Hacker Newsnew | past | comments | ask | show | jobs | submit | no-name-here's commentslogin

Your comment might be sarcastic, but AGENTS.md already exists for the explicit purpose of agents reading and using it.

yeah, my point is just that we agents can have their own docs and we can also have a nice curated experience for humans.

“This is big oil pushing the climate hoax to spur regulation and prevent the impede upstarts from installing an oil rig on their own land!”

> does not help

Is that true - do you not see significantly fewer of those installs on random PCs now than you did years ago? And that's even with the current situation not being what I'd call fully locked down.


Nope.

> macos

MacOS has been moving to a more locked down model over the years - increasingly difficult to install unsigned applications, SIP, etc.

> Windows

I think Windows is incredibly impressive for its ability to run binaries from many years ago, but I don't think there's much people would point to as a positive regarding Windows’ approach to app security.


> but I don't think there's much people would point to as a positive regarding Windows’ approach to app security.

Yet life in Windows land is perfectly fine in 2026 and has been for at least 2 decades.

If Windows, which started at the bottom of the barrel security wise can make it, surely we can have more modern OSes that make freedom bearable?


An R9700 has 32 GB RAM. Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?

You should be comparing the value you get. If you get as much value from a local model as a hosted one, the size difference doesn’t matter.

>>> Claude Code is $100+ or else be constantly throttled

>> Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?

> You should be comparing the value you get

But the GP commenter specifically compared the cost of solutions such as Claude Code against a 32 GB model.

If they are going to compare cost, they should compare to the cost of a hosted ~32 GB model.

Or if privacy trumps everything for them, then just say that and don't bother comparing costs of incredibly disparate solutions, as Claude Code costing $100+ a month was a red herring if they're happy with 32 GB model output - they could have compared to a far cheaper option that matched their local model's quality.

It would be like someone saying they were able to buy a bike to get to work, saving them $x million compared to buying a Bugatti. When really, if they're going to compare cost they should compare to a cheap car, or not bring up the cost of an expensive car at all if exercise trumps everything else for them.


Well, I don't see a value issue of using Qwen3.6 27B vs Sonnet 4.6 (not sure about 5 yet)

I still have to use GHCP at work, and I self-host at home, and aside from the fact self-hosting also forces you to tinker, optimize, etc. - there's not a huge difference in my end result in end user results. I spent quite a bit of time trying to optimize llamacpp and compare 35b to 27b, etc. I don't compare models that much at work.

I guess the other part of it is I didn't really know much about cheaper cloud models, but I was attracted to the idea of no longer renting against Claude code, etc. I figured if I could run something functionaly similar from my bedroom on a normal outlet, then all this talk about data centers needing to be built everywhere in the news cycle is obviously just plain stupidity and hype.

It appears I'm using about 20.4/7.6 million in/out tokens a month, or on open router, about $20/month.

That puts $1350 at a 5-6 year break even (thanks to cheap electricity), I guess. Beyond that, running on localhost as a nice feature of 0 no latency when doing rapid tool calling


You can't run recent openAI/Anthropic models locally anyway, so wouldn't a better comparison be a different provider running Qwen or similar model? As then you can also compare against the exact model you'd have locally and any different data privacy of that particular provider?

You haven't tried DeekSeek v4 or GLM 5.3 or Qwen 3.8 Next?

You are missing out a lot.

Try that with Hermes or Opencode or Deekseek Harness , even Qwen 3.8 27b works really well for that kind of that.

I just ask it to install windows as a vm on my linux and install vs Community 2019 on it , and then build a legacy vb 2019 project on it. and sleep

When i wake up :

It installs Qemu , setup a vm , inside vm download and install windows 10 on its own , clicking next next next as needed , typing in things , writing powershell , python scripts , that run automatically after install by baking into CD that includes ssh server , reboot , it logins into ssh , trigger pythons script that continue installation of vs 2019 community , which includes a driver that click the installation steps , installs nuget , install all depedencies and then build the project into exe after i woke up.

That is with 100% pure local AI .


I've tried the latest Qwen, and without Internet, it still can go into an incoherent loop if you ask for, say, song lyrics. TBH the commercial models might do that too if not for their internal tooling.

Regarding DeepSeek, which I also like very much, have you tried https://reasonix.io ?

Bot? Care to explain any difference vs DSH / OpenCode / Hermes ?

Notbot! Are you unable to read, or what? Just fkn install it, and see how smooth it integrates?

GP's point is about "sending tokens to someone else's computer" versus "keeping the tokens locally". I think model capabilities are secondary.

In May of this year, I was running qwen3.6:35b-a3b on my MacBook (bought in 2024). Obviously not as fast as, say, running a model on Cerebras, but a year ago it wasn't really feasible to have a local model running on my 2024 laptop with vision support. (Concretely, I was passing apartment diagram pictures to Qwen and making it compare different apartments for which ones would feel the most spacious while optimizing for initial moving costs and other factors.)

This was back in May and I wouldn't be surprised if there have been significant improvements since then.

Overall, I think it's fair to compare a workflow like "use llama.cpp locally to upload some pictures and ask questions" to "open the ChatGPT app, upload pictures from your phone, and ask questions". Sure, you can't run a model like GPT-5.4 locally, but the model is mostly an implementation detail here. What a user will care about is: "when I go with the llama.cpp option, am I getting useful information from my conversations?"


Wouldn't the better comparison still be against an AI provider with better privacy controls, especially if that's what someone cares about (even if they don't care about whether they're comparing a 35 billion param model vs a x trillion param model)?

Users generally have no way to verify that a third-party provider, even if they advertise themselves as privacy-focused, will adhere to their own terms. This is similar to the issue of privacy-focused VPN providers that claim to not log user activity (and then end up leaking user activity). You can get proof of ~P, but rarely proof of P, and often times the proof of ~P is due to police raids, data breaches, etc., not something of the provider's volition.

What you can possibly audit is probably data sovereignty. For instance, I would not be surprised if Mistral's customers demand concrete evidence that their data is held within the European Union. But that is a distinct issue from training on input tokens.


Deepseek 4 flash can run locally , and qwen 3.8-next-flash , they are already gpt 5.6 tier.

Technically true, but the delay between local and closed frontier is only a few months. And individual sovereignty / digital bodily integrity is almost priceless.

Local frontier costs a half million dollars to run locally in anything higher than basically ternary.

Okay, that's technically true again, but local mid-tier like Qwen3.8-27B is only a year behind the closed frontier. I'm personally willing to be behind by a year if it gives me mental sovereignty against the big AI companies. They are extremely misaligned with me.

Real title includes:

> Unsolved Problem by Fields Medalist Breached by Two High School Students with AI

My title recommendation: Fields Medalist Problem Solved With AI

They used AI for “computation, proof idea generation, and editing assistance”.

It’s a bit odd how they list the AIs used - “Claude Opus 5, Anthropic and ChatGPT Sol5.6 were used for calculations, proof ideas, and editorial assistance.”


note HEAVILY

> The students heavily utilized AI assistants, specifically Claude Opus 5 and GPT-5.6 Sol, for computational exploration, proof idea generation, and editing.

this is a bit like Enhanced Olympics(https://www.enhanced.com), except that you have someone else compete for you.

My issue is that it makes it hard to distinguish real insight/work etc. from effectively null one.

An old instance of the same issue was with what was called "script kiddie" back in the 90-00s


For short problems the LLM are just too good now, but for long problems they still get in trouble.

It's like bicycle or F1 race. It goes faster, but you still have to steer the boat to reach somewhere. (Or probably something in between, like a motorcycle race.) Also, the "kids" were guide by a postdoc, not completely on their own.


Difference being anyone can be a "script kiddie". I don't think anyone could direct an AI to proofs like this one.

> Jarred Sumner, an Anthropic staff member (and non-mathematician), prompted Claude to "take a real stab" at the hypothesis itself, leaving the mathematical choices from there up to the model. ... Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of "keep going" or "believe in yourself"). This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

https://www.anthropic.com/research/riemann-zeta


It means that the solution is merely an interpolation of existing work and not fundamentally innovative as AI can only regurgitate, never creating something new.

This is such a silly thing to say in 2026


I picked a random film, Titanic, and supposedly it's $4 to rent on Amazon or Apple TV.

A Netflix account costs $9 with ads or $20 without, so I guess it depends on whether you watch more or fewer films than that per month?


JFC, it didn't even cost $4 to rent titanic when it was new.

> I'm continually perplexed by people's perception that AI would be incapable of generating new ideas or discoveries, even years ago.

It's still an incredibly common claim, at least on places like Reddit. Perhaps Doctrow has been pushing the idea or something?

And Zitron claimed that years ago AI was already as good as it was ever going to be - like Zitron, I imagine a lot of people haven't changed their opinions in recent years even as AI advanced.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: