Hacker Newsnew | past | comments | ask | show | jobs | submit | GoToRO's commentslogin

It will be just in time for when AI will run very well on consumer hardware and the need for data centers will collapse.

There will be demand for both

Some demand. The question is if the capacity will meet it, be under or be over.

What if the AI companies start selling the hardware with their models?

I.e., you get a locked down device with access to their AI and maybe a way to run third party apps. Of course the developers of those apps will have to pay a percentage of their revenue.

Sound familiar?


More likely they would sell a chip with the model etched in silicon for the dramatically higher token/s.

A chip that can only be accessed through their software. With a monthly subscription fee, and other enshittification surprises.

Most likely those models would be extracted from the hardware in no time.

I suspect no, because most of the tokens would be used for its internal reasoning loop, and a cheaper model could be used to obfuscate the output.

That's never going to happen. By the time you can run current frontier models on your $10k desktop the frontier will have massively advanced and people will want those models instead.

I’m not so sure about that. Already AI vendors are back to cutting prices to try and keep customers from cutting back on their usage. My own employer is working hard at pivoting to much smaller fine-tuned models for established use cases, and seeing model performance improvement in addition to large inference cost reductions. Being able to run them locally hasn’t exactly been a disaster for devex, either.

It may turn out that demand for SOTA frontier models isn’t so limitless after all.


Every product follows demand curves. At a price of 0 you could find infinite usage. This has nearly zero relation to how much it costs to provide the product.

>At a price of 0 you could find infinite usage

Infinite demand isn't a thing. Even if they'd offer free compute forever (not likely possible) I and many others would still use local models that we have full control over and that do not harvest our personal data.


Except of course it relates. All else being equal, we will prefer $X COGS over $2X COGS because that helps us with both profit margins and price competition.

It relates in the sense there's a minimum cost of production without losses, not the actual price people are willing to pay.

Framing it in terms of the price people might be willing to pay for a single product in isolation frames the point I was making, which was about price competition, right out of the picture.

Maybe I'd be willing to pay $10 for product A if I had other options. But if there's a product B for $3 that's not quite as nice but still ticks all my boxes, then product instantly becomes a lot less attractive.


This + your ROI in $10k device will be always lower than busy datacenter, its literally math. They sell free compute to others when you dont use it, you will never sell at that level or even you magically sell home compute, you will not compete at price

That greater efficiency only benefits the LLM SaaS providers as long as the hardware manufacturers, probably especially the VRAM manufacturers, remain supply constrained, since the high-efficiency users are the ones who can pay top dollar. But the hardware guys' dream is presumably to get parts into millions of laptops which remain on standby for 19 hours a day, not to bargain with SaaS providers who obsessively optimise their memory consumption.

I'm not sure if this prediction will hold true.

We're not seeing the progress in those "frontier models" that we have previously seen. There's certainly still gas left in tank tank, but we're way into the diminishing returns by now.

Cloud inference still beats hardware investments by orders of magnitude of course, but that's only if your data doesn't really matter to you.


We are certainly not in the diminishing returns phase for LLM progress. No sign of that yet.

I’ll grant that for specialized applications like coding agents and mathematics, but even there I suspect that most the real gains are actually taking place in the harness.

But I suspect returns may have already diminished into negative territory for at least some other use cases. One of my least favorite job responsibilities in this brave new era is figuring out how to avoid performance and behavior regressions when an older model were using for some application reaches end of life. It’s getting uncommon for me to look at our benchmark results and say, “Oh, good, it does better on one of the newer models!”


>suspect that most the real gains are actually taking place in the harness.

Part of the reason harnesses work well is you can run a lot of agents in parallel. That doesn't slow down demand.


That is true, but the eventual realization that more machines doing more coin flips in parallel does not mean "more work gets done" might.

LLMs are amazing tech, but they're terrible without oversight. More agents faster just makes reality collapse on them quicker.

But yeah, you're right, temporarily, this will still push demand. But the topic was about "diminishing returns" as in "tech getting better". Not as in "customer spending".


It's kind of weird because more machines working together does mean more work gets done. Coin flips and weighted coin flips are totally different things. Any biases weights towards reality push you closer to reality when you use them.

New models keep being able to use more and more agents on longer time frames. Your hypothesis doesn't look like what we're measuring.


Who is we?

The people mapping AI capabilities.

Oh cool, so that we includes me! :)

Maybe turn on your light when you use a ruler? Not sure what else to say.

But high demand for LLM time isn't sufficient to keep customers at the frontier LLM SaaS providers. That demand can be satisfied locally or at non-frontier outlets, absent hardware shortages at least. The Tier 1 providers (and the would-be Tier 1s) presumably need to open up a much bigger lead in model quality, one that doesn't simply get distilled away this time, and/or continue to be protected by ongoing (or worsening!) hardware shortages. (And that's overlooking the revenue shortfalls which OpenAI and Anthropic seem to be facing already.)

I had actually been thinking more about all the non-LLM functionality that go into the harnesses. I'm not going to name names and I haven't done any rigorous testing, but my general impression is that choice of harness matters more than choice of model. In terms of basic task completion success specifically, not code aesthetics.

A perfect harness will not extract gold from a dumb model. It's a system that builds on each other, though we've not probed that frontier much to have a good intuition on what effects what.

One thing that I really want to know - the better models from today vs a year ago - what has changed. They have already pre-trained on all available public data. Scooping up the last percentage of archaic texts which were never digitized is not going to move the needle.

Is it just that the providers are generating tons of synthetic datasets on coding tasks so that the models get more exposure to the right thing to do? Every time someone points out an LLM stupidity they add some training data to patch over the weakness (trivial to generate "there are two 'l's in llama")?


Google scholar has a flood of papers showing LLM diminishing returns on pretty much every facet

https://scholar.google.com/scholar?hl=en&as_sdt=0%2C23&q=llm...


The first title I see there is "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs":

https://proceedings.iclr.cc/paper_files/paper/2026/hash/3b4e...


And you apparently only read the title and ignored the content.

It’s also poor reasoning to you only pick one thing you think supports your view, and ignore vastly more things not supporting it, all from the same useful criteria.

Now if only you’d carefully read the report you chose, and spend equal time looking at ample presented evidence, you’d develop a more accurate understanding.


Well I mean if I wanted to be extra pedantic, I would argue that we've been in that phase since LLMs were first introduced.

Before that, we had 0. After that, we had more than 1.

A leap as far as that is hard to recreate.

But that wasn't my point. That's just trolling.

The actual point is that LLMs aren't gaining new capabilities anymore. They just get more reliable at the ones they already have; turning what was a coin flip to some higher probability.

That's (intuitively speaking, not strictly mathematically speaking) kinda the mathematical definition of diminishing returns.


It’s a constant tension in computing that has been around since mainframes and clients… Neither is going to disappear. My general feeling is normal people care more about how thin and light something is than their privacy, so if data center powered LLMs will have a strong future.

Hmm I'm not 100% sure about that, given that edge is very viable, and the geopolitical climate has changed quite significantly.

I agree that datacenters are not going to go away, but I have doubts that the buildup that has happened is really going to pay off for most operators.


they really dont want to hear this bro lol

I can see that by those reddit-style vote swings, but who are "they", exactly?

Who is so emotionally invested into random comment sections being purely positive about their pet.. uuuuuuuh.. tech?

Very weird.


On the contrary I would love it if AI stagnates. I don't want to be out of a job.

But I also don't believe things just because I want them to be true.


It's going to happen very soon, which is why these frontier labs are scrambling to shut down open source language models. There's an existential risk threatening their obscene returns.

It is for that reason they are being archived and torrented as a very large middle finger.

I feel like cost competitiveness of local has been going down, if anything, not up. API providers can use hardware more and have scale efficiencies. Do you see any reason this will reverse?

> It's going to happen very soon

Why? You can't just assert it. There are very good reasons to think it won't happen soon, and you've given no reasons to think it will happen soon.


Because everything is converging on a backlog of huge efficiency gains established in research, waiting to be combined. Looped transformers, a whole host of diffusion techniques and new quantization techniques, maturation of ternary distillation and new ways to separate logic from stuff that can be looked up. It would surprise me if most frontier models were actually even that big at that point in terms of active params. I highly doubt it.

You asserted that it was never going to happen first

But he gave a reason for that. "Before that happens the frontier will move". Why do you think it will happen anyway? Do you think the frontier will not move fast enough that local models are unable to catch up, or do you think people will prefer local models at a point. Or something else?

You don’t need frontier-level performance for every task. That’s why companies hire both junior and senior developers. I suspect a decent percentage of people using Fable would probably be fine with Opus.

Also, the frontier can’t keep advancing at this rate forever. Eventually the low-hanging fruit will all be gone and advances will slow down.


Or even more amusing in time for the bubble to burst, I hope the big RAM makers end up holding the whole bag for that. Greedy bastards.

Nearly the entire reason we're in this mess is because RAM/SSD/HDD makers are terrified of AI bubble bursting.

Long-term contracts, not building new fabs - they're in full on hedging mode right now.


Seemingly the AI bulls are either not sufficiently optimistic, or not sufficiently wealthy(?!), to build new fabs as joint ventures with the manufacturers in which they agree to assume most of the downside risk?

They signed long term contracts for the hardware delivery over years. AI investment is approaching 1 trillion per year and expected to grow to well over 1 trillion per year. [1]

Even over their many years, projects like the Manhattan Project, the Apollo Program, or the U.S. Interstate Highway System never added up to that. [2]

Is your argument that there is insufficient AI spending at present as they aren’t also taking on building their own fabs?

[1] https://www.pwc.com/gx/en/news-room/press-releases/2026/glob...

[2] https://www.aljazeera.com/news/2026/2/19/visualising-ai-spen...


If the long term contact are with any AI companies, then good luck getting anything back in bankrupcy proceedings!

If the AI bubble bursts, the DRAM manufacturers switch back to making conventional DRAM and pent-up business and consumer demand results in a surge of purchases, cushioning the fall considerably.

Look into the time frames a bit more, this isn’t a switch they flip. If the bubble bursts today it will be years before those downstream effects add up.

Looking forward to see them burn.

Since you’re not greedy when there is a memory glut I’m sure you’ll be willing to pay extra to make it fair.

Not at all, I’m going to relax with a cooling drink and watch the consequences of this unbelievably destructive and wasteful venture implode. I also adore the idea that it's on the customer to be "fair" to a company that dumped us a group in favor of chasing B2B money.

Their choice, their consequences.


I think they're talking about the past: memory has often been boom/bust. Memory manufacturers lose money during gluts, and did so long before AI. People benefited from cheap prices at that time.

Internet is a fad, it will implode at any moment.

In fact as a business opportunity it did just that because the massive investment in the mid-late 1990's was without anything like a connection to profitability or sustainable demand. It took years after that crash for the concept of e-commerce to begin both a recovery and evolution into what we see today.

I expect something similar for AI. It's useful tech... just not very profitable tech, and certainly not to the tune of trillions of dollars worth of public demand. The promises of superintelligence, replacing everyone, and the rest of the hype will die with the companies who made the promises, but the tech will survive and thrive.


>The promises of superintelligence, replacing everyone, and the rest of the hype will die with the companies who made the promises

The hype will die but there are many reasons why super intelligence and robots are a separate entity from said hype.

Look, back a few decades and tell someone our gdp would be in the trillions and it's likely they'd have a hard time believing you. A huge portion of our products that we use day to day would be complete science fiction to them.

All we're negotiating at this point is the timescale it will take.


Unpredictable things happen routinely, therefore these specific unpredictable things will happen?

These things aren't really unpredictable. Work long enough on a robot with more dexterity and you will get just that. The timetable on when it happens is a bit more up in the air, economic downturns and wars can have huge impacts on it.

But in my mind robots and superintelligence are inevitable unless there is a massive setback in humanity before then. Nature already did it once. We're not inventing something totally new. Add in every new invention and bit of intelligence we build on and actualize pushes us that much closer to the goal. Information technology allows this to speed up even further with easy sharing of information and testing.

Superintelligence is not like faster than light travel. We have many frameworks that show us FTL is impossible. I don't believe there is a single widely accepted framework that shows there is some limit to intelligence and we're near it. Intelligence isn't even a singular thing. My calculator is a super intelligent adder compared to me. It would seem highly improbable that somehow nature random walked into making brains the most efficient general intelligence device in all dimensions and scales.


Autism looks to me similar to what a guy had when we were children. His mother was known for drinking even durring pregnancy. I'm pretty sure that we are disrupting some development mechanism today with all the substances we consume, even without knowing. Food, cosmetics, clothes, the air and so on. I think this onslaught on our bodies is what disrupts normal body functioning, in unpredictable ways. Combine that with pregnancy and the disruptions get written in "hardware". So no real cure, maybe some improvements. At government levels I think this is known, but they don't know what to do about it. Giving people a clean environment again, will have to undo most of the "progress" we got. Anyway, this is just my hunch but at individual level is better not to wait for 100% scientific confirmation.

Japan ambassador in Romania said just that in an interview.

Russia and the US seek to expand by invasion, at the same time there are NINE countries with a stated ambition to join the European Union, and now there seem to be even more countries who want to "more closely align" with the EU (whatever that means).

Whenever countries consider their needs and desires, and what they want their country to be, they look at the countries of Europe - not at Russia, not at the US, and not at China.


I don't like those headlights also. I think it's a case of when founders bring too much of their personality into the product. Otherwise great product and this self drive system at least has a sound start (track objects across video frames)

Why look foolish if you don't have to?

"You are right, that was a red light!"

At first I was skeptical, but reading the article, they do track object across video frames which for me it's the minimum a system like this should do. All previous FSD systems are just Fake self-driving.

"That proprietary AI driver identifies features in images and point clouds, groups them into objects, and tracks them across frames, time-stamped to the millisecond to account for differing frame rates. The AI thus builds confidence over time, acting on object detections that persist across several frames, rather than, say, slamming the brakes due to a camera blip on a single frame."


>they do track object across video frames which for me it's the minimum a system like this should do. All previous FSD systems are just Fake self-driving.

I'm pretty sure Waymo does this. It's hard to imagine not using this information, it's basically 50% of what I do when driving (e.g. that kid that was running on the sidewalk suddenly disappeared from view in an unexpected way; I wonder if he's about to cut across the street, etc).


I wander what it does for different objects.

Oups a ball crossed the street 15m ahead. Let's slowdown to a crawl and watch for the kid is what a human would do.

I suppose they may have special cased the ball case.


> The AI thus builds confidence over time, acting on object detections that persist across several frames, rather than, say, slamming the brakes due to a camera blip on a single frame.

Nice in theory but in practice they still need a ton of training data. The bitter lesson is very bitter. Tesla's end to end FSD uses occupancy networks and brake stabbing still hasn't been fully solved.


This is very basic for any self driving car system you can see on the road today or even 10 years ago.

> slamming the brakes due to a camera blip on a single frame

Ironically, I'm sure most humans have done the eyeball equivalent a handful of times, despite our knowledge of object permeance.


Where did he said that? A new experiment can be done, responsibly, right?

The difficulty in getting to an agreement on what "responsible" means, and on how a "failure" should be handled means that said experiment will never occur. Nothing is ever safe enough and no failure is acceptable -- so no test which can fail can be run.

A situation which would be in obvious conflict (negative expected value) with the goal of saving the man hours and lives currently being cost by manual driving.


Add facebook too. No problem for them to accept money from websites that were reported as fraud. They pause it a little then they are right back at it.

This is not a game. Choosing from wrong answers only is not a game.

First you introduce the tax, then you increase it. Minimum public anger. They messed up because it's a recycling tax that I presume does not apply to... everything? So it feels punitive. Also it's big. I pay way less recycling tax on bigger and harder to recycle items.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: