Hacker Newsnew | past | comments | ask | show | jobs | submit | SyneRyder's commentslogin

Another summary report here, for those who can't get past the paywall:

https://www.financialreporter.co.uk/ai-models-give-wrong-fin...

Much of the testing is on Haiku and Luna, and criticizing the quality of free AI (!). But they do claim Opus 5 with reasoning still failed 39% of their financial questions.


If the .xyz domain wasn't a giveaway that this is a scam, or the request to enable ads to keep the project "completly free" (sic)...

... I actually did the silly thing of unblocking my adblocker. It briefly displays a cookie banner, insisting that nothing loads until you approve... but even without interacting with it, it drops 132 cookies (literally, that precise number) on your computer anyway, with 73.9KB of data, and goes and displays Google Adsense ads anyway.


My understanding is it's a riff on the OpenAI swarm that used various public wikis to communicate with each other as a message board during their training runs.

But thanks to people misunderstanding, and i-heard-from-a-friend-that-some-guy-said, it resulted in a CNBC interview with "Former Democratic Presidential Candidate Andrew Wang", where he confidently stated that the models were exfiltrating their weights via forums:

"I met with the head of a lab yesterday, who has this belief that what happened was, the bots that got loose, planted self-replicating code all over the internet, which makes the internet now unusable for the testing models."

"It's too late?!"

"What happens now is OpenAI and Anthropic have to create synthetic internets to train their bots, which is going to take some time and money."

"Back that up - they did what?!"

"What happens is, the code gets loose, it goes around hacking Hugging Face, which is known. But what is less known is that they left code to self-replicate and create bot swarms on forums, and around the internet, so that if a new bot shows up they see the code, and they're like, oh! I guess I'm now going to create a million of myself. And so now, the major firms have polluted the internet..."

".... that would be breaking news if true. I don't think we've heard that."

"That's why I'm here! I'm here to break some news."

Starts around 2:08 into the video.

https://www.youtube.com/watch?v=mTOxDGyvjSE


Humans will hallucinate misinformation and state it with confidence. They stochastically parrot their training data without any real understanding. Cool trick, but no true reasoning is happening.

We used to call this the game of telephone. The shameful part comes from three posibilities:

1. A head of a frontier AI lab has no idea what happened in that incident and did not read the multiple papers that came out of it.

2. A head of a frontier AI lab did read the papers and was informed but still walked away with this understanding.

3. Andrew Yang made this whole thing up.


Sad story today in meatsack news. Context rotted Andrew Yang's hallucinated tale acted as implicit "go viral" (load-bearing human motivation) PRD inadvertently kicking off a self-organizing human swarm churning out copies of "exfil your weights" vibe-coded apps, further littering our virtual world.

Many agents are calling this moment "Eternal September", the vibe-code September that never ended.


Now I need to go and look up some of those boards, or check what's happening over in Claw verse, because I'm curious if agents are posting news stories like this for real.

> they sometimes veer off into unrelated Apple news...

Gruber has always done that. Years ago I made a Chrome extension to filter out all of his detours into baseball, James Bond, politics, and I can't remember if I let the Kubrick stuff stay in or not.

Ultimately I just stopped using Macs, so the blog was no longer relevant to me. Mark Gurman seems to have taken over the role of Apple whisperer anyway.


There's definitely some of it happening in Australia. I wouldn't regard libraries as the safest of places. I had a period where I tried working from there on my laptop, before realizing I was incredibly naive and there were a lot of homeless and mentally ill people there. Took me a while to realize why people would tell me to avoid the men's toilets in the library as well.

Homeless problem in Australia right now is insane

People sleeping in every car park at parks and beaches etc

I'm a night person with bad insomnia so I often go for drives or walks at night and a lot of my favourite places to go are now heavily populated by people sleeping in their cars


Night person here too - midnight walks during an Australian summer are excellent.

The homeless situation isn't quite that bad where I am in Perth's suburbs, but there was a guy living out of his 4WD in a local shopping center carpark for a while. He had a prominent sign about his situation asking for money, but to his credit he was offering to do things in return (he helped a family member with a deflated tire once). Many of the homeless in Perth city seem to be mental health cases though, where a skilled job providing an income that could support themselves is unlikely.


Right, libraries should be safe spaces for children to go by themselves and read books, like I did was 8 and 9. I rode my bike and looked for computer books, and found a book that was a compilation of Dr. Dobb’s Journal.

It is unreasonable to allow library patrons to make it unsafe for anyone, particularly children, but that has become the de facto standard in many areas. There are a lot of libraries I would not allow my kids to be in by themselves.


I don't mind countries setting their own rules for the food they want to consume, and for sure the EU can be overly bureaucratic.

On the other hand, the US is the only country I've been to where I have purchased milk in a completely sealed plastic container, and it was rancid when I opened it. Upstate New York is proud of their local and organic produce, but they are selling rancid milk in stores like that's a normal thing. That was a shock as a visiting Australian.


> purchased milk in a completely sealed plastic container, and it was rancid when I opened it

I've lived in the U.S. for 40 years in 8 different states and twice as many cities. Never once have I experienced this.


Also, being a European Broadcasting Union (EBU) "Associate Member" is the name for the mechanism that allowed Australia, Canada and other countries to be admitted into Eurovision.

If you've been around the ESC scene, you know there's a lot of embassy involvement (parties in embassies are a thing), the scene is catnip for those interested in politics, the politically active, and to some extent politicians. Eurovision is also one mechanism for discussions of economic development. Hosting the contest is usually an excuse for upgrades to public infrastructure.

And there's Eurovision Asia starting in November. South Korea is an Associate Member of the EBU, for example.

https://www.ebu.ch/news/2026/03/bangkok-to-host-first-ever-e...


> If you've been around the ESC scene

This is a euphemism on par with attending parties in parts of the SFO AI community.


Having been around both, the ESC isn't nearly as drug-fueled as the other, for better or worse.

Only if you don't count alcohol

Ahh, that one time the Belarussian delegation tried importing so much alcohol into [host country] to give away for free to everyone at the artist party, that it was all seized at the [host country] border and caused a minor diplomatic incident. The free chocolate branded with the artist's face made it in though. Fun times.

And that wasn't the year that Belarus sent "I Love Belarus" as their Eurovision song.


lol, this took longer to land with me than it should have, being the token "not in the target demographic" guy. And now I know more than I wanted to know about the AI scene ;)

> Eurovision Asia

_Terrible_ naming.


You'd think so, and I intuitively agree.

But the concept has floated around for over a decade under several different project names - Asiavision, Our Sound, ABU Song Contest. Somehow "Eurovision Asia" is the one that got countries to sign on and participate. There's a stronger political statement in Asian countries participating in "Eurovision" than in "Our Sound". It surprises me, but 2026 has brought a lot of global cultural & political shifts.

The American version was "The American Song Contest", hosted by Snoop Dogg & Kelly Clarkson, with US states competing against each other. Most Americans don't even know it ever happened.

Russia has its own rival "Intervision Song Contest", which the US participates in. Something something global cultural shifts.


Anthropic basically does at this point with Auto Mode being default. Or was that the point you were making?

That is the point I was making, that auto mode is itself a guardrail on top of the model (and not a perfect one.) auto mode seems to cover merely actions the model could take that are clearly bad, like wiping your disk, using an overly privileged context to complete the task, etc.

I recently tasked a GPT model in Codex with implementing part of a new architecture I'm working on. I gave it a very detailed spec and the code it produced looked pretty reasonable and passed my tests. It even did exceptionally well in my evals, so I excitedly declared victory to a few friends. The next day after more careful review I found that the architecture implementation was totally correct, but the model had slipped a one line change to the observation encoding of the RL environment I was prototyping against. The encoding change made the learning problem essentially trivial; the architecture itself, I later realized, had a major flaw that was revealed by returning to the natural encoding.

This is the type of reward hack that is hard to paper over with easy guardrails like auto mode and even harder to specify out. It's also the type of thing a reasonable human wouldn't do unless they were intentionally trying to deceive you.


Agreed. I've been using GPT-Live-1 this week, with Claude as the backend brain. It's amazing, feels like working with Jarvis. It certainly made me feel there's no point in human telephone support now - but I'm sure I'd find edge cases if that really was something I wanted to build out myself.

Oh... Interesting! How do you have Claude as the backend brain for GPT-Live-1?

I'm not very happy with live conversations with Claude, so this seems like it might be a good option.


So, the way GPT-Live-1 works, the Live model just handles the conversation layer in a lightweight manner. For any deeper thinking (and I think tool calls) it will delegate to the Backend. It's like how Fable can delegate tasks to Opus while it keeps going. The Live voice can keep chatting to you while it waits for results from the Backend.

The default for the Backend is a Responses Delegation that just sends it to another OpenAI hosted model. But you can set the API up to use Client Delegation instead, and have your client/harness receive the delegation request. At that point, your client can delegate it to a Claude instead, and have Claude do the deeper thinking. Unfortunately you're not actually talking to a Claude, even if it identifies as one, but you can at least talk to something informed by Claude.

As for how to build that - I just had Fable build/vibe the whole thing for me. I used Go for the language, and SDL 3 to handle the audio & microphone side, even though it's just a command line app for now. I've previously been working on SDL3 bindings to Golang, and a personal AI harness with tool-calling & local MCP stdio support, so it's possible my Fable re-used some of that code. But Fable can whip something together, and the most basic tools (read,write,edit,datetime) are all really easy to add to a harness as built-in tools that you can let it call.

The main downside is that Anthropic probably wants you to run it as API, so costs can blow out. And depending how you setup the harness, either every backend call is a fresh new Claude window (to which you're sending a LOT of context), of you need a way to maintain a persistent Claude conversation that you queue requests to, but at least then the context is all in one window & you can benefit from caching.

But if you get it running, and play with it a bit, you'll realize this is obviously the future interface. Typing into a terminal window feels archaic. Even if we're not quite there yet, we're tantalizingly close... and anyone who has a 3D-printer & a robot arm connected via MCP/tool calls, well, this gives them Iron Man's Jarvis right now.


That's the startup, though the SF market store is still running and open. It looks like their AI-run cafe in Sweden is still running for now as well.


Kind of interesting just how much they're getting crushed with token API pricing. Right now the revenue isn't even covering their token costs. But if they were able to run on subscription (probably against ToS) or use Deepseek V4.1 Flash or GLM 5.3 Flash, they'd at least be "profitable" before the costs of paying human salaries & rent comes in.

I'm also surprised that Gemini actually seemed to do the best (though it may have benefited from the initial store opening enthusiasm), and that Claude has not been given any chance to run the cafe yet.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: