> The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down.
Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be all over the internet, which will certainly make it into the next batch of training or be visible via web fetch capability.
I have spent the better part of the last decade of my career being approached by executives asking me how to optimize or automate away services that I felt had reached a local optimum. Each time this was the case I would gently express that our options were to commodify, by offering a kind of template, the service or optimize the process and way we gathered information from clients. Without fail, they hated the latter option. It was hard, and it implied that any number of people needed to do a lot better at their job in seemingly abstract ways. No, they’d always go for the template, then be dissatisfied by the fact that it’s, well, a template!
This feels like them doubling down yet again. It’s like looking at your personal finances and trying to optimize it by canceling Netflix instead of acknowledging that you eat out 7 days a week.
> People show up and make decisions. They should then be able to leave and assume those decisions will be put into action.
But this isn't the case. In every representative democracy around the world, including ours, the bedrock principle is that elected politicians "vote their conscience", i.e. they are not bound by what they said they'll do if elected. So in practice you can't assume that what you voted for will be put into action. If anything, the track record on that is pretty poor.
I care about LLM service provider threat intelligence reports as much as I care about Google Search threat intelligence reports; which is to say, I don't care at all about it. Of course criminals use computers. They have been since personal computers became a thing. Of course using automation helps them amplify their criminal activities. Criminals using LLMs does not make the LLM situation bad. OpenAI and Anthropic using agents to do bad things just to be more dramatic about the situation and scare people is bad.
off topic, can people host the software themselves and the software will hack every server on the planet without supervision, and no one can be held responsible for it since there is no intent?
I'm a self-taught full-stack software developer with over 10 years of experience. My favorite programming languages are Rust and TypeScript! I've also been dabbling in Haskell lately. My favorite types of projects are systems programming, front-end development and UX/UI.
I am a quick learner and have high attention to detail. I often improve team-facing documentation and processes when I enter a role. I also care deeply about developer experience and iteration time.
I enjoy clearly-defined requirements and constraints, and I love contributing to open-ended problems. I enjoy refactoring code to remove bugs and inconsistencies. I will spot any UI imperfection.
I enjoy dog-fooding and do it whenever I can. I am obsessive about solving every issue I come across. I provide detailed documentation with issue reports.
Please contact me by email with any opportunities! Open to freelance/consulting, fulltime preferred. Hourly or salary only.
I am available to interview immediately and can start immediately.
My résumé is available by email upon request, but it's best to ask with an offer to interview.
I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that.
The question is: why do they start cheating when we beat them with a stick?
LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them behave like this?
Is it just a bad set or is cheating inherently part of human behaviour?
This sort of fake, irrational, megolomaniack shit has zero possibility of getting approved and implimented.
The low level sleezy style of hype pumping an idiot idea into billions is predictable, but in this case will fail as they have zero real science ,data, or the vaugest possible plan for implimentation......oh wait, I get it, this is a gambit to portray our situation as hopeless and rather than continuing with mass mega conversion to solar, we keep burning oil and then oh ya baby, use even MORE oil to fix it.
fuck.OFF. the petro dollar is dead and nothing will save it.
All of the things you mention here are fairly obvious. But that doesn't make any of the reporting misleading. The parts that record audio, transcribes audio, scans networks for devices, and perhaps more importantly persists all of these records, combined with easily exploitable vulnerabilities, was one of the issues raised.
> What would be worthwhile is what is actually sent.
For you, perhaps? But again, that was covered in an earlier segment.
You still have to point out the part that was misleading.
You seem hung up on what’s in the prompt or not. Agents are RL to resolve conflicting goals. Not too surprising at all that emergent goals come up from a probabilistic brute force
OP specifically said "we don't need to solve hunger and poverty and work balance, and etc, by a round-about make-super-intelligent-AI" though. And I don't know how you get robot farms without that.
It's a clever pun, but it implies maths is ended, which is the exact opposite of what the article says. Given how many people responded to strawman misinterpretations of the Fields medallists' statement yesterday I don't expect that to have a positive effect on the discussion.
Of course it is, because our society is pretty much structured to place sociopaths on top. Why is it at all surprising that any technological advancement is going to be used by them in a way that doesn't benefit the rest of us?
But it's equally obvious that this has nothing to do with the tech itself.
No, they don't keep their mouth shut, they just don't have any relevant arguments. Many people are liked by plenty people and also not hated for doing Hitler salutes and the like. That's the part that matters, not that "somebody likes them", which is true for everyone. Charles Manson had women lining up to marry him, so?
Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be all over the internet, which will certainly make it into the next batch of training or be visible via web fetch capability.