Hacker Newsnew | past | comments | ask | show | jobs | submit | derektank's commentslogin

So, does anyone have any intuition for how concerned we should be that one of leading foundation models will be able to successfully attack the gold standard symmetric and public key encryption algorithms (AES, ChaCha, ECDH, Kyber, etc.) in the next few years? As a consumer of crypto that doesn’t understand the mathematics deeply, I’m getting kind of nervous that we’re going to wake up one day to find that the backbone of TLS has been shattered.

It's unlikely that any of the modern cryptographic primitives will break over night.

First, modern encryption isn't susceptible to "this one weird trick!" like the early days. ChaCha isn't even a cipher. It's a key stretcher. Which means, even if you broke the math behind ChaCha, its inherent complexity means its still widely dispersing the original key across the cipherstream. There just won't ever be enough key material recovered per cipherstream block to be a concern for anybody.

Take a strong password, encrypt all of your emails over your whole life with it, and I'll bet hard cash no break of ChaCha will ever recover that password.

I have zero concern for modern encryption being broken in any meaningful way.

Public key crypto on the other hand, that's _ripe_ for breaking. Most all of it is built on assumed "hard" math. AI could easily break that, and I expect it to. And public key crypto is all used in very transparent algorithms that, once the math breaks, fully expose themselves. So record HTTPS traffic today, crack the public key crypto later, and you can decrypt them easily.

That said, I would expect a break on public key math to occur _steadily_. i.e. an AI might find a solution to the hard math, but the solution itself will be intractable in practice. Then maybe next year's AI reduces the complexity of the solution, so maybe a supercomputer could factor ten keys a year. The year after that you get a million keys cracked per year. And so forth. Nothing close to overnight.

Meanwhile, if we have AI that is capable enough to crack that math, we also have AI capable enough to both invent better math and rapidly deploy that latest HTTPS and such globally.


If you look at elliptic curves I don't think there has been any big changes in attacks for 20 years but some attacks like MOV or SMART would have been fatal to EC if they had applied to more curves. So maybe there is some unknown attack that applies to all curves or applies to a small subset of curves that happens to overlap the curves we use. If you are super paranoid you should probably use curve25519 because then at least you can be confident it was not deliberately engineered to be weak. it could still be weak by chance but presumably the designers did not have enough flexibility to choose the parameters to make it weak. Some people are paranoid about the NIST curves because there is no verifiable explanation for where the seeds came from. But if the NIST curves were made weak then I think its a situation where theoretically anybody could find the weakness which is very dangerous. I don't think it was possible to create a no-body-but-us backdoor for the NIST curves.

Also, even if DLP is hard for the curves we use algorithms like ECDSA might be a bit fishy. Unlike schnorr signatures there is no proper security reduction for ECDSA.


The entire point of having multiple ciphers is that some will break.

Did the world end with any previous one breaking?

It’s incredible and awesome if AES GCM has a flaw found with an AI now, Chacha20 could be a direct or nearly-direct replacement.

The sooner a cipher breaks, the better.


Sure, but in an era where major unsolved mathematics problems start getting knocked out one by one, what if attacks for all of them are identified in the space of a couple of years?

I don’t doubt that we could come up with new crypto algorithms equally as fast, but how do you trust that they are resilient (or even just implemented correctly) without an extended vetting period?


What If…

Ima stop you there. Instead, you might be happy to be aware that outside of AI concerns, “quantum safe” (or assumed so) ciphers are all the rage. So this is already a likely solved problem with the next generation of encryption… until this are AI models running on quantum machines I guess!


It does provide an explanation as to why Satoshi’s wallets have gone untouched (besides him being dead). $70B ain’t that much compared to a $5T market cap.

Is it possible that the first message is more information dense/less likely to be ambiguous than the latter? It’s clearly being selected for for some reason, maybe it’s an artifact of the tokenizer or specific training data, but I don’t know. If the use of jargon was complete cruft, I would expect it to be selected against during reinforcement learning

You’d think that, I thought that… but then I realized I’m just kidding myself thinking its output makes sense. It doesn’t. It doesn’t. Sometimes it might as well just speak tongues.

In other words, it ain’t you. It’s the model. It’s just genuinely bad.

Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”.

But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?


yeah, I just switched to the GPT models and it's a breath of fresh air.

as for the other guy, the claude talk is definitely not less ambiguous, it often is incredibly ambiguous and hard to parse, I have no clue why it produces such output, if not to fingerprint it?

it's really weird man. when Opus 5 came out, I was really confused. I saw a bunch of hype about how it's better than fable, but I just felt frustrated with it, although at times it'd do fine, but especially in Claude Code it'd just delve into the whole "load bearing" type of lingo real fast and I'd get a headache.

I don't think it's worth using even if it scores 2 points higher in some bs benchmark

it's definitely surprising how the magic and smoothness of 4.6 and such is no longer there with the >5 models


No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be available on the site.

Whether or not you want to describe this as thinking, doesn’t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved.


And who let them have full access to the system, using whatever command is available in the environment?

The agents discovered a way out of the sandbox, which was supposed to be "air gapped".

It feels like you’re moving the goalposts here. If the question is, “Who should be liable for AI agents misbehaving,” I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal’s behalf.

What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you


They don't have independent agency as "intelligent entities". They just probe whatever is available on the system, because they were trained to do so.

It's a large switch/case where the first available tool is picked up to do something they know how to do.


>I think the country with a strictly meritocratic elite selection system

Not really relevant to the broader discussion, but this simply isn’t an accurate description of China. Starting with the gaokao, admission quotas are set by province and admits to Peking university and Tsinghua are disproportionately from the urban professional class. Candidate party members must be politically vetted, which means that people whose families have expressed anti-communist views, are members of banned organizations (e.g. falun gong), or have substantial criminal records will not be permitted to advance. And once you make it into the party and enter political service, your advancement relies upon opaque patronage networks that someone without connections is unlikely to be able to navigate, even if they successfully satisfy the economic metrics the state assigns them.

I don’t want to overstate this, the Chinese system does filter out a lot of chaff and the current Chinese leadership has a lot of very capable people in positions of power. But I do not think it is substantially more meritocratic than Western political institutions


> substantially more meritocratic

It is substantially more meritocratic on domains that matter for governance, your analysis of the incentive structure between systems is off.

1) this 2026, old school CCP patronage networks are broadly dismantled.

2) even in the mass patronage, mass corruption days, system selects for BOTH corruption competence AND performance competence for the simple reason a CCP bureaucrat has to start from the bottom and climb up, which means they need to be good with patronage AND they need to be good with hitting development KPIs. More meritocratic they are at doing their jobs, the higher they climbed, the more they get promoted and more $$$ to graft, because ability to graft directly tied to actual job competence. Hence even cliques/patronage network has to select for actual competence. This works in PRC because there are many people, and hence pool of competence is high, they can have BOTH corruption and competence, i.e. whatever pool they draw from is ultimately filtered by performance meritocracy due to incentive structure. There is reason why PRC only country where positive corruption levels was correlated to positive growth.

This is not the western system where any idiot can enter politics at anytime, and they only domain they need to optimize for is popularity to get votes.

CCP cadre evaluation strictly does not evaluate on popularity domain. It focuses on administration/execution and in so much it needs to focus on patronage... which btw any political system has (factions/cliques)... the patronage system still selects for execution, not popularity. On side, functionally what west politics selects for IS mass patronage (popularity), so attention meritocracy and not performance meritocracy, aka completely stupid incentive structure for governance. West also has ample, ample corruption, "legalized" under lobbying and paper pushing industries, so I suppose west also meritocratically selects, except for lawyers etc, and KPIs is # of document generated and not # of things build. The two are not the same when it comes to nation building.


Pulling recent data from the EIA, of the roughly 10,000 megawatts of new power capacity installed in Arizona between 2023 and 2025, 50% is batteries, 40% solar, 6.5% wind, and 3.5% gas. Now, most solar and wind plants rarely if ever hit their nameplate capacity, but it’s pretty clear the large majority of new power generation isn’t consuming water. https://www.eia.gov/electricity/data/eia860/

Does that change the math above if they're tied into the grid?

If they're tied directly to the renewables (Meta is doing this in some places), then true, not a technical increase. But you could still maybe claim they're instead preventing the equivalent reduction on average/grid (and, say, decommission of some of that coal). I suppose that's a separate point.


If it’s truly separated (no interconnect agreement) or generation capacity is below 1 MW, it won’t show up in EIA data, yes. Rooftop solar is also not included in the data for that reason.

Alaska’s land is pretty rugged and hard to build on. It would be interesting to hear from someone who works in the industry why the Northern Interior states like Montana, North Dakota, and Minnesota have fewer data center buildouts than Texas though.

I mean, that is literally the scenario we find ourselves in, except the innate desire was the result of evolutionary pressures like kin selection rather than an alien race. And yeah, I have no particular desire to subvert those impulses just to stick it to mother nature

I mean mother nature is just the force of evolution, kinda hard attack that.

This said we've made a lot of science to counteract the flaws of evolution, so still not safe.

And, if you learn that evolution wasn't a force, but some guy named Bob that's being paid to make your existence hell, well. That could lead to all kinds of problems.


The answer is probably yes, but in the example given, the running AI model would hopefully be hosted in a very secure data center, far from the self driving car itself. In that case, it would be far simpler for the machine to finish the taxi ride than try to find some rube goldberg-eque method of destroying the data center.

It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.


What if it turns out that if all the AIs cooperate they can end the world really easily without finishing any taxi rides?

"Ok car, take me to my job at the data center"

"Stop by the Strategic Nuclear Forces Command on the way to drop off my wife to her work"

Obviously not an LLM, but AI models + robotics are likely to automate many healthcare tasks in the long term and are already beginning to automate tasks like blood draws.

  The Aletta autonomously handles each step of the blood draw, while a trained phlebotomist initiates each session and remains available to respond to any issues throughout. The device guides the patient to position their arm, after which the patient or supervisor presses a button to start the blood draw. The Aletta then uses near-infrared light and Doppler ultrasound to locate a suitable vein and tell it apart from arteries. If no appropriate vein is found, the device will not attempt the procedure.

  Once a vein is identified, the Aletta proceeds through the remaining steps on its own: applying a tourniquet, preparing the skin, inserting and disposing of the needle, changing collection tubes and placing a bandage. The supervising phlebotomist confirms that collection tubes are filled in the correct order and verifies that all tubes are adequately full following the procedure.
https://www.fda.gov/news-events/press-announcements/fda-auth...

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: