It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine.
It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).
On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".
It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.
Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend an infinite amount of money.
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine.
It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).
On that note, I actually had an overall harness (for experimenting) that was essentially like this:
"for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".
It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.
Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend infinite amount of money.
This is one part where it should really just be done by AI for 99% of the cases. Just have a security focused AI model that is biased towards flagging things a bit more conservatively.
Only apps that get flagged by the AI reviewer then have to go through a separate, slower human review process.
With latest AI models, we will most likely see shift back to just fully Native apps even for smaller teams.
This is type of stuff that AI models do really well, if you get a particular feature or even a bigger app change done in one platform, you can just tell Opus or Fable to replicate to the other platforms and in almost all cases it will just do it as well as most human teams would have done it.
Well, if in future we do end up with a magical tool that can solve any formal mathematical problem on a whim, we really won’t need field of mathematics anymore as it is today.
There would be no need to deliver new mathematical insights by solving problems. You would just have a magical math problem solving machine and that’s it.
What do you mean “we won’t need mathematics as it is today”?
To further human understanding is itself a goal that single-handedly justifies our efforts.
Jumping straight to the “answer” and therefore missing both the understanding of the actual problem, and any useful discoveries along the way is a waste at best, and actively harmful at worst.
I'm not the one you replied to, but I think I get their point. If you have a machine that can solve math problems at the push of a button, you don't really need to amass discoveries anymore just in case someone might need it years down the line.
Someone in the thread mentioned Minkowski’s math work being instrumental to Einstein’s physics. If Einstein had had a math machine that can spit out the result at will, he wouldn't have needed Minkowski.
Nevertheless, I also see the opposing point: if Minkowski hadn't already published his results, it's possible Einstein would not have even had the inspiration to derive general relativity from it. I think this is the human aspect most critics are worried about.
Physics alone is more than enough to “further human understanding”. All current mathematicians can move to other sciences, closest being fields in physics, and it will all continue to progress just fine.
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
While AI companies have almost infinite money, they still don’t want to blow million dollar budgets on problems if there isn’t high likelihood that it will be successful.
But within next 10 years as costs drop significantly and even more improvements are made, yes it is very likely that almost every single existing math problem will get a serious AI cracking done on it
You are glossing over the fact that there is reasonable suspicion that OpenAI used privileged information to do “the scoop”. I.e. they used the fact that researcher used OpenAI tools to get advantage .
Imagine if OpenAI opened up a high frequency trading arm and suddenly stole all the prompts and research that other HfT firms are doing through OpenAI tools and start making bank based on that . Wouldn’t that be straight up insane?
This Tristan guy's statement reads like something a normal, reasonable human being would write.
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win".
https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
It is so disappointing that we can't have such a monumental moment in history without the controversy. OpenAI leadership clearly doesn't seem to care too much about ethics. Is it a requirement to completely lack integrity to have a ground breaking company?
The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.
As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.
It is simultaneously the best and the worst time to be a mathematician right now.
On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".
It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.
Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend an infinite amount of money.
reply