For every important match problem solved by AI, without mathematicians we wouldn't know about the existence and importance of the problem.
Famous mathematical conjectures are social constructs, formed by decades of even centuries of attention given to them by members of the math community. Without it, the danger is that future math "progress" will be reduced to generating tables of Lean statements and a probable/unprovable bit generated by AI.
Could just be a sampling bias. Humanity had something like 3000 years to make famous conjectures, whereas AI mathematicians have been around for a month or so. Give them time, I'm sure they'll start formulating highly consequential unsolved problems soon enough.
> AI mathematicians have been around for a month or so.
LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).
That is what I'm saying. There's a gold rush to extract what can be done with LLMs, but no indication that it'll continue. Maybe it will, but the evidence for its continuation seems to be entirely hype based on the fact that we've suddenly thought of new things to try. It's not because 'the capabilities are improving at a staggering rate omg the singularity is here guys'; it's mostly because no one bothered to spend enough time and money (and currently absurd amounts are required) to give it a serious shot until recently.
I like the test proposed by Demis Hassabis. Something like: train an LLM on pre-Einstein physics and see if it can rediscover what Einstein did. Despite the incessant noise online, we seem to be no closer to this. If there's material evidence that we are, I'd sincerely like to hear about it!
LLMs have improved at a staggering rate over the years. From spewing incoherent gibberish to barely managing low digit arithmetic to only tackling high school level problems.
2 years ago, a DeepMind system achieved IMO Silver. This system was not a pure LLM, nor were the problems solved end to end in language. In fact the role of the LLM was to translate the problem into lean and generate strategies that could be auto verified. Some problems took up to 3 days to compute.
1 year ago, both OpenAI and DeepMind announced internal models had achieved IMO Gold. Unlike the cobbled together Alphaproof and AlphaGeometry, these results were purely LLMs, solved end to end in language, with any lean proofs generated by the LLM afterwards. They achieved this result in the time limits alloted to humans (4.5 hours)
Today, we have publicly available models capable of solving well known and open standing conjectures. And in less than the time it took Alphaproof/Geometry to attain IMO Silver, internal models can even solve a millenium prize problem.
Comments like yours just make it plain you aren't paying any attention at all. The indication that it will continue is the fact that it has continued despite some assuring us we're at the precipice of some plateau at every single moment. You think it's because of 'breathless hype'? Only if you're blind.
> internal models can even solve a millenium prize problem.
OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.
What do you think has enabled the 'staggering rate' of improvement? Scaling up? More data? It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached. Anthropic are already reportedly buying up all of the world's second-hand books and destructively (their words) scanning them in a desperate rush for ever more training data, so I don't think this theory is unfounded.
What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
>OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.
No they did not. Tristan (with heavy aid from LLMs btw) solved a sub-problem via a path Open AI's full solution didn't take (forced vs unforced). To say they stole his work would be silly.
And they explicitly say - “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
>What do you think has enabled the 'staggering rate' of improvement?
Compute and data
>It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached.
Data used to train the models has become increasingly synthetic. The reinforcement learning getting them superhuman at math isn't tied to human data. If this is what you're betting on to make things stall, best to move on.
>Anthropic are already reportedly buying up all of the world's second-hand books
That is not what they are doing.
>What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
There's not enough pre-einstein text digitzed to make a fair shot at such an experiment so neat idea but kind of a waste of time. And it's kind of meaningless. We're already in the midst of AI becoming superhuman in one domain. We don't need any what ifs for something that is happening right now, in front of us. If some want to stick their hands in the sands, shouting 'la la humans are special' then there's nothing to do about that.
There is an important asymmetry, though: Buckmaster is a mathematician in it for fun (as almost all mathematicians are, because there certainly isn’t much money in it), and OpenAI is a company desperately scrambling for anything they can present as evidence that the singularity is here and you must invest your money now. Buckmaster abides by the unwritten code of mathematicians (alluded to in said interview); OpenAI abides by absolutely no such code other than ‘move fast and break things’.
I don’t think it’s a stretch to say that each side’s claim isn’t equally trustworthy.
But there is strong ego involved on part of Mr Buckmaster, egos tend to cloud judgement and bias a person. I would not speculate on either side and I would be careful to avoid people who simply take one side for granted.
Listening to him, I didn’t get that impression. I get that academia at large is full of egomaniacs (or so I’ve heard and read online), but the mathematics community is a weirdly non-toxic oasis of reasonable enthusiasts — genuinely, surprising though that may be to those with no experience!
So I have my reasons to weigh his statements more heavily. I certainly don’t take anything for granted.
I don't mean he must specifically be an egotistic person. But if you feel slighted and feel a problem you have been working hard for months taken from under you (quite regardless of the truth of that statement), any person would have a egotistic bias.
unddoch · · focus · HN ↗
Famous mathematical conjectures are social constructs, formed by decades of even centuries of attention given to them by members of the math community. Without it, the danger is that future math "progress" will be reduced to generating tables of Lean statements and a probable/unprovable bit generated by AI.
Xirdus · · focus · HN ↗
xanderlewis · · focus · HN ↗
LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).
Xirdus · · focus · HN ↗
xanderlewis · · focus · HN ↗
I like the test proposed by Demis Hassabis. Something like: train an LLM on pre-Einstein physics and see if it can rediscover what Einstein did. Despite the incessant noise online, we seem to be no closer to this. If there's material evidence that we are, I'd sincerely like to hear about it!
famouswaffles · · focus · HN ↗
2 years ago, a DeepMind system achieved IMO Silver. This system was not a pure LLM, nor were the problems solved end to end in language. In fact the role of the LLM was to translate the problem into lean and generate strategies that could be auto verified. Some problems took up to 3 days to compute.
1 year ago, both OpenAI and DeepMind announced internal models had achieved IMO Gold. Unlike the cobbled together Alphaproof and AlphaGeometry, these results were purely LLMs, solved end to end in language, with any lean proofs generated by the LLM afterwards. They achieved this result in the time limits alloted to humans (4.5 hours)
Today, we have publicly available models capable of solving well known and open standing conjectures. And in less than the time it took Alphaproof/Geometry to attain IMO Silver, internal models can even solve a millenium prize problem.
Comments like yours just make it plain you aren't paying any attention at all. The indication that it will continue is the fact that it has continued despite some assuring us we're at the precipice of some plateau at every single moment. You think it's because of 'breathless hype'? Only if you're blind.
xanderlewis · · focus · HN ↗
OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.
What do you think has enabled the 'staggering rate' of improvement? Scaling up? More data? It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached. Anthropic are already reportedly buying up all of the world's second-hand books and destructively (their words) scanning them in a desperate rush for ever more training data, so I don't think this theory is unfounded.
What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
famouswaffles · · focus · HN ↗
No they did not. Tristan (with heavy aid from LLMs btw) solved a sub-problem via a path Open AI's full solution didn't take (forced vs unforced). To say they stole his work would be silly.
And they explicitly say - “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
<a href="https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html" rel="nofollow">https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
>What do you think has enabled the 'staggering rate' of improvement?
Compute and data
>It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached.
Data used to train the models has become increasingly synthetic. The reinforcement learning getting them superhuman at math isn't tied to human data. If this is what you're betting on to make things stall, best to move on.
>Anthropic are already reportedly buying up all of the world's second-hand books
That is not what they are doing.
>What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
There's not enough pre-einstein text digitzed to make a fair shot at such an experiment so neat idea but kind of a waste of time. And it's kind of meaningless. We're already in the midst of AI becoming superhuman in one domain. We don't need any what ifs for something that is happening right now, in front of us. If some want to stick their hands in the sands, shouting 'la la humans are special' then there's nothing to do about that.
xanderlewis · · focus · HN ↗
He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?
> There's not enough pre-einstein text digitzed to make a fair shot at such an experiment
Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.
hardbass · · focus · HN ↗
xanderlewis · · focus · HN ↗
I don’t think it’s a stretch to say that each side’s claim isn’t equally trustworthy.
hardbass · · focus · HN ↗
xanderlewis · · focus · HN ↗
So I have my reasons to weigh his statements more heavily. I certainly don’t take anything for granted.
hardbass · · focus · HN ↗
famouswaffles · · focus · HN ↗
He can say whatever he wants. I'm not obligated to take it at face value, especially when it seems to fall apart with some scrutiny.
>Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.
Billions of years of evolution probably has something to do with it.