If youre going to spend 10 million dollars on 10000 agents trying to solve some important maths problem I think it's reasonable to ask for some grants to help digest whatever they came up with.
Or you could hire mathematicians and do it in-house, but I guarantee you grants to PhD students are cheaper than silicon value salaries.
Not for now, but it's AI is getting better. A big step is "refactoring" a very long proof into a few intermediate lemas and theorems that are more inteligible and useful for other proof. It may take a few years or decades in some cases.
Anyway, I expect AI to be better at "refactoring", but for now a centaur is better.
Actually, for what you are mentioning, it is getting worse. There was a sweet spot somewhere around the release of GPT-o3, and ever since, the LLMs have been getting more accurate at solving problems, but worse at explaining how, and to hone in on what is interesting. This isn't surprising, as RL strategies shifted from RLHF to RLVR, so priorities during learning changed. I don't expect AI labs to reverse course on this. We can expect AI proofs to become increasingly incomprehensible over time.
I still remember Garry Kasparov vs. Deep Blue. Now, Magnus vs Stockfish is not even funny. I've seen AlphaGo and AlphaStar, and how their communities reacted...
I remember when Mathematica only could tell the answers and you had to type the formula correctly. Now there are apps that solve the exercise from a photograph with all the intermediate steps. We reminded the T.A. to be more alert during the midterms becuse we already had problems. I'm very worry about the magical glasses now, but it's important to be not overreact and be polite with the students.
Back to refactoring unintelligible long math proof: Let's talk again in 2031.
The paper is secondary to developing a general understanding of new ideas. The AI is not going to do that. All of the AI bros in this thread are going to be addicted to super AI Netflix and the math community will be dead.
The paper, both writing it and reading it, are a means by which we develop our understanding of new ideas. We can surely come up with a better means if we try, but "stop trying and leave actual understanding to machines" ain't gonna cut it.
They aren't buying tokens, more likely it costs them 1m to run the 10000 agents for 88 hours. Still leaves room for the PhD ofc.
I'm trying to map it on other fields and cant help but think it is absurd. If some company spends 1m developing a new alloy, should metallurgists be upset when they publish the recipe?
AI companies are already paying human mathematicians, to help train their models. Including to understand AI-generated stuff. The pay is kind of shit but still quite a lot better than what we were paid as phd students, especially if you do it 40 hours a week (which, tbf, was very short week in grad school).
I know this because I spent a lot of social time around a Math dept in grad school and many of them are moonlighting working for subcontractors on production and evaluation of training data [1].
Anyways, mathematics is a bit of a funny discipline... some buckets:
There are lots of obvious cases where progress has obvious applications for incremental progress, and where you can ask for novel math from the perspective of the application or ask for applications from the perspective of novel math. That seems like the sort of thing that you could throw $10M at get something valuable without much human input. I do this, at much much much smaller dollar amounts, pretty regularly, and with open models, so nothing secret/sota/etc.
But there are also lots of cases where some progress gets made on something very pure and esoteric and the implications aren't really clear. Those are cases where just asking a bunch of nerds to marinate on it while interacting with a messy world in the full social generality of a modern university could make a lot of sense. The connection between the four color theorem and register allocation, for example, feels very... hard to do without humans free-associating in a diverse academic environment.
And, at last, there are cases where there are subfields or problems that are big and important in the way that driving a rare car or wearing a particularly fascinating rolex is important. Some of modern mathematics feels like it has its cultural roots in intellectual gamesmanship amongst wealthy gentleman and/or those under their patronage. A lot of the... less well-argued... objections come from this set.
Like I said, math is kind of a uniquely weird discipline. It's got aspects of practical gritty science, noblesse oblige/high-culture, fashion/taste, craftsmanship, tutelage, and so on.
--
[1] Some enterprising journalist should track down all the subcontractors used by OpenAI for mathematics data, then talk to all of those people and figure out how much breakthrough-adjacent data work was happening... pure conspiracy but I'd be curious for someone to explore the hypothesis that there were maybe people helping find fragments here and there and it wasn't all purely mechanical monkeying.
Animats · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
[deleted] · · focus · HN ↗
[deleted]
unddoch · · focus · HN ↗
ChickeNES · · focus · HN ↗
eli_gottlieb · · focus · HN ↗
gus_massa · · focus · HN ↗
Anyway, I expect AI to be better at "refactoring", but for now a centaur is better.
hodgehog11 · · focus · HN ↗
gus_massa · · focus · HN ↗
I remember when Mathematica only could tell the answers and you had to type the formula correctly. Now there are apps that solve the exercise from a photograph with all the intermediate steps. We reminded the T.A. to be more alert during the midterms becuse we already had problems. I'm very worry about the magical glasses now, but it's important to be not overreact and be polite with the students.
Back to refactoring unintelligible long math proof: Let's talk again in 2031.
GPerson · · focus · HN ↗
eli_gottlieb · · focus · HN ↗
GPerson · · focus · HN ↗
6510 · · focus · HN ↗
I'm trying to map it on other fields and cant help but think it is absurd. If some company spends 1m developing a new alloy, should metallurgists be upset when they publish the recipe?
GPerson · · focus · HN ↗
fultonn · · focus · HN ↗
I know this because I spent a lot of social time around a Math dept in grad school and many of them are moonlighting working for subcontractors on production and evaluation of training data [1].
Anyways, mathematics is a bit of a funny discipline... some buckets:
There are lots of obvious cases where progress has obvious applications for incremental progress, and where you can ask for novel math from the perspective of the application or ask for applications from the perspective of novel math. That seems like the sort of thing that you could throw $10M at get something valuable without much human input. I do this, at much much much smaller dollar amounts, pretty regularly, and with open models, so nothing secret/sota/etc.
But there are also lots of cases where some progress gets made on something very pure and esoteric and the implications aren't really clear. Those are cases where just asking a bunch of nerds to marinate on it while interacting with a messy world in the full social generality of a modern university could make a lot of sense. The connection between the four color theorem and register allocation, for example, feels very... hard to do without humans free-associating in a diverse academic environment.
And, at last, there are cases where there are subfields or problems that are big and important in the way that driving a rare car or wearing a particularly fascinating rolex is important. Some of modern mathematics feels like it has its cultural roots in intellectual gamesmanship amongst wealthy gentleman and/or those under their patronage. A lot of the... less well-argued... objections come from this set.
Like I said, math is kind of a uniquely weird discipline. It's got aspects of practical gritty science, noblesse oblige/high-culture, fashion/taste, craftsmanship, tutelage, and so on.
--
[1] Some enterprising journalist should track down all the subcontractors used by OpenAI for mathematics data, then talk to all of those people and figure out how much breakthrough-adjacent data work was happening... pure conspiracy but I'd be curious for someone to explore the hypothesis that there were maybe people helping find fragments here and there and it wasn't all purely mechanical monkeying.