‹ BackHN Continuity

Thread

Responsible Release of AI-Generated Mathematics

123 points · 223 comments · aureianimus

  1. unddoch · · focus · HN ↗
    For every important match problem solved by AI, without mathematicians we wouldn't know about the existence and importance of the problem.

    Famous mathematical conjectures are social constructs, formed by decades of even centuries of attention given to them by members of the math community. Without it, the danger is that future math "progress" will be reduced to generating tables of Lean statements and a probable/unprovable bit generated by AI.

    1. Xirdus · · focus · HN ↗
      Could just be a sampling bias. Humanity had something like 3000 years to make famous conjectures, whereas AI mathematicians have been around for a month or so. Give them time, I'm sure they'll start formulating highly consequential unsolved problems soon enough.
      1. amoss · · focus · HN ↗
        Somewhat tiring that as alway any criticism is reduced to "but have you tried this on the latest model".
        1. auggierose · · focus · HN ↗
          Maybe tiring, but that is the reality. Note that mathematicians are only upset now that "have you tried this on the latest model" works for so many of their problems now, but didn't for the model before that.
        2. Xirdus · · focus · HN ↗
          Those are the two extremes, both are just as bad. Don't excuse shortcomings of the current models with promises of future improvements. But also don't demand literal miracles in 21 business days.
      2. eru · · focus · HN ↗
        Not just sampling bias, but also human bias.

        In a sense, how do you know whether the problem your AI has just solved is important? A simple proxy is to just check whether humans have thought it's important.

        That's also why famous open problems are a good benchmark or proxy: you don't need to convince the rest of the world that the problem your lab's new AI just solved is actually useful or hard.

        1. xanderlewis · · focus · HN ↗
          Not all of the 'famous' open problems 'mathematicians have failed to solve for decades' are actually that famous though. In some cases they've remained open because no one cared or had even heard of them. Those outside the mathematics community seem to think that all mathematicians can, and do, work on essentially all problems (for example, most mathematicians have in mind the millennium problems as a goal), but this is very far from the case. Most serious problem statements are not even particularly understandable to most mathematicians, let alone workable on.

          There's also a difference between important and hard. There are important problems that turn out to be easy, and hard problems that turn out to be useless.

          As usual, I think everyone would agree that something like "curing cancer" would be both hard and important!

          I think the AI labs' current obsession with showing off mathematical results to uninformed outsiders is a cheap trick. If they were really interested in 'enabling human flourishing' (rather than just wowing, by any means possible, investors with more money than sense), they'd be showing off cures to diseases rather than solving obscure problems in combinatorics only previously considered by three Russians fifty years ago and then declaring that The Singularity is here.

      3. xanderlewis · · focus · HN ↗
        > AI mathematicians have been around for a month or so.

        LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).

        1. Xirdus · · focus · HN ↗
          Yes, and they've done absolutely fuck all with this knowledge until very, very, very recently.
          1. xanderlewis · · focus · HN ↗
            That is what I'm saying. There's a gold rush to extract what can be done with LLMs, but no indication that it'll continue. Maybe it will, but the evidence for its continuation seems to be entirely hype based on the fact that we've suddenly thought of new things to try. It's not because 'the capabilities are improving at a staggering rate omg the singularity is here guys'; it's mostly because no one bothered to spend enough time and money (and currently absurd amounts are required) to give it a serious shot until recently.

            I like the test proposed by Demis Hassabis. Something like: train an LLM on pre-Einstein physics and see if it can rediscover what Einstein did. Despite the incessant noise online, we seem to be no closer to this. If there's material evidence that we are, I'd sincerely like to hear about it!

            1. famouswaffles · · focus · HN ↗
              LLMs have improved at a staggering rate over the years. From spewing incoherent gibberish to barely managing low digit arithmetic to only tackling high school level problems.

              2 years ago, a DeepMind system achieved IMO Silver. This system was not a pure LLM, nor were the problems solved end to end in language. In fact the role of the LLM was to translate the problem into lean and generate strategies that could be auto verified. Some problems took up to 3 days to compute.

              1 year ago, both OpenAI and DeepMind announced internal models had achieved IMO Gold. Unlike the cobbled together Alphaproof and AlphaGeometry, these results were purely LLMs, solved end to end in language, with any lean proofs generated by the LLM afterwards. They achieved this result in the time limits alloted to humans (4.5 hours)

              Today, we have publicly available models capable of solving well known and open standing conjectures. And in less than the time it took Alphaproof/Geometry to attain IMO Silver, internal models can even solve a millenium prize problem.

              Comments like yours just make it plain you aren't paying any attention at all. The indication that it will continue is the fact that it has continued despite some assuring us we're at the precipice of some plateau at every single moment. You think it's because of 'breathless hype'? Only if you're blind.

              1. xanderlewis · · focus · HN ↗
                > internal models can even solve a millenium prize problem.

                OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.

                What do you think has enabled the 'staggering rate' of improvement? Scaling up? More data? It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached. Anthropic are already reportedly buying up all of the world's second-hand books and destructively (their words) scanning them in a desperate rush for ever more training data, so I don't think this theory is unfounded.

                What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.

                1. famouswaffles · · focus · HN ↗
                  >OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.

                  No they did not. Tristan (with heavy aid from LLMs btw) solved a sub-problem via a path Open AI's full solution didn't take (forced vs unforced). To say they stole his work would be silly.

                  And they explicitly say - “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

                  <a href="https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;10&#x2F;science&#x2F;tristan-buckmaster-openai-math-navier-stokes.html" rel="nofollow">https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;10&#x2F;science&#x2F;tristan-buckmaste...

                  &gt;What do you think has enabled the &#x27;staggering rate&#x27; of improvement?

                  Compute and data

                  &gt;It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they&#x27;d have nothing), once all the data has been eaten, some sort of plateau will indeed be reached.

                  Data used to train the models has become increasingly synthetic. The reinforcement learning getting them superhuman at math isn&#x27;t tied to human data. If this is what you&#x27;re betting on to make things stall, best to move on.

                  &gt;Anthropic are already reportedly buying up all of the world&#x27;s second-hand books

                  That is not what they are doing.

                  &gt;What do you think of the Hassabis Test? Until it passes, I don&#x27;t see how humans are replaced.

                  There&#x27;s not enough pre-einstein text digitzed to make a fair shot at such an experiment so neat idea but kind of a waste of time. And it&#x27;s kind of meaningless. We&#x27;re already in the midst of AI becoming superhuman in one domain. We don&#x27;t need any what ifs for something that is happening right now, in front of us. If some want to stick their hands in the sands, shouting &#x27;la la humans are special&#x27; then there&#x27;s nothing to do about that.

                  1. xanderlewis · · focus · HN ↗
                    &gt; To say they stole his work would be silly.

                    He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?

                    &gt; There&#x27;s not enough pre-einstein text digitzed to make a fair shot at such an experiment

                    Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.

                    1. hardbass · · focus · HN ↗
                      He has the right to present his side. Without analyzing both sides claims, no comment can be made.
                      1. xanderlewis · · focus · HN ↗
                        There is an important asymmetry, though: Buckmaster is a mathematician in it for fun (as almost all mathematicians are, because there certainly isn’t much money in it), and OpenAI is a company desperately scrambling for anything they can present as evidence that the singularity is here and you must invest your money now. Buckmaster abides by the unwritten code of mathematicians (alluded to in said interview); OpenAI abides by absolutely no such code other than ‘move fast and break things’.

                        I don’t think it’s a stretch to say that each side’s claim isn’t equally trustworthy.

                        1. hardbass · · focus · HN ↗
                          But there is strong ego involved on part of Mr Buckmaster, egos tend to cloud judgement and bias a person. I would not speculate on either side and I would be careful to avoid people who simply take one side for granted.
                          1. xanderlewis · · focus · HN ↗
                            Listening to him, I didn’t get that impression. I get that academia at large is full of egomaniacs (or so I’ve heard and read online), but the mathematics community is a weirdly non-toxic oasis of reasonable enthusiasts — genuinely, surprising though that may be to those with no experience!

                            So I have my reasons to weigh his statements more heavily. I certainly don’t take anything for granted.

                            1. hardbass · · focus · HN ↗
                              I don&#x27;t mean he must specifically be an egotistic person. But if you feel slighted and feel a problem you have been working hard for months taken from under you (quite regardless of the truth of that statement), any person would have a egotistic bias.
                    2. famouswaffles · · focus · HN ↗
                      &gt;He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?

                      He can say whatever he wants. I&#x27;m not obligated to take it at face value, especially when it seems to fall apart with some scrutiny.

                      &gt;Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.

                      Billions of years of evolution probably has something to do with it.

    2. jillesvangurp · · focus · HN ↗
      You make a good point about LLMs being (so far) great at identifying solutions to complex puzzles but not yet coming up with their own theories, questions, etc. The business of finding interesting questions to answer rather than answering them is the essence of what scientists do. Good science identifies more questions than it answers.

      With Fermat&#x27;s theorem, the genius was in the original theorem. Which then caused generations of mathematicians to break their heads over trying to prove it correct. Fermat didn&#x27;t write down a proof. His theorem was famously just a scribble in a side line of a book. Probably it was something that he had an hunch about that he couldn&#x27;t quickly falsify.

      I don&#x27;t think AI is being used much for coming up with new problems yet. But I don&#x27;t see why that would not be possible either. It&#x27;s the obvious next frontier after AI clears the backlog of existing theorems. I imagine scientists are already using AI to find new interesting problems to work on and generally explore the problem space. But fundamentally, the reflex of asking or imagining &quot;if this is true, what else could be true&quot; is something that distinguishes people from AIs. For now at least.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.