‹ BackHN Continuity

Thread

Responsible Release of AI-Generated Mathematics

123 points · 223 comments · aureianimus

  1. unddoch · · focus · HN ↗
    For every important match problem solved by AI, without mathematicians we wouldn't know about the existence and importance of the problem.

    Famous mathematical conjectures are social constructs, formed by decades of even centuries of attention given to them by members of the math community. Without it, the danger is that future math "progress" will be reduced to generating tables of Lean statements and a probable/unprovable bit generated by AI.

    1. Xirdus · · focus · HN ↗
      Could just be a sampling bias. Humanity had something like 3000 years to make famous conjectures, whereas AI mathematicians have been around for a month or so. Give them time, I'm sure they'll start formulating highly consequential unsolved problems soon enough.
      1. xanderlewis · · focus · HN ↗
        > AI mathematicians have been around for a month or so.

        LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).

        1. Xirdus · · focus · HN ↗
          Yes, and they've done absolutely fuck all with this knowledge until very, very, very recently.
          1. xanderlewis · · focus · HN ↗
            That is what I'm saying. There's a gold rush to extract what can be done with LLMs, but no indication that it'll continue. Maybe it will, but the evidence for its continuation seems to be entirely hype based on the fact that we've suddenly thought of new things to try. It's not because 'the capabilities are improving at a staggering rate omg the singularity is here guys'; it's mostly because no one bothered to spend enough time and money (and currently absurd amounts are required) to give it a serious shot until recently.

            I like the test proposed by Demis Hassabis. Something like: train an LLM on pre-Einstein physics and see if it can rediscover what Einstein did. Despite the incessant noise online, we seem to be no closer to this. If there's material evidence that we are, I'd sincerely like to hear about it!

            1. famouswaffles · · focus · HN ↗
              LLMs have improved at a staggering rate over the years. From spewing incoherent gibberish to barely managing low digit arithmetic to only tackling high school level problems.

              2 years ago, a DeepMind system achieved IMO Silver. This system was not a pure LLM, nor were the problems solved end to end in language. In fact the role of the LLM was to translate the problem into lean and generate strategies that could be auto verified. Some problems took up to 3 days to compute.

              1 year ago, both OpenAI and DeepMind announced internal models had achieved IMO Gold. Unlike the cobbled together Alphaproof and AlphaGeometry, these results were purely LLMs, solved end to end in language, with any lean proofs generated by the LLM afterwards. They achieved this result in the time limits alloted to humans (4.5 hours)

              Today, we have publicly available models capable of solving well known and open standing conjectures. And in less than the time it took Alphaproof/Geometry to attain IMO Silver, internal models can even solve a millenium prize problem.

              Comments like yours just make it plain you aren't paying any attention at all. The indication that it will continue is the fact that it has continued despite some assuring us we're at the precipice of some plateau at every single moment. You think it's because of 'breathless hype'? Only if you're blind.

              1. xanderlewis · · focus · HN ↗
                > internal models can even solve a millenium prize problem.

                OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.

                What do you think has enabled the 'staggering rate' of improvement? Scaling up? More data? It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached. Anthropic are already reportedly buying up all of the world's second-hand books and destructively (their words) scanning them in a desperate rush for ever more training data, so I don't think this theory is unfounded.

                What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.

                1. famouswaffles · · focus · HN ↗
                  >OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.

                  No they did not. Tristan (with heavy aid from LLMs btw) solved a sub-problem via a path Open AI's full solution didn't take (forced vs unforced). To say they stole his work would be silly.

                  And they explicitly say - “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

                  <a href="https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;10&#x2F;science&#x2F;tristan-buckmaster-openai-math-navier-stokes.html" rel="nofollow">https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;10&#x2F;science&#x2F;tristan-buckmaste...

                  &gt;What do you think has enabled the &#x27;staggering rate&#x27; of improvement?

                  Compute and data

                  &gt;It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they&#x27;d have nothing), once all the data has been eaten, some sort of plateau will indeed be reached.

                  Data used to train the models has become increasingly synthetic. The reinforcement learning getting them superhuman at math isn&#x27;t tied to human data. If this is what you&#x27;re betting on to make things stall, best to move on.

                  &gt;Anthropic are already reportedly buying up all of the world&#x27;s second-hand books

                  That is not what they are doing.

                  &gt;What do you think of the Hassabis Test? Until it passes, I don&#x27;t see how humans are replaced.

                  There&#x27;s not enough pre-einstein text digitzed to make a fair shot at such an experiment so neat idea but kind of a waste of time. And it&#x27;s kind of meaningless. We&#x27;re already in the midst of AI becoming superhuman in one domain. We don&#x27;t need any what ifs for something that is happening right now, in front of us. If some want to stick their hands in the sands, shouting &#x27;la la humans are special&#x27; then there&#x27;s nothing to do about that.

                  1. xanderlewis · · focus · HN ↗
                    &gt; To say they stole his work would be silly.

                    He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?

                    &gt; There&#x27;s not enough pre-einstein text digitzed to make a fair shot at such an experiment

                    Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.

                    1. famouswaffles · · focus · HN ↗
                      &gt;He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?

                      He can say whatever he wants. I&#x27;m not obligated to take it at face value, especially when it seems to fall apart with some scrutiny.

                      &gt;Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.

                      Billions of years of evolution probably has something to do with it.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.