‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. andrepd · · focus · HN ↗
      > "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that).

      In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!

      1. hatthew · · focus · HN ↗
        I'm not aware of anyone disputing that AI solved NS. As far as I'm aware, the controversies are about how useful of a result forced blowup is, whether the model built off of unpublished work by Buckmaster, and the ethics of essentially trying to scoop him. All of those are very valid concerns, and none of them affect the fact that AI solved a difficult open math problem.

        If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.

        1. sokoloff · · focus · HN ↗
          “AI conclusively did this, but we’re uncertain whether it relied on the unpublished work of a human while doing it.”

          If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.

          1. hatthew · · focus · HN ↗
            If building off of someone's work means you didn't solve it yourself, then nobody has solved anything themselves since some caveman counting piles of rocks tens of thousands of year ago. A more generous interpretation of what you're saying is: the AI's contributions to the solution were not meaningful enough to count as it "solving the problem". That's certainly defensible, but I would still disagree. Could you clarify your point?
            1. mvc · · focus · HN ↗
              I dunno. Didn't it still need to be driven by a team of experienced Mathematicians? I don't believe that two months ago, you or I could've just typed "Solve Navier Stokes. Make no mistakes" into claude and come back some time later and expect to see a solution.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.