‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. andrepd · · focus · HN ↗
      > "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that).

      In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!

      1. hatthew · · focus · HN ↗
        I'm not aware of anyone disputing that AI solved NS. As far as I'm aware, the controversies are about how useful of a result forced blowup is, whether the model built off of unpublished work by Buckmaster, and the ethics of essentially trying to scoop him. All of those are very valid concerns, and none of them affect the fact that AI solved a difficult open math problem.

        If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.

        1. sokoloff · · focus · HN ↗
          “AI conclusively did this, but we’re uncertain whether it relied on the unpublished work of a human while doing it.”

          If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.

          1. hatthew · · focus · HN ↗
            If building off of someone's work means you didn't solve it yourself, then nobody has solved anything themselves since some caveman counting piles of rocks tens of thousands of year ago. A more generous interpretation of what you're saying is: the AI's contributions to the solution were not meaningful enough to count as it "solving the problem". That's certainly defensible, but I would still disagree. Could you clarify your point?
            1. mvc · · focus · HN ↗
              I dunno. Didn't it still need to be driven by a team of experienced Mathematicians? I don't believe that two months ago, you or I could've just typed "Solve Navier Stokes. Make no mistakes" into claude and come back some time later and expect to see a solution.
            2. sokoloff · · focus · HN ↗
              At some point in time, perhaps Jan 2023, there was an open question about Navier-Stokes.

              We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.

              Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.

              Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.

              At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.

              Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).

              In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.

              In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.

              What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”

              1. jasode · · focus · HN ↗
                >, we fork the universe in 3. In one, Bob continues his work and solves the open question.

                To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.

                >In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.

                Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.

                We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.

              2. hatthew · · focus · HN ↗
                In addition to jasode's comment, which I endorse, I'd still argue that Charlie solved the problem. At some point in time, Bob and Charlie both had the same information (Bob's notes). From that same information, Charlie got to the final solution and Bob didn't, despite Bob having the advantage of years of familiarity with the problem. I feel like Charlie deserves quite a bit of credit. Depending on the circumstances, I might argue that Bob deserves a greater share of the credit than Charlie, but I think it would not be right to claim that Charlie didn't solve the problem (I'm intentionally leaving off any explicit qualifiers of amount of help, because my original statement was "solve an open math problem" which also has no such qualifiers).

                Now of course Alice has the advantage of massive parallelization, but that's not a reason to discredit her, that's a genuine advantage that she has, applicable to all problems.

                To respond to a couple specific points:

                > whether Bob’s contributions were required to Alice’s final step

                Even if Bob's contributions were required, I don't think that's a reason to fully discredit Alice.

                > Alice is capable of solving a second open math question

                Given that AI has already solved dozens of open math questions, I don't think this is up for debate.

          2. keeda · · focus · HN ↗
            We are no longer uncertain. Even if you want to dismiss OpenAI’s categorical denial there is the little matter of hundreds of longstanding open problems also solved by AI without controversy.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.