‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. hexapus · · focus · HN ↗
      You should add an additional item: Be able to relay the contents of this exam accurately.
    2. andrepd · · focus · HN ↗
      > "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that).

      In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!

      1. hatthew · · focus · HN ↗
        I'm not aware of anyone disputing that AI solved NS. As far as I'm aware, the controversies are about how useful of a result forced blowup is, whether the model built off of unpublished work by Buckmaster, and the ethics of essentially trying to scoop him. All of those are very valid concerns, and none of them affect the fact that AI solved a difficult open math problem.

        If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.

        1. sokoloff · · focus · HN ↗
          “AI conclusively did this, but we’re uncertain whether it relied on the unpublished work of a human while doing it.”

          If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.

          1. hatthew · · focus · HN ↗
            If building off of someone's work means you didn't solve it yourself, then nobody has solved anything themselves since some caveman counting piles of rocks tens of thousands of year ago. A more generous interpretation of what you're saying is: the AI's contributions to the solution were not meaningful enough to count as it "solving the problem". That's certainly defensible, but I would still disagree. Could you clarify your point?
            1. mvc · · focus · HN ↗
              I dunno. Didn't it still need to be driven by a team of experienced Mathematicians? I don't believe that two months ago, you or I could've just typed "Solve Navier Stokes. Make no mistakes" into claude and come back some time later and expect to see a solution.
            2. sokoloff · · focus · HN ↗
              At some point in time, perhaps Jan 2023, there was an open question about Navier-Stokes.

              We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.

              Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.

              Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.

              At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.

              Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).

              In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.

              In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.

              What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”

              1. jasode · · focus · HN ↗
                >, we fork the universe in 3. In one, Bob continues his work and solves the open question.

                To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.

                >In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.

                Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.

                We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.

              2. hatthew · · focus · HN ↗
                In addition to jasode's comment, which I endorse, I'd still argue that Charlie solved the problem. At some point in time, Bob and Charlie both had the same information (Bob's notes). From that same information, Charlie got to the final solution and Bob didn't, despite Bob having the advantage of years of familiarity with the problem. I feel like Charlie deserves quite a bit of credit. Depending on the circumstances, I might argue that Bob deserves a greater share of the credit than Charlie, but I think it would not be right to claim that Charlie didn't solve the problem (I'm intentionally leaving off any explicit qualifiers of amount of help, because my original statement was "solve an open math problem" which also has no such qualifiers).

                Now of course Alice has the advantage of massive parallelization, but that's not a reason to discredit her, that's a genuine advantage that she has, applicable to all problems.

                To respond to a couple specific points:

                > whether Bob’s contributions were required to Alice’s final step

                Even if Bob's contributions were required, I don't think that's a reason to fully discredit Alice.

                > Alice is capable of solving a second open math question

                Given that AI has already solved dozens of open math questions, I don't think this is up for debate.

          2. keeda · · focus · HN ↗
            We are no longer uncertain. Even if you want to dismiss OpenAI’s categorical denial there is the little matter of hundreds of longstanding open problems also solved by AI without controversy.
    3. intelkishan · · focus · HN ↗
      Could you share your other 5 tests, if they are public?
      1. hatthew · · focus · HN ↗
        Here&#x27;s my previous comment: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902

        I didn&#x27;t really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name &quot;humanity&#x27;s last exam&quot;. So take it with a grain of salt.

        1. andai · · focus · HN ↗
          Thanks for sharing.

          I think non-RLHF&#x27;d LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don&#x27;t know if anyone has tested them for that. (Also I&#x27;m not sure how to come by base models without post-training crap, even the &quot;base&quot; models of recent releases start spamming assistant-type text constantly, i.e. they&#x27;re clearly putting it in the pretraining data.)

          If I&#x27;m right on that then we hit that benchmark like five years ago.

          The egg thing, probably 2030-ish.

          1. Kim_Bruning · · focus · HN ↗
            Do consider the answer given by the Claude model family to be good enough?

            The Claude answer is &#x27;neutral&#x27;, which is sure to anger people at either extreme of the AI debate (and does).

          2. lostmsu · · focus · HN ↗
            &gt; non-RLHF&#x27;d LLMs sound natural enough to pass the Turing test

            No, they will entirely forget things you ask them to remember in the beginning of a conversation. That requires at least an ability to compact context.

        2. dingdongditchme · · focus · HN ↗
          I like your list of &quot;challenges&quot; but have a personal issue with two of them:

          1. &quot;improve uniteds&#x27; plane shedule&quot; -&gt; the word &quot;improve&quot; does a lot of heavy lifting there..

          2. Turing test for me is solved: &quot;AI expert&quot; is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can&#x27;t communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol&#x27; website. In a standard llm session I don&#x27;t really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.

          1. lovich · · focus · HN ↗
            You can buy handwriting machines on Amazon.

            Can’t imagine it’d be too much work to hook up and LLMs output to one.

            <a href="https:&#x2F;&#x2F;www.amazon.com&#x2F;s?k=handwriting+machine+for+letters" rel="nofollow">https:&#x2F;&#x2F;www.amazon.com&#x2F;s?k=handwriting+machine+for+letters

          2. hatthew · · focus · HN ↗
            1. Yeah I don&#x27;t really have a good clarification for this. My thought process went only as far as &quot;flight scheduling and routing optimization is a very difficult problem, would be impressive if an LLM could optimize it&quot;

            2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn&#x27;t trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM &quot;tells&quot;.

            1. CamperBob2 · · focus · HN ↗
              I was setting new goalposts: essentially, solve all known obvious LLM &quot;tells&quot;.

              I don&#x27;t think those are problems to be solved. They are deliberate misfeatures added by the labs through RLHF to keep the models from doing the equivalent of passing a Turing test.

              They don&#x27;t want another GPT-4o, where people threatened to burn down the building, jump off of bridges, etc. when they unplugged it.

              1. tavavex · · focus · HN ↗
                The 4o situation was caused by the model acting like a yes-man and showering users in what they saw as affirmation and support. Its output was still very obviously AI-generated. It wasn&#x27;t too good at being an LLM, OpenAI just went too hard with pumping it full of tricks and behaviors that increase user retention and addiction.

                We can safely say that the AI labs aren&#x27;t deliberately holding back. There&#x27;s too many different companies making their own models, and an LLM that doesn&#x27;t feel like an LLM is too lucrative of an opportunity for one of them not to defect. They are obviously going all out with this and still can&#x27;t get it. I think this is why OP&#x27;s benchmark is so interesting, because it seems that there are a bunch of persistent LLM defects that can&#x27;t be solved definitively. They can try to squeeze it by making these defects less likely, but actually resolving what&#x27;s causing them probably requires a new breakthrough in the field.

                1. CamperBob2 · · focus · HN ↗
                  We&#x27;ll have to agree to disagree on that. The models say &quot;It&#x27;s not X, it&#x27;s Y&quot; because they were trained specifically to say that and not something else. Everything we think of as a surefire slop signal is there because that&#x27;s what the lab wanted.
                  1. tavavex · · focus · HN ↗
                    To clarify, is your opinion that these defects first came up organically and then the AI labs kept artificially reinforcing it to make it obvious that their models produce AI output, but they could fix it perfectly if they wanted to?

                    It seems very hard for me to believe, because everything motivates them to do the opposite. Can you imagine the flood of people and businesses that would come down on a model that actually sounds like a human being? The ability to impersonate a person without any tells would be a dream come true for businesses, marketers and scammers alike. They could put in zero effort and get what looks like normal human behavior in return. That is just too tempting for any of the labs to pass up. Besides, all of them have picked profit over sanity 10 times out of 10, so the argument that they drew just this single line in the sand and none of them ever crossed it on purpose seems unconvincing, especially with the sheer number of models out there, corporate and open.

                    1. CamperBob2 · · focus · HN ↗
                      I see it as the verbal equivalent of the piss filter that OpenAI added to the later Dall-E models. It absolutely was not there in the very beginning, if you look back at the published papers or remember how it worked immediately after they first released it.

                      At least to me, the first Dall-E 3 images were more realistic in some respects than anything that has shipped since. Likewise, I don&#x27;t think it&#x27;s a coincidence that the conversational capabilities of today&#x27;s frontier models aren&#x27;t much better than they were a year or two ago, even though other aspects and capabilities have improved massively. If they wanted a Turing test-capable model with no superficial tells, they&#x27;d have it, so I have to assume they don&#x27;t want it.

                      (Elsewhere in the thread someone else suggests that the conversational degradation&#x2F;lack of improvement might be the result of increased training on synthetic data, and that&#x27;s another theory that sounds reasonable to me. If so, it&#x27;s another thing they could fix if they wanted to.)

                      1. hatthew · · focus · HN ↗
                        Genuine question, not trying to be difficult: how much have you worked with training&#x2F;finetuning generative models? In my experience, models frequently converge to a semi-arbitrary style. I wouldn&#x27;t be surprised if some of the quirks we see in recent gen AI models are artifacts of a particular training run or amplification of unintentional dataset bias, rather than artifacts of an architecture or human guidance.
                        1. CamperBob2 · · focus · HN ↗
                          Not much beyond the Zero-to-Hero videos, TBH. I do know that once bias is identified and labeled, it can be post-trained out... or in.
          3. Dylan16807 · · focus · HN ↗
            Making the judge an expert might be moving the goal posts. But what you&#x27;re describing is not at all serving the purpose of a Turing test. Depending on the situation, having plausibly human-sounding work-related communication with a robot has been possible for decades. It&#x27;s not a Turing test unless the main goal of the conversation is figuring out human versus AI. And it has to be a proper conversation going back and forth many many times.
        3. HanClinto · · focus · HN ↗
          &quot;A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects.&quot;

          -- Robert A. Heinlein

          1. _superposition_ · · focus · HN ↗
            This is literally my linked in bio.
            1. HanClinto · · focus · HN ↗
              It&#x27;s not a complete list, but it really reminds me of the parent comment&#x27;s criteria.

              If AI isn&#x27;t achieving superhuman performance in all of these areas, I&#x27;m not sure we can actually call it &quot;Humanity&#x27;s Last Exam&quot; -- it feels like a bit of an overextension.

    4. pfortuny · · focus · HN ↗
      The Jacobian Conjecture settles the &quot;solve a math problem&quot; and is controversy-free.
    5. an0malous · · focus · HN ↗
      How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers? I find it mind boggling that HN just accepts these shenanigans with no transparency. At the very least, they could share the conversation &#x2F; thinking trace easily and if their claims are true there shouldn’t be anything controversial or negative for their company in the trace.
      1. senordevnyc · · focus · HN ↗
        Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.

        I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).

        Looking forward to the cope when the next big problem falls.

        1. dormento · · focus · HN ↗
          &gt; Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.

          I hope I&#x27;m not being too blunt, but the other alternative is to &quot;just trust me bro&quot; the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don&#x27;t think this is the way.

          1. pessimizer · · focus · HN ↗
            AI can&#x27;t be a religion if you don&#x27;t attack with fire the people for whom &quot;just trust me bro&quot; isn&#x27;t good enough. &quot;Just Trust Them, Bros!&quot; is the central tenet.
          2. senordevnyc · · focus · HN ↗
            Then don&#x27;t trust them?
            1. someonebaggy · · focus · HN ↗
              That&#x27;s what they&#x27;re asking for - by getting the tokens
        2. eviks · · focus · HN ↗
          Yeah, they could indeed easily just share the tokens. And you can ask your favorite AI to explain that it doesn&#x27;t have to be a public release to all the competitors if you can&#x27;t come up with a better alternative yourself. There is a big range between no-one and every-one
        3. _superposition_ · · focus · HN ↗
          Yeah they don&#x27;t want to share the tokens because I guarantee it&#x27;s a bunch of nonsense trial and error leaning on lean for verification. Those tokens will show the lack of intelligence, not the presence of it.
          1. someonebaggy · · focus · HN ↗
            A machine that solves open problems by brute force trial and error is still pretty cool.
      2. off_with_their_ · · focus · HN ↗

        [dead]

      3. olmo23 · · focus · HN ↗
        navier stokes is not the only example of an open problem in maths that was solved by AI (eg the counterexample to the jacobian)
        1. tsunamifury · · focus · HN ↗
          In the other hand why does anyone find it surprising at all that a computer solved a math problem.
          1. stronglikedan · · focus · HN ↗
            Because up until then, computers had not solved math problems. Humans solved them, often using computers as a tool.
            1. tsunamifury · · focus · HN ↗
              I&#x27;m pretty sure computers have been solving math problems for a very very long time through various techniques.
          2. ben_w · · focus · HN ↗
            To ask that question suggests you are unfamiliar with the difference between arithmetic (which computers are good at) and pure mathematics (which is like metaprogramming combined with formal methods, and Gödel&#x27;s incompleteness theorem is worse than the Halting problem)?

            Simply put: For the same reason computers have not already solved all problems in mathematics.

            More concretely:

            Consider the Collatz conjecture. It&#x27;s a very simple rule to write down:

              Take some positive integer: If the number is even, divide it by two; otherwise triple it and add one. With enough repetition, do all positive integers converge to 1?
            
            Trivial to write a program to test numbers starting at 1 and going up. We know it holds up to at least 2.36e21 (according to Wikipedia), but to prove it is true with such a program requires testing all of the infinite set of positive integers.

            But you may notice some things about the rules, that they suggest a subset of numbers will trivially always converge to 1, so that you don&#x27;t need to even test them: any integer 2^n where n is also a positive integer.

            You may find other easy wins, or ways to simplify the test, e.g. once you know the numbers up to m will converge to 1, you can terminate your loop early if you test m+1 and it ever has an intermediate value less than or equal to m. You can combine that with applying one of the rules in reverse, and know that all even numbers between m and 2m will on their first move be halved, making them smaller than m, which means you know they&#x27;ll eventually converge.

            But actually proving this is fully general? Nobody knows. You can&#x27;t just throw arithmetic at the problem directly, you have to figure out patterns that would let you prove that it always holds, no matter what.

            Or, you may find many such patterns and directly calculate some number not in any of them, to find one which doesn&#x27;t converge to 1.

      4. Jtariiiii · · focus · HN ↗
        &gt;How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers?

        The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.

        1. someonebaggy · · focus · HN ↗
          If I want to solve an open problem, stealing a solution to half of it surely makes it much easier. Can I claim i did it myself? Well... Maybe.
      5. 3uruiueijjj · · focus · HN ↗
        Man I can&#x27;t believe mathematicians were just about to solve dozens of significant problems all in the same year, and that&#x27;s exactly the year LLMs come around to steal their solutions.

        Talk about bad luck!

        1. irthomasthomas · · focus · HN ↗
          No one ever spent tens of millions of dollars trying before. To evaluate the achievement we must see all the logs, know how much they spent, and how much help they had from humans.
    6. ralfd · · focus · HN ↗
      &gt; I expected it to be more like &quot;write a 500 page novel that a publisher accepts&quot;,

      It only has 272 pages, but there is a big scandal right now about France&#x27;s most prestigous literary award remobing a critically aclaimed(!) bestseller(!!) novel because it likely was &quot;almost entirely written by AI&quot;

      <a href="https:&#x2F;&#x2F;www.theguardian.com&#x2F;books&#x2F;2026&#x2F;sep&#x2F;25&#x2F;thelyson-orelien-goncourt-prize-france" rel="nofollow">https:&#x2F;&#x2F;www.theguardian.com&#x2F;books&#x2F;2026&#x2F;sep&#x2F;25&#x2F;thelyson-oreli...

      1. mikeocool · · focus · HN ↗
        Note that the author denies the claim. And it also appears that AI detectors are split on the matter, and likely don’t have a rich training corpus of Haitian-French to draw on.

        His publisher also said the original draft was submitted in 2019, though they’ve also become lukewarm in their backing of the author lately (he was also just accused of classic plagiarism for an unrelated short story).

        So hard to say this one is settled.

      2. someonebaggy · · focus · HN ↗
        Note that &quot;bestseller&quot; doesn&#x27;t mean anything and hasn&#x27;t for a long time as it&#x27;s been manipulated to all hell with stuff like wash sales. Neither does &quot;critically acclaimed&quot; which never did.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.