‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. intelkishan · · focus · HN ↗
      Could you share your other 5 tests, if they are public?
      1. hatthew · · focus · HN ↗
        Here&#x27;s my previous comment: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902

        I didn&#x27;t really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name &quot;humanity&#x27;s last exam&quot;. So take it with a grain of salt.

        1. andai · · focus · HN ↗
          Thanks for sharing.

          I think non-RLHF&#x27;d LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don&#x27;t know if anyone has tested them for that. (Also I&#x27;m not sure how to come by base models without post-training crap, even the &quot;base&quot; models of recent releases start spamming assistant-type text constantly, i.e. they&#x27;re clearly putting it in the pretraining data.)

          If I&#x27;m right on that then we hit that benchmark like five years ago.

          The egg thing, probably 2030-ish.

          1. Kim_Bruning · · focus · HN ↗
            Do consider the answer given by the Claude model family to be good enough?

            The Claude answer is &#x27;neutral&#x27;, which is sure to anger people at either extreme of the AI debate (and does).

          2. lostmsu · · focus · HN ↗
            &gt; non-RLHF&#x27;d LLMs sound natural enough to pass the Turing test

            No, they will entirely forget things you ask them to remember in the beginning of a conversation. That requires at least an ability to compact context.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.