‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. intelkishan · · focus · HN ↗
      Could you share your other 5 tests, if they are public?
      1. hatthew · · focus · HN ↗
        Here&#x27;s my previous comment: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902

        I didn&#x27;t really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name &quot;humanity&#x27;s last exam&quot;. So take it with a grain of salt.

        1. dingdongditchme · · focus · HN ↗
          I like your list of &quot;challenges&quot; but have a personal issue with two of them:

          1. &quot;improve uniteds&#x27; plane shedule&quot; -&gt; the word &quot;improve&quot; does a lot of heavy lifting there..

          2. Turing test for me is solved: &quot;AI expert&quot; is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can&#x27;t communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol&#x27; website. In a standard llm session I don&#x27;t really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.

          1. Dylan16807 · · focus · HN ↗
            Making the judge an expert might be moving the goal posts. But what you&#x27;re describing is not at all serving the purpose of a Turing test. Depending on the situation, having plausibly human-sounding work-related communication with a robot has been possible for decades. It&#x27;s not a Turing test unless the main goal of the conversation is figuring out human versus AI. And it has to be a proper conversation going back and forth many many times.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.