‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. intelkishan · · focus · HN ↗
      Could you share your other 5 tests, if they are public?
      1. hatthew · · focus · HN ↗
        Here&#x27;s my previous comment: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=42809902

        I didn&#x27;t really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name &quot;humanity&#x27;s last exam&quot;. So take it with a grain of salt.

        1. andai · · focus · HN ↗
          Thanks for sharing.

          I think non-RLHF&#x27;d LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don&#x27;t know if anyone has tested them for that. (Also I&#x27;m not sure how to come by base models without post-training crap, even the &quot;base&quot; models of recent releases start spamming assistant-type text constantly, i.e. they&#x27;re clearly putting it in the pretraining data.)

          If I&#x27;m right on that then we hit that benchmark like five years ago.

          The egg thing, probably 2030-ish.

          1. Kim_Bruning · · focus · HN ↗
            Do consider the answer given by the Claude model family to be good enough?

            The Claude answer is &#x27;neutral&#x27;, which is sure to anger people at either extreme of the AI debate (and does).

          2. lostmsu · · focus · HN ↗
            &gt; non-RLHF&#x27;d LLMs sound natural enough to pass the Turing test

            No, they will entirely forget things you ask them to remember in the beginning of a conversation. That requires at least an ability to compact context.

        2. dingdongditchme · · focus · HN ↗
          I like your list of &quot;challenges&quot; but have a personal issue with two of them:

          1. &quot;improve uniteds&#x27; plane shedule&quot; -&gt; the word &quot;improve&quot; does a lot of heavy lifting there..

          2. Turing test for me is solved: &quot;AI expert&quot; is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can&#x27;t communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol&#x27; website. In a standard llm session I don&#x27;t really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.

          1. lovich · · focus · HN ↗
            You can buy handwriting machines on Amazon.

            Can’t imagine it’d be too much work to hook up and LLMs output to one.

            <a href="https:&#x2F;&#x2F;www.amazon.com&#x2F;s?k=handwriting+machine+for+letters" rel="nofollow">https:&#x2F;&#x2F;www.amazon.com&#x2F;s?k=handwriting+machine+for+letters

          2. hatthew · · focus · HN ↗
            1. Yeah I don&#x27;t really have a good clarification for this. My thought process went only as far as &quot;flight scheduling and routing optimization is a very difficult problem, would be impressive if an LLM could optimize it&quot;

            2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn&#x27;t trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM &quot;tells&quot;.

            1. CamperBob2 · · focus · HN ↗
              I was setting new goalposts: essentially, solve all known obvious LLM &quot;tells&quot;.

              I don&#x27;t think those are problems to be solved. They are deliberate misfeatures added by the labs through RLHF to keep the models from doing the equivalent of passing a Turing test.

              They don&#x27;t want another GPT-4o, where people threatened to burn down the building, jump off of bridges, etc. when they unplugged it.

              1. tavavex · · focus · HN ↗
                The 4o situation was caused by the model acting like a yes-man and showering users in what they saw as affirmation and support. Its output was still very obviously AI-generated. It wasn&#x27;t too good at being an LLM, OpenAI just went too hard with pumping it full of tricks and behaviors that increase user retention and addiction.

                We can safely say that the AI labs aren&#x27;t deliberately holding back. There&#x27;s too many different companies making their own models, and an LLM that doesn&#x27;t feel like an LLM is too lucrative of an opportunity for one of them not to defect. They are obviously going all out with this and still can&#x27;t get it. I think this is why OP&#x27;s benchmark is so interesting, because it seems that there are a bunch of persistent LLM defects that can&#x27;t be solved definitively. They can try to squeeze it by making these defects less likely, but actually resolving what&#x27;s causing them probably requires a new breakthrough in the field.

                1. CamperBob2 · · focus · HN ↗
                  We&#x27;ll have to agree to disagree on that. The models say &quot;It&#x27;s not X, it&#x27;s Y&quot; because they were trained specifically to say that and not something else. Everything we think of as a surefire slop signal is there because that&#x27;s what the lab wanted.
                  1. tavavex · · focus · HN ↗
                    To clarify, is your opinion that these defects first came up organically and then the AI labs kept artificially reinforcing it to make it obvious that their models produce AI output, but they could fix it perfectly if they wanted to?

                    It seems very hard for me to believe, because everything motivates them to do the opposite. Can you imagine the flood of people and businesses that would come down on a model that actually sounds like a human being? The ability to impersonate a person without any tells would be a dream come true for businesses, marketers and scammers alike. They could put in zero effort and get what looks like normal human behavior in return. That is just too tempting for any of the labs to pass up. Besides, all of them have picked profit over sanity 10 times out of 10, so the argument that they drew just this single line in the sand and none of them ever crossed it on purpose seems unconvincing, especially with the sheer number of models out there, corporate and open.

                    1. CamperBob2 · · focus · HN ↗
                      I see it as the verbal equivalent of the piss filter that OpenAI added to the later Dall-E models. It absolutely was not there in the very beginning, if you look back at the published papers or remember how it worked immediately after they first released it.

                      At least to me, the first Dall-E 3 images were more realistic in some respects than anything that has shipped since. Likewise, I don&#x27;t think it&#x27;s a coincidence that the conversational capabilities of today&#x27;s frontier models aren&#x27;t much better than they were a year or two ago, even though other aspects and capabilities have improved massively. If they wanted a Turing test-capable model with no superficial tells, they&#x27;d have it, so I have to assume they don&#x27;t want it.

                      (Elsewhere in the thread someone else suggests that the conversational degradation&#x2F;lack of improvement might be the result of increased training on synthetic data, and that&#x27;s another theory that sounds reasonable to me. If so, it&#x27;s another thing they could fix if they wanted to.)

                      1. hatthew · · focus · HN ↗
                        Genuine question, not trying to be difficult: how much have you worked with training&#x2F;finetuning generative models? In my experience, models frequently converge to a semi-arbitrary style. I wouldn&#x27;t be surprised if some of the quirks we see in recent gen AI models are artifacts of a particular training run or amplification of unintentional dataset bias, rather than artifacts of an architecture or human guidance.
                        1. CamperBob2 · · focus · HN ↗
                          Not much beyond the Zero-to-Hero videos, TBH. I do know that once bias is identified and labeled, it can be post-trained out... or in.
          3. Dylan16807 · · focus · HN ↗
            Making the judge an expert might be moving the goal posts. But what you&#x27;re describing is not at all serving the purpose of a Turing test. Depending on the situation, having plausibly human-sounding work-related communication with a robot has been possible for decades. It&#x27;s not a Turing test unless the main goal of the conversation is figuring out human versus AI. And it has to be a proper conversation going back and forth many many times.
        3. HanClinto · · focus · HN ↗
          &quot;A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects.&quot;

          -- Robert A. Heinlein

          1. _superposition_ · · focus · HN ↗
            This is literally my linked in bio.
            1. HanClinto · · focus · HN ↗
              It&#x27;s not a complete list, but it really reminds me of the parent comment&#x27;s criteria.

              If AI isn&#x27;t achieving superhuman performance in all of these areas, I&#x27;m not sure we can actually call it &quot;Humanity&#x27;s Last Exam&quot; -- it feels like a bit of an overextension.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.