‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. ErrantX · · focus · HN ↗
    What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.

    And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.

    But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.

    That alone tells you a lot IMO

    1. ianjbutler · · focus · HN ↗
      Sigh, the whole "obviously the turing test is solved" meme is annoying.

      Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?

      More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.

      The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.

      1. johnsmith1840 · · focus · HN ↗
        I was just thinking how anyone still thought AI didn't pass turing already. There's been literal papers proving average people cannot tell reliably.
        1. Kotlopou · · focus · HN ↗
          AFAIK people refer to this paper [0]. I think it only proves very little, because a typical conversation they studied looks like this:

          Q: do you like doing psych studies and why?

          A: theyre chill, easy money tbh

          Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?

          A: nah i just get the box mix lol

          Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?

          A: axolotl, theyre weirdly cute

          And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.

          IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.

          [0]: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2503.23674" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2503.23674 (now published at <a href="https:&#x2F;&#x2F;www.pnas.org&#x2F;doi&#x2F;epdf&#x2F;10.1073&#x2F;pnas.2524472123" rel="nofollow">https:&#x2F;&#x2F;www.pnas.org&#x2F;doi&#x2F;epdf&#x2F;10.1073&#x2F;pnas.2524472123). This is the top result in Google Scholar for &quot;Turing test&quot; from 2025 onwards.

          [1]: <a href="https:&#x2F;&#x2F;turingtest.live&#x2F;" rel="nofollow">https:&#x2F;&#x2F;turingtest.live&#x2F;

          [2]: &quot;Dull Rigid Human meets Ace Mechanical Translator&quot; (<a href="https:&#x2F;&#x2F;www.cambridge.org&#x2F;core&#x2F;books&#x2F;abs&#x2F;once-and-future-turing&#x2F;dull-rigid-human-meets-ace-mechanical-translator&#x2F;9E307F63E1D8FD23E0D3447EF4BCFD75" rel="nofollow">https:&#x2F;&#x2F;www.cambridge.org&#x2F;core&#x2F;books&#x2F;abs&#x2F;once-and-future-tur... or alternative access methods thereof)

          1. johnsmith1840 · · focus · HN ↗
            I mean that example is literally passing turing to me. It talks exactly like a human, what do you want it to do?

            Turing test does not mean perfectly human it just means you can talk to one without knowing that has been passed for a long time now.

            I have not been able to tell for a long time now especially if I directly give it human like writing instructions for outbound content.

            1. suddenlybananas · · focus · HN ↗
              &gt;It talks exactly like a human, what do you want it to do?

              Eliza can sometimes pass the Turing test!

              1. avadodin · · focus · HN ↗
                Tell me more
                1. Kotlopou · · focus · HN ↗
                  Read the linked [0] paper! The detection rate for ELIZA was far below 100%.
                  1. sokoloff · · focus · HN ↗
                    “Tell me more” was a stock ELIZA prompt, making this an interesting meta-Turing test exchange.
                  2. zahlman · · focus · HN ↗
                    …I think the comment you&#x27;re replying to was intended as satire: demonstrating how ELIZA might respond.
                    1. Kotlopou · · focus · HN ↗
                      :D Oops! Time to log off...
            2. Kotlopou · · focus · HN ↗
              I don&#x27;t mean that this example is somehow robotic, only that it&#x27;s absurd for the Turing Test to be four lines long and with no adversarial attempts. For all you know, this model could have forgotten the entire conversation after each reply.

              Conversations with strangers can be hard to get going, but they aren&#x27;t this bad.

              1. johnsmith1840 · · focus · HN ↗
                The difference is it could be 4 lines on essantially any topic a human knows. It&#x27;s turing test passing an inch deep and a million miles wide and it&#x27;s getting deeper all the time.
        2. Gtex555 · · focus · HN ↗
          Sure but you havnt addressed his main point, why are people still complaining about AI slop post or AI slop emails if the turning test has been solved. Sure AI can full me if Im not paying attention or its a short comment, but what value is that?
          1. BeetleB · · focus · HN ↗
            It seems reasoning skills are declining rapidly here.

            That some models with some system prompts don&#x27;t pass the Turing test doesn&#x27;t mean other models with other prompts can&#x27;t.

          2. johnsmith1840 · · focus · HN ↗
            So the goal post is that every instance of AI must pass turing test?

            I didn&#x27;t respond to it because it&#x27;s a bad argument.

        3. Barrin92 · · focus · HN ↗
          &gt;There&#x27;s been literal papers proving average people cannot tell reliably.

          the average person reads at a 7th grade level and can&#x27;t tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:

          <a href="https:&#x2F;&#x2F;pastebin.com&#x2F;NjfCLSXa" rel="nofollow">https:&#x2F;&#x2F;pastebin.com&#x2F;NjfCLSXa

          1. artisin · · focus · HN ↗
            this gave me a good chuckle.

            &gt; Oh, right. Aramaic. My mistake—I somehow read &quot;Italian&quot; and didn&#x27;t question it. &gt; I can try, although my Aramaic is very rusty. Do you mean Classical&#x2F;Syriac Aramaic, or one of the modern varieties?

          2. kelseyfrog · · focus · HN ↗
            We have to remember that one of these average people is also participating as a player; the bar for AI is also that low.
            1. zahlman · · focus · HN ↗
              You&#x27;re missing the point, which is that the AI is completely failing to lower its presented capability to a human level, in the areas where it has an advantage. The fact that &quot;these average people&quot; would be even more hopeless at repeating themselves in arbitrarily chosen languages (including dead ones) than highly intelligent and&#x2F;or well-studied people, is exactly the point. Whereas translating between human languages is a task you&#x27;d naturally expect an LLM to be especially well-suited to.
          3. woobar · · focus · HN ↗
            Interesting. Just tried similar convo (i did warn Claude I will trick it) and it did not fail so spectacularly

            <a href="https:&#x2F;&#x2F;pastebin.com&#x2F;sVKA1tyw" rel="nofollow">https:&#x2F;&#x2F;pastebin.com&#x2F;sVKA1tyw

            1. woobar · · focus · HN ↗

              [dead]

        4. Mikhail_Edoshin · · focus · HN ↗
          I once sat in a barbershop and talked with the barber. At first I listened to her sympathetically but then realized she was mad; at least she had noticeable psychical problems. We cannot quickly conclude a person is mad, can we? Even specialists cannot. AI is similar. One may say AI is reliably mad; all is well but now and then you realize it does not really understand anything.
        5. ianjbutler · · focus · HN ↗
          Average people cannot tell reliably isn’t an appropriate or interesting test tho, otherwise Eliza and markov models etc. the framing that matters is explicitly adversarial. Play like your life depends on it instead of rooting for the machine, and you can’t win?

          One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.

          Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!

          Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.

          1. johnsmith1840 · · focus · HN ↗
            I found a cladder example that hits 97.7% passing on that benchmark? And it&#x27;s like an insanely small dumb model that hit it.

            I don&#x27;t think philosophy has any real value in assesment here. I&#x27;m bias but even before AI I thought it wasn&#x27;t accurate model of how thought works and I think AI has reinforced that.

            Reminds me of the 4 humors of medicine in medieval europe. It has some truth but it&#x27;s not really accurate.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.