‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. ErrantX · · focus · HN ↗
    What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.

    And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.

    But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.

    That alone tells you a lot IMO

    1. ianjbutler · · focus · HN ↗
      Sigh, the whole "obviously the turing test is solved" meme is annoying.

      Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?

      More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.

      The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.

      1. johnsmith1840 · · focus · HN ↗
        I was just thinking how anyone still thought AI didn't pass turing already. There's been literal papers proving average people cannot tell reliably.
        1. ianjbutler · · focus · HN ↗
          Average people cannot tell reliably isn’t an appropriate or interesting test tho, otherwise Eliza and markov models etc. the framing that matters is explicitly adversarial. Play like your life depends on it instead of rooting for the machine, and you can’t win?

          One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.

          Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!

          Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.

          1. johnsmith1840 · · focus · HN ↗
            I found a cladder example that hits 97.7% passing on that benchmark? And it's like an insanely small dumb model that hit it.

            I don't think philosophy has any real value in assesment here. I'm bias but even before AI I thought it wasn't accurate model of how thought works and I think AI has reinforced that.

            Reminds me of the 4 humors of medicine in medieval europe. It has some truth but it's not really accurate.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.