Vote on which of Hacker News' challenges for AI have been met
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Vote on which of Hacker News' challenges for AI have been met
Unofficial Hacker News client; not affiliated with Y Combinator.
ErrantX · · focus · HN ↗
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
ianjbutler · · focus · HN ↗
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
johnsmith1840 · · focus · HN ↗
Kotlopou · · focus · HN ↗
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
[0]: <a href="https://arxiv.org/pdf/2503.23674" rel="nofollow">https://arxiv.org/pdf/2503.23674 (now published at <a href="https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123" rel="nofollow">https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.
[1]: <a href="https://turingtest.live/" rel="nofollow">https://turingtest.live/
[2]: "Dull Rigid Human meets Ace Mechanical Translator" (<a href="https://www.cambridge.org/core/books/abs/once-and-future-turing/dull-rigid-human-meets-ace-mechanical-translator/9E307F63E1D8FD23E0D3447EF4BCFD75" rel="nofollow">https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)
johnsmith1840 · · focus · HN ↗
Turing test does not mean perfectly human it just means you can talk to one without knowing that has been passed for a long time now.
I have not been able to tell for a long time now especially if I directly give it human like writing instructions for outbound content.
suddenlybananas · · focus · HN ↗
Eliza can sometimes pass the Turing test!
avadodin · · focus · HN ↗
Kotlopou · · focus · HN ↗
sokoloff · · focus · HN ↗