‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. WhitneyLand · · focus · HN ↗
      1. It’s hard to trust a 2026 paper that’s showing results for such old models.

      2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.

      3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

      1. carodgers · · focus · HN ↗
        &gt; Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

        A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would &quot;destroy any human at chess.&quot;

        Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?

        1. nimbleal · · focus · HN ↗
          Maybe practically it doesn’t matter? Perhaps AGI is not the model but the model plus everything it’s got access to. If we’re modelling intelligence in the way we seem to have to to have any coherent definition of AGI, it seems to me &lt;model + everything it can access&gt; is always going to be more “intelligent” than &lt;model&gt; alone.
          1. Planktonne · · focus · HN ↗
            That would mean we should consider any human with coding knowledge a chess grandmaster, which is obviously not the case.
            1. empath75 · · focus · HN ↗
              If the goal is merely to &quot;win at chess&quot;, then yes, an LLM using stockfish is better than any human alone at performing the task. When you are talking about what AI agents are capable of doing, there is no such thing as &quot;cheating&quot;. They are as capable as the tools they can use effectively. The entire history of human civilization was driven by effectively using tools to achieve goals.
              1. Planktonne · · focus · HN ↗
                But then the human should also get Stockfish, and we&#x27;re back at a stalemate.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.