‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. consensus1 · · focus · HN ↗
      This isn&#x27;t how intelligence works. The LLM may not be able to play chess directly through inference, but it can write a program to do it and execute that program. Same as how human intelligence works. We can&#x27;t fly, but we can build planes.
      1. nefarious_ends · · focus · HN ↗
        Thanks for saying this, feels like everyone has gone insane over this stuff.
        1. what · · focus · HN ↗
          Humans don’t code a $game engine to play $game, they can just play it. It seems like you are the one that has gone insane.
          1. hackinthebochs · · focus · HN ↗
            And how many years of direct play and study does it take for a human to get good at chess or any other game? Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. That&#x27;s just not how the brain works. If LLMs could do that they would truly be superintelligence.
            1. sph · · focus · HN ↗
              No, learning is definitely not a sign of super intelligence. I know words don’t mean anything anymore, but that is simply general intelligence, despite the claims we have reached this milestone.
              1. hackinthebochs · · focus · HN ↗
                No, but superhuman capabilities derived from learning is, which is what the parent comment described.
            2. lelanthran · · focus · HN ↗
              &gt; And how many years of direct play and study does it take for a human to get good at chess or any other game?

              Time is irrelevant to training; the more relevant comparison is &quot;how many games does a human need to play to get diminishing returns&quot;.

              1. hackinthebochs · · focus · HN ↗
                Yes, obviously. The point was simply that human&#x27;s don&#x27;t one-shot chess so why does anyone expect an LLM to?
            3. orwin · · focus · HN ↗
              A week. My brother learned and was above 1100 online within 12 hours, after a few hundred games.
              1. hackinthebochs · · focus · HN ↗
                We&#x27;re obviously using different meanings for &quot;good&quot; here. But aside from that, it took 100&#x27;s to 1000&#x27;s of reinforcement iterations for your brother to play competently. While certainly impressive, that is still an entirely different category from piecing together disparate facts learned during training (LLMs aren&#x27;t analyzing a board as they&#x27;re learning the rules or ingesting PNG files), to executing a competent performance in one shot.
            4. mtlmtlmtlmtl · · focus · HN ↗
              &gt; Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess.

              Maybe not, but you&#x27;d be surprised how little it takes.

              A six year old child can learn the rules of chess well enough to be able to play legal moves only in a single day. And they can improve their game at a pace which is almost frightening to behold. I have taught children, and I&#x27;ve witnessed significant improvement materialise in a single game. LLMs have probably thousands of chess books, games, videos, etc in their training data, yet they are unable to even follow the rules.

              This is, at the very least, interesting. It illustrates many of the things brains can do, which current ML systems in general, and LLMs in particular, can&#x27;t.

              1. hackinthebochs · · focus · HN ↗
                It is interesting, but people are drawing the wrong conclusion from it. For one, LLMs don&#x27;t go through a &quot;chess learning phase&quot;. They&#x27;re not analyzing a board as they&#x27;re learning the rules or studying games to create a coherent model of chess. They&#x27;re just imbibing raw relationships as disparate fragments of information. The fact that they can&#x27;t unify this into a coherent model of chess playing in one shot and execute a competent game of chess says nothing interesting about the limits of their intelligence. If you give frontier models the rules of chess in their context window, could they perform only legal moves? I bet they could, excepting trickier scenarios like pins and failing to respond to a check. But those kinds of scenarios have to be reinforced in any human player as well.
      2. thesmtsolver2 · · focus · HN ↗
        Human beings can play chess directly without coding up a tool.
        1. consensus1 · · focus · HN ↗
          Very poorly compared to the tools we have built. Similar to the LLM.
          1. shimman · · focus · HN ↗
            Poorly in what sense? I think human chess leagues are way more popular and fun than just playing a computer by yourself. Human oriented communities are always a vastly better experience than their digital counterparts.

            There&#x27;s more to games than simply winning you know.

          2. thesmtsolver2 · · focus · HN ↗
            Comparing to raw LLMs? Much much better.
        2. qarl · · focus · HN ↗
          If they wanted to train an LLM to play chess they could easily do so.

          But nobody wants that.

        3. jstanley · · focus · HN ↗
          Asking an LLM to play chess by writing algebraic notation is like asking a human to play chess blindfolded.

          Yes some people can do it but most people can&#x27;t even if they&#x27;re unusually intelligent.

          You really need to be giving the LLM a board representation.

          EDIT: I see that they actually were giving the LLMs a board representation and they still played badly. Fair enough then.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.