‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. MattCruikshank · · focus · HN ↗
      What happens when you ask those same frontier models to write a chess-playing program?

      I feel like, this is a huge stumbling block that many people have. They&#x27;ll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn&#x27;t fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.

      It&#x27;s really neat to see what a frontier model can do itself. No doubt.

      But &quot;play chess by hand&quot; is a frankly awful metric. It&#x27;s kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.

      1. cbolton · · focus · HN ↗
        It&#x27;s a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it&#x27;s noteworthy that they underperform on that test.

        Letting the model execute a chess program (that it wrote) would make sense if you&#x27;re measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.

        1. empath75 · · focus · HN ↗
          &gt; It&#x27;s a great test of cognitive abilities.

          It isn&#x27;t. Stockfish running on your laptop can beat every human being on earth easily at chess. It&#x27;s not intelligent _at all_ in any sense that matters.

          1. cbolton · · focus · HN ↗
            Well it&#x27;s not a perfect test so you need a bit of care in how you use it. If you have no idea what the subject is doing, then you don&#x27;t know if you&#x27;re measuring cognitive ability or something else (like cheating ability, or algorithmic sophistication or whatever). But failing the test is a pretty clear sign of certain cognitive abilities being poor.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.