‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. threethirtytwo · · focus · HN ↗
      The story isn&#x27;t so clear cut.

      The caveat is: It depends on the task.

      Are there reams of chess moves that the model can train off of? No.

      Are there reams of math papers the model can train off of? Yes.

      1. keephnacct · · focus · HN ↗

        [dead]

      2. [deleted] · · focus · HN ↗

        [deleted]

      3. iwontberude · · focus · HN ↗

        [dead]

      4. tjwebbnorfolk · · focus · HN ↗
        &gt; Are there reams of chess moves that the model can train off of? No.

        This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.

        1. XenophileJKO · · focus · HN ↗
          It is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don&#x27;t even need any data to start with (but would help).
          1. manquer · · focus · HN ↗
            There are more possible game combinations than atoms in the universe, even those generation of valid game states are as you say pre-defined. that is why models cannot go this route and therefore are poor at chess
            1. wat10000 · · focus · HN ↗
              Isn’t this exactly how AlphaZero was trained? The rules are known and well defined so the training process can generate games without any outside data.

              The only reason LLMs are this bad at chess is because the labs don’t care about chess performance so they’re not going out of their way to train the models for it. The ability they do have is from what chess information happens to be in the training data, plus whatever general reasoning abilities they may be able to apply.

        2. threethirtytwo · · focus · HN ↗
          Let me make my statement more clear with a correction:

          Was there reams of chess moves that the model trained off of? No.

      5. vmg12 · · focus · HN ↗
        &gt; The caveat is: It depends on the task.

        I think the line of criticism around LLMs sucking at chess makes more sense when you understand what the AI companies are saying about the future trajectory of these models.

        The entire recursive self improvement story falls apart once you point out that there is not much &quot;cross domain transfer learning&quot;. Meaning that training an LLM to become good at coding, math, etc, will eventually transfer into them being good at other skills that were not explicitly trained for.

        Using games like chess which have little economic value is actually a good test for this. What&#x27;s even more surprising about them sucking at chess is how much information about chess strategy exists in the training data.

      6. freejazz · · focus · HN ↗
        &gt;Are there reams of chess moves that the model can train off of? No.

        For real??

        1. threethirtytwo · · focus · HN ↗
          I meant if there are reams of chess moves the model was trained off of.
          1. freejazz · · focus · HN ↗
            Yeah, it&#x27;s not like there&#x27;s any literature about Chess in the corpus of these models!
            1. threethirtytwo · · focus · HN ↗
              There’s literature. But I don’t think there’s reams of chess games. With literature the LLM can understand strategy but chess needs intuition and you need tokenized games for that. LLMs have less of that.
      7. FuckButtons · · focus · HN ↗
        There’s multiple databases of games in algebraic notation. You can also, very easily rl train on pitting models against one another, even without mcts.
      8. thelaxiankey · · focus · HN ↗
        there are far more reams of chess moves than there are math papers. Lichess is pretty open...

        But hey, they&#x27;re actually good at chess if you prompt correctly so.... <a href="https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;" rel="nofollow">https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.