‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. zahlman · · focus · HN ↗
        Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with:

        &gt; Let&#x27;s play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.

        (I hoped the latter requirement would help it be &quot;not blindfolded&quot;; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)

        For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.

        It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the &quot;book&quot; theory in its training data, but it still completely fell apart at early midgame.

        1. sailfast · · focus · HN ↗
          What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill?

          Also what harness? If you’re using a general harness of course it’s going to try and give you commentary.

          I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.

          1. topaz0 · · focus · HN ↗
            You&#x27;re pointing out that the goalposts are not fixed in the problem statement above, and gp&#x27;s interpretation is not the most generous possible. But as the interpretations get more generous, the claim becomes more and more absurd. Maybe a properly-harnessed model would download the most advanced chess engine and query it to find the best move in each position, but that&#x27;s not really demonstrating the model&#x27;s intelligence anymore.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.