‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. zahlman · · focus · HN ↗
        Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with:

        &gt; Let&#x27;s play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.

        (I hoped the latter requirement would help it be &quot;not blindfolded&quot;; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)

        For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.

        It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the &quot;book&quot; theory in its training data, but it still completely fell apart at early midgame.

        1. titzer · · focus · HN ↗

          [dead]

          1. lirolero · · focus · HN ↗

            [dead]

          2. danpalmer · · focus · HN ↗
            Sure, but installing a chess program is child&#x2F;teen level general ability, and playing chess well is highly trained expert level ability. Which one are we sold AI as being?
            1. alpinisme · · focus · HN ↗
              I think we are being sold AI as expert only when given tools (although that is not emphasized). The (quasi?) miracle of AI right now is that you can get an agent to accomplish the task of a team of intelligent but not exceptional humans at speeds far exceeding what the human could do. Which makes it “cheap” to throw (effectively) dozens of teams at a problem for the equivalent of hundreds of man hours.

              That may not be the AI of sci fi fantasy but it’s still a game changing reality.

          3. kavok · · focus · HN ↗
            I often don’t see agents reaching for available or potential tools&#x2F;libraries unless explicitly told to.

            Sometimes they’ll even manually search or write bespoke code to search json instead of using something like jq.

          4. HarHarVeryFunny · · focus · HN ↗
            A Transformer has a massive amount of state - it&#x27;s entire KV cache, in addition to the user asking it to draw the state after every move, which is really unnecessary.

            A human, at least a trained human (for fairer comparison to an LLM whose training data contained a ton of chess games) can absolutely do this - have you never seen demonstrations of expert players playing a dozen or more games while blindfolded?

            A Transformer&#x2F;LLM is not a human of course, and the way it will by default play chess is by prediction, not reasoning. An LLM actually does surprisingly well if you only give it the most recent 20 moves of a game where 40 moves have been played so far, since the moves NOT played tell it just as much as the ones that were played, letting it effectively infer a lot of what is on the board.

            1. zahlman · · focus · HN ↗
              I just want to make sure it&#x27;s clear: the reason I was asking it to redraw the board is because last time I tried (which was like a month ago), I didn&#x27;t ask for that, and basically as soon as the opening was &quot;out of book&quot; it started trying to make illegal moves and made false statements about the position in its running commentary (and after being corrected on these points, started dropping pieces for no reason).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.