‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. zahlman · · focus · HN ↗
        Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with:

        &gt; Let&#x27;s play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.

        (I hoped the latter requirement would help it be &quot;not blindfolded&quot;; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)

        For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.

        It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the &quot;book&quot; theory in its training data, but it still completely fell apart at early midgame.

        1. sailfast · · focus · HN ↗
          What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill?

          Also what harness? If you’re using a general harness of course it’s going to try and give you commentary.

          I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.

          1. datsci_est_2015 · · focus · HN ↗
            Why does a 6 year old not need any of these guardrails?

            Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking &#x2F; influencing about.

            Maybe not all AGIs have a path to digital singularity. Maybe our current era of intelligence modeling has fundamental flaws and we are in a local minimum of the artificial intelligence space.

            To note, I would bet with a good amount of certainty that we have enough compute power and automation to DDOS the internet out of existence with botnets. That doesn’t make the frontier models intelligent, that just makes their handlers reckless.

            1. trio8453 · · focus · HN ↗
              &gt; Why does a 6 year old not need any of these guardrails?

              They&#x27;re not guardrails, they&#x27;re a different input&#x2F;output environment.

            2. solenoid0937 · · focus · HN ↗
              Ask a 6 year old to draw a chess board from scratch every turn and they too will make mistakes.
              1. datsci_est_2015 · · focus · HN ↗
                A 6 year old will figure out how to ask you to help them after they get it wrong.
              2. freejazz · · focus · HN ↗
                No one has spent the past three years telling me that a 6 year old will take my job!!!
                1. claytongulick · · focus · HN ↗
                  And the 6 year old doesn&#x27;t cost more than the GDP of a medium sized country.
                  1. sailfast · · focus · HN ↗
                    [delayed]
              3. wavemode · · focus · HN ↗
                [delayed]
            3. gf000 · · focus · HN ↗
              Well, would a dissected frontal lobe in and of itself be intelligence?

              I think the same goes for LLMs, they may be a core part of an LLM harness, but you may still need a couple other components (e.g. it may itself write itself a deterministic function to validate steps).

              In and of itself intelligence is an ill-defined and badly understood concept.

            4. themgt · · focus · HN ↗
              Why does a 6 year old not need any of these guardrails?

              Why does a bird not need jet engines or regular professional maintenance?

          2. topaz0 · · focus · HN ↗
            You&#x27;re pointing out that the goalposts are not fixed in the problem statement above, and gp&#x27;s interpretation is not the most generous possible. But as the interpretations get more generous, the claim becomes more and more absurd. Maybe a properly-harnessed model would download the most advanced chess engine and query it to find the best move in each position, but that&#x27;s not really demonstrating the model&#x27;s intelligence anymore.
          3. zahlman · · focus · HN ↗
            &gt; but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.

            This isn&#x27;t just about judging LLM capability. This is about pointing out that these capabilities are not &quot;AGI&quot;. If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.

            Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it&#x27;s hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.