‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. empath75 · · focus · HN ↗
      I want you to consider how relevant this is in any practical sense.

      First -- most _people_ cannot do this, without having a physical board in front of them.

      Second -- Claude Code is perfectly capable of downloading and running stockfish. People focus too much on LLMs by themselves as the entity of concern instead of the entire harness and all of it&#x27;s capabilities together.

      1. Capricorn2481 · · focus · HN ↗
        Because they are obviously testing for general intelligence. If you want a thread about how cool the harness is, that&#x27;s down the street.

        We don&#x27;t really consider humans downloading stockfish to beat people at chess as noteworthy endeavors.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.