‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. sobellian · · focus · HN ↗
        I tested both myself and a weak bot against Astra xhigh, <a href="https:&#x2F;&#x2F;lichess.org&#x2F;study&#x2F;27lCQqDa" rel="nofollow">https:&#x2F;&#x2F;lichess.org&#x2F;study&#x2F;27lCQqDa. It&#x27;s still pretty bad at chess, though it takes longer to devolve into illegal moves.
        1. losvedir · · focus · HN ↗
          &gt; though it takes longer to devolve into illegal moves

          Is this because the context is being saturated? How did you set it up?

          Was the prompt something like &quot;Here&#x27;s the state of the board, you&#x27;re white, your move, what do you do?&quot; and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (thanks for sharing!) to my mental model and understanding.

          I&#x27;d be curious how it would work if it started fresh each time. My guess is it would never make an illegal move, although it may not actually play all that well.

          1. sobellian · · focus · HN ↗
            You can see the entire conversation for my game at <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aaac17b-1384-83e8-98fd-4350a0ef69cd" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aaac17b-1384-83e8-98fd-4350a0ef69....
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.