‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. tossandthrow · · focus · HN ↗
      Llm systems are not really build for adhering to a grammar (other than &quot;a string og tokens&quot;).

      It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents.

      Certainly,a harness can easily correct for it.

      1. vkazanov · · focus · HN ↗
        By the promise of it, llms should be able to both adhere to grammars, or go free form where necessary. I mean, doing math is supposed to be strict but in practice it&#x27;s a somewhat educated random walk in the space of correct lean theorems.

        Harnesses do correct things, sure.

        1. tossandthrow · · focus · HN ↗
          You are right. I am imprecise.

          Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.

          Games and programming languages (including lean) does not allow this flexibility.

          A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.

          Certainly it must be like that, otherwise reviews in math was rendered moot.

          Do we blame research mathematicians for not adhering to the grammar?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.