‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. famouswaffles · · focus · HN ↗
      Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there&#x27;s a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn&#x27;t make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.
      1. bigstrat2003 · · focus · HN ↗
        &gt; Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.

        If the models were actually intelligent, the way that the boosters claim, they wouldn&#x27;t need to be tuned to play chess in order to be good at it. That&#x27;s kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

        1. skydhash · · focus · HN ↗
          Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That&#x27;s the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use&#x2F;play something. Mastery can be gained overtime.
          1. willmarch · · focus · HN ↗
            If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn).

            Your assumptions&#x2F;intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).

            1. what · · focus · HN ↗
              &gt; considering LLMs currently play better than a brand new human player would

              They’ve ingested all the literature on playing chess, a brand new human player has not.

              1. willmarch · · focus · HN ↗
                Yes, but my point is that humans can’t even do the thing that the above comments are claiming humans can do (read a book or two and be decent at chess), and then they complain that LLMs can’t do the same thing (that humans can’t do either).

                We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.

                1. skydhash · · focus · HN ↗
                  &gt; The only way you get better at chess is by playing a lot of games and learning from mistakes

                  How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice.

                  This kickstarting then gradual refinement is how most people learn. And the foundational knowledge stays. Even a basic player knows to not do illegal moves.

                  1. willmarch · · focus · HN ↗
                    Reading can kickstart the process, but you can also make random moves guided by some sort of system (such as a computer GUI) or learn by watching other players play. The overall point is that you learn through observation and lots of trial and error (whether you are a human or a computer). And beginners in chess often make illegal moves even after learning the rules, it&#x27;s fairly common.

                    It feels like you&#x27;re trying to say that humans never make illegal moves while learning chess, which doesn&#x27;t match with my experience. I&#x27;m trying to understand your overall point.

                    1. skydhash · · focus · HN ↗
                      &gt; The overall point is that you learn through observation and lots of trial and error

                      That’s the most inefficient way and people usually avoid doing that. Instead they find someone that knows how to do the thing and ask him to be a teacher. Or use a proxy like a book or videos.

                      &gt; It feels like you&#x27;re trying to say that humans never make illegal moves while learning chess, which doesn&#x27;t match with my experience. I&#x27;m trying to understand your overall point

                      There’s learning the basic stuff (which is done after a few games) and there’s mastery. The thread started with the observation that even with all that knowledge (through content ingested in training), LLMs still makes illegal moves. Humans can be erratic, but they can constrain themselves to the rules for the task at hand after learning them.

                      1. willmarch · · focus · HN ↗
                        The only way to master anything is lots of trial and error. Coaches and teachers can help guide you towards more focused trial and error paths but the student still has to do the lessons and put in the work of learning, and learning only truly happens through doing.

                        Humans are not perfect and make mistakes in learning even when they have memorized the rules. A simple example is new players will often move a piece, exposing their king to check, and a more experienced player must point out to them that they have made an illegal move (because a new player often has not encoded that pattern for looking for exposed checks because they&#x27;re more focused on how the pieces move, not what that piece exposes.)

                        We&#x27;re just going to have to agree to disagree here.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.