‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. famouswaffles · · focus · HN ↗
      Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there&#x27;s a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn&#x27;t make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.
      1. bigstrat2003 · · focus · HN ↗
        &gt; Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.

        If the models were actually intelligent, the way that the boosters claim, they wouldn&#x27;t need to be tuned to play chess in order to be good at it. That&#x27;s kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

        1. skydhash · · focus · HN ↗
          Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That&#x27;s the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use&#x2F;play something. Mastery can be gained overtime.
          1. willmarch · · focus · HN ↗
            If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn).

            Your assumptions&#x2F;intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).

            1. what · · focus · HN ↗
              &gt; considering LLMs currently play better than a brand new human player would

              They’ve ingested all the literature on playing chess, a brand new human player has not.

              1. willmarch · · focus · HN ↗
                Yes, but my point is that humans can’t even do the thing that the above comments are claiming humans can do (read a book or two and be decent at chess), and then they complain that LLMs can’t do the same thing (that humans can’t do either).

                We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.

                1. skydhash · · focus · HN ↗
                  &gt; The only way you get better at chess is by playing a lot of games and learning from mistakes

                  How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice.

                  This kickstarting then gradual refinement is how most people learn. And the foundational knowledge stays. Even a basic player knows to not do illegal moves.

                  1. willmarch · · focus · HN ↗
                    Reading can kickstart the process, but you can also make random moves guided by some sort of system (such as a computer GUI) or learn by watching other players play. The overall point is that you learn through observation and lots of trial and error (whether you are a human or a computer). And beginners in chess often make illegal moves even after learning the rules, it&#x27;s fairly common.

                    It feels like you&#x27;re trying to say that humans never make illegal moves while learning chess, which doesn&#x27;t match with my experience. I&#x27;m trying to understand your overall point.

                    1. skydhash · · focus · HN ↗
                      &gt; The overall point is that you learn through observation and lots of trial and error

                      That’s the most inefficient way and people usually avoid doing that. Instead they find someone that knows how to do the thing and ask him to be a teacher. Or use a proxy like a book or videos.

                      &gt; It feels like you&#x27;re trying to say that humans never make illegal moves while learning chess, which doesn&#x27;t match with my experience. I&#x27;m trying to understand your overall point

                      There’s learning the basic stuff (which is done after a few games) and there’s mastery. The thread started with the observation that even with all that knowledge (through content ingested in training), LLMs still makes illegal moves. Humans can be erratic, but they can constrain themselves to the rules for the task at hand after learning them.

                      1. willmarch · · focus · HN ↗
                        The only way to master anything is lots of trial and error. Coaches and teachers can help guide you towards more focused trial and error paths but the student still has to do the lessons and put in the work of learning, and learning only truly happens through doing.

                        Humans are not perfect and make mistakes in learning even when they have memorized the rules. A simple example is new players will often move a piece, exposing their king to check, and a more experienced player must point out to them that they have made an illegal move (because a new player often has not encoded that pattern for looking for exposed checks because they&#x27;re more focused on how the pieces move, not what that piece exposes.)

                        We&#x27;re just going to have to agree to disagree here.

            2. rsfern · · focus · HN ↗
              the discussion isn’t really about whether language models can become strong chess players though, the point is they seem to struggle to consistently make valid moves. Most humans don’t need to read two books to pick that up, just a couple lines of basic instructions
              1. willmarch · · focus · HN ↗
                That has not been my experience with new players, they regularly make invalid or incorrect moves even after detailed instructions especially in novel situations.
                1. rsfern · · focus · HN ↗
                  Maybe it depends on the person? My six year old isn’t great at strategy but they can pretty consistently make valid moves. Sometimes they ask for confirmation on a move which is also not a trait I see in language models (at least unprompted)
                  1. willmarch · · focus · HN ↗
                    Your child never messed up en passant (or had trouble understanding it in a real game), castled through or into a check, didn&#x27;t see a discovered check after moving a piece, never got confused by how stalemate works?
            3. orwin · · focus · HN ↗
              That&#x27;s quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).
              1. willmarch · · focus · HN ↗
                So your brother tried to make illegal moves while learning the game and it took your brother hundreds of games to get to be a decent player? I don&#x27;t see how this contradicts anything I said...
                1. orwin · · focus · HN ↗
                  The _only_ illegal move a human might make as a beginner is a failed en passant or a bad castle. And yes, a few hundred games is all it takes to be better than any publicly available LLM at the moment.
                  1. willmarch · · focus · HN ↗
                    ...and mistakes like not seeing discovered checks, castling into check, trying to castle out of check, castling through a check, misunderstanding en passant, missing checks when promoting a piece, misunderstanding how stalemate works, etc.

                    If you compared a human after hundreds of games to a SOTA LLM that was also trained on the output of hundreds of chess games that it played, I suspect you would notice similar improvements.

                    1. orwin · · focus · HN ↗
                      Honestly, if we&#x27;re only talking about SOTA llms, in sandbox mode without harness? No shot. Without harness LLMs have no memory of previous moves. I can&#x27;t make the last version of chatgpt remember more than 5 movements. I guarantee you, if you do not add flags in the harness with &#x27;left rook moved&#x27; or &#x27;right rook moved&#x27;, it will try to illegally castle 100% of the time it&#x27;s in the situation. Llm+harness, just make it call stockfish tbh.
          2. hi_im_greg_h · · focus · HN ↗
            1. The LLMs have surely ingested hundreds if not thousands of books on chess.

            2. The study (along with other posters here) show the models can’t even stick to following the rules of the game

        2. famouswaffles · · focus · HN ↗
          If humans were actually intelligent, they wouldn&#x27;t need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?
          1. diehunde · · focus · HN ↗
            Except all these LLMs were already trained with hundreds of chess book and game databases and they still suck
            1. famouswaffles · · focus · HN ↗
              If all you do is read chess books, you&#x27;ll be a shit player. Training and practice is what it takes to be great.
              1. sdf32dsf · · focus · HN ↗
                WTF even is this post?
              2. diehunde · · focus · HN ↗
                Oh right. But if all you do is reading programming books you are an amazing programmer? Where is all the training and practice LLMs did to become so good at coding?
                1. famouswaffles · · focus · HN ↗
                  LLMs (and Humans) don&#x27;t get good from programming books lol. The training and practice is the actual code they predict and learn from in the process of predicting.
                  1. diehunde · · focus · HN ↗
                    Oh I see. So if someone just reads books AND actual code then they can become experts, got it. And by the way LLMs are also trained with probably hundreds of thousands of actual games not just books
                2. hackinthebochs · · focus · HN ↗
                  &gt;Where is all the training and practice LLMs did to become so good at coding?

                  Coding is a matter of translating between the natural language description of a problem to the code specification while keeping the semantics fixed (and filling in the missing semantics reasonably well). It is not considerably more difficult than translating between two dissimilar natural languages. Chess isn&#x27;t a matter of language translation, but a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Chess takes directed practice and reinforcement whereas language translation does not.

                3. brindleth · · focus · HN ↗
                  It&#x27;s called post-training, typically through some form of reinforcement learning, and is a significant part of modern LLM development.

                  You have the first stage, pre-training, which is learning from next token prediction. That&#x27;s where the model memorises a lot of facts about things and generally gets good at forms of writing. It&#x27;s like reading a lot of books on programming and reading through a lot of source code. It&#x27;s learning how to autocomplete code, essentially. Doing that requires a developing a reasonable understanding of code, but it&#x27;s also learning how to autocomplete bad code as well as good, and won&#x27;t make it a &quot;good&quot; programmer.

                  Pre-training uses a method called Cross-Entropy Loss to update the weights of the network.

                  Then comes post-training. This is where the model is trained against huge sets of example problems, like fixing a bug, adding a new feature based on a spec, etc. They are set the task and try to complete it inside a training environment. Once they&#x27;re done, their complete solution is evaluated (either by humans, or by some separate evaluation model that was developed based on human feedback) and they are updated based on whether the solution was good or not.

                  Post-training uses a different method called Proximal policy optimization to update the weights of the network.

                  So these really are very different forms of learning, and mainstream LLMs are not post-trained to be good at chess. They could be. You could easily create a reinforcement learning environment that evaluated and improved their ability to play and win at chess. The result would be a very strong chess playing AI, something we know is possible because the strongest chess playing programs we have are neural network based, but it is not a priority for AI companies.

            2. WarmWash · · focus · HN ↗
              Contrary to popular belief, you need a lot of training on something for an LLM to be good and consistent with it.

              People think that if one mention exists in the training set, then the LLM is perfect at it.

              1. diehunde · · focus · HN ↗
                Not one mention. Hundreds of books, articles and databases of games.
          2. bigstrat2003 · · focus · HN ↗
            OpenAI making the next model good at chess is not analogous to a human training to get good at chess. It is analogous to God creating Human 2.0 which now has increased chess playing ability. If LLMs were intelligent the way humans are, then the models that exist right now would be able to spend time improving themselves at chess and become good at it. They can&#x27;t do this because they are not, in fact, intelligent.
            1. cindyllm · · focus · HN ↗

              [dead]

        3. Gregkion · · focus · HN ↗
          Thats just absolutly not true.

          A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

          And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

          1. foldr · · focus · HN ↗
            Humans don&#x27;t need a lot of training and finite tuning to make only legal moves.

            An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.

            An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).

            1. thom · · focus · HN ↗
              How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?
              1. foldr · · focus · HN ↗
                I don&#x27;t know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.
                1. thom · · focus · HN ↗
                  I&#x27;ve not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.
                  1. foldr · · focus · HN ↗
                    I haven&#x27;t tried it myself, but people seem to report that the illegal moves surface eventually. It just takes longer: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751

                    Nothing is forcing the LLM to play &#x27;blind&#x27;. If it&#x27;s smart, it should be able to create its own representation of the chess board and update it with every move, just like a human could.

                    1. thom · · focus · HN ↗
                      A human wouldn&#x27;t do that, they&#x27;d look at the board. I&#x27;m not disagreeing that to demonstrate clear superhuman ability the LLM should be able to do this, but it plays better than most humans blindfolded, and with fair prompts seems very good otherwise.
                      1. foldr · · focus · HN ↗
                        That&#x27;s what a human will do if they have a physical board to look at. But if someone, say, posed you a chess move exam question via FEN notation, you&#x27;d sketch a visual representation of the board off your own initiative to help you answer the question. There is nothing in principle to stop the LLM creating its own board representations in a format that enable it to easily keep track of legal and illegal moves. If it fails to do so, that&#x27;s a sign of its own limitations.
                        1. thom · · focus · HN ↗
                          I maintain that the amount of effort to teach a human to do this vastly outweighs the amount of effort to teach an LLM to do this unless you&#x27;re deliberately trying to make them fail. I honestly have no bigger point than that, I just think this isn&#x27;t a very good thing by which to evaluate LLM capabilities. If there&#x27;s no argument you&#x27;ll accept, I am happy to move on.
                          1. foldr · · focus · HN ↗
                            You don’t need to teach a human anything except the rules of chess and the details of a particular chess notation. No special skill or training is required to make a sketch of a chess board. Surely there is no chess player who, if confronted with a sequence of chess moves in algebraic notation, would not think to construct a representation of the chess board in order to understand what was going on.

                            &gt; I just think this isn&#x27;t a very good thing by which to evaluate LLM capabilities

                            I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)

                            &gt; If there&#x27;s no argument you&#x27;ll accept

                            It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your comments so far. I could equally well say the same thing to you!

                            1. thom · · focus · HN ↗
                              I&#x27;m just going to keep repeating: it is utterly trivial to get an LLM to play chess without making illegal moves. Easier than teaching a human. Sorry this doesn&#x27;t happen out of the box, but it shouldn&#x27;t budge your priors about LLM intelligence one bit.
                2. zahlman · · focus · HN ↗
                  More importantly, beginner human players don&#x27;t exhibit that tendency. The history of the position doesn&#x27;t bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.
                  1. thom · · focus · HN ↗
                    Humans do make these errors when playing blindfolded. If you even the playing field and give the LLM the position at each turn, it does not make mistakes.
                    1. zahlman · · focus · HN ↗
                      &gt; If you even the playing field and give the LLM the position at each turn, it does not make mistakes.

                      It absolutely still makes mistakes if you ask it to draw the board each turn, which should be equivalent to giving it the position because it only has to update one move at a time and then it has the position in the context window.

                      1. thom · · focus · HN ↗
                        Yes, we can come up with all sorts of weird situations where you can get it to be confused. But what I&#x27;m saying is it&#x27;s _trivial_ to give it a simple prompt that prevents it from ever making any errors, and so I don&#x27;t think it&#x27;s this big LLM gotcha (of which there are many!)
              2. Capricorn2481 · · focus · HN ↗
                The actual question is backwards: how do we keep the prompt and context small enough so the LLM doesn&#x27;t start hallucinating basic rules of chess.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.