‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. famouswaffles · · focus · HN ↗
      Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there&#x27;s a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn&#x27;t make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.
      1. bigstrat2003 · · focus · HN ↗
        &gt; Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.

        If the models were actually intelligent, the way that the boosters claim, they wouldn&#x27;t need to be tuned to play chess in order to be good at it. That&#x27;s kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

        1. Gregkion · · focus · HN ↗
          Thats just absolutly not true.

          A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

          And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

          1. foldr · · focus · HN ↗
            Humans don&#x27;t need a lot of training and finite tuning to make only legal moves.

            An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.

            An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).

            1. thom · · focus · HN ↗
              How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?
              1. foldr · · focus · HN ↗
                I don&#x27;t know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.
                1. thom · · focus · HN ↗
                  I&#x27;ve not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.
                  1. foldr · · focus · HN ↗
                    I haven&#x27;t tried it myself, but people seem to report that the illegal moves surface eventually. It just takes longer: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751

                    Nothing is forcing the LLM to play &#x27;blind&#x27;. If it&#x27;s smart, it should be able to create its own representation of the chess board and update it with every move, just like a human could.

                    1. thom · · focus · HN ↗
                      A human wouldn&#x27;t do that, they&#x27;d look at the board. I&#x27;m not disagreeing that to demonstrate clear superhuman ability the LLM should be able to do this, but it plays better than most humans blindfolded, and with fair prompts seems very good otherwise.
                      1. foldr · · focus · HN ↗
                        That&#x27;s what a human will do if they have a physical board to look at. But if someone, say, posed you a chess move exam question via FEN notation, you&#x27;d sketch a visual representation of the board off your own initiative to help you answer the question. There is nothing in principle to stop the LLM creating its own board representations in a format that enable it to easily keep track of legal and illegal moves. If it fails to do so, that&#x27;s a sign of its own limitations.
                        1. thom · · focus · HN ↗
                          I maintain that the amount of effort to teach a human to do this vastly outweighs the amount of effort to teach an LLM to do this unless you&#x27;re deliberately trying to make them fail. I honestly have no bigger point than that, I just think this isn&#x27;t a very good thing by which to evaluate LLM capabilities. If there&#x27;s no argument you&#x27;ll accept, I am happy to move on.
                          1. foldr · · focus · HN ↗
                            You don’t need to teach a human anything except the rules of chess and the details of a particular chess notation. No special skill or training is required to make a sketch of a chess board. Surely there is no chess player who, if confronted with a sequence of chess moves in algebraic notation, would not think to construct a representation of the chess board in order to understand what was going on.

                            &gt; I just think this isn&#x27;t a very good thing by which to evaluate LLM capabilities

                            I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)

                            &gt; If there&#x27;s no argument you&#x27;ll accept

                            It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your comments so far. I could equally well say the same thing to you!

                            1. thom · · focus · HN ↗
                              I&#x27;m just going to keep repeating: it is utterly trivial to get an LLM to play chess without making illegal moves. Easier than teaching a human. Sorry this doesn&#x27;t happen out of the box, but it shouldn&#x27;t budge your priors about LLM intelligence one bit.
                2. zahlman · · focus · HN ↗
                  More importantly, beginner human players don&#x27;t exhibit that tendency. The history of the position doesn&#x27;t bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.
                  1. thom · · focus · HN ↗
                    Humans do make these errors when playing blindfolded. If you even the playing field and give the LLM the position at each turn, it does not make mistakes.
                    1. zahlman · · focus · HN ↗
                      &gt; If you even the playing field and give the LLM the position at each turn, it does not make mistakes.

                      It absolutely still makes mistakes if you ask it to draw the board each turn, which should be equivalent to giving it the position because it only has to update one move at a time and then it has the position in the context window.

                      1. thom · · focus · HN ↗
                        Yes, we can come up with all sorts of weird situations where you can get it to be confused. But what I&#x27;m saying is it&#x27;s _trivial_ to give it a simple prompt that prevents it from ever making any errors, and so I don&#x27;t think it&#x27;s this big LLM gotcha (of which there are many!)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.