‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. 21asdffdsa12 · · focus · HN ↗
        So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.
        1. user43928 · · focus · HN ↗
          There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1.

          The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier.

          That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.

          1. Topfi · · focus · HN ↗
            Fortunately, a fellow commenter was so kind and did it with Astra. Didn&#x27;t do that well either [0]. I&#x27;m sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...

            I&#x27;ll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. &quot;Just&quot; having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn&#x27;t even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.

            [0] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751

            1. user43928 · · focus · HN ↗
              Doesn&#x27;t look impressive, although I&#x27;m hearing a marked improvement in choosing legal moves, compared to early 2025.

              Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better?

              I would not be surprised if OpenAI released a model that beats humans at chess this year.

              1. Topfi · · focus · HN ↗
                I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish.

                Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models have gained in actually capability that is in the training data, but not RLHFd to hell, so to speak. Tracking the state of pieces, I suspect given similar in Sudoku [0], is what these models struggle with in game settings, whilst tracking the state of code changes can be reliable over 250k tokens. Essentially, for the latter they were trained in the specific manner that lead them to abstract the capability, but that doesn&#x27;t track to the former, which is a massive difference between LLMs data focused training and human learning.

                So yeah, GPT-7 or any upcoming&#x2F;present LLM could do massively better in Chess than GPT-6 Astra, but not because the approach was emergent out of pure data. Rather, it requires a very specific training data type and stack for a model to gain capabilities that track a specific task long enough to adhere to the rules of a game such as chess.

                [0] <a href="https:&#x2F;&#x2F;logicalintelligence.com&#x2F;blog&#x2F;energy-based-model-sudoku-demo" rel="nofollow">https:&#x2F;&#x2F;logicalintelligence.com&#x2F;blog&#x2F;energy-based-model-sudo...

                1. user43928 · · focus · HN ↗
                  I&#x27;m wondering if instructing it to track the board state in a file would make a significant difference then.

                  It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new context compaction method increased the performance dramatically. However, I think that is not applicable here.

                2. 21asdffdsa12 · · focus · HN ↗
                  So what is the supposed leap? One agent per option to change, evaluating the board state that there move would create, by having a army evaluate the remaining piece options and average over that? Wee-Free-Man as a hierarchical army ? Pet-LLMs trained on one thing?
                  1. Topfi · · focus · HN ↗
                    Honestly, for intelligence I don&#x27;t know and I doubt anyone can claim to know. Maybe JEPA, there is potential concerning some shortcomings inherent to LLMs but it has its own, maybe scaling up the electron microscope stuff Google just did (though the connections are inferred), maybe future implementations of autoregressive and diffusion LLMs can at some point address its issues after all, maybe something else entirely.

                    All I know is, AGI, as in actual intelligence, is quite a massive accomplishment to claim and we shouldn&#x27;t loose sight of that fact, especially as &quot;not being intelligent&quot; does not make these models any less impressive, fascinating to work on or useful in many tasks. Personally, the only thing I am fairly convinced on is that if we were to find a way to create actual intelligence, it likely wouldn&#x27;t start out as useful as todays LLMs are and may thus be dismissed early. But again, pure speculation on that front.

                    If for leap you just mean more utility from LLMs as they are, then I&#x27;ll pretty confidently put my money on higher quality, not more, training data for a wide range of verifiable tasks. What makes maths, coding, etc. comparatively easy to make gains in (though less verifiable tasks can also make similar as seen with the writing in Kimi K2).

              2. datsci_est_2015 · · focus · HN ↗
                Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the concentration of wealth even further and fully realize the American dream of eliminating the middle class.
                1. user43928 · · focus · HN ↗
                  I&#x27;ve seen some of his videos, and got the impression he didn&#x27;t understand how GPT-Live delegates to the more powerful regular model with reasoning.

                  The regular model generally does not suffer the same issues he is demonstrating with the real time audio version.

                  In my view the investment into datacenters is well justified by the current demand, and progress has been very impressive.

                  1. freejazz · · focus · HN ↗
                    Really? It was being sold as a total replacement for jobs like software engineering and being an attorney, but its looking a lot more that its just going to be a tool those professions use and doesn&#x27;t actually seem to be taking jobs away.
                2. Quinner · · focus · HN ↗
                  I find it amusing that you&#x27;re describing a huge misallocation of capital and a society enabling such, and that is the optimisitic scenario (in my mind anyway).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.