‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. MattCruikshank · · focus · HN ↗
      What happens when you ask those same frontier models to write a chess-playing program?

      I feel like, this is a huge stumbling block that many people have. They&#x27;ll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn&#x27;t fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.

      It&#x27;s really neat to see what a frontier model can do itself. No doubt.

      But &quot;play chess by hand&quot; is a frankly awful metric. It&#x27;s kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.

      1. cbolton · · focus · HN ↗
        It&#x27;s a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it&#x27;s noteworthy that they underperform on that test.

        Letting the model execute a chess program (that it wrote) would make sense if you&#x27;re measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.

        1. MattCruikshank · · focus · HN ↗
          &gt; The fact that the human would have a much harder time writing a useful program is irrelevant.

          Why?

          There&#x27;s a box.

          You give it a problem, and it comes up with a solution.

          Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes?

          Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess &quot;in its head&quot;, or if it has to use scratch paper?

          1. lionkor · · focus · HN ↗
            The question is to what end? This is a benchmark task, because playing chess, or solving other well-understood problems is more of a party trick than it is useful.

            If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.

            1. MattCruikshank · · focus · HN ↗
              Do you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back?

              More to my point, I think it&#x27;s stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs.

              I&#x27;m advising people that they should think about this distinction, themselves, when they have data and want answers.

              1. cbolton · · focus · HN ↗
                Neither. As I said I want to measure cognitive abilities.

                Your &quot;ability of the box&quot; is like &quot;economic potential&quot; in my previous comment. If that&#x27;s what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.

                1. MattCruikshank · · focus · HN ↗
                  I agree that it&#x27;s a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had.

                  That said, it&#x27;s really weird to me when people use (and judge) LLMs one way... and won&#x27;t try using them another way.

                  Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.

                  Otherwise, it&#x27;s like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I&#x27;ve ever seen: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=gDy9AUQJ3Fg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=gDy9AUQJ3Fg

                  This lamp, without a working power outlet? It really doesn&#x27;t do anything...

                  1. cbolton · · focus · HN ↗
                    I completely agree.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.