‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. WhitneyLand · · focus · HN ↗
      1. It’s hard to trust a 2026 paper that’s showing results for such old models.

      2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.

      3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

      1. paimapi · · focus · HN ↗
        so prove it! get a public repo out there, have it play against some open source engines

        also I think the operative letter in AGI is the G - and if the G is short for &#x27;variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill&#x27; then its not really G at all, is it?

        1. BobbyJo · · focus · HN ↗
          I suck at chess. Are you saying I can&#x27;t be intelligent?
          1. paimapi · · focus · HN ↗
            is that what I&#x27;m saying? or am I talking about AGI? perhaps there&#x27;s some irony here to be explored when it comes to basic reading comprehension gaps
            1. jibal · · focus · HN ↗
              That&#x27;s a polite way to put it. :-)
            2. BobbyJo · · focus · HN ↗
              My point was you are misunderstanding G, or at least applying it erroneously here. Being good at chess is not a generalization of any other body of knowledge, it is a rigorous set of rules. The only way to be good at chess is to practice chess, or to apply deep calculations. The latter is the model writing code.

              The illegal move aspect has more to do with a failure of online&#x2F;in-context learning, which would support your point. I tend to think it is a byproduct of reasoning in language, which newer architectures would fix, but we shall see.

              1. paimapi · · focus · HN ↗
                chess is not a rigorous set of rules, it is rules as foundation. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun. knowledge for chess, specifically, is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. it is not at all different from any other body of knowledge - it only &#x27;feels&#x27; different to us humans because it is so logic-based and it takes a long time before our inferential, pattern-recognition kicks in and starts seeing the board in a naturally, systemic way. for an AGI, that should be a cakewalk, trained as it were to surpass human capability in any and every domain (thus the G for &#x27;general&#x27; and not &#x27;H&#x27; for &#x27;hyperspecific&#x27;)
                1. BobbyJo · · focus · HN ↗
                  &gt; an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain.

                  AGI != ASI. You are confusing the two.

                  1. paimapi · · focus · HN ↗
                    I&#x27;m not. AGI is almost necessarily closer to ASI than it is to human intelligence. you take the concept of domain knowledge transferring to other areas. presumably, an &#x27;AGI&#x27; that is generally as good as a really good human at every task under-the-sun will already be much better than most humans because it can incorporate cross-domain knowledge and apply it in a reasonable fashion. it&#x27;s like the parable of Newton and the apple - the domain knowledge that an apple falls according to certain rules observed before igniting the creative spark that led to universal gravitation
                    1. BobbyJo · · focus · HN ↗
                      &gt; presumably, an &#x27;AGI&#x27; that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion.

                      I disagree with this definition of AGI, and I disagree that chess skills significantly benefit from generalizing non-chess knowledge, outside of computing moves probabilistically.

                      AGI has historically been defined as human level or better, with generality to new domains. I think blurring it with ASI makes the terminology confusing to use.

                      Chess is learned rules and the ability to apply those rules. Strategy as a whole is applying a set of rules to circumstances, that&#x27;s how it is taught: &quot;here are examples of circumstances and actions, try to pattern match to future circumstance and apply commensurate action.&quot;

                      If you make the point that chess is a large part of the training data, or that LLMs are unable to learn chess well, I&#x27;ll accept that as refuting that LLMs are AGI, but these other points I disagree with.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.