‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. sigmoid10 · · focus · HN ↗
        The actual current frontier plays somewhere around GM level.

        <a href="https:&#x2F;&#x2F;chessbench-ai.github.io&#x2F;#leaderboard" rel="nofollow">https:&#x2F;&#x2F;chessbench-ai.github.io&#x2F;#leaderboard

        It&#x27;s also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I&#x27;m sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

        1. minraws · · focus · HN ↗
          I know HN readers and posters just read numbers and can&#x27;t be bothered to read, but please read the methodology before making any claims.

          &gt; About their ELO ratings from their own website:

          &gt; A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

          I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

          Please folks at least use your AIs to read stuff before making claims.

          AI is not GM level, it&#x27;s not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

          A GM is 2600 they can beat me in under 20 moves...

          Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

          Maybe I should stop doing that will be a happier life, don&#x27;t think just believe in the AGI.

          1. echelon · · focus · HN ↗
            The AI can write a chess bot program that will beat you.

            You&#x27;re thinking about this the wrong way. The system is built and delivered as it is because that&#x27;s how the providers make the most money. If they cared to have it perform well in chess games, you&#x27;d see a different shape and behavior.

            We shouldn&#x27;t ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing&#x27;s flight guidance system to do so.

            1. minraws · · focus · HN ↗
              So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here.

              Delusion runs deep in HN circles.

              I say that as someone heavily invested in AI startups and projects and as someone working in the field.

              I think most people on HN should touch grass and find real human contact. Lmao

              Incredible reasoning all around here.

              1. diehunde · · focus · HN ↗
                AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI

                Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

                1. hackinthebochs · · focus · HN ↗
                  &gt;LLM can’t beat an avg chess player.

                  Why should that matter?

                  1. janalsncm · · focus · HN ↗
                    If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this.

                    So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

                    1. Gregkion · · focus · HN ↗
                      So we humans are not a general intelligence then?

                      And the stuff i&#x27;m using LLMs daily is just fake?

                      I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

                      1. dosisking · · focus · HN ↗
                        &gt; And the stuff i&#x27;m using LLMs daily is just fake?

                        It simply means that LLMs are smarter than you, but not smarter than the average person

                      2. zahlman · · focus · HN ↗
                        &gt; So we humans are not a general intelligence then?

                        No, because we can, in fact, generally read the rules of a game and then follow them. It&#x27;s actually a hobby for many of us.

                        &gt; And the stuff i&#x27;m using LLMs daily is just fake?

                        This misses the point completely.

                        1. hackinthebochs · · focus · HN ↗
                          &gt; generally read the rules of a game and then follow them

                          How many times do you think chess.com prevents illegal moves from being executed? Even Super GM&#x27;s fall for mate-in-1&#x27;s occasionally, which is functionally equivalent to missing a fork or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn&#x27;t pass the smell test.

                          1. diehunde · · focus · HN ↗
                            Do you play chess ? Do you even know what an illegal move is ?
                            1. hackinthebochs · · focus · HN ↗
                              If you have something to contribute to the discussion, just say it
                          2. zahlman · · focus · HN ↗
                            Chess.com has to accommodate people who haven&#x27;t learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they&#x27;re often outlining sequences of multiple &quot;pre-moves&quot; during the opponent&#x27;s turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.