‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. sigmoid10 · · focus · HN ↗
        The actual current frontier plays somewhere around GM level.

        <a href="https:&#x2F;&#x2F;chessbench-ai.github.io&#x2F;#leaderboard" rel="nofollow">https:&#x2F;&#x2F;chessbench-ai.github.io&#x2F;#leaderboard

        It&#x27;s also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I&#x27;m sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

        1. minraws · · focus · HN ↗
          I know HN readers and posters just read numbers and can&#x27;t be bothered to read, but please read the methodology before making any claims.

          &gt; About their ELO ratings from their own website:

          &gt; A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

          I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

          Please folks at least use your AIs to read stuff before making claims.

          AI is not GM level, it&#x27;s not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

          A GM is 2600 they can beat me in under 20 moves...

          Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

          Maybe I should stop doing that will be a happier life, don&#x27;t think just believe in the AGI.

          1. YeGoblynQueenne · · focus · HN ↗
            &gt;&gt; I know HN readers and posters just read numbers and can&#x27;t be bothered to read, but please read the methodology before making any claims.

            This is unfair to HN readers all of whom but one did not post the comment you replied to. You can&#x27;t just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.

            1. minraws · · focus · HN ↗
              How many posts if I link that do the same thing will you agree this is the norm here.

              Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine.

              I don&#x27;t even know if there is critical thought or we believe what we read&#x2F;shared&#x2F;etc

              1. dezsiszabi · · focus · HN ↗
                50% + 1 of all comments
              2. YeGoblynQueenne · · focus · HN ↗
                No, I don&#x27;t agree it&#x27;s the norm. There is though a general tendency to opine with strong views on subjects posters have no expertise on. I think that&#x27;s because many are software engineers (or equivalent) and they are used to being expected to &quot;wing it&quot; on whatever technical subject comes up. On the other hand you can always find informed comments by users who have specialist knowledge.

                And there&#x27;s plenty of pushback on here about the chess thing besides your very valid points.

                EDIT: anyway if I can offer a bit of unsolicited advice, it won&#x27;t do you or anyone any good to accuse everyone who doesn&#x27;t agree with you of laziness, even if you can see e.g. they haven&#x27;t really read an article. Just say the thing you wan to say and let them figure it out. Most people will appreciate that much better and you will feel better about yourself for acting like a mature adult.

                It&#x27;s even in the site guidelines:

                Please don&#x27;t comment on whether someone read an article. &quot;Did you even read the article? It mentions that&quot; can be shortened to &quot;The article mentions that&quot;.

                1. minraws · · focus · HN ↗
                  It&#x27;s not been my personal experience on this website in the last 2-3 years atleast, pre-covid perhaps.

                  But despite that you aren&#x27;t wrong and the only reason I even visit this website is because people sometimes did&#x2F;do take time to reflect on things based on their experience and knowledge.

                  And in hindsight pointing out that hn has issues wasn&#x27;t even the point but I feel frustrated when everyone is readily agreeing to things on here without reading. When that in this moment feels like the one thing that separates humans from machines that we get to think and learn.

                  I possibly should just drop reading this place until we have most noisy people go away. I have for one tried to always only comment on things where I could be a value add, this one does feel like I could I have done better.

                  In the moment I probably thought if they are GM level and I can beat them, is this some interesting find, my disappointment honestly led me to making a rather incorrect call on this one.

                  Either way I still do think HN as a whole has devolved into mindless herd follower mindset, I can point to more than a few posts that just say adopt the hacker mindset aka move fast don&#x27;t care about the consequences.

                  And I for one find this laughable even though that&#x27;s the reality of my job&#x2F;work as well.

                  1. YeGoblynQueenne · · focus · HN ↗
                    [delayed]
                    1. minraws · · focus · HN ↗
                      &gt;&gt; Sorry, I didn&#x27;t get this? What was the incorrect call you made?

                      Talking about people&#x27;s inability to read rather than just pointing out that the article pointed at something else.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.