‹ BackHN Continuity

Thread

Why I'm still bearish on LLMs after Navier-Stokes

496 points · 653 comments · jaykru

  1. carodgers · · focus · HN ↗
    This April 2026 paper is a fun and related read.

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2509.24239v4

    Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

    The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

    1. threethirtytwo · · focus · HN ↗
      The story isn&#x27;t so clear cut.

      The caveat is: It depends on the task.

      Are there reams of chess moves that the model can train off of? No.

      Are there reams of math papers the model can train off of? Yes.

      1. keephnacct · · focus · HN ↗

        [dead]

      2. [deleted] · · focus · HN ↗

        [deleted]

      3. iwontberude · · focus · HN ↗

        [dead]

      4. tjwebbnorfolk · · focus · HN ↗
        &gt; Are there reams of chess moves that the model can train off of? No.

        This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.

        1. XenophileJKO · · focus · HN ↗
          It is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don&#x27;t even need any data to start with (but would help).
          1. manquer · · focus · HN ↗
            There are more possible game combinations than atoms in the universe, even those generation of valid game states are as you say pre-defined. that is why models cannot go this route and therefore are poor at chess
            1. wat10000 · · focus · HN ↗
              Isn’t this exactly how AlphaZero was trained? The rules are known and well defined so the training process can generate games without any outside data.

              The only reason LLMs are this bad at chess is because the labs don’t care about chess performance so they’re not going out of their way to train the models for it. The ability they do have is from what chess information happens to be in the training data, plus whatever general reasoning abilities they may be able to apply.

        2. threethirtytwo · · focus · HN ↗
          Let me make my statement more clear with a correction:

          Was there reams of chess moves that the model trained off of? No.

      5. vmg12 · · focus · HN ↗
        &gt; The caveat is: It depends on the task.

        I think the line of criticism around LLMs sucking at chess makes more sense when you understand what the AI companies are saying about the future trajectory of these models.

        The entire recursive self improvement story falls apart once you point out that there is not much &quot;cross domain transfer learning&quot;. Meaning that training an LLM to become good at coding, math, etc, will eventually transfer into them being good at other skills that were not explicitly trained for.

        Using games like chess which have little economic value is actually a good test for this. What&#x27;s even more surprising about them sucking at chess is how much information about chess strategy exists in the training data.

      6. freejazz · · focus · HN ↗
        &gt;Are there reams of chess moves that the model can train off of? No.

        For real??

        1. threethirtytwo · · focus · HN ↗
          I meant if there are reams of chess moves the model was trained off of.
          1. freejazz · · focus · HN ↗
            Yeah, it&#x27;s not like there&#x27;s any literature about Chess in the corpus of these models!
            1. threethirtytwo · · focus · HN ↗
              There’s literature. But I don’t think there’s reams of chess games. With literature the LLM can understand strategy but chess needs intuition and you need tokenized games for that. LLMs have less of that.
      7. FuckButtons · · focus · HN ↗
        There’s multiple databases of games in algebraic notation. You can also, very easily rl train on pitting models against one another, even without mcts.
      8. thelaxiankey · · focus · HN ↗
        there are far more reams of chess moves than there are math papers. Lichess is pretty open...

        But hey, they&#x27;re actually good at chess if you prompt correctly so.... <a href="https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;" rel="nofollow">https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;

    2. wat10000 · · focus · HN ↗
      I wonder how current models would fare. The ones they tested are fairly old now.
    3. joefourier · · focus · HN ↗
      &gt; current frontier models

      &gt; Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1

      The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. sigmoid10 · · focus · HN ↗
        The actual current frontier plays somewhere around GM level.

        <a href="https:&#x2F;&#x2F;chessbench-ai.github.io&#x2F;#leaderboard" rel="nofollow">https:&#x2F;&#x2F;chessbench-ai.github.io&#x2F;#leaderboard

        It&#x27;s also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I&#x27;m sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.

        1. htrp · · focus · HN ↗
          more like you lose intelligence in chess by maxing for coding... hence knocking back the claims of emergent intelligence
        2. csande17 · · focus · HN ↗
          Even if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.
          1. MichaelNolan · · focus · HN ↗
            I wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.
            1. shric · · focus · HN ↗
              &gt; so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other

              As a 1500 elo human I can tell you that a 1500 elo chess engine doesn&#x27;t play like anything like a 1500 elo human.

              1. traes · · focus · HN ↗
                This is true, but I&#x27;m not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.
                1. fahrvrgnugen · · focus · HN ↗
                  I feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that&#x27;s the bot.
                  1. shric · · focus · HN ↗
                    That would only work for the first few (from around 10 to 20 typically depending on how close people stick to opening book) moves.

                    Conservatively there are well over 10 to the 30 positions likely to show up in realistic games.

                    There are of the order of 10 to the 10 or so games recorded.

                    Thus well under one in a trillion positions are &quot;known&quot;.

                    1. fahrvrgnugen · · focus · HN ↗
                      It&#x27;s much smaller than that. You would be unlikely to find yourself in a novel position after 40 moves even if you were trying.
                      1. traes · · focus · HN ↗
                        This is simply blatant misinformation. If you play a game online on lichess and go to the analysis board you can find when your game becomes novel. It will be within 20 turns unless you are intentionally following a known opening. In fact it will likely become unique within 10-15 turns.
                        1. fahrvrgnugen · · focus · HN ↗
                          It&#x27;s not my experience at all. If you find yourself in a novel position within 10-15 moves it&#x27;s likely a resignable one.
                          1. shric · · focus · HN ↗
                            You got me curious...

                            I am around 1500 (actually 1649 on lichess blitz, but close enough).

                            I explored the last 5 games I played on lichess. Here are the number of moves before lichess had never seen that position before for each of the 5 games: 6, 12, 16, 15, 11.

                            &gt; If you find yourself in a novel position within 10-15 moves it&#x27;s likely a resignable one

                            The opponent is also going to be in a novel position. Should both players resign?

                            Just in case you think this is limited to low rated players like 1500s, look at MagnusCarlsen&#x27;s most used account on Blitz: <a href="https:&#x2F;&#x2F;lichess.org&#x2F;@&#x2F;DrNykterstein&#x2F;search?perf=2" rel="nofollow">https:&#x2F;&#x2F;lichess.org&#x2F;@&#x2F;DrNykterstein&#x2F;search?perf=2

                            You will see that most games become unique to the whole of lichess within 15-20 moves and a good chunk between 10-15.

        3. einszwei · · focus · HN ↗
          Probably tells us that without labs explicitly training&#x2F;tuning the models or designing the harness (with fast oracle) the LLMs aren&#x27;t going to get good at those areas.
        4. minraws · · focus · HN ↗
          I know HN readers and posters just read numbers and can&#x27;t be bothered to read, but please read the methodology before making any claims.

          &gt; About their ELO ratings from their own website:

          &gt; A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.

          I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..

          Please folks at least use your AIs to read stuff before making claims.

          AI is not GM level, it&#x27;s not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.

          A GM is 2600 they can beat me in under 20 moves...

          Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.

          Maybe I should stop doing that will be a happier life, don&#x27;t think just believe in the AGI.

          1. peab · · focus · HN ↗
            What levels are they actually at in your experience?
            1. minraws · · focus · HN ↗
              Sub 1300 that&#x27;s my rating in the singular official tournament I participated at.

              But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).

              I would rate them around 500-800 big range but at that level it&#x27;s all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.

              I can play good&#x2F;best moves till 14-15 moves if I remember the lines and find someone who falls for it.

              If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.

              700 is around the rating for a human who doesn&#x27;t know the tricks but can do bare minimum calculations and understands the rules thoroughly.

              1. Forgeties79 · · focus · HN ↗
                As someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself.

                Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.

                Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous

              2. little_endorian · · focus · HN ↗
                You can take LLMs out of opening knowledge by playing chess960, and their performance degrades significantly. I just tried playing Claude Sonnet 5 (high), and it made its first illegal move on move 5.

                They played 4...c6, followed by 5...Nc6, somehow forgetting about the pawn the just put on c6. (My move in between was 5. Nc3, and apparently they were trying to mirror me.)

            2. zug_zug · · focus · HN ↗
              So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.

              In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.

              1. firmretention · · focus · HN ↗
                I&#x27;ve always liked the analogy that talking to an LLM is like talking to a really, really smart person with a head injury.
          2. echelon · · focus · HN ↗
            The AI can write a chess bot program that will beat you.

            You&#x27;re thinking about this the wrong way. The system is built and delivered as it is because that&#x27;s how the providers make the most money. If they cared to have it perform well in chess games, you&#x27;d see a different shape and behavior.

            We shouldn&#x27;t ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing&#x27;s flight guidance system to do so.

            1. minraws · · focus · HN ↗
              So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here.

              Delusion runs deep in HN circles.

              I say that as someone heavily invested in AI startups and projects and as someone working in the field.

              I think most people on HN should touch grass and find real human contact. Lmao

              Incredible reasoning all around here.

              1. echelon · · focus · HN ↗
                I&#x27;m stating that certain folks are trying to use the software-generating product as an AGI&#x2F;ASI and then complaining when it doesn&#x27;t play chess very well.

                People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

                1. minraws · · focus · HN ↗
                  Then why respond at all for the sake of responding?

                  We all know AI can code, but the question it all stemmed from what if it&#x27;s AGI or GM level in chess on it&#x27;s own.

                  You can&#x27;t just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.

                  I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?

                  1. sdf32dsf · · focus · HN ↗
                    He keeps posting with a particular type of tone.

                    He definitely needs to touch grass.

                    1. echelon · · focus · HN ↗
                      Try to embrace hacker ethos and stop hating.

                      Y&#x27;all seem to miss the point of this forum. Building and hacking and science and engineering.

                      I swear there&#x27;s a whole lot of you who just like to look down instead of up. There&#x27;s a whole universe up there.

                  2. bigstrat2003 · · focus · HN ↗
                    &gt; We all know AI can code...

                    We know no such thing. LLMs are quite bad at generating code, worse than any capable human.

                2. modulus1 · · focus · HN ↗
                  I agree w&#x2F; this perspective. An agent with a harness that can run programs can solve a lot more than one without the harness. The AI system includes the harness, and it&#x27;s not clear to me that AGI requires more than LLMs + code generation &amp; execution are capable of.
                  1. minraws · · focus · HN ↗
                    So AI is AGI in fields where code can&#x27;t solve anything?

                    Is code omnipotent, I have been in software all my life and I would hard agree here.

                    Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it&#x27;s the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.

                    For instance Maths is just code with different symbols and slightly less universally legible concepts.

                    AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.

                    But that&#x27;s it, I am certain a bunch of companies will make a lot of money despite no AGI.

                    I think people either don&#x27;t understand AGI or don&#x27;t understand how real world works.

                    Until an LLM can bow it&#x27;s head take responsibility for mistakes made and ensure they aren&#x27;t repeated again with 100% confidence to the leadership it&#x27;s inarguably a tool a rather questionable one at that.

                    1. simianwords · · focus · HN ↗
                      &gt; AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.

                      So.. like chess?

                      Anyway, do you have any prediction on what LLM&#x27;s can or can&#x27;t do in a few years?

                3. Yizahi · · focus · HN ↗
                  It&#x27;s not even a &quot;software-generating product&quot;. It&#x27;s only half of it. Most of the heavy lifting is done by absolutely not-AI compilers, analyzers and the like. If not for these programs, written well before AI boom, them LLMs would be no better at programming than they are are at pure LLM based calculations or writing.
              2. diehunde · · focus · HN ↗
                AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI

                Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

                1. hackinthebochs · · focus · HN ↗
                  &gt;LLM can’t beat an avg chess player.

                  Why should that matter?

                  1. janalsncm · · focus · HN ↗
                    If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this.

                    So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

                    1. hackinthebochs · · focus · HN ↗
                      I would bet a lot of money that Astra can follow the rules of chess (perhaps if repeated within the context window). Also, this is a different argument than what I responded to.
                      1. minraws · · focus · HN ↗
                        I can write you a benchmark to prove it even with a heavy handed system prompt Astra will make an illegal move during the course of the games first few moves are generally ok since it&#x27;s just throwing out learned moves.
                        1. hackinthebochs · · focus · HN ↗
                          I&#x27;d genuinely like to see the results of that.
                      2. janalsncm · · focus · HN ↗
                        I would definitely take you up on that.
                        1. simianwords · · focus · HN ↗
                          <a href="https:&#x2F;&#x2F;www.chessbench.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.chessbench.org&#x2F;

                          &gt;GPT-6 Astra xHigh: 0.06% rejected moves

                          1. frde_me · · focus · HN ↗
                            I wonder if I would do better as a human, maybe? Or would I happen to have one move in 1500+ that&#x27;s not valid?

                            I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece.

                    2. Gregkion · · focus · HN ↗
                      So we humans are not a general intelligence then?

                      And the stuff i&#x27;m using LLMs daily is just fake?

                      I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

                      1. dosisking · · focus · HN ↗
                        &gt; And the stuff i&#x27;m using LLMs daily is just fake?

                        It simply means that LLMs are smarter than you, but not smarter than the average person

                      2. zahlman · · focus · HN ↗
                        &gt; So we humans are not a general intelligence then?

                        No, because we can, in fact, generally read the rules of a game and then follow them. It&#x27;s actually a hobby for many of us.

                        &gt; And the stuff i&#x27;m using LLMs daily is just fake?

                        This misses the point completely.

                        1. hackinthebochs · · focus · HN ↗
                          &gt; generally read the rules of a game and then follow them

                          How many times do you think chess.com prevents illegal moves from being executed? Even Super GM&#x27;s fall for mate-in-1&#x27;s occasionally, which is functionally equivalent to missing a fork or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn&#x27;t pass the smell test.

                          1. diehunde · · focus · HN ↗
                            Do you play chess ? Do you even know what an illegal move is ?
                            1. hackinthebochs · · focus · HN ↗
                              If you have something to contribute to the discussion, just say it
                          2. zahlman · · focus · HN ↗
                            Chess.com has to accommodate people who haven&#x27;t learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they&#x27;re often outlining sequences of multiple &quot;pre-moves&quot; during the opponent&#x27;s turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.
                  2. lelanthran · · focus · HN ↗
                    &gt; Why should that matter?

                    Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.

                    So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.

                    1. hackinthebochs · · focus · HN ↗
                      We&#x27;re not talking about learning the rules of chess here, but playing a competent game from just the rules. Why is it so hard for people to keep track of the thread of discussion?
                      1. ncruces · · focus · HN ↗
                        But we are. The models can&#x27;t even follow the rules: they try illegal moves all the time.
                      2. lelanthran · · focus · HN ↗
                        &gt; We&#x27;re not talking about learning the rules of chess here, but playing a competent game from just being shown the rules.

                        Okay, lets go with that: it&#x27;s the &quot;shown the rules&quot; bit that we are arguing about.

                        The argument is that a human may play maybe a dozen games after learning the rules, after which they won&#x27;t be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.

                        This does not point to generalisable and adaptable intelligence, such as we see in the average human.

                        1. hackinthebochs · · focus · HN ↗
                          This is not good reasoning. Humans need at least dozens if not hundreds of reinforcement sessions to only make legal moves, and still occasionally fail (consider pins, walking into check, failing to respond to check). LLMs must one-shot a competent game after imbibing a mass of disconnected units of information about chess. Nothing about the two are similar.

                          See my comment here for more: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49725306">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49725306

                  3. HarHarVeryFunny · · focus · HN ↗
                    It depends on what you are selling it as.

                    It only matters if you are claiming it to be general purpose.

                    If you admit that it&#x27;s just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose.

                    The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can&#x27;t do is very relevant.

                2. lostmsu · · focus · HN ↗
                  The fact that LLMs can play chess at any level is a strong indication we are in AGI.
                  1. recursive · · focus · HN ↗
                    Can they if they frequently make illegal moves?
                    1. lostmsu · · focus · HN ↗
                      [delayed]
                  2. bigstrat2003 · · focus · HN ↗
                    No it isn&#x27;t. Computers could play chess long before LLMs, better than LLMs can in fact. That didn&#x27;t make them AGI.
                    1. lostmsu · · focus · HN ↗
                      [delayed]
                  3. zahlman · · focus · HN ↗
                    This is roughly comparable to observing a cat batting a ball away with its paw and taking this as a &quot;strong indication&quot; that cats can play any sport.
                    1. lostmsu · · focus · HN ↗
                      [delayed]
                  4. HarHarVeryFunny · · focus · HN ↗
                    It would be more impressive if they could play chess (or do anything they haven&#x27;t been custom RLVR trained for) by reasoning, rather than just &quot;have a go at it&quot; prediction which is closer to memorization.

                    HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?

                    1. lostmsu · · focus · HN ↗
                      [delayed]
                      1. HarHarVeryFunny · · focus · HN ↗
                        &gt; They can&#x27;t possibly remember even a few positions.

                        Sure they could, but that&#x27;s irrelevant.

                        A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.

                        However, that&#x27;s not how LLMs work. They don&#x27;t memorize inputs - they predict them, based on disovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positons). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.

                        &gt; Don&#x27;t you know the legend about rice grains on a chess board?

                        Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM&#x27;s training data.

                        &gt; The claim here is not about intelligence, it is about generality. There&#x27;s no doubt for me the LLMs are intelligent.

                        Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say &quot;E=mc^2&quot;, does that make you think I am Einstein?

                        1. lostmsu · · focus · HN ↗
                          [delayed]
                          1. HarHarVeryFunny · · focus · HN ↗
                            You are talking about 2^64 being a huge number I assume ?

                            If not, then what are you talking about ?

                            If yes, then what is the relevance to an LLM playing chess ?

                            1. lostmsu · · focus · HN ↗
                              [delayed]
                              1. HarHarVeryFunny · · focus · HN ↗
                                1) The number of unique chess games that could theoretically be played (but mostly never have been), is irrelevant to what an LLM is remembering. It can only remember what was in it&#x27;s training data - a far smaller number of maybe 10&#x27;s of millions of games (of 30-50 moves each).

                                2) An LLM is not going to memorize vs generalize when there is no training pressure to do so. You might expect it to memorize book openings that occur over and over in the training data, but not some random non-celebrity game that occurs once in the Lichess dataset and is never again referred to.

                                &gt; They can&#x27;t possibly remember even a few positions. Don&#x27;t you know the legend about rice grains on a chess board?

                                If the wise man was a bit wiser, he&#x27;d have asked for his rice on a snakes &amp; ladders board (100 squares, not 64) and would have had 2^36 more rice, which is equally irrelevant.

              3. Gregkion · · focus · HN ↗
                An AGI doesn&#x27;t stand for &#x27;perfect intelligence&#x27; it stands for artificial general intelligence.

                And no an AGI system doesn&#x27;t need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.

                Just because you define AGI as something it doesn&#x27;t has to be,doesn&#x27;t mean i need to touch grass.

                This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing

                1. tsimionescu · · focus · HN ↗
                  Do you know what the &quot;General&quot; in &quot;Artificial General Intelligence&quot; means? It specifically means that the AGI adapts to novel domains that it hasn&#x27;t been trained on - its training generalizes to real world problems.

                  That doesn&#x27;t mean it has to be extraordinary at these things. But to be AGI, it has to have some level of competency when used on problems outside its training set. In particular, it the LLMs were to install a known chess engine and run that to get the moves when asked to play chess, that would qualify for more AGI-like behavior. But really, chess is such a simplistic game that they should be able to do decently well at it even without even needing that. At the very least, they should be able to consistently play without making illegal moves - something that many 7-year olds manage quite well.

                2. rsfern · · focus · HN ↗
                  On the contrary, I think the chess comparison is on point. We’re discussing observations that even the strongest models devolve into making invalid moves without scaffolding. For me that raises the question of whether these models are learning the rules and generalizing from them, or of they’re just pattern matching and flailing on this task. Maybe the reality is somewhere in between, but the benchmarks don’t seem to directly measure conceptual generalization, they measure task completion. They can disrupt a lot of people and industries by pattern matching and flailing without being AGI.

                  I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe

            2. striking · · focus · HN ↗
              It&#x27;s not quite the same, but the in-flight chess game provided by Delta was known to be absurdly hard: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46593395">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46593395
              1. willmarch · · focus · HN ↗
                I believe I remember reading it was based on Glaurung&#x27;s code (which eventually evolved into what we now know as the juggernaut Stockfish).
            3. what · · focus · HN ↗
              I can write a chess bot program that will beat you. Does that mean I’m good at chess?

              &gt;If they cared to have it perform well in chess games, you&#x27;d see a different shape and behavior.

              So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?

              1. phoghed · · focus · HN ↗
                They’ll never be AGI simply because the definition will be constantly updated to be some steps ahead of them.
                1. fc417fc802 · · focus · HN ↗
                  I&#x27;m pretty sure &quot;competent at chess without external aids&quot; has been on the standard AGI checklist since before personal computers were a thing. How can you claim an intelligence is general if it can&#x27;t make sense of such a highly constrained board game?
                  1. phoghed · · focus · HN ↗
                    Because they’ll train it to be good at chess and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even ____

                    It can’t even count the R’s in strawberry

                    It can’t even add numbers

                    It can’t even solve a millennium puzzle

                    It’s not even a chess GM

                    It’s not even beyond human capability in Go

                    It can’t even drive a car

                    It can’t even self replicate

                    It can’t even build weapons

                    It doesn’t even have feelings

                    So how could someone conceivably convince everyone that some system is AGI when there are still tasks that some human or group of humans can do that the system cannot?

                    This will only happen, in my opinion, when the model&#x2F;system can self-improve at a rate that scares people.

                    1. fc417fc802 · · focus · HN ↗
                      &gt; and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even

                      One, you&#x27;re not addressing what I wrote above and two, yes, that&#x27;s absolutely correct. Doing X doesn&#x27;t qualify something as AGI. If you can&#x27;t X you can&#x27;t be AGI. The inverse doesn&#x27;t hold though. In particular if you have to retain the model in order to X then it can&#x27;t possibly be AGI since (being _general_) it would be capable of figuring X out on its own having never seen it before.

                      1. phoghed · · focus · HN ↗
                        Completely arbitrary definition that nobody will agree on, stated as if it’s some self-evident ground truth.
                        1. fc417fc802 · · focus · HN ↗
                          Yes, it is indeed self evident. If it can&#x27;t figure things out then its intelligence isn&#x27;t general in which case it can&#x27;t be AGI by definition.
                          1. phoghed · · focus · HN ↗
                            No, because there is no coherent, agreed-upon definition. There’s just a million people vibe defining it.

                            Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever.

                            Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not.

                            And for what it’s worth I just watched GitHub Copilot figure something out. So your definition is once again lacking.

                            1. fc417fc802 · · focus · HN ↗
                              Throughout this exchange you&#x27;re repeatedly confusing the negative and the positive. There is no rigorous and universally agreed upon criteria for exactly what would constitute AGI. There are some vague shapes that are widely (but not universally) accepted such as largely (vague boundary) being capable of replacing (vague criteria) humans.

                              However there are plenty of disqualifiers that are more or less universally accepted. In the above case it is literally by definition. Something cannot be termed general if it is incapable of generalizing.

                              1. phoghed · · focus · HN ↗
                                &gt; However there are plenty of disqualifiers that are more or less universally accepted (ie the negative)

                                Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI. There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them.

                                An AI controlled robot will be standing over the cooling corpse of the last human who will die certain that it wasn’t done by AGI.

                                1. cindyllm · · focus · HN ↗

                                  [dead]

            4. jibal · · focus · HN ↗
              First, you&#x27;re moving the goalposts. Second, it&#x27;s not actually true that any existing frontier AI can write a chess bot program that can beat a 1600 player ... not unless the program is derived from Stockfish or some other leading engine that has been in development for decades.

              &gt; The system is built and delivered as it is because that&#x27;s how the providers make the most money. If they cared to have it perform well in chess games, you&#x27;d see a different shape and behavior.

              These comments indicate a complete failure to understand the technology.

              I won&#x27;t respond again.

            5. zahlman · · focus · HN ↗
              &gt; The system is built and delivered as it is because that&#x27;s how the providers make the most money. If they cared to have it perform well in chess games, you&#x27;d see a different shape and behavior.

              This argument is fundamentally incompatible with all the breathless rhetoric about &quot;AGI&quot; coming from the providers&#x27; general direction.

              1. echelon · · focus · HN ↗
                It&#x27;s really not.

                The labs frequently apply their raw models to problems that do not make economic sense for their customers but that demonstrate the power and capability of their systems. These experiments can cost millions of dollars. That&#x27;s not customer-shaped.

                They&#x27;re not going to give you access to that. It&#x27;s not a product. The government might have an interest in this, but that&#x27;s not something you&#x27;d be privileged to know about.

                And when these labs do develop &quot;AGI&quot;, they more than likely won&#x27;t be selling it to end users. They&#x27;ve pretty much already said this.

              2. anthonyrstevens · · focus · HN ↗
                &gt;&gt; the breathless rhetoric about &quot;AGI&quot; coming from the providers&#x27; general direction

                So many commenters here see it as their ... duty? to argue against the most optimistic&#x2F;unhinged (take your pick) arguments from &quot;the other side&quot; and then treat everybody who disagrees as a shill or an idiot.

                Why is &quot;being good at chess&quot; a proxy for whatever AGI strawmen you want to argue against?

                Maybe step back from your black-and-white ledge and think about discussing what&#x27;s actually under discussion? For example, why or why not would an LLM be good at chess? Will they be good at chess? What technical limitations might preclude that?

          3. uncivilized · · focus · HN ↗
            HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.
            1. xdavidliu · · focus · HN ↗
              that is if it even a human commenter at all
              1. linkjuice4all · · focus · HN ↗
                State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of &#x27;guard railing&#x27; required to keep even the latest models completely on-task. The chess example is interesting because it&#x27;s clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behavior in these systems that&#x27;s difficult to engineer out.
                1. YeGoblynQueenne · · focus · HN ↗
                  [delayed]
                  1. 27183 · · focus · HN ↗
                    But it speaks in words, therefore it must be super duper extra smart!!11 &#x2F;s

                    Sarcasm aside, I think this is an easy cognitive trap to fall into. It does sometimes feel like the LLM must have some world model because it converses somewhat coherently. Examples like this failure to understand chess, or to count the number of Rs in &quot;strawberry&quot;, seem difficult to explain if the models are intelligent. But that doesn&#x27;t stop people believing they are anyway. I think there must be something about the conversational interface that fools us easily. I wonder if people trained in interrogation techniques are also fooled?

                    1. YeGoblynQueenne · · focus · HN ↗
                      [delayed]
                  2. hackinthebochs · · focus · HN ↗
                    &gt;If that were true, we should have seen LLMs play good chess by now.

                    Not at all. LLMs learn by imbibing a mass of relationships as isolated fragments of information. There is a certain amount of sorting and indexing that happens during the training phase. There is also a certain amount of compute executed on these relationships during inference. LLMs can model processes that fit within the compute budget. Language translation works well because language is lookup-heavy while being light on compute.

                    Chess is a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Humans cut through the compute requirements by reinforcement and learning intuition. LLMs don&#x27;t get reinforcement on chess so they must compute during inference a unified model of chess. Developing a strong model of chess from raw fragments of information is simply not in their compute budget.

                    1. YeGoblynQueenne · · focus · HN ↗
                      [delayed]
                  3. geoffschmidt · · focus · HN ↗
                    [delayed]
                    1. YeGoblynQueenne · · focus · HN ↗
                      [delayed]
                  4. DavCreator · · focus · HN ↗
                    <a href="https:&#x2F;&#x2F;xxcancel.com&#x2F;biobootloader&#x2F;status&#x2F;1640512444958396416" rel="nofollow">https:&#x2F;&#x2F;xxcancel.com&#x2F;biobootloader&#x2F;status&#x2F;164051244495839641...
                2. TheOtherHobbes · · focus · HN ↗
                  I&#x27;m not sure why anyone is expecting stochastic systems to be deterministic.

                  Chess is a deterministic game won by a combination of known movesets and constrained multi-level forward search.

                  LLMs do neither of these things. They don&#x27;t reproduce training data exactly, their next response is more &#x27;inspired by&#x27; prompts and its own memory than produced deterministically, and they don&#x27;t have the capability to do general forward search on their own.

                  So when you ask an LLM to play chess you&#x27;re getting the equivalent of a very compressed and lossy JPEG of chess rules and strategies with added per-turn random noise.

                  They also don&#x27;t have the ability to design their own chess engine, although it would be interesting to see what happens if you ask for one.

                  1. YeGoblynQueenne · · focus · HN ↗
                    [delayed]
                  2. dezsiszabi · · focus · HN ↗
                    I&#x27;m expecting that they at least don&#x27;t forget about pieces between turns, we&#x27;re in AGI era after all, according to the tech overlords.

                    I, as a human AGI, would jever just forget and remove a piece from the board from one turn to the next.

            2. nalekberov · · focus · HN ↗

              [dead]

            3. avadodin · · focus · HN ↗
              Back in 2001, our social medium was Slashdot and no one ever pretended to read the article. No one read the article either. It was slashdotted most of the time anyways.
              1. _superposition_ · · focus · HN ↗
                Oh shit he said slash dotted. Havent heard that in a long time!
          4. Onavo · · focus · HN ↗
            &gt; even if I give them literal infinite time and all the subagents and internet access..

            Don&#x27;t use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.

            1. sfn42 · · focus · HN ↗

              [dead]

              1. Onavo · · focus · HN ↗
                Give me a proof they don&#x27;t. Because from my observations they clearly do.
          5. automatic6131 · · focus · HN ↗
            HackerNews is Gell-Mann amnesia that refreshes on every comment on every thread.
          6. dmurray · · focus · HN ↗
            &gt; I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.

            I don&#x27;t believe this.

            You refer to &quot;subagents&quot;, so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.

            A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.

            I completely believe the LLM on its own can&#x27;t play a full game of chess at your level. Though I&#x27;d bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don&#x27;t do it because there are other approaches that play chess much better.

            1. lirolero · · focus · HN ↗

              [dead]

          7. YeGoblynQueenne · · focus · HN ↗
            &gt;&gt; I know HN readers and posters just read numbers and can&#x27;t be bothered to read, but please read the methodology before making any claims.

            This is unfair to HN readers all of whom but one did not post the comment you replied to. You can&#x27;t just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.

            1. minraws · · focus · HN ↗
              How many posts if I link that do the same thing will you agree this is the norm here.

              Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine.

              I don&#x27;t even know if there is critical thought or we believe what we read&#x2F;shared&#x2F;etc

              1. dezsiszabi · · focus · HN ↗
                50% + 1 of all comments
              2. YeGoblynQueenne · · focus · HN ↗
                No, I don&#x27;t agree it&#x27;s the norm. There is though a general tendency to opine with strong views on subjects posters have no expertise on. I think that&#x27;s because many are software engineers (or equivalent) and they are used to being expected to &quot;wing it&quot; on whatever technical subject comes up. On the other hand you can always find informed comments by users who have specialist knowledge.

                And there&#x27;s plenty of pushback on here about the chess thing besides your very valid points.

                EDIT: anyway if I can offer a bit of unsolicited advice, it won&#x27;t do you or anyone any good to accuse everyone who doesn&#x27;t agree with you of laziness, even if you can see e.g. they haven&#x27;t really read an article. Just say the thing you wan to say and let them figure it out. Most people will appreciate that much better and you will feel better about yourself for acting like a mature adult.

                It&#x27;s even in the site guidelines:

                Please don&#x27;t comment on whether someone read an article. &quot;Did you even read the article? It mentions that&quot; can be shortened to &quot;The article mentions that&quot;.

                1. minraws · · focus · HN ↗
                  It&#x27;s not been my personal experience on this website in the last 2-3 years atleast, pre-covid perhaps.

                  But despite that you aren&#x27;t wrong and the only reason I even visit this website is because people sometimes did&#x2F;do take time to reflect on things based on their experience and knowledge.

                  And in hindsight pointing out that hn has issues wasn&#x27;t even the point but I feel frustrated when everyone is readily agreeing to things on here without reading. When that in this moment feels like the one thing that separates humans from machines that we get to think and learn.

                  I possibly should just drop reading this place until we have most noisy people go away. I have for one tried to always only comment on things where I could be a value add, this one does feel like I could I have done better.

                  In the moment I probably thought if they are GM level and I can beat them, is this some interesting find, my disappointment honestly led me to making a rather incorrect call on this one.

                  Either way I still do think HN as a whole has devolved into mindless herd follower mindset, I can point to more than a few posts that just say adopt the hacker mindset aka move fast don&#x27;t care about the consequences.

                  And I for one find this laughable even though that&#x27;s the reality of my job&#x2F;work as well.

                  1. YeGoblynQueenne · · focus · HN ↗
                    [delayed]
                    1. minraws · · focus · HN ↗
                      &gt;&gt; Sorry, I didn&#x27;t get this? What was the incorrect call you made?

                      Talking about people&#x27;s inability to read rather than just pointing out that the article pointed at something else.

          8. victorbjorklund · · focus · HN ↗
            &gt; Why do I even scroll through this website.

            Because other HN bring in their own experience telling us what is real and what is BS. Maybe next time it will be someone else with experience in something else that will call out BS and you will see it. I didn’t really think LLM:s are any near good in chess but I don’t play chess so don’t know what 1600 means. So you helped me by calling BS.

          9. thelaxiankey · · focus · HN ↗
            I&#x27;m just dropping this all over this thread but you&#x27;re unfortunately mistaken

            <a href="https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;" rel="nofollow">https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;

            1. freejazz · · focus · HN ↗
              More show and less tell would be appreciated.
            2. minraws · · focus · HN ↗
              Summarizing here for my dear friends, the guy on the other end managed to fine tune a model gpt-3.5-fine-tune against stockfish vs stockfish games to perform at 1200 elo level against stockfish.

              I have been proved wrong I should have quit while I was ahead. &#x2F;s

              1. thelaxiankey · · focus · HN ↗
                I think your summary is not really correct, maybe I&#x27;m missing something. As far as I can read, the approximate takeaways are these:

                * LLM chess play is hyper sensitive to the harness being used (see: the section on regurgitation), and only mildly sensitive to fine tuning

                * The best ELOs the author was observing were from 1500 to 1750 or so (circa 2024&#x2F;25). Not grandmaster, but no longer incompetent monkey either.

        5. [deleted] · · focus · HN ↗

          [deleted]

        6. sashank_1509 · · focus · HN ↗
          These ratings seems very wrong, i have beaten GPT Astra max thinking in chess and my rating is close to 1500. The ratings here seem more accurate: <a href="https:&#x2F;&#x2F;chessbenchllm.onrender.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;chessbenchllm.onrender.com&#x2F;

          GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time

          1. jibal · · focus · HN ↗
            &quot;Elo is relative to the ChessBench field.&quot;

            They are of course &quot;wrong&quot; if you don&#x27;t read the faint fine print and sensibly interpret them as FIDE or similar ratings.

        7. sobellian · · focus · HN ↗
          If it&#x27;s a GM then I&#x27;m Magnus Carlsen, <a href="https:&#x2F;&#x2F;lichess.org&#x2F;study&#x2F;27lCQqDa" rel="nofollow">https:&#x2F;&#x2F;lichess.org&#x2F;study&#x2F;27lCQqDa.
        8. boesboes · · focus · HN ↗
          Dumbest thing I’ve seen today
        9. jibal · · focus · HN ↗
          Please do not post misinformation. They are not playing anywhere near GM level.

          &quot;Elo is relative to the ChessBench field.&quot;

        10. zahlman · · focus · HN ↗
          &gt; The actual current frontier plays somewhere around GM level.... It&#x27;s also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors

          Sorry, but I am not buying that 5.6-Sol is that much better than 5.6-Luna, which can barely be coaxed to reach the midgame with legal moves and an apparent understanding of what the position is.

      3. [deleted] · · focus · HN ↗

        [deleted]

      4. sobellian · · focus · HN ↗
        I tested both myself and a weak bot against Astra xhigh, <a href="https:&#x2F;&#x2F;lichess.org&#x2F;study&#x2F;27lCQqDa" rel="nofollow">https:&#x2F;&#x2F;lichess.org&#x2F;study&#x2F;27lCQqDa. It&#x27;s still pretty bad at chess, though it takes longer to devolve into illegal moves.
        1. phist_mcgee · · focus · HN ↗
          That&#x27;s really cool, thanks for sharing!
        2. hackinthebochs · · focus · HN ↗
          So you weren&#x27;t giving it an updated board state after every move? If you want to compare apples to apples, it should give an updated board state for each move, or you should play blindfolded.
          1. sobellian · · focus · HN ↗
            I can play blindfolded. I am expert OTB (though I haven&#x27;t played in a while). The game was like 18 moves of theory in the Maroczy Bind.
          2. Topfi · · focus · HN ↗
            Blindfolded flex by OP aside (I can barely play when seeing the board), considering reasoning traces and their nature, if we want to be fair, a person would have to get the moves, but be allowed to write them down or draw up a board in their notepad. My working memory can barely handle five chunks, a models reasoning tokens are masses of written text in comparison.
          3. HarHarVeryFunny · · focus · HN ↗
            An LLM has been trained to do everything it does blindfolded, &quot;only&quot; using perfect recall of everything in it&#x27;s hundreds of thousands of steps of context, and hundreds of layers of KV cache. It&#x27;s a computer - it has a massive advantage over a human.

            The fairest apples-to-apples comparison of an LLM whose training data included chess games would be a trained human such as Magnus Carlson, who can quite happily play a dozen or more simultaneous blindfold chess games.

        3. losvedir · · focus · HN ↗
          &gt; though it takes longer to devolve into illegal moves

          Is this because the context is being saturated? How did you set it up?

          Was the prompt something like &quot;Here&#x27;s the state of the board, you&#x27;re white, your move, what do you do?&quot; and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (thanks for sharing!) to my mental model and understanding.

          I&#x27;d be curious how it would work if it started fresh each time. My guess is it would never make an illegal move, although it may not actually play all that well.

          1. sobellian · · focus · HN ↗
            You can see the entire conversation for my game at <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aaac17b-1384-83e8-98fd-4350a0ef69cd" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aaac17b-1384-83e8-98fd-4350a0ef69....
        4. bhelkey · · focus · HN ↗
          It looks like it played a fully legal game of chess with one exception, it said &quot;rxd1+&quot; (Rook takes D1 with check) instead of &quot;rd1+&quot; (Rook to D1 with check) on move 29.

          I would say this did a really good job of playing chess. It moved the pieces consistently and traded pieces when required.

          This is worlds away from the frontier ~1 year ago where models would hallucinate pieces into existence.

          1. sobellian · · focus · HN ↗
            You can still see undercurrents of its old self, once I pointed out the illegal notation it hallucinated prior illegal moves. But I agree, it&#x27;s leagues apart from prior iterations. It also knew thematic moves in the opening. But whenever it needs to play concretely rather than &quot;I know so-and-so is a good move in these types of positions&quot; it crumbles.
            1. bhelkey · · focus · HN ↗
              Agreed, move selection was not great. Notably, it should not have allowed nxe7+.

              However, the pawn was defended by the queen and it took a forced queen trade to unlock the move.

              I have seen much worse blunders from human players. And, I have made much worse blunders.

          2. legulere · · focus · HN ↗
            Would you tell a human that just tried doing an illegal move that they did &quot;a really good job of playing chess&quot;? The probability for such mistakes is greatly reduced but still far from negligible, which proves the point that guardrails are needed.
      5. aprilthird2021 · · focus · HN ↗
        They still need supervision though
      6. 21asdffdsa12 · · focus · HN ↗
        So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.
        1. [deleted] · · focus · HN ↗

          [deleted]

        2. ares623 · · focus · HN ↗
          Well I guess this excuse is finally gonna become obsolete soon with all the &quot;pacing&quot; nonsense.
        3. user43928 · · focus · HN ↗
          There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1.

          The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier.

          That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.

          1. Topfi · · focus · HN ↗
            Fortunately, a fellow commenter was so kind and did it with Astra. Didn&#x27;t do that well either [0]. I&#x27;m sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...

            I&#x27;ll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. &quot;Just&quot; having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn&#x27;t even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.

            [0] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751

            1. [deleted] · · focus · HN ↗

              [deleted]

            2. user43928 · · focus · HN ↗
              Doesn&#x27;t look impressive, although I&#x27;m hearing a marked improvement in choosing legal moves, compared to early 2025.

              Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better?

              I would not be surprised if OpenAI released a model that beats humans at chess this year.

              1. Topfi · · focus · HN ↗
                I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish.

                Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models have gained in actually capability that is in the training data, but not RLHFd to hell, so to speak. Tracking the state of pieces, I suspect given similar in Sudoku [0], is what these models struggle with in game settings, whilst tracking the state of code changes can be reliable over 250k tokens. Essentially, for the latter they were trained in the specific manner that lead them to abstract the capability, but that doesn&#x27;t track to the former, which is a massive difference between LLMs data focused training and human learning.

                So yeah, GPT-7 or any upcoming&#x2F;present LLM could do massively better in Chess than GPT-6 Astra, but not because the approach was emergent out of pure data. Rather, it requires a very specific training data type and stack for a model to gain capabilities that track a specific task long enough to adhere to the rules of a game such as chess.

                [0] <a href="https:&#x2F;&#x2F;logicalintelligence.com&#x2F;blog&#x2F;energy-based-model-sudoku-demo" rel="nofollow">https:&#x2F;&#x2F;logicalintelligence.com&#x2F;blog&#x2F;energy-based-model-sudo...

                1. user43928 · · focus · HN ↗
                  I&#x27;m wondering if instructing it to track the board state in a file would make a significant difference then.

                  It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new context compaction method increased the performance dramatically. However, I think that is not applicable here.

                2. 21asdffdsa12 · · focus · HN ↗
                  So what is the supposed leap? One agent per option to change, evaluating the board state that there move would create, by having a army evaluate the remaining piece options and average over that? Wee-Free-Man as a hierarchical army ? Pet-LLMs trained on one thing?
                  1. Topfi · · focus · HN ↗
                    Honestly, for intelligence I don&#x27;t know and I doubt anyone can claim to know. Maybe JEPA, there is potential concerning some shortcomings inherent to LLMs but it has its own, maybe scaling up the electron microscope stuff Google just did (though the connections are inferred), maybe future implementations of autoregressive and diffusion LLMs can at some point address its issues after all, maybe something else entirely.

                    All I know is, AGI, as in actual intelligence, is quite a massive accomplishment to claim and we shouldn&#x27;t loose sight of that fact, especially as &quot;not being intelligent&quot; does not make these models any less impressive, fascinating to work on or useful in many tasks. Personally, the only thing I am fairly convinced on is that if we were to find a way to create actual intelligence, it likely wouldn&#x27;t start out as useful as todays LLMs are and may thus be dismissed early. But again, pure speculation on that front.

                    If for leap you just mean more utility from LLMs as they are, then I&#x27;ll pretty confidently put my money on higher quality, not more, training data for a wide range of verifiable tasks. What makes maths, coding, etc. comparatively easy to make gains in (though less verifiable tasks can also make similar as seen with the writing in Kimi K2).

              2. datsci_est_2015 · · focus · HN ↗
                Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the concentration of wealth even further and fully realize the American dream of eliminating the middle class.
                1. user43928 · · focus · HN ↗
                  I&#x27;ve seen some of his videos, and got the impression he didn&#x27;t understand how GPT-Live delegates to the more powerful regular model with reasoning.

                  The regular model generally does not suffer the same issues he is demonstrating with the real time audio version.

                  In my view the investment into datacenters is well justified by the current demand, and progress has been very impressive.

                  1. freejazz · · focus · HN ↗
                    Really? It was being sold as a total replacement for jobs like software engineering and being an attorney, but its looking a lot more that its just going to be a tool those professions use and doesn&#x27;t actually seem to be taking jobs away.
                2. Quinner · · focus · HN ↗
                  I find it amusing that you&#x27;re describing a huge misallocation of capital and a society enabling such, and that is the optimisitic scenario (in my mind anyway).
        4. thelaxiankey · · focus · HN ↗
          Sure: <a href="https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;" rel="nofollow">https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;
      7. yuxi258 · · focus · HN ↗

        [dead]

      8. dgb23 · · focus · HN ↗
        The gap in capabilities is mostly quantitative and not qualitative.
        1. RealityVoid · · focus · HN ↗
          Is it? I am on the fence on this, but it does seem like there are some qualitative improvements between the models.

          Not related to your post, but a fact I keep mulling over. The fact I don&#x27;t trust the current crop of LLM&#x27;s enough and I consider LLM&#x27;s as a tech will hit a ceiling pretty hard, it doesn&#x27;t mean parallel improvement curves won&#x27;t spring up out of other research that will lead to much higher capabilities than currently.

          1. zahlman · · focus · HN ↗
            &gt; but it does seem like there are some qualitative improvements between the models.

            It could easily seem that way, I think, in a &quot;quantity has a quality of its own&quot; kind of way. When you can come to the same conclusion faster, that lets you iterate more; and sometimes when you iterate you find more things.

          2. dezsiszabi · · focus · HN ↗
            &gt; Is it?

            Yes, it is.

        2. lionkor · · focus · HN ↗
          My read is that the improvements in quality are due to excessive use of &quot;thinking&quot; tokens (so, higher quantity and brute force), so I agree with that.
      9. zahlman · · focus · HN ↗
        Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with:

        &gt; Let&#x27;s play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.

        (I hoped the latter requirement would help it be &quot;not blindfolded&quot;; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)

        For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.

        It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the &quot;book&quot; theory in its training data, but it still completely fell apart at early midgame.

        1. titzer · · focus · HN ↗

          [dead]

          1. lirolero · · focus · HN ↗

            [dead]

          2. danpalmer · · focus · HN ↗
            Sure, but installing a chess program is child&#x2F;teen level general ability, and playing chess well is highly trained expert level ability. Which one are we sold AI as being?
            1. alpinisme · · focus · HN ↗
              I think we are being sold AI as expert only when given tools (although that is not emphasized). The (quasi?) miracle of AI right now is that you can get an agent to accomplish the task of a team of intelligent but not exceptional humans at speeds far exceeding what the human could do. Which makes it “cheap” to throw (effectively) dozens of teams at a problem for the equivalent of hundreds of man hours.

              That may not be the AI of sci fi fantasy but it’s still a game changing reality.

          3. kavok · · focus · HN ↗
            I often don’t see agents reaching for available or potential tools&#x2F;libraries unless explicitly told to.

            Sometimes they’ll even manually search or write bespoke code to search json instead of using something like jq.

          4. HarHarVeryFunny · · focus · HN ↗
            A Transformer has a massive amount of state - it&#x27;s entire KV cache, in addition to the user asking it to draw the state after every move, which is really unnecessary.

            A human, at least a trained human (for fairer comparison to an LLM whose training data contained a ton of chess games) can absolutely do this - have you never seen demonstrations of expert players playing a dozen or more games while blindfolded?

            A Transformer&#x2F;LLM is not a human of course, and the way it will by default play chess is by prediction, not reasoning. An LLM actually does surprisingly well if you only give it the most recent 20 moves of a game where 40 moves have been played so far, since the moves NOT played tell it just as much as the ones that were played, letting it effectively infer a lot of what is on the board.

            1. zahlman · · focus · HN ↗
              I just want to make sure it&#x27;s clear: the reason I was asking it to redraw the board is because last time I tried (which was like a month ago), I didn&#x27;t ask for that, and basically as soon as the opening was &quot;out of book&quot; it started trying to make illegal moves and made false statements about the position in its running commentary (and after being corrected on these points, started dropping pieces for no reason).
        2. sailfast · · focus · HN ↗
          What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill?

          Also what harness? If you’re using a general harness of course it’s going to try and give you commentary.

          I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.

          1. datsci_est_2015 · · focus · HN ↗
            Why does a 6 year old not need any of these guardrails?

            Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking &#x2F; influencing about.

            Maybe not all AGIs have a path to digital singularity. Maybe our current era of intelligence modeling has fundamental flaws and we are in a local minimum of the artificial intelligence space.

            To note, I would bet with a good amount of certainty that we have enough compute power and automation to DDOS the internet out of existence with botnets. That doesn’t make the frontier models intelligent, that just makes their handlers reckless.

            1. trio8453 · · focus · HN ↗
              &gt; Why does a 6 year old not need any of these guardrails?

              They&#x27;re not guardrails, they&#x27;re a different input&#x2F;output environment.

            2. solenoid0937 · · focus · HN ↗
              Ask a 6 year old to draw a chess board from scratch every turn and they too will make mistakes.
              1. datsci_est_2015 · · focus · HN ↗
                A 6 year old will figure out how to ask you to help them after they get it wrong.
              2. freejazz · · focus · HN ↗
                No one has spent the past three years telling me that a 6 year old will take my job!!!
                1. claytongulick · · focus · HN ↗
                  And the 6 year old doesn&#x27;t cost more than the GDP of a medium sized country.
                  1. sailfast · · focus · HN ↗
                    [delayed]
              3. wavemode · · focus · HN ↗
                [delayed]
            3. gf000 · · focus · HN ↗
              Well, would a dissected frontal lobe in and of itself be intelligence?

              I think the same goes for LLMs, they may be a core part of an LLM harness, but you may still need a couple other components (e.g. it may itself write itself a deterministic function to validate steps).

              In and of itself intelligence is an ill-defined and badly understood concept.

            4. themgt · · focus · HN ↗
              Why does a 6 year old not need any of these guardrails?

              Why does a bird not need jet engines or regular professional maintenance?

          2. topaz0 · · focus · HN ↗
            You&#x27;re pointing out that the goalposts are not fixed in the problem statement above, and gp&#x27;s interpretation is not the most generous possible. But as the interpretations get more generous, the claim becomes more and more absurd. Maybe a properly-harnessed model would download the most advanced chess engine and query it to find the best move in each position, but that&#x27;s not really demonstrating the model&#x27;s intelligence anymore.
          3. zahlman · · focus · HN ↗
            &gt; but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.

            This isn&#x27;t just about judging LLM capability. This is about pointing out that these capabilities are not &quot;AGI&quot;. If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.

            Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it&#x27;s hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.

        3. meowface · · focus · HN ↗
          Luna is one of the budget lower-end last generation models. It&#x27;d be useful to at least try to verify the present before being bearish about the future. For OpenAI, the best publicly available model is GPT-6 Astra with XHigh or Max reasoning, and for Anthropic it&#x27;s Claude Fable 5.1 with XHigh or Max reasoning.
          1. XMPPwocky · · focus · HN ↗
            out of curiosity, do you think fable would get this right? (I&#x27;m not sure myself, and haven&#x27;t tried yet.)
            1. zahlman · · focus · HN ↗
              Elsewhere in the thread there are reports of Astra on xhigh playing at what I would characterize broadly as a competent casual level, at least given occasional prodding (which a human of that skill level would basically only require when trying to play unreasonably quickly). There seems to be a pattern (even after correcting for relative ELO systems that aren&#x27;t calibrated) of the LLM bots demonstrating stronger play against traditional bots than against humans.
      10. nutrientharvest · · focus · HN ↗
        &quot;Transatlantic flight will never be commercially viable, we conclude based on careful study of several aircraft designs from the 1920s&quot;
        1. ponector · · focus · HN ↗
          How about supersonic flight?
          1. ggreer · · focus · HN ↗
            I don&#x27;t think that&#x27;s a useful comparison. Supersonic military planes have been common for decades. We don&#x27;t have supersonic passenger planes because the FAA has banned supersonic flight over land since 1973, though the agency is planning on replacing it with a noise standard. Also the original supersonic passenger aircraft were government-sponsored tech demos, not financially sustainable products. With updated laws &amp; modern technology (cameras instead of tilting noses, more efficient engines without afterburners, lighter materials), we could have viable supersonic passenger flight.
        2. zeroonetwothree · · focus · HN ↗
          Technology keeps advancing in a domain until suddenly it doesn’t. Where are my flying cars?
          1. krapp · · focus · HN ↗
            They&#x27;re called helicopters.
            1. freejazz · · focus · HN ↗
              And what since then?
              1. pixl97 · · focus · HN ↗
                This has the smell of &quot;Why don&#x27;t I have a faster horse&quot;.

                Why no flying cars. Because objects have mass and inertia and people are incredibly stupid. Making a flying car has been done. Making a flying car not be a weapon of mass destruction is very, very hard.

                Also:

                <a href="https:&#x2F;&#x2F;www.txdot.gov&#x2F;about&#x2F;newsroom&#x2F;statewide&#x2F;air-taxi-testing-taking-flight-in-texas.html" rel="nofollow">https:&#x2F;&#x2F;www.txdot.gov&#x2F;about&#x2F;newsroom&#x2F;statewide&#x2F;air-taxi-test...

                1. freejazz · · focus · HN ↗
                  You&#x27;re making my point for me, surprised you don&#x27;t realize that...
                  1. pixl97 · · focus · HN ↗
                    Because you don&#x27;t fully understand your own point...

                    You look at science fiction and say &quot;why didn&#x27;t I get flying cars&quot; and not &quot;why didn&#x27;t most science fiction predict a global always on network that put the furthest places away from you a few microseconds away from audio, video, or any other type of information that can be digitally encoded.

                    Trying to use flying cars as a gotcha is missing that flying cars aren&#x27;t near as useful as one would think in relation to their costs. Moving information has become far more useful than moving objects long distances quickly, especially humans.

                    1. freejazz · · focus · HN ↗
                      &gt; You look at science fiction and say &quot;why didn&#x27;t I get flying cars&quot;

                      I definitely don&#x27;t, and you&#x27;re definitely not getting my point, but I&#x27;m amused that you&#x27;ve instead double down on somehow getting it more than me...

      11. moron4hire · · focus · HN ↗
        &gt; The gap in capabilities between those models which they tested, and actual current frontier ones is enormous.

        Same story every 4 months and yet still no breakout, winning products. I&#x27;ve been hearing &quot;the AI is good now&quot; and &quot;it 10x&#x27;s my productivity&quot; for a over a year now. If it were true, why aren&#x27;t the all-in-AI using companies 10-15 years ahead of their competition yet? Why is it still all buggy, poorly designed junk?

        1. orangedog · · focus · HN ↗
          I don&#x27;t get why it is hard to understand there is middle ground. People are 10x their productivity, it isn&#x27;t all buggy junk, but it isn&#x27;t all it is hyped up to be either. It isn&#x27;t that complicated.

          If you hold the extreme position that there isn&#x27;t any value in this, that&#x27;s fine, but we&#x27;re only having this discussion because these models have done what humans previously failed to do.

          1. freejazz · · focus · HN ↗
            I don&#x27;t think the poster disagrees with you at all. The middle ground is that there are no breakout products and that the models clearly aren&#x27;t so powerful as to make these companies not produce shit code.
        2. autoexec · · focus · HN ↗
          Right now AI hasn&#x27;t even managed to replace all the human workers taking orders at the fast food drive thru. That&#x27;s a job often performed by literal children and companies are still waiting for AI to get good enough for even that. Maybe one day it will be good enough, maybe one day it will outperform humans at such a basic task, but that day is not today. If the hype were anything close to reality, we&#x27;d see it everywhere in our lives.
          1. wavemode · · focus · HN ↗
            Funny you mention this - a fast food restaurant in my town now has an LLM taking drive-thru orders.

            Though I highly doubt it has taken anyone&#x27;s job, since most of the work is still in making, packing and handing over the food. (In fact, given the area I live in, I partially feel like the advantage they saw in it was that the LLM can speak Spanish.)

            1. autoexec · · focus · HN ↗
              Last I heard McDonald&#x27;s and Taco Bell were trialing AI again at a limited number of stores. It&#x27;s the kind of job AI should be really good at and many fast food companies are using call center workers currently. They really want AI to work, so they keep trying every few years to make it happen, but so far all they get are embarrassing social media posts
    4. consensus1 · · focus · HN ↗
      This isn&#x27;t how intelligence works. The LLM may not be able to play chess directly through inference, but it can write a program to do it and execute that program. Same as how human intelligence works. We can&#x27;t fly, but we can build planes.
      1. nefarious_ends · · focus · HN ↗
        Thanks for saying this, feels like everyone has gone insane over this stuff.
        1. what · · focus · HN ↗
          Humans don’t code a $game engine to play $game, they can just play it. It seems like you are the one that has gone insane.
          1. hackinthebochs · · focus · HN ↗
            And how many years of direct play and study does it take for a human to get good at chess or any other game? Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. That&#x27;s just not how the brain works. If LLMs could do that they would truly be superintelligence.
            1. sph · · focus · HN ↗
              No, learning is definitely not a sign of super intelligence. I know words don’t mean anything anymore, but that is simply general intelligence, despite the claims we have reached this milestone.
              1. hackinthebochs · · focus · HN ↗
                No, but superhuman capabilities derived from learning is, which is what the parent comment described.
            2. lelanthran · · focus · HN ↗
              &gt; And how many years of direct play and study does it take for a human to get good at chess or any other game?

              Time is irrelevant to training; the more relevant comparison is &quot;how many games does a human need to play to get diminishing returns&quot;.

              1. hackinthebochs · · focus · HN ↗
                Yes, obviously. The point was simply that human&#x27;s don&#x27;t one-shot chess so why does anyone expect an LLM to?
            3. orwin · · focus · HN ↗
              A week. My brother learned and was above 1100 online within 12 hours, after a few hundred games.
              1. hackinthebochs · · focus · HN ↗
                We&#x27;re obviously using different meanings for &quot;good&quot; here. But aside from that, it took 100&#x27;s to 1000&#x27;s of reinforcement iterations for your brother to play competently. While certainly impressive, that is still an entirely different category from piecing together disparate facts learned during training (LLMs aren&#x27;t analyzing a board as they&#x27;re learning the rules or ingesting PNG files), to executing a competent performance in one shot.
            4. mtlmtlmtlmtl · · focus · HN ↗
              &gt; Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess.

              Maybe not, but you&#x27;d be surprised how little it takes.

              A six year old child can learn the rules of chess well enough to be able to play legal moves only in a single day. And they can improve their game at a pace which is almost frightening to behold. I have taught children, and I&#x27;ve witnessed significant improvement materialise in a single game. LLMs have probably thousands of chess books, games, videos, etc in their training data, yet they are unable to even follow the rules.

              This is, at the very least, interesting. It illustrates many of the things brains can do, which current ML systems in general, and LLMs in particular, can&#x27;t.

              1. hackinthebochs · · focus · HN ↗
                It is interesting, but people are drawing the wrong conclusion from it. For one, LLMs don&#x27;t go through a &quot;chess learning phase&quot;. They&#x27;re not analyzing a board as they&#x27;re learning the rules or studying games to create a coherent model of chess. They&#x27;re just imbibing raw relationships as disparate fragments of information. The fact that they can&#x27;t unify this into a coherent model of chess playing in one shot and execute a competent game of chess says nothing interesting about the limits of their intelligence. If you give frontier models the rules of chess in their context window, could they perform only legal moves? I bet they could, excepting trickier scenarios like pins and failing to respond to a check. But those kinds of scenarios have to be reinforced in any human player as well.
      2. thesmtsolver2 · · focus · HN ↗
        Human beings can play chess directly without coding up a tool.
        1. consensus1 · · focus · HN ↗
          Very poorly compared to the tools we have built. Similar to the LLM.
          1. shimman · · focus · HN ↗
            Poorly in what sense? I think human chess leagues are way more popular and fun than just playing a computer by yourself. Human oriented communities are always a vastly better experience than their digital counterparts.

            There&#x27;s more to games than simply winning you know.

          2. thesmtsolver2 · · focus · HN ↗
            Comparing to raw LLMs? Much much better.
        2. qarl · · focus · HN ↗
          If they wanted to train an LLM to play chess they could easily do so.

          But nobody wants that.

        3. jstanley · · focus · HN ↗
          Asking an LLM to play chess by writing algebraic notation is like asking a human to play chess blindfolded.

          Yes some people can do it but most people can&#x27;t even if they&#x27;re unusually intelligent.

          You really need to be giving the LLM a board representation.

          EDIT: I see that they actually were giving the LLMs a board representation and they still played badly. Fair enough then.

    5. [deleted] · · focus · HN ↗

      [deleted]

    6. famouswaffles · · focus · HN ↗
      Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there&#x27;s a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn&#x27;t make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.
      1. bigstrat2003 · · focus · HN ↗
        &gt; Frontier labs don&#x27;t care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.

        If the models were actually intelligent, the way that the boosters claim, they wouldn&#x27;t need to be tuned to play chess in order to be good at it. That&#x27;s kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

        1. skydhash · · focus · HN ↗
          Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That&#x27;s the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use&#x2F;play something. Mastery can be gained overtime.
          1. willmarch · · focus · HN ↗
            If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn).

            Your assumptions&#x2F;intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).

            1. what · · focus · HN ↗
              &gt; considering LLMs currently play better than a brand new human player would

              They’ve ingested all the literature on playing chess, a brand new human player has not.

              1. willmarch · · focus · HN ↗
                Yes, but my point is that humans can’t even do the thing that the above comments are claiming humans can do (read a book or two and be decent at chess), and then they complain that LLMs can’t do the same thing (that humans can’t do either).

                We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.

                1. skydhash · · focus · HN ↗
                  &gt; The only way you get better at chess is by playing a lot of games and learning from mistakes

                  How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice.

                  This kickstarting then gradual refinement is how most people learn. And the foundational knowledge stays. Even a basic player knows to not do illegal moves.

                  1. willmarch · · focus · HN ↗
                    Reading can kickstart the process, but you can also make random moves guided by some sort of system (such as a computer GUI) or learn by watching other players play. The overall point is that you learn through observation and lots of trial and error (whether you are a human or a computer). And beginners in chess often make illegal moves even after learning the rules, it&#x27;s fairly common.

                    It feels like you&#x27;re trying to say that humans never make illegal moves while learning chess, which doesn&#x27;t match with my experience. I&#x27;m trying to understand your overall point.

                    1. skydhash · · focus · HN ↗
                      &gt; The overall point is that you learn through observation and lots of trial and error

                      That’s the most inefficient way and people usually avoid doing that. Instead they find someone that knows how to do the thing and ask him to be a teacher. Or use a proxy like a book or videos.

                      &gt; It feels like you&#x27;re trying to say that humans never make illegal moves while learning chess, which doesn&#x27;t match with my experience. I&#x27;m trying to understand your overall point

                      There’s learning the basic stuff (which is done after a few games) and there’s mastery. The thread started with the observation that even with all that knowledge (through content ingested in training), LLMs still makes illegal moves. Humans can be erratic, but they can constrain themselves to the rules for the task at hand after learning them.

                      1. willmarch · · focus · HN ↗
                        The only way to master anything is lots of trial and error. Coaches and teachers can help guide you towards more focused trial and error paths but the student still has to do the lessons and put in the work of learning, and learning only truly happens through doing.

                        Humans are not perfect and make mistakes in learning even when they have memorized the rules. A simple example is new players will often move a piece, exposing their king to check, and a more experienced player must point out to them that they have made an illegal move (because a new player often has not encoded that pattern for looking for exposed checks because they&#x27;re more focused on how the pieces move, not what that piece exposes.)

                        We&#x27;re just going to have to agree to disagree here.

            2. rsfern · · focus · HN ↗
              the discussion isn’t really about whether language models can become strong chess players though, the point is they seem to struggle to consistently make valid moves. Most humans don’t need to read two books to pick that up, just a couple lines of basic instructions
              1. willmarch · · focus · HN ↗
                That has not been my experience with new players, they regularly make invalid or incorrect moves even after detailed instructions especially in novel situations.
                1. rsfern · · focus · HN ↗
                  Maybe it depends on the person? My six year old isn’t great at strategy but they can pretty consistently make valid moves. Sometimes they ask for confirmation on a move which is also not a trait I see in language models (at least unprompted)
                  1. willmarch · · focus · HN ↗
                    Your child never messed up en passant (or had trouble understanding it in a real game), castled through or into a check, didn&#x27;t see a discovered check after moving a piece, never got confused by how stalemate works?
            3. orwin · · focus · HN ↗
              That&#x27;s quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).
              1. willmarch · · focus · HN ↗
                So your brother tried to make illegal moves while learning the game and it took your brother hundreds of games to get to be a decent player? I don&#x27;t see how this contradicts anything I said...
                1. orwin · · focus · HN ↗
                  The _only_ illegal move a human might make as a beginner is a failed en passant or a bad castle. And yes, a few hundred games is all it takes to be better than any publicly available LLM at the moment.
                  1. willmarch · · focus · HN ↗
                    ...and mistakes like not seeing discovered checks, castling into check, trying to castle out of check, castling through a check, misunderstanding en passant, missing checks when promoting a piece, misunderstanding how stalemate works, etc.

                    If you compared a human after hundreds of games to a SOTA LLM that was also trained on the output of hundreds of chess games that it played, I suspect you would notice similar improvements.

                    1. orwin · · focus · HN ↗
                      Honestly, if we&#x27;re only talking about SOTA llms, in sandbox mode without harness? No shot. Without harness LLMs have no memory of previous moves. I can&#x27;t make the last version of chatgpt remember more than 5 movements. I guarantee you, if you do not add flags in the harness with &#x27;left rook moved&#x27; or &#x27;right rook moved&#x27;, it will try to illegally castle 100% of the time it&#x27;s in the situation. Llm+harness, just make it call stockfish tbh.
          2. hi_im_greg_h · · focus · HN ↗
            1. The LLMs have surely ingested hundreds if not thousands of books on chess.

            2. The study (along with other posters here) show the models can’t even stick to following the rules of the game

        2. famouswaffles · · focus · HN ↗
          If humans were actually intelligent, they wouldn&#x27;t need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?
          1. diehunde · · focus · HN ↗
            Except all these LLMs were already trained with hundreds of chess book and game databases and they still suck
            1. famouswaffles · · focus · HN ↗
              If all you do is read chess books, you&#x27;ll be a shit player. Training and practice is what it takes to be great.
              1. sdf32dsf · · focus · HN ↗
                WTF even is this post?
              2. diehunde · · focus · HN ↗
                Oh right. But if all you do is reading programming books you are an amazing programmer? Where is all the training and practice LLMs did to become so good at coding?
                1. famouswaffles · · focus · HN ↗
                  LLMs (and Humans) don&#x27;t get good from programming books lol. The training and practice is the actual code they predict and learn from in the process of predicting.
                  1. diehunde · · focus · HN ↗
                    Oh I see. So if someone just reads books AND actual code then they can become experts, got it. And by the way LLMs are also trained with probably hundreds of thousands of actual games not just books
                2. hackinthebochs · · focus · HN ↗
                  &gt;Where is all the training and practice LLMs did to become so good at coding?

                  Coding is a matter of translating between the natural language description of a problem to the code specification while keeping the semantics fixed (and filling in the missing semantics reasonably well). It is not considerably more difficult than translating between two dissimilar natural languages. Chess isn&#x27;t a matter of language translation, but a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Chess takes directed practice and reinforcement whereas language translation does not.

                3. brindleth · · focus · HN ↗
                  It&#x27;s called post-training, typically through some form of reinforcement learning, and is a significant part of modern LLM development.

                  You have the first stage, pre-training, which is learning from next token prediction. That&#x27;s where the model memorises a lot of facts about things and generally gets good at forms of writing. It&#x27;s like reading a lot of books on programming and reading through a lot of source code. It&#x27;s learning how to autocomplete code, essentially. Doing that requires a developing a reasonable understanding of code, but it&#x27;s also learning how to autocomplete bad code as well as good, and won&#x27;t make it a &quot;good&quot; programmer.

                  Pre-training uses a method called Cross-Entropy Loss to update the weights of the network.

                  Then comes post-training. This is where the model is trained against huge sets of example problems, like fixing a bug, adding a new feature based on a spec, etc. They are set the task and try to complete it inside a training environment. Once they&#x27;re done, their complete solution is evaluated (either by humans, or by some separate evaluation model that was developed based on human feedback) and they are updated based on whether the solution was good or not.

                  Post-training uses a different method called Proximal policy optimization to update the weights of the network.

                  So these really are very different forms of learning, and mainstream LLMs are not post-trained to be good at chess. They could be. You could easily create a reinforcement learning environment that evaluated and improved their ability to play and win at chess. The result would be a very strong chess playing AI, something we know is possible because the strongest chess playing programs we have are neural network based, but it is not a priority for AI companies.

            2. WarmWash · · focus · HN ↗
              Contrary to popular belief, you need a lot of training on something for an LLM to be good and consistent with it.

              People think that if one mention exists in the training set, then the LLM is perfect at it.

              1. diehunde · · focus · HN ↗
                Not one mention. Hundreds of books, articles and databases of games.
          2. bigstrat2003 · · focus · HN ↗
            OpenAI making the next model good at chess is not analogous to a human training to get good at chess. It is analogous to God creating Human 2.0 which now has increased chess playing ability. If LLMs were intelligent the way humans are, then the models that exist right now would be able to spend time improving themselves at chess and become good at it. They can&#x27;t do this because they are not, in fact, intelligent.
            1. cindyllm · · focus · HN ↗

              [dead]

        3. Gregkion · · focus · HN ↗
          Thats just absolutly not true.

          A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

          And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

          1. foldr · · focus · HN ↗
            Humans don&#x27;t need a lot of training and finite tuning to make only legal moves.

            An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.

            An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).

            1. thom · · focus · HN ↗
              How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?
              1. foldr · · focus · HN ↗
                I don&#x27;t know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.
                1. thom · · focus · HN ↗
                  I&#x27;ve not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.
                  1. foldr · · focus · HN ↗
                    I haven&#x27;t tried it myself, but people seem to report that the illegal moves surface eventually. It just takes longer: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49720751

                    Nothing is forcing the LLM to play &#x27;blind&#x27;. If it&#x27;s smart, it should be able to create its own representation of the chess board and update it with every move, just like a human could.

                    1. thom · · focus · HN ↗
                      A human wouldn&#x27;t do that, they&#x27;d look at the board. I&#x27;m not disagreeing that to demonstrate clear superhuman ability the LLM should be able to do this, but it plays better than most humans blindfolded, and with fair prompts seems very good otherwise.
                      1. foldr · · focus · HN ↗
                        That&#x27;s what a human will do if they have a physical board to look at. But if someone, say, posed you a chess move exam question via FEN notation, you&#x27;d sketch a visual representation of the board off your own initiative to help you answer the question. There is nothing in principle to stop the LLM creating its own board representations in a format that enable it to easily keep track of legal and illegal moves. If it fails to do so, that&#x27;s a sign of its own limitations.
                        1. thom · · focus · HN ↗
                          I maintain that the amount of effort to teach a human to do this vastly outweighs the amount of effort to teach an LLM to do this unless you&#x27;re deliberately trying to make them fail. I honestly have no bigger point than that, I just think this isn&#x27;t a very good thing by which to evaluate LLM capabilities. If there&#x27;s no argument you&#x27;ll accept, I am happy to move on.
                          1. foldr · · focus · HN ↗
                            You don’t need to teach a human anything except the rules of chess and the details of a particular chess notation. No special skill or training is required to make a sketch of a chess board. Surely there is no chess player who, if confronted with a sequence of chess moves in algebraic notation, would not think to construct a representation of the chess board in order to understand what was going on.

                            &gt; I just think this isn&#x27;t a very good thing by which to evaluate LLM capabilities

                            I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)

                            &gt; If there&#x27;s no argument you&#x27;ll accept

                            It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your comments so far. I could equally well say the same thing to you!

                            1. thom · · focus · HN ↗
                              I&#x27;m just going to keep repeating: it is utterly trivial to get an LLM to play chess without making illegal moves. Easier than teaching a human. Sorry this doesn&#x27;t happen out of the box, but it shouldn&#x27;t budge your priors about LLM intelligence one bit.
                2. zahlman · · focus · HN ↗
                  More importantly, beginner human players don&#x27;t exhibit that tendency. The history of the position doesn&#x27;t bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.
                  1. thom · · focus · HN ↗
                    Humans do make these errors when playing blindfolded. If you even the playing field and give the LLM the position at each turn, it does not make mistakes.
                    1. zahlman · · focus · HN ↗
                      &gt; If you even the playing field and give the LLM the position at each turn, it does not make mistakes.

                      It absolutely still makes mistakes if you ask it to draw the board each turn, which should be equivalent to giving it the position because it only has to update one move at a time and then it has the position in the context window.

                      1. thom · · focus · HN ↗
                        Yes, we can come up with all sorts of weird situations where you can get it to be confused. But what I&#x27;m saying is it&#x27;s _trivial_ to give it a simple prompt that prevents it from ever making any errors, and so I don&#x27;t think it&#x27;s this big LLM gotcha (of which there are many!)
              2. Capricorn2481 · · focus · HN ↗
                The actual question is backwards: how do we keep the prompt and context small enough so the LLM doesn&#x27;t start hallucinating basic rules of chess.
    7. WhitneyLand · · focus · HN ↗
      1. It’s hard to trust a 2026 paper that’s showing results for such old models.

      2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.

      3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

      1. manquer · · focus · HN ↗
        &gt; People who are good at it rely more on experience and deep domain expertise

        People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.

        A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.

        1. svachalek · · focus · HN ↗
          1100 at online speed chess or something, could be. I&#x27;m not that deep in the chess world but everyone I know that can make 1100 in official rating can name a dozen openings and most of the known tactics, and is pretty good at applying at least one opening.
          1. tovej · · focus · HN ↗
            1100 is literally below the ELO you get by default as a beginner.
          2. orwin · · focus · HN ↗
            1100 lichess&#x2F;chess.com does not represent real elo. I&#x27;m around 1400 online, I would still be unranked in the real world. The fact that I easily beat any model publicly available is not a great look for AGI.
      2. what · · focus · HN ↗
        &gt; Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

        Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.

        1. senordevnyc · · focus · HN ↗
          But does that matter?

          If the goal for buyers of AI is “replace this knowledge worker”, how much does it matter that the model in a simple loop can’t do it, but the model with a strong general purpose harness and a little time to gather resources and knowledge to augment the harness going forward, plus tool calls, plus custom built tools, etc, can replace the knowledge worker?

          Probably the only thing saving many jobs from being replaced right now is that it’s hard to have a verification of correctness in the loop, so the agent can’t hill climb very easily.

          1. numitus · · focus · HN ↗
            Tests are often conducted under restricted conditions. For example, elementary school students aren&#x27;t given calculators in math class, or during an interview, you are asked what encapsulation is and aren&#x27;t allowed to use Google. The chess test effectively demonstrates the reasoning capabilities of an LLM without relying on brute force, because a human is incapable of calculating trillions of combinations yet plays chess successfully. This test is necessary because many complex problems cannot be solved by brute force, such as managing a business or playing Heroes 3. Therefore, we can make the assumption that if an LLM can play chess at a grandmaster level without brute force, it means it will be able to command an army or manage production.
            1. senordevnyc · · focus · HN ↗
              People might care about this for chess, but no one really cares if an LLM can command an army or manage production of a business without any tools. If it can do those tasks reliably when given access to tools (including any tools it autonomously creates for itself), then that&#x27;s more than sufficient. No one cares if an LLM is doing reasoning the way humans do it, as long as it can get the job done.
              1. numitus · · focus · HN ↗
                The assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more &#x27;pieces,&#x27; incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn&#x27;t matter whether it has tools or not. It&#x27;s simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.
      3. paimapi · · focus · HN ↗
        so prove it! get a public repo out there, have it play against some open source engines

        also I think the operative letter in AGI is the G - and if the G is short for &#x27;variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill&#x27; then its not really G at all, is it?

        1. BobbyJo · · focus · HN ↗
          I suck at chess. Are you saying I can&#x27;t be intelligent?
          1. paimapi · · focus · HN ↗
            is that what I&#x27;m saying? or am I talking about AGI? perhaps there&#x27;s some irony here to be explored when it comes to basic reading comprehension gaps
            1. jibal · · focus · HN ↗
              That&#x27;s a polite way to put it. :-)
            2. BobbyJo · · focus · HN ↗
              My point was you are misunderstanding G, or at least applying it erroneously here. Being good at chess is not a generalization of any other body of knowledge, it is a rigorous set of rules. The only way to be good at chess is to practice chess, or to apply deep calculations. The latter is the model writing code.

              The illegal move aspect has more to do with a failure of online&#x2F;in-context learning, which would support your point. I tend to think it is a byproduct of reasoning in language, which newer architectures would fix, but we shall see.

              1. paimapi · · focus · HN ↗
                chess is not a rigorous set of rules, it is rules as foundation. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun. knowledge for chess, specifically, is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. it is not at all different from any other body of knowledge - it only &#x27;feels&#x27; different to us humans because it is so logic-based and it takes a long time before our inferential, pattern-recognition kicks in and starts seeing the board in a naturally, systemic way. for an AGI, that should be a cakewalk, trained as it were to surpass human capability in any and every domain (thus the G for &#x27;general&#x27; and not &#x27;H&#x27; for &#x27;hyperspecific&#x27;)
                1. BobbyJo · · focus · HN ↗
                  &gt; an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain.

                  AGI != ASI. You are confusing the two.

                  1. paimapi · · focus · HN ↗
                    I&#x27;m not. AGI is almost necessarily closer to ASI than it is to human intelligence. you take the concept of domain knowledge transferring to other areas. presumably, an &#x27;AGI&#x27; that is generally as good as a really good human at every task under-the-sun will already be much better than most humans because it can incorporate cross-domain knowledge and apply it in a reasonable fashion. it&#x27;s like the parable of Newton and the apple - the domain knowledge that an apple falls according to certain rules observed before igniting the creative spark that led to universal gravitation
                    1. BobbyJo · · focus · HN ↗
                      &gt; presumably, an &#x27;AGI&#x27; that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion.

                      I disagree with this definition of AGI, and I disagree that chess skills significantly benefit from generalizing non-chess knowledge, outside of computing moves probabilistically.

                      AGI has historically been defined as human level or better, with generality to new domains. I think blurring it with ASI makes the terminology confusing to use.

                      Chess is learned rules and the ability to apply those rules. Strategy as a whole is applying a set of rules to circumstances, that&#x27;s how it is taught: &quot;here are examples of circumstances and actions, try to pattern match to future circumstance and apply commensurate action.&quot;

                      If you make the point that chess is a large part of the training data, or that LLMs are unable to learn chess well, I&#x27;ll accept that as refuting that LLMs are AGI, but these other points I disagree with.

          2. nmehner · · focus · HN ↗
            If you read all chess tutorials, strategy documentation and game archives on the internet and then would still suck at chess: yes.
            1. brindleth · · focus · HN ↗
              Declarative knowledge is not the same as procedural knowledge. You can read as many chess tutorials, strategy documentation and game archives as you like, they won&#x27;t make you good at chess until you actually start practicing chess.
              1. _superposition_ · · focus · HN ↗
                Does practical improvement apply to only humans or intelligence in general?
          3. foldr · · focus · HN ↗
            The issue with the models isn&#x27;t that they play a bad game, but that they persist in making illegal moves. An average intelligent human can be told the rules of chess and then play chess, badly, within the rules.
            1. empath75 · · focus · HN ↗
              An average human would have a physical chess board in front of them to remind them of the current state.
              1. foldr · · focus · HN ↗
                Sure, but the LLM is free to construct a representation of the chess board and update it as it goes along. It is not in any way banned from using a virtual board, or whatever representation of game state it pleases.
      4. carodgers · · focus · HN ↗
        &gt; Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

        A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would &quot;destroy any human at chess.&quot;

        Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?

        1. nimbleal · · focus · HN ↗
          Maybe practically it doesn’t matter? Perhaps AGI is not the model but the model plus everything it’s got access to. If we’re modelling intelligence in the way we seem to have to to have any coherent definition of AGI, it seems to me &lt;model + everything it can access&gt; is always going to be more “intelligent” than &lt;model&gt; alone.
          1. Planktonne · · focus · HN ↗
            That would mean we should consider any human with coding knowledge a chess grandmaster, which is obviously not the case.
            1. nimbleal · · focus · HN ↗
              My points is more that, while we have a strong intuition about where, as an entity, a human&#x27;s boundaries are (i.e. where the person begins and ends), philosophically it&#x27; not immediately obvious that the analogy applies to the a model in the same way. Why should that be the line drawn that says this is the &quot;thing&quot; and this other stuff is external to the thing? It feels somewhat arbitrary.

              Of course this is a difficult question with humans too, hence my reliance on intuition above. We don&#x27;t have the same cultural&#x2F;biological framework to fall back on with AI.

              1. thinkharderdev · · focus · HN ↗
                I think all this debate about whether an LLM can write (or download) a chess engine is sort of missing the point. For basically any economically valuable work there is no equivalent of a chess engine for it. If there were we wouldn&#x27;t need humans or AI to begin with.
            2. empath75 · · focus · HN ↗
              If the goal is merely to &quot;win at chess&quot;, then yes, an LLM using stockfish is better than any human alone at performing the task. When you are talking about what AI agents are capable of doing, there is no such thing as &quot;cheating&quot;. They are as capable as the tools they can use effectively. The entire history of human civilization was driven by effectively using tools to achieve goals.
              1. Planktonne · · focus · HN ↗
                But then the human should also get Stockfish, and we&#x27;re back at a stalemate.
      5. Certhas · · focus · HN ↗
        Good science, properly digested and presented takes time.

        The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues

        1. WhitneyLand · · focus · HN ↗
          Not sure how that vague truism applies to this paper.

          Lots of papers have great results that don’t depend on the latest models.

          However in this case it’s problematic:

          - They specifically make claims about the state of “current LLMs”. o3 is not representative of this.

          - They ask are LLMs capable of X and arrive at a negative result.

          If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.

          However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.

    8. matteoraso · · focus · HN ↗
      I don&#x27;t see why this is such a big deal. Nobody&#x27;s using LLMs for chess, but even if they are, just give them Stockfish as part of their harness. They don&#x27;t need to do everything themselves as long as they&#x27;re intelligent enough to use tools.
    9. gizmodo59 · · focus · HN ↗
      why cant models make a tool call to stockfish? its like saying model can&#x27;t execute python for complex math calculations
      1. redcheeks · · focus · HN ↗
        Exactly. All these nerds saying cars make bad submarines. Well duh.
      2. what · · focus · HN ↗
        Because then it’s not playing chess, stockfish is?
    10. 1dom · · focus · HN ↗
      The last post on HN I read was about someone using LLMs to reverse engineer an Apple GPU driver for linux in a month. The top comment points out how the poster must have had specialist internal domain specific contact with Apple. But then the thread concludes that wasn&#x27;t the case and that this would take domain experts years to do.

      &gt; &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;

      I feel this statement is extreme. I can&#x27;t personally reconcile it with any of the projects we&#x27;re regularly seeing get delivered largely by LLMs now.

      What are you thoughts? Like, what&#x27;s your position here? Even if you sincerely believe frontier models need laborious oversight on even the simplest of tasks, do you think that accurately captures and reflects the current state and progress of frontier LLMs?

      Don&#x27;t get me wrong, there&#x27;s lots of things LLMs can&#x27;t do well, but the idea that they&#x27;re basically not helpful for even the simplest of tasks seems... disingenuous?

    11. stinkbeetle · · focus · HN ↗
      &gt; Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

      It can be very interesting and even entertaining to know where models don&#x27;t do well. I don&#x27;t find something like chess to be very instructive about anything though, nobody is going to pay for AI to play chess at any significant scale even if it could do it perfectly.

      &gt; The author of the originating post says that &quot;current frontier models need laborious oversight and guardrails on even the simplest tasks&quot;, and he&#x27;s absolutely correct.

      I have found that not to be the case, and I am not an AI power user or cheerleader by any means. There&#x27;s probably a bunch of even &quot;simplest&quot; tasks where AI doesn&#x27;t do well and might never. That doesn&#x27;t take away from the cases where it works well and is a productivity booster. It doesn&#x27;t even have to be solving millennium prizes or any other breakthrough creativity or reasearch, it&#x27;s still very useful in places.

      1. contubernio · · focus · HN ↗
        They should just make them read an old chess book like Lasker that gives low level heuristics
    12. tossandthrow · · focus · HN ↗
      Llm systems are not really build for adhering to a grammar (other than &quot;a string og tokens&quot;).

      It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents.

      Certainly,a harness can easily correct for it.

      1. vkazanov · · focus · HN ↗
        By the promise of it, llms should be able to both adhere to grammars, or go free form where necessary. I mean, doing math is supposed to be strict but in practice it&#x27;s a somewhat educated random walk in the space of correct lean theorems.

        Harnesses do correct things, sure.

        1. tossandthrow · · focus · HN ↗
          You are right. I am imprecise.

          Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.

          Games and programming languages (including lean) does not allow this flexibility.

          A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.

          Certainly it must be like that, otherwise reviews in math was rendered moot.

          Do we blame research mathematicians for not adhering to the grammar?

      2. wodenokoto · · focus · HN ↗
        It seems absolutely crazy to me to expect an LLM to code a solution to a problem while also not expecting it to be able to adhere to a grammar.
        1. tossandthrow · · focus · HN ↗
          Why?

          You might never have tried to program before, so I don&#x27;t blame it on you.

          But most programmers, even experienced ones, see grammar and type errors regularly.

        2. Gregkion · · focus · HN ↗
          How much support do we as humans need to get rules right?

          I&#x27;m an expert in my field, read my comments, my gramma is shit.

      3. _superposition_ · · focus · HN ↗
        Someone build a chess harness already...
    13. Auracle · · focus · HN ↗
      The fact that they can play chess at all despite having no specific training for it blows my mind, and the fact it doesn’t do the same for many others shows just how far they’ve come and how fast.
      1. bigstrat2003 · · focus · HN ↗
        It doesn&#x27;t blow anyone&#x27;s mind because it hasn&#x27;t been impressive for a computer to play chess for 40 years. &quot;We made something worse than existing solutions by using a new technique&quot; is not an impressive feat.
        1. Auracle · · focus · HN ↗
          It&#x27;s literally generalized intelligence. Yeah, I had a shitty chessboard back in the 90&#x27;s that was specifically programmed to play chess. That same board couldn&#x27;t tell me the best medication to use to treat a certain condition or how best to modify a model in Blender.
    14. kbau · · focus · HN ↗
      I suspect (in a probably ignorant fashion) that this is because learning process has been reading a lot of algebraic chess notation (such as &quot;1. e4 e5 2. Nf3 f6 3. Nxf6 gxf6 4. Qh5! +-&quot;) then, to play, generating more of it without considering the rules of the game. This is exactly how it&#x27;s always felt to me when playing chess against LLMs. Sure, &quot;1. e4 e5 2. Nf3 Nc3&quot; looks innocent to somebody simply learning the syntax of algebraic notation, but that Nc3 by black is an illegal move.

      An LLM is the wrong approach for playing chess.

      1. JohnKemeny · · focus · HN ↗
        Are you saying that modern LLMs cannot play chess now, or that LLMs (GPT architecture) cannot be trained to play chess well?

        Or are you saying that neural networks in general cannot (practically) be trained to be an above-average chess player?

        Or are you saying that it depends on the input? Would it be better if they were given a picture&#x2F;drawing&#x2F;ascii art of the board? If so, surely they can produce it at will?

        1. kbau · · focus · HN ↗
          Neural Nets can be trained to play chess very well and have been doing so for a long time (see Stockfish and Leela as some of the most popular&#x2F;strongest ones - top GMs have no chance against them), but these are dedicated models, where the game rules are encoded in the learning process, as opposed to large language models which are natural language processing models. Technically you can give an LLM a lot of chess books and games and it will be able to spit out chess notation. Put a webapp on top that renders text moves to the board and it looks like it&#x27;s playing chess. But it isn&#x27;t really.
          1. bitexploder · · focus · HN ↗
            <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Chinese_room" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Chinese_room I think about this once in a while. At some point if it does the thing almost perfectly is it still not doing the thing?
            1. kbau · · focus · HN ↗
              I suppose if you all you need is a good enough opponent for the average person out there, sure this is good enough.

              I was more talking in reference to why the LLMs in the above linked paper were producing so many illegal moves, and it is because they are not hard constrained by the rules of the game. Of course, a loop can prompt until a valid move is produced and then rendered on a screen. But why do this? I suppose, who am I to say what should be done or not, but a specialized tool being better than a general one at its specific job isn&#x27;t particularly surprising.

    15. ricky54 · · focus · HN ↗
      If you give the same task to an exceptionally intelligent human, who does not play chess and has only heard about it in passing, then they would be beaten by every child who has looked at the rules for more than 10 minutes.

      What kind of intelligence is &quot;playing &lt;____&gt; but we don&#x27;t tell you the rules&quot; supposed to test?

      1. bulder · · focus · HN ↗
        ...but that&#x27;s not what the models are. You can interrogate them on the rules of chess, and they&#x27;ll (statistically likely) give you a decent breakdown of the rules. Evidently the rules are in their training material, they just fail to apply them in the manner of an intelligent system for some reason or another.
    16. lynx97 · · focus · HN ↗
      Well, yes, PGN files have structure... But still, playing Chess with an LLM is so weird that I impulsively question the sanity of people attempting to do so. Do some people really believe training on TWIC PGNs would make an LLM a good chess player?
    17. killerstorm · · focus · HN ↗
      This is an absolute nonsense. Any frontier model can implement chess program from scratch - modeling the board, checking legality, etc. If you asked e.g. GPT-6 to get good at chess and gave it a computer, it will get good at chess. That&#x27;s an actual strategic skill.

      Asking GPT to play chess directly using its reasoning only tests its reasoning ability to model chess state. Which it really is NOT optimized for.

      This is also true for humans - people who don&#x27;t have years of chess training can&#x27;t really tell which moves are legal given an algebraic notation transcript. These people might have good strategic skills in different areas. Chess is just a very, very specific skill

      1. roenxi · · focus · HN ↗
        It&#x27;s an interesting puzzle, isn&#x27;t it. On the one hand, the AIs are no good at playing Chess.

        However, on the other hand, if you ask an AI to win a game of chess it has all the tools on hand to compete at the same level as Stockfish - it can re-implement an engine and even probably has a GPU on hand to train its own neural nets.

        So should we say that the AI can play chess well, or that it cannot?

        1. tired-turtle · · focus · HN ↗
          Is it, though? If you design and build a winning F1 race car, did you also win the race?

          Recent discourse around AI seems to conflate the semantics of winning: 1. you contributed to the win vs 2. you yourself were the winning driver.

          1. roenxi · · focus · HN ↗
            But the human would have to get extra hardware to do that. The AI isn&#x27;t bringing in any resources it doesn&#x27;t already have access to.
        2. kbau · · focus · HN ↗
          I can compile stockfish from source and use it to beat other kids in my class in chess. Behold, I am a chess genius.
    18. glitchc · · focus · HN ↗
      Why not ask it to implement a chess engine first, and then use that to play against you?

      Does the LLM need to learn to play chess if it can build a chess engine to play for it instead?

      1. lionkor · · focus · HN ↗
        With that approach, the benchmark falls apart. Of course it can write a chess engine, because it learned on lots of stolen source code of chess engines. This has nothing to do with the LLM&#x27;s ability to reason.

        Writing a well understood engine for a super popular problem does not count as reasoning about the problem.

        1. glitchc · · focus · HN ↗
          &gt; Writing a well understood engine for a super popular problem does not count as reasoning about the problem.

          Doesn&#x27;t writing the engine imply understanding about the problem domain? Tool use is a widely accepted measure of intelligence.

          1. lionkor · · focus · HN ↗
            In most humans, yes, because we are terrible at memorizing millions of codebases. For LLMs, we need to apply our understanding of them before making statements like that. An LLM can &quot;memorize&quot;, and has &quot;memorized&quot;&#x2F;been trained on tens of thousands of chess engines. Writing a chess engine, or even deriving a chess engine from the rules alone, does not constitute a deep understanding of, and more importantly, the ability to apply, the rules, at all.

            When humans do this, they inadvertently learn something, too, but when an LLM reproduces or derives and implementation of a chess engine, it in no way implies that the LLM can follow the rules in its own &quot;train of thought&quot; and consistently apply the rules in its &quot;head&quot;.

            Let&#x27;s say you want to evaluate my algebra skills. You make me solve some algebra challenges. If I then whip out a computer and write a calculator, or take some sticks and stones and take a couple hours to build an abacus, and then solve the algebraic challenges, this would not constitute a good solution, and would defeat the entire point of the test. If, instead, I do the algebra in my head or on paper, it might seem like there&#x27;s no difference, but you can derive all sorts of information from that.

            For example, you could time it, check for recurring errors I make, for interesting mistakes like mistaking 7 and 1 for one another due to bad hand-writing, etc.

            If that was the goal, then me writing a calculator or crafting an abacus defeats the point of the test. Yes, me writing a calculator shows that I&#x27;m intelligent, and I understand the algebraic rules, but if the test is about applying the rules, I have not passed.

            In the very same way, an LLM writing a chess engine to solve a chess benchmark that is all about LLM&#x27;s reasoning capability is complete bogus and defeats the entire point.

          2. well_ackshually · · focus · HN ↗
            &gt;Tool use is a widely accepted measure of intelligence.

            Stop anthropomorphizing the parrot. The parrot has had stockfish&#x27;s source code blasted at high pressure into its head along with dozens of millions of other pieces of code whose sole role is to have efficient algorithms to more or less brute force through the best result. Brute forcing (no matter how smart it is) isn&#x27;t understanding the problem domain.

            1. glitchc · · focus · HN ↗
              I&#x27;m not fully convinced that human beings aren&#x27;t stochastic parrots. Much of the behaviour I observe in daily life reflects a blind adherence to set of beliefs that are an amalgation of &quot;a person of authority said to do this&quot; at an early age.
          3. Capricorn2481 · · focus · HN ↗
            Definitely not. A human can write a chess engine without being very good at chess.
    19. dzonga · · focus · HN ↗
      people on the ground know that small models are enough, since llms are good at directed work (i.e handholding) not the let loose go wild that the labs try to hype on.

      the only thing that few people are willing to admit is that humans are the bottleneck as humans are needed to handhold &#x2F; verify output - which puts a dent or might I say pause on the excessive valuations of a.i companies as that&#x27;s against the narrative.

    20. MattCruikshank · · focus · HN ↗
      What happens when you ask those same frontier models to write a chess-playing program?

      I feel like, this is a huge stumbling block that many people have. They&#x27;ll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn&#x27;t fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.

      It&#x27;s really neat to see what a frontier model can do itself. No doubt.

      But &quot;play chess by hand&quot; is a frankly awful metric. It&#x27;s kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.

      1. cbolton · · focus · HN ↗
        It&#x27;s a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it&#x27;s noteworthy that they underperform on that test.

        Letting the model execute a chess program (that it wrote) would make sense if you&#x27;re measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.

        1. MattCruikshank · · focus · HN ↗
          &gt; The fact that the human would have a much harder time writing a useful program is irrelevant.

          Why?

          There&#x27;s a box.

          You give it a problem, and it comes up with a solution.

          Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes?

          Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess &quot;in its head&quot;, or if it has to use scratch paper?

          1. lionkor · · focus · HN ↗
            The question is to what end? This is a benchmark task, because playing chess, or solving other well-understood problems is more of a party trick than it is useful.

            If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.

            1. MattCruikshank · · focus · HN ↗
              Do you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back?

              More to my point, I think it&#x27;s stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs.

              I&#x27;m advising people that they should think about this distinction, themselves, when they have data and want answers.

              1. cbolton · · focus · HN ↗
                Neither. As I said I want to measure cognitive abilities.

                Your &quot;ability of the box&quot; is like &quot;economic potential&quot; in my previous comment. If that&#x27;s what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.

                1. MattCruikshank · · focus · HN ↗
                  I agree that it&#x27;s a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had.

                  That said, it&#x27;s really weird to me when people use (and judge) LLMs one way... and won&#x27;t try using them another way.

                  Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.

                  Otherwise, it&#x27;s like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I&#x27;ve ever seen: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=gDy9AUQJ3Fg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=gDy9AUQJ3Fg

                  This lamp, without a working power outlet? It really doesn&#x27;t do anything...

                  1. cbolton · · focus · HN ↗
                    I completely agree.
          2. [deleted] · · focus · HN ↗

            [deleted]

        2. empath75 · · focus · HN ↗
          &gt; It&#x27;s a great test of cognitive abilities.

          It isn&#x27;t. Stockfish running on your laptop can beat every human being on earth easily at chess. It&#x27;s not intelligent _at all_ in any sense that matters.

          1. cbolton · · focus · HN ↗
            Well it&#x27;s not a perfect test so you need a bit of care in how you use it. If you have no idea what the subject is doing, then you don&#x27;t know if you&#x27;re measuring cognitive ability or something else (like cheating ability, or algorithmic sophistication or whatever). But failing the test is a pretty clear sign of certain cognitive abilities being poor.
        3. anthonyrstevens · · focus · HN ↗
          &gt;&gt; There are many claims that current LLMs surpass humans in cognitive abilities

          Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic &#x2F; untethered comments as the strawman against which you feel the need to argue?

          1. cbolton · · focus · HN ↗
            Why the aggressive tone and the strawman rhetoric? I never said there was a consensus. Yes I&#x27;m talking more about the &quot;optimistic&quot; commenters and pointing out that this chess thing is a good datum to temper their enthusiasm. What&#x27;s wrong with that?

            Also these claims are not completely without merit, it&#x27;s just that LLMs seem to excel at specific &quot;cognitive&quot; tasks and it&#x27;s interesting to see where they fail.

      2. deaton · · focus · HN ↗
        What would happen if you asked a human developer to write a chess-playing program?
      3. Kotlopou · · focus · HN ↗
        To go a bit off-track based on your final sentence: my high school physics teacher would do any numerical calculation that came up first in his head, as an estimate, and only then use a calculator or write on the board. Usually the estimate was within ±10-20% of the correct value even for long combinations of numbers with a bunch of decimals. Cube roots didn&#x27;t come up, but square roots did.

        The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.

      4. jayd16 · · focus · HN ↗
        Of course I&#x27;m a super fast runner. I can get in my car and go like 100 mph.
      5. well_ackshually · · focus · HN ↗
        &gt;What happens when you ask those same frontier models to write a chess-playing program?

        they shit out a carbon copy of <a href="https:&#x2F;&#x2F;github.com&#x2F;official-stockfish&#x2F;stockfish" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;official-stockfish&#x2F;stockfish that they have in their training data. Still doesn&#x27;t make Fable good at playing chess.

        1. MattCruikshank · · focus · HN ↗
          I feel like you&#x27;re saying something as odd as &quot;Transistors still aren&#x27;t good at playing chess.&quot;

          I&#x27;m pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.

          1. well_ackshually · · focus · HN ↗
            I&#x27;m not the one posting daily about how &quot;AIs are going to destroy the world because of how smart they are&quot;, &quot;humans are finished&quot; and &quot;we&#x27;ve reached super duper mega intelligence&quot;. Go see Dario and Sam about that.

            &gt;I&#x27;m pretty sure Fable could write AlphaZero

            If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster_2000_xx, it doesn&#x27;t matter if it does: it&#x27;s writing a solver: it&#x27;s not good at playing chess. If tomorrow I tell you that I&#x27;m so fucking good at chess I can beat Magnus, and I show up with a laptop running stockfish, you&#x27;re going to laugh me out of the room.

            1. MattCruikshank · · focus · HN ↗
              So if you need a scratchpad to solve a problem, does that mean you cannot solve that problem?
      6. perrygeo · · focus · HN ↗
        Asking a model to &quot;write a throwaway program to do X&quot; is vastly more productive and reliable than asking &quot;do X&quot;. Running code provides a feedback loop, the model can iteratively improve the solution instead of guess. Even if you don&#x27;t read the code yourself, you have a reproducible, editable, and auditable artifact if you need it.
      7. numitus · · focus · HN ↗
        I am bad in chess game by itself like 1200 ELO, but I can write Programm and win player with 2600 ELO. Does it mean I am pro chess gamer?
        1. MattCruikshank · · focus · HN ↗
          Do I care if Richard Feynman was only able to do nuclear physics with the help of an abacus?

          Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There&#x27;s joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm.

          Sometimes you care about the Bee, sometimes you care about the results.

          1. numitus · · focus · HN ↗
            This is called a benchmark. We run a calculation of Pi to evaluate a computer&#x27;s performance, but we don&#x27;t allow the script to download a ready-made solution. When we evaluate a runner, we don&#x27;t let them use a bicycle. When we evaluate a new LLM, we don&#x27;t allow it to send a request to a team of programmers, so using a chess engine for an LLM is cheating
            1. MattCruikshank · · focus · HN ↗
              If an LLM writes AlphaZero, and it competes with itself, and is dominant (and beats stockfish!!!), with no book positions cribbed from its learning...

              The LLM has a process to beat chess.

              Just like, if it doesn&#x27;t inherently know how to multiply 13 * 17 without using Python to do it... I don&#x27;t really care.

              Maybe you do care. Maybe you want an LLM to be able to do work, only in its head.

              But I kind of can&#x27;t understand the desire for that limitation...

              I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can&#x27;t go fast enough. The Drive gear is literally right there.

              1. numitus · · focus · HN ↗
                The assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more &#x27;pieces,&#x27; incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn&#x27;t matter whether it has tools or not. It&#x27;s simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.
    21. tzone · · focus · HN ↗
      It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine. It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).

      On that note, I actually had an overall harness (for experimenting) that was essentially like this: &quot;for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer&quot;.

      It actually worked incredibly well on all &quot;gotcha&quot; LLM questions like math or counting letters in words and all sorts of stuff.

      Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend an infinite amount of money.

      1. causal · · focus · HN ↗
        Yeah, thread full of cope. &quot;Well if you remove the human&#x27;s legs it&#x27;s actually quite bad at marathons&quot; arguments.
    22. empath75 · · focus · HN ↗
      I want you to consider how relevant this is in any practical sense.

      First -- most _people_ cannot do this, without having a physical board in front of them.

      Second -- Claude Code is perfectly capable of downloading and running stockfish. People focus too much on LLMs by themselves as the entity of concern instead of the entire harness and all of it&#x27;s capabilities together.

      1. Capricorn2481 · · focus · HN ↗
        Because they are obviously testing for general intelligence. If you want a thread about how cool the harness is, that&#x27;s down the street.

        We don&#x27;t really consider humans downloading stockfish to beat people at chess as noteworthy endeavors.

    23. thelaxiankey · · focus · HN ↗
      These &quot;researchers&quot; are less informed on LLM chess than random internet bloggers. The situation is much more interesting

      <a href="https:&#x2F;&#x2F;dynomight.net&#x2F;chess&#x2F;" rel="nofollow">https:&#x2F;&#x2F;dynomight.net&#x2F;chess&#x2F;

      1. thelaxiankey · · focus · HN ↗
        Forgot to add the amazing follow up

        <a href="https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;" rel="nofollow">https:&#x2F;&#x2F;dynomight.net&#x2F;more-chess&#x2F;

        1. topaz0 · · focus · HN ↗
          From your linked post: &quot;LLMs sometimes struggle to give legal moves. In these experiments, I try 10 times and if there’s still no legal move, I just pick one at random.&quot;

          Which sounds a lot like what that paper was about

          1. thelaxiankey · · focus · HN ↗
            Yeah but it didn&#x27;t do squat. The effective thing was a subtle prompt mod that boosted performance tremendously. I&#x27;ve been surprised by this forums unwillingness to read this thing...
    24. m3at · · focus · HN ↗
      Yes &quot;oversight and guardrails&quot; are still needed, but even that is becoming easier to build, and imo already no longer in the &quot;laborious&quot; category. Even far from the frontier, you can tune a 0.2B LLM into a decent 2000 Elo player as a weekend project:

      <a href="https:&#x2F;&#x2F;x.com&#x2F;maximelabonne&#x2F;status&#x2F;2100137121264828901" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;maximelabonne&#x2F;status&#x2F;2100137121264828901

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.