Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
What happens when you ask those same frontier models to write a chess-playing program?
I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn't fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.
It's really neat to see what a frontier model can do itself. No doubt.
But "play chess by hand" is a frankly awful metric. It's kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.
It's a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it's noteworthy that they underperform on that test.
Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.
> The fact that the human would have a much harder time writing a useful program is irrelevant.
Why?
There's a box.
You give it a problem, and it comes up with a solution.
Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes?
Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess "in its head", or if it has to use scratch paper?
The question is to what end? This is a benchmark task, because playing chess, or solving other well-understood problems is more of a party trick than it is useful.
If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
Neither. As I said I want to measure cognitive abilities.
Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.
I agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had.
That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way.
Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.
Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: <a href="https://www.youtube.com/watch?v=gDy9AUQJ3Fg" rel="nofollow">https://www.youtube.com/watch?v=gDy9AUQJ3Fg
This lamp, without a working power outlet? It really doesn't do anything...
> It's a great test of cognitive abilities.
It isn't. Stockfish running on your laptop can beat every human being on earth easily at chess. It's not intelligent _at all_ in any sense that matters.
Well it's not a perfect test so you need a bit of care in how you use it. If you have no idea what the subject is doing, then you don't know if you're measuring cognitive ability or something else (like cheating ability, or algorithmic sophistication or whatever). But failing the test is a pretty clear sign of certain cognitive abilities being poor.
>> There are many claims that current LLMs surpass humans in cognitive abilities
Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic / untethered comments as the strawman against which you feel the need to argue?
Why the aggressive tone and the strawman rhetoric? I never said there was a consensus. Yes I'm talking more about the "optimistic" commenters and pointing out that this chess thing is a good datum to temper their enthusiasm. What's wrong with that?
Also these claims are not completely without merit, it's just that LLMs seem to excel at specific "cognitive" tasks and it's interesting to see where they fail.
To go a bit off-track based on your final sentence: my high school physics teacher would do any numerical calculation that came up first in his head, as an estimate, and only then use a calculator or write on the board. Usually the estimate was within ±10-20% of the correct value even for long combinations of numbers with a bunch of decimals. Cube roots didn't come up, but square roots did.
The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.
>What happens when you ask those same frontier models to write a chess-playing program?
they shit out a carbon copy of <a href="https://github.com/official-stockfish/stockfish" rel="nofollow">https://github.com/official-stockfish/stockfish that they have in their training data. Still doesn't make Fable good at playing chess.
I'm not the one posting daily about how "AIs are going to destroy the world because of how smart they are", "humans are finished" and "we've reached super duper mega intelligence". Go see Dario and Sam about that.
>I'm pretty sure Fable could write AlphaZero
If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster_2000_xx, it doesn't matter if it does: it's writing a solver: it's not good at playing chess. If tomorrow I tell you that I'm so fucking good at chess I can beat Magnus, and I show up with a laptop running stockfish, you're going to laugh me out of the room.
Asking a model to "write a throwaway program to do X" is vastly more productive and reliable than asking "do X". Running code provides a feedback loop, the model can iteratively improve the solution instead of guess. Even if you don't read the code yourself, you have a reproducible, editable, and auditable artifact if you need it.
Do I care if Richard Feynman was only able to do nuclear physics with the help of an abacus?
Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm.
Sometimes you care about the Bee, sometimes you care about the results.
This is called a benchmark. We run a calculation of Pi to evaluate a computer's performance, but we don't allow the script to download a ready-made solution. When we evaluate a runner, we don't let them use a bicycle. When we evaluate a new LLM, we don't allow it to send a request to a team of programmers, so using a chess engine for an LLM is cheating
If an LLM writes AlphaZero, and it competes with itself, and is dominant (and beats stockfish!!!), with no book positions cribbed from its learning...
The LLM has a process to beat chess.
Just like, if it doesn't inherently know how to multiply 13 * 17 without using Python to do it... I don't really care.
Maybe you do care. Maybe you want an LLM to be able to do work, only in its head.
But I kind of can't understand the desire for that limitation...
I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can't go fast enough. The Drive gear is literally right there.
The assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more 'pieces,' incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn't matter whether it has tools or not. It's simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
MattCruikshank · · focus · HN ↗
I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn't fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.
It's really neat to see what a frontier model can do itself. No doubt.
But "play chess by hand" is a frankly awful metric. It's kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.
cbolton · · focus · HN ↗
Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.
MattCruikshank · · focus · HN ↗
Why?
There's a box.
You give it a problem, and it comes up with a solution.
Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes?
Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess "in its head", or if it has to use scratch paper?
lionkor · · focus · HN ↗
If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
MattCruikshank · · focus · HN ↗
More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs.
I'm advising people that they should think about this distinction, themselves, when they have data and want answers.
cbolton · · focus · HN ↗
Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.
MattCruikshank · · focus · HN ↗
That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way.
Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.
Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: <a href="https://www.youtube.com/watch?v=gDy9AUQJ3Fg" rel="nofollow">https://www.youtube.com/watch?v=gDy9AUQJ3Fg
This lamp, without a working power outlet? It really doesn't do anything...
cbolton · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
empath75 · · focus · HN ↗
It isn't. Stockfish running on your laptop can beat every human being on earth easily at chess. It's not intelligent _at all_ in any sense that matters.
cbolton · · focus · HN ↗
anthonyrstevens · · focus · HN ↗
Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic / untethered comments as the strawman against which you feel the need to argue?
cbolton · · focus · HN ↗
Also these claims are not completely without merit, it's just that LLMs seem to excel at specific "cognitive" tasks and it's interesting to see where they fail.
deaton · · focus · HN ↗
Kotlopou · · focus · HN ↗
The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.
jayd16 · · focus · HN ↗
well_ackshually · · focus · HN ↗
they shit out a carbon copy of <a href="https://github.com/official-stockfish/stockfish" rel="nofollow">https://github.com/official-stockfish/stockfish that they have in their training data. Still doesn't make Fable good at playing chess.
MattCruikshank · · focus · HN ↗
I'm pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.
well_ackshually · · focus · HN ↗
>I'm pretty sure Fable could write AlphaZero
If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster_2000_xx, it doesn't matter if it does: it's writing a solver: it's not good at playing chess. If tomorrow I tell you that I'm so fucking good at chess I can beat Magnus, and I show up with a laptop running stockfish, you're going to laugh me out of the room.
MattCruikshank · · focus · HN ↗
perrygeo · · focus · HN ↗
numitus · · focus · HN ↗
MattCruikshank · · focus · HN ↗
Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm.
Sometimes you care about the Bee, sometimes you care about the results.
numitus · · focus · HN ↗
MattCruikshank · · focus · HN ↗
The LLM has a process to beat chess.
Just like, if it doesn't inherently know how to multiply 13 * 17 without using Python to do it... I don't really care.
Maybe you do care. Maybe you want an LLM to be able to do work, only in its head.
But I kind of can't understand the desire for that limitation...
I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can't go fast enough. The Drive gear is literally right there.
numitus · · focus · HN ↗