Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
The AI can write a chess bot program that will beat you.
You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
This is roughly comparable to observing a cat batting a ball away with its paw and taking this as a "strong indication" that cats can play any sport.
It would be more impressive if they could play chess (or do anything they haven't been custom RLVR trained for) by reasoning, rather than just "have a go at it" prediction which is closer to memorization.
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
> They can't possibly remember even a few positions.
Sure they could, but that's irrelevant.
A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.
However, that's not how LLMs work. They don't memorize inputs - they predict them, based on disovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positons). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.
> Don't you know the legend about rice grains on a chess board?
Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?
1) The number of unique chess games that could theoretically be played (but mostly never have been), is irrelevant to what an LLM is remembering. It can only remember what was in it's training data - a far smaller number of maybe 10's of millions of games (of 30-50 moves each).
2) An LLM is not going to memorize vs generalize when there is no training pressure to do so. You might expect it to memorize book openings that occur over and over in the training data, but not some random non-celebrity game that occurs once in the Lichess dataset and is never again referred to.
> They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board?
If the wise man was a bit wiser, he'd have asked for his rice on a snakes & ladders board (100 squares, not 64) and would have had 2^36 more rice, which is equally irrelevant.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
sigmoid10 · · focus · HN ↗
<a href="https://chessbench-ai.github.io/#leaderboard" rel="nofollow">https://chessbench-ai.github.io/#leaderboard
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
minraws · · focus · HN ↗
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
echelon · · focus · HN ↗
You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
minraws · · focus · HN ↗
Delusion runs deep in HN circles.
I say that as someone heavily invested in AI startups and projects and as someone working in the field.
I think most people on HN should touch grass and find real human contact. Lmao
Incredible reasoning all around here.
diehunde · · focus · HN ↗
Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
lostmsu · · focus · HN ↗
recursive · · focus · HN ↗
lostmsu · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
lostmsu · · focus · HN ↗
zahlman · · focus · HN ↗
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
Sure they could, but that's irrelevant.
A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.
However, that's not how LLMs work. They don't memorize inputs - they predict them, based on disovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positons). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.
> Don't you know the legend about rice grains on a chess board?
Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
If not, then what are you talking about ?
If yes, then what is the relevance to an LLM playing chess ?
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
2) An LLM is not going to memorize vs generalize when there is no training pressure to do so. You might expect it to memorize book openings that occur over and over in the training data, but not some random non-celebrity game that occurs once in the Lichess dataset and is never again referred to.
> They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board?
If the wise man was a bit wiser, he'd have asked for his rice on a snakes & ladders board (100 squares, not 64) and would have had 2^36 more rice, which is equally irrelevant.