Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
I tested both myself and a weak bot against Astra xhigh, <a href="https://lichess.org/study/27lCQqDa" rel="nofollow">https://lichess.org/study/27lCQqDa. It's still pretty bad at chess, though it takes longer to devolve into illegal moves.
So you weren't giving it an updated board state after every move? If you want to compare apples to apples, it should give an updated board state for each move, or you should play blindfolded.
Blindfolded flex by OP aside (I can barely play when seeing the board), considering reasoning traces and their nature, if we want to be fair, a person would have to get the moves, but be allowed to write them down or draw up a board in their notepad. My working memory can barely handle five chunks, a models reasoning tokens are masses of written text in comparison.
An LLM has been trained to do everything it does blindfolded, "only" using perfect recall of everything in it's hundreds of thousands of steps of context, and hundreds of layers of KV cache. It's a computer - it has a massive advantage over a human.
The fairest apples-to-apples comparison of an LLM whose training data included chess games would be a trained human such as Magnus Carlson, who can quite happily play a dozen or more simultaneous blindfold chess games.
> though it takes longer to devolve into illegal moves
Is this because the context is being saturated? How did you set it up?
Was the prompt something like "Here's the state of the board, you're white, your move, what do you do?" and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (thanks for sharing!) to my mental model and understanding.
I'd be curious how it would work if it started fresh each time. My guess is it would never make an illegal move, although it may not actually play all that well.
You can see the entire conversation for my game at <a href="https://chatgpt.com/share/6aaac17b-1384-83e8-98fd-4350a0ef69cd" rel="nofollow">https://chatgpt.com/share/6aaac17b-1384-83e8-98fd-4350a0ef69....
It looks like it played a fully legal game of chess with one exception, it said "rxd1+" (Rook takes D1 with check) instead of "rd1+" (Rook to D1 with check) on move 29.
I would say this did a really good job of playing chess. It moved the pieces consistently and traded pieces when required.
This is worlds away from the frontier ~1 year ago where models would hallucinate pieces into existence.
You can still see undercurrents of its old self, once I pointed out the illegal notation it hallucinated prior illegal moves. But I agree, it's leagues apart from prior iterations. It also knew thematic moves in the opening. But whenever it needs to play concretely rather than "I know so-and-so is a good move in these types of positions" it crumbles.
Would you tell a human that just tried doing an illegal move that they did "a really good job of playing chess"? The probability for such mistakes is greatly reduced but still far from negligible, which proves the point that guardrails are needed.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
sobellian · · focus · HN ↗
phist_mcgee · · focus · HN ↗
hackinthebochs · · focus · HN ↗
sobellian · · focus · HN ↗
Topfi · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
The fairest apples-to-apples comparison of an LLM whose training data included chess games would be a trained human such as Magnus Carlson, who can quite happily play a dozen or more simultaneous blindfold chess games.
losvedir · · focus · HN ↗
Is this because the context is being saturated? How did you set it up?
Was the prompt something like "Here's the state of the board, you're white, your move, what do you do?" and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (thanks for sharing!) to my mental model and understanding.
I'd be curious how it would work if it started fresh each time. My guess is it would never make an illegal move, although it may not actually play all that well.
sobellian · · focus · HN ↗
bhelkey · · focus · HN ↗
I would say this did a really good job of playing chess. It moved the pieces consistently and traded pieces when required.
This is worlds away from the frontier ~1 year ago where models would hallucinate pieces into existence.
sobellian · · focus · HN ↗
bhelkey · · focus · HN ↗
However, the pawn was defended by the queen and it took a forced queen trade to unlock the move.
I have seen much worse blunders from human players. And, I have made much worse blunders.
legulere · · focus · HN ↗