Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
Even if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.
I wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.
This is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.
This is simply blatant misinformation. If you play a game online on lichess and go to the analysis board you can find when your game becomes novel. It will be within 20 turns unless you are intentionally following a known opening. In fact it will likely become unique within 10-15 turns.
I am around 1500 (actually 1649 on lichess blitz, but close enough).
I explored the last 5 games I played on lichess. Here are the number of moves before lichess had never seen that position before for each of the 5 games: 6, 12, 16, 15, 11.
> If you find yourself in a novel position within 10-15 moves it's likely a resignable one
The opponent is also going to be in a novel position. Should both players resign?
Just in case you think this is limited to low rated players like 1500s, look at MagnusCarlsen's most used account on Blitz: <a href="https://lichess.org/@/DrNykterstein/search?perf=2" rel="nofollow">https://lichess.org/@/DrNykterstein/search?perf=2
You will see that most games become unique to the whole of lichess within 15-20 moves and a good chunk between 10-15.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
sigmoid10 · · focus · HN ↗
<a href="https://chessbench-ai.github.io/#leaderboard" rel="nofollow">https://chessbench-ai.github.io/#leaderboard
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
csande17 · · focus · HN ↗
MichaelNolan · · focus · HN ↗
shric · · focus · HN ↗
As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.
traes · · focus · HN ↗
fahrvrgnugen · · focus · HN ↗
shric · · focus · HN ↗
Conservatively there are well over 10 to the 30 positions likely to show up in realistic games.
There are of the order of 10 to the 10 or so games recorded.
Thus well under one in a trillion positions are "known".
fahrvrgnugen · · focus · HN ↗
traes · · focus · HN ↗
fahrvrgnugen · · focus · HN ↗
shric · · focus · HN ↗
I am around 1500 (actually 1649 on lichess blitz, but close enough).
I explored the last 5 games I played on lichess. Here are the number of moves before lichess had never seen that position before for each of the 5 games: 6, 12, 16, 15, 11.
> If you find yourself in a novel position within 10-15 moves it's likely a resignable one
The opponent is also going to be in a novel position. Should both players resign?
Just in case you think this is limited to low rated players like 1500s, look at MagnusCarlsen's most used account on Blitz: <a href="https://lichess.org/@/DrNykterstein/search?perf=2" rel="nofollow">https://lichess.org/@/DrNykterstein/search?perf=2
You will see that most games become unique to the whole of lichess within 15-20 moves and a good chunk between 10-15.