Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
1. It’s hard to trust a 2026 paper that’s showing results for such old models.
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
Not sure how that vague truism applies to this paper.
Lots of papers have great results that don’t depend on the latest models.
However in this case it’s problematic:
- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.
- They ask are LLMs capable of X and arrive at a negative result.
If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.
However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
WhitneyLand · · focus · HN ↗
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
Certhas · · focus · HN ↗
The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues
WhitneyLand · · focus · HN ↗
Lots of papers have great results that don’t depend on the latest models.
However in this case it’s problematic:
- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.
- They ask are LLMs capable of X and arrive at a negative result.
If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.
However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.