Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
With that approach, the benchmark falls apart. Of course it can write a chess engine, because it learned on lots of stolen source code of chess engines. This has nothing to do with the LLM's ability to reason.
Writing a well understood engine for a super popular problem does not count as reasoning about the problem.
In most humans, yes, because we are terrible at memorizing millions of codebases. For LLMs, we need to apply our understanding of them before making statements like that. An LLM can "memorize", and has "memorized"/been trained on tens of thousands of chess engines. Writing a chess engine, or even deriving a chess engine from the rules alone, does not constitute a deep understanding of, and more importantly, the ability to apply, the rules, at all.
When humans do this, they inadvertently learn something, too, but when an LLM reproduces or derives and implementation of a chess engine, it in no way implies that the LLM can follow the rules in its own "train of thought" and consistently apply the rules in its "head".
Let's say you want to evaluate my algebra skills. You make me solve some algebra challenges. If I then whip out a computer and write a calculator, or take some sticks and stones and take a couple hours to build an abacus, and then solve the algebraic challenges, this would not constitute a good solution, and would defeat the entire point of the test. If, instead, I do the algebra in my head or on paper, it might seem like there's no difference, but you can derive all sorts of information from that.
For example, you could time it, check for recurring errors I make, for interesting mistakes like mistaking 7 and 1 for one another due to bad hand-writing, etc.
If that was the goal, then me writing a calculator or crafting an abacus defeats the point of the test. Yes, me writing a calculator shows that I'm intelligent, and I understand the algebraic rules, but if the test is about applying the rules, I have not passed.
In the very same way, an LLM writing a chess engine to solve a chess benchmark that is all about LLM's reasoning capability is complete bogus and defeats the entire point.
>Tool use is a widely accepted measure of intelligence.
Stop anthropomorphizing the parrot. The parrot has had stockfish's source code blasted at high pressure into its head along with dozens of millions of other pieces of code whose sole role is to have efficient algorithms to more or less brute force through the best result. Brute forcing (no matter how smart it is) isn't understanding the problem domain.
I'm not fully convinced that human beings aren't stochastic parrots. Much of the behaviour I observe in daily life reflects a blind adherence to set of beliefs that are an amalgation of "a person of authority said to do this" at an early age.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
glitchc · · focus · HN ↗
Does the LLM need to learn to play chess if it can build a chess engine to play for it instead?
lionkor · · focus · HN ↗
Writing a well understood engine for a super popular problem does not count as reasoning about the problem.
glitchc · · focus · HN ↗
Doesn't writing the engine imply understanding about the problem domain? Tool use is a widely accepted measure of intelligence.
lionkor · · focus · HN ↗
When humans do this, they inadvertently learn something, too, but when an LLM reproduces or derives and implementation of a chess engine, it in no way implies that the LLM can follow the rules in its own "train of thought" and consistently apply the rules in its "head".
Let's say you want to evaluate my algebra skills. You make me solve some algebra challenges. If I then whip out a computer and write a calculator, or take some sticks and stones and take a couple hours to build an abacus, and then solve the algebraic challenges, this would not constitute a good solution, and would defeat the entire point of the test. If, instead, I do the algebra in my head or on paper, it might seem like there's no difference, but you can derive all sorts of information from that.
For example, you could time it, check for recurring errors I make, for interesting mistakes like mistaking 7 and 1 for one another due to bad hand-writing, etc.
If that was the goal, then me writing a calculator or crafting an abacus defeats the point of the test. Yes, me writing a calculator shows that I'm intelligent, and I understand the algebraic rules, but if the test is about applying the rules, I have not passed.
In the very same way, an LLM writing a chess engine to solve a chess benchmark that is all about LLM's reasoning capability is complete bogus and defeats the entire point.
well_ackshually · · focus · HN ↗
Stop anthropomorphizing the parrot. The parrot has had stockfish's source code blasted at high pressure into its head along with dozens of millions of other pieces of code whose sole role is to have efficient algorithms to more or less brute force through the best result. Brute forcing (no matter how smart it is) isn't understanding the problem domain.
glitchc · · focus · HN ↗
Capricorn2481 · · focus · HN ↗