Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with:
> Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.
(I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)
For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.
It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the "book" theory in its training data, but it still completely fell apart at early midgame.
Sure, but installing a chess program is child/teen level general ability, and playing chess well is highly trained expert level ability. Which one are we sold AI as being?
I think we are being sold AI as expert only when given tools (although that is not emphasized). The (quasi?) miracle of AI right now is that you can get an agent to accomplish the task of a team of intelligent but not exceptional humans at speeds far exceeding what the human could do. Which makes it “cheap” to throw (effectively) dozens of teams at a problem for the equivalent of hundreds of man hours.
That may not be the AI of sci fi fantasy but it’s still a game changing reality.
A Transformer has a massive amount of state - it's entire KV cache, in addition to the user asking it to draw the state after every move, which is really unnecessary.
A human, at least a trained human (for fairer comparison to an LLM whose training data contained a ton of chess games) can absolutely do this - have you never seen demonstrations of expert players playing a dozen or more games while blindfolded?
A Transformer/LLM is not a human of course, and the way it will by default play chess is by prediction, not reasoning. An LLM actually does surprisingly well if you only give it the most recent 20 moves of a game where 40 moves have been played so far, since the moves NOT played tell it just as much as the ones that were played, letting it effectively infer a lot of what is on the board.
I just want to make sure it's clear: the reason I was asking it to redraw the board is because last time I tried (which was like a month ago), I didn't ask for that, and basically as soon as the opening was "out of book" it started trying to make illegal moves and made false statements about the position in its running commentary (and after being corrected on these points, started dropping pieces for no reason).
What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill?
Also what harness? If you’re using a general harness of course it’s going to try and give you commentary.
I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.
Why does a 6 year old not need any of these guardrails?
Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about.
Maybe not all AGIs have a path to digital singularity. Maybe our current era of intelligence modeling has fundamental flaws and we are in a local minimum of the artificial intelligence space.
To note, I would bet with a good amount of certainty that we have enough compute power and automation to DDOS the internet out of existence with botnets. That doesn’t make the frontier models intelligent, that just makes their handlers reckless.
Well, would a dissected frontal lobe in and of itself be intelligence?
I think the same goes for LLMs, they may be a core part of an LLM harness, but you may still need a couple other components (e.g. it may itself write itself a deterministic function to validate steps).
In and of itself intelligence is an ill-defined and badly understood concept.
You're pointing out that the goalposts are not fixed in the problem statement above, and gp's interpretation is not the most generous possible. But as the interpretations get more generous, the claim becomes more and more absurd. Maybe a properly-harnessed model would download the most advanced chess engine and query it to find the best move in each position, but that's not really demonstrating the model's intelligence anymore.
> but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.
This isn't just about judging LLM capability. This is about pointing out that these capabilities are not "AGI". If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.
Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it's hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.
Luna is one of the budget lower-end last generation models. It'd be useful to at least try to verify the present before being bearish about the future. For OpenAI, the best publicly available model is GPT-6 Astra with XHigh or Max reasoning, and for Anthropic it's Claude Fable 5.1 with XHigh or Max reasoning.
Elsewhere in the thread there are reports of Astra on xhigh playing at what I would characterize broadly as a competent casual level, at least given occasional prodding (which a human of that skill level would basically only require when trying to play unreasonably quickly). There seems to be a pattern (even after correcting for relative ELO systems that aren't calibrated) of the LLM bots demonstrating stronger play against traditional bots than against humans.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
zahlman · · focus · HN ↗
> Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.
(I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)
For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.
It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the "book" theory in its training data, but it still completely fell apart at early midgame.
titzer · · focus · HN ↗
[dead]
lirolero · · focus · HN ↗
[dead]
danpalmer · · focus · HN ↗
alpinisme · · focus · HN ↗
That may not be the AI of sci fi fantasy but it’s still a game changing reality.
kavok · · focus · HN ↗
Sometimes they’ll even manually search or write bespoke code to search json instead of using something like jq.
HarHarVeryFunny · · focus · HN ↗
A human, at least a trained human (for fairer comparison to an LLM whose training data contained a ton of chess games) can absolutely do this - have you never seen demonstrations of expert players playing a dozen or more games while blindfolded?
A Transformer/LLM is not a human of course, and the way it will by default play chess is by prediction, not reasoning. An LLM actually does surprisingly well if you only give it the most recent 20 moves of a game where 40 moves have been played so far, since the moves NOT played tell it just as much as the ones that were played, letting it effectively infer a lot of what is on the board.
zahlman · · focus · HN ↗
sailfast · · focus · HN ↗
Also what harness? If you’re using a general harness of course it’s going to try and give you commentary.
I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.
datsci_est_2015 · · focus · HN ↗
Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about.
Maybe not all AGIs have a path to digital singularity. Maybe our current era of intelligence modeling has fundamental flaws and we are in a local minimum of the artificial intelligence space.
To note, I would bet with a good amount of certainty that we have enough compute power and automation to DDOS the internet out of existence with botnets. That doesn’t make the frontier models intelligent, that just makes their handlers reckless.
trio8453 · · focus · HN ↗
They're not guardrails, they're a different input/output environment.
solenoid0937 · · focus · HN ↗
datsci_est_2015 · · focus · HN ↗
freejazz · · focus · HN ↗
claytongulick · · focus · HN ↗
sailfast · · focus · HN ↗
wavemode · · focus · HN ↗
gf000 · · focus · HN ↗
I think the same goes for LLMs, they may be a core part of an LLM harness, but you may still need a couple other components (e.g. it may itself write itself a deterministic function to validate steps).
In and of itself intelligence is an ill-defined and badly understood concept.
themgt · · focus · HN ↗
Why does a bird not need jet engines or regular professional maintenance?
topaz0 · · focus · HN ↗
zahlman · · focus · HN ↗
This isn't just about judging LLM capability. This is about pointing out that these capabilities are not "AGI". If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.
Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it's hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.
meowface · · focus · HN ↗
XMPPwocky · · focus · HN ↗
zahlman · · focus · HN ↗