Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there's a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn't make any illegal moves even after a few thousand moves.
Frontier labs care deeply about automating knowledge work and computer use. They are working hard on getting models better and better, and they are succeeding. Astra is a step change on that front. So good luck i guess, if chess performance is your barometer.
> Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player.
If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.
Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That's the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use/play something. Mastery can be gained overtime.
If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn).
Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).
That's quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).
So your brother tried to make illegal moves while learning the game and it took your brother hundreds of games to get to be a decent player? I don't see how this contradicts anything I said...
The _only_ illegal move a human might make as a beginner is a failed en passant or a bad castle. And yes, a few hundred games is all it takes to be better than any publicly available LLM at the moment.
...and mistakes like not seeing discovered checks, castling into check, trying to castle out of check, castling through a check, misunderstanding en passant, missing checks when promoting a piece, misunderstanding how stalemate works, etc.
If you compared a human after hundreds of games to a SOTA LLM that was also trained on the output of hundreds of chess games that it played, I suspect you would notice similar improvements.
Honestly, if we're only talking about SOTA llms, in sandbox mode without harness? No shot. Without harness LLMs have no memory of previous moves. I can't make the last version of chatgpt remember more than 5 movements. I guarantee you, if you do not add flags in the harness with 'left rook moved' or 'right rook moved', it will try to illegally castle 100% of the time it's in the situation. Llm+harness, just make it call stockfish tbh.
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
famouswaffles · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.
skydhash · · focus · HN ↗
willmarch · · focus · HN ↗
Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).
orwin · · focus · HN ↗
willmarch · · focus · HN ↗
orwin · · focus · HN ↗
willmarch · · focus · HN ↗
If you compared a human after hundreds of games to a SOTA LLM that was also trained on the output of hundreds of chess games that it played, I suspect you would notice similar improvements.
orwin · · focus · HN ↗