Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
From your linked post: "LLMs sometimes struggle to give legal moves. In these experiments, I try 10 times and if there’s still no legal move, I just pick one at random."
Yeah but it didn't do squat. The effective thing was a subtle prompt mod that boosted performance tremendously. I've been surprised by this forums unwillingness to read this thing...
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
thelaxiankey · · focus · HN ↗
<a href="https://dynomight.net/chess/" rel="nofollow">https://dynomight.net/chess/
thelaxiankey · · focus · HN ↗
<a href="https://dynomight.net/more-chess/" rel="nofollow">https://dynomight.net/more-chess/
topaz0 · · focus · HN ↗
Which sounds a lot like what that paper was about
thelaxiankey · · focus · HN ↗