Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
This has the smell of "Why don't I have a faster horse".
Why no flying cars. Because objects have mass and inertia and people are incredibly stupid. Making a flying car has been done. Making a flying car not be a weapon of mass destruction is very, very hard.
Because you don't fully understand your own point...
You look at science fiction and say "why didn't I get flying cars" and not "why didn't most science fiction predict a global always on network that put the furthest places away from you a few microseconds away from audio, video, or any other type of information that can be digitally encoded.
Trying to use flying cars as a gotcha is missing that flying cars aren't near as useful as one would think in relation to their costs. Moving information has become far more useful than moving objects long distances quickly, especially humans.
> You look at science fiction and say "why didn't I get flying cars"
I definitely don't, and you're definitely not getting my point, but I'm amused that you've instead double down on somehow getting it more than me...
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
nutrientharvest · · focus · HN ↗
zeroonetwothree · · focus · HN ↗
krapp · · focus · HN ↗
freejazz · · focus · HN ↗
pixl97 · · focus · HN ↗
Why no flying cars. Because objects have mass and inertia and people are incredibly stupid. Making a flying car has been done. Making a flying car not be a weapon of mass destruction is very, very hard.
Also:
<a href="https://www.txdot.gov/about/newsroom/statewide/air-taxi-testing-taking-flight-in-texas.html" rel="nofollow">https://www.txdot.gov/about/newsroom/statewide/air-taxi-test...
freejazz · · focus · HN ↗
pixl97 · · focus · HN ↗
You look at science fiction and say "why didn't I get flying cars" and not "why didn't most science fiction predict a global always on network that put the furthest places away from you a few microseconds away from audio, video, or any other type of information that can be digitally encoded.
Trying to use flying cars as a gotcha is missing that flying cars aren't near as useful as one would think in relation to their costs. Moving information has become far more useful than moving objects long distances quickly, especially humans.
freejazz · · focus · HN ↗
I definitely don't, and you're definitely not getting my point, but I'm amused that you've instead double down on somehow getting it more than me...