Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
> The gap in capabilities between those models which they tested, and actual current frontier ones is enormous.
Same story every 4 months and yet still no breakout, winning products. I've been hearing "the AI is good now" and "it 10x's my productivity" for a over a year now. If it were true, why aren't the all-in-AI using companies 10-15 years ahead of their competition yet? Why is it still all buggy, poorly designed junk?
I don't get why it is hard to understand there is middle ground. People are 10x their productivity, it isn't all buggy junk, but it isn't all it is hyped up to be either. It isn't that complicated.
If you hold the extreme position that there isn't any value in this, that's fine, but we're only having this discussion because these models have done what humans previously failed to do.
I don't think the poster disagrees with you at all. The middle ground is that there are no breakout products and that the models clearly aren't so powerful as to make these companies not produce shit code.
Right now AI hasn't even managed to replace all the human workers taking orders at the fast food drive thru. That's a job often performed by literal children and companies are still waiting for AI to get good enough for even that. Maybe one day it will be good enough, maybe one day it will outperform humans at such a basic task, but that day is not today. If the hype were anything close to reality, we'd see it everywhere in our lives.
Funny you mention this - a fast food restaurant in my town now has an LLM taking drive-thru orders.
Though I highly doubt it has taken anyone's job, since most of the work is still in making, packing and handing over the food. (In fact, given the area I live in, I partially feel like the advantage they saw in it was that the LLM can speak Spanish.)
Last I heard McDonald's and Taco Bell were trialing AI again at a limited number of stores. It's the kind of job AI should be really good at and many fast food companies are using call center workers currently. They really want AI to work, so they keep trying every few years to make it happen, but so far all they get are embarrassing social media posts
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
moron4hire · · focus · HN ↗
Same story every 4 months and yet still no breakout, winning products. I've been hearing "the AI is good now" and "it 10x's my productivity" for a over a year now. If it were true, why aren't the all-in-AI using companies 10-15 years ahead of their competition yet? Why is it still all buggy, poorly designed junk?
orangedog · · focus · HN ↗
If you hold the extreme position that there isn't any value in this, that's fine, but we're only having this discussion because these models have done what humans previously failed to do.
freejazz · · focus · HN ↗
autoexec · · focus · HN ↗
wavemode · · focus · HN ↗
Though I highly doubt it has taken anyone's job, since most of the work is still in making, packing and handing over the food. (In fact, given the area I live in, I partially feel like the advantage they saw in it was that the LLM can speak Spanish.)
autoexec · · focus · HN ↗