Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
The AI can write a chess bot program that will beat you.
You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well.
People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.
Then why respond at all for the sake of responding?
We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own.
You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.
I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?
I agree w/ this perspective. An agent with a harness that can run programs can solve a lot more than one without the harness. The AI system includes the harness, and it's not clear to me that AGI requires more than LLMs + code generation & execution are capable of.
So AI is AGI in fields where code can't solve anything?
Is code omnipotent, I have been in software all my life and I would hard agree here.
Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it's the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.
For instance Maths is just code with different symbols and slightly less universally legible concepts.
AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.
But that's it, I am certain a bunch of companies will make a lot of money despite no AGI.
I think people either don't understand AGI or don't understand how real world works.
Until an LLM can bow it's head take responsibility for mistakes made and ensure they aren't repeated again with 100% confidence to the leadership it's inarguably a tool a rather questionable one at that.
> AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.
So.. like chess?
Anyway, do you have any prediction on what LLM's can or can't do in a few years?
It's not even a "software-generating product". It's only half of it. Most of the heavy lifting is done by absolutely not-AI compilers, analyzers and the like. If not for these programs, written well before AI boom, them LLMs would be no better at programming than they are are at pure LLM based calculations or writing.
If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this.
So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
I would bet a lot of money that Astra can follow the rules of chess (perhaps if repeated within the context window). Also, this is a different argument than what I responded to.
I can write you a benchmark to prove it even with a heavy handed system prompt Astra will make an illegal move during the course of the games first few moves are generally ok since it's just throwing out learned moves.
I wonder if I would do better as a human, maybe? Or would I happen to have one move in 1500+ that's not valid?
I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece.
> generally read the rules of a game and then follow them
How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a fork or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
Chess.com has to accommodate people who haven't learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they're often outlining sequences of multiple "pre-moves" during the opponent's turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.
Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.
So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.
We're not talking about learning the rules of chess here, but playing a competent game from just the rules. Why is it so hard for people to keep track of the thread of discussion?
> We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules.
Okay, lets go with that: it's the "shown the rules" bit that we are arguing about.
The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.
This does not point to generalisable and adaptable intelligence, such as we see in the average human.
This is not good reasoning. Humans need at least dozens if not hundreds of reinforcement sessions to only make legal moves, and still occasionally fail (consider pins, walking into check, failing to respond to check). LLMs must one-shot a competent game after imbibing a mass of disconnected units of information about chess. Nothing about the two are similar.
See my comment here for more: <a href="https://news.ycombinator.com/item?id=49725306">https://news.ycombinator.com/item?id=49725306
It only matters if you are claiming it to be general purpose.
If you admit that it's just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose.
The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can't do is very relevant.
This is roughly comparable to observing a cat batting a ball away with its paw and taking this as a "strong indication" that cats can play any sport.
It would be more impressive if they could play chess (or do anything they haven't been custom RLVR trained for) by reasoning, rather than just "have a go at it" prediction which is closer to memorization.
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
> They can't possibly remember even a few positions.
Sure they could, but that's irrelevant.
A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.
However, that's not how LLMs work. They don't memorize inputs - they predict them, based on disovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positons). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.
> Don't you know the legend about rice grains on a chess board?
Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?
1) The number of unique chess games that could theoretically be played (but mostly never have been), is irrelevant to what an LLM is remembering. It can only remember what was in it's training data - a far smaller number of maybe 10's of millions of games (of 30-50 moves each).
2) An LLM is not going to memorize vs generalize when there is no training pressure to do so. You might expect it to memorize book openings that occur over and over in the training data, but not some random non-celebrity game that occurs once in the Lichess dataset and is never again referred to.
> They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board?
If the wise man was a bit wiser, he'd have asked for his rice on a snakes & ladders board (100 squares, not 64) and would have had 2^36 more rice, which is equally irrelevant.
An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.
And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.
Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.
This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing
Do you know what the "General" in "Artificial General Intelligence" means? It specifically means that the AGI adapts to novel domains that it hasn't been trained on - its training generalizes to real world problems.
That doesn't mean it has to be extraordinary at these things. But to be AGI, it has to have some level of competency when used on problems outside its training set. In particular, it the LLMs were to install a known chess engine and run that to get the moves when asked to play chess, that would qualify for more AGI-like behavior. But really, chess is such a simplistic game that they should be able to do decently well at it even without even needing that. At the very least, they should be able to consistently play without making illegal moves - something that many 7-year olds manage quite well.
On the contrary, I think the chess comparison is on point. We’re discussing observations that even the strongest models devolve into making invalid moves without scaffolding. For me that raises the question of whether these models are learning the rules and generalizing from them, or of they’re just pattern matching and flailing on this task. Maybe the reality is somewhere in between, but the benchmarks don’t seem to directly measure conceptual generalization, they measure task completion. They can disrupt a lot of people and industries by pattern matching and flailing without being AGI.
I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe
It's not quite the same, but the in-flight chess game provided by Delta was known to be absurdly hard: <a href="https://news.ycombinator.com/item?id=46593395">https://news.ycombinator.com/item?id=46593395
I'm pretty sure "competent at chess without external aids" has been on the standard AGI checklist since before personal computers were a thing. How can you claim an intelligence is general if it can't make sense of such a highly constrained board game?
Because they’ll train it to be good at chess and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even ____
It can’t even count the R’s in strawberry
It can’t even add numbers
It can’t even solve a millennium puzzle
It’s not even a chess GM
It’s not even beyond human capability in Go
It can’t even drive a car
It can’t even self replicate
It can’t even build weapons
It doesn’t even have feelings
So how could someone conceivably convince everyone that some system is AGI when there are still tasks that some human or group of humans can do that the system cannot?
This will only happen, in my opinion, when the model/system can self-improve at a rate that scares people.
> and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even
One, you're not addressing what I wrote above and two, yes, that's absolutely correct. Doing X doesn't qualify something as AGI. If you can't X you can't be AGI. The inverse doesn't hold though. In particular if you have to retain the model in order to X then it can't possibly be AGI since (being _general_) it would be capable of figuring X out on its own having never seen it before.
Yes, it is indeed self evident. If it can't figure things out then its intelligence isn't general in which case it can't be AGI by definition.
No, because there is no coherent, agreed-upon definition. There’s just a million people vibe defining it.
Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever.
Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not.
And for what it’s worth I just watched GitHub Copilot figure something out. So your definition is once again lacking.
Throughout this exchange you're repeatedly confusing the negative and the positive. There is no rigorous and universally agreed upon criteria for exactly what would constitute AGI. There are some vague shapes that are widely (but not universally) accepted such as largely (vague boundary) being capable of replacing (vague criteria) humans.
However there are plenty of disqualifiers that are more or less universally accepted. In the above case it is literally by definition. Something cannot be termed general if it is incapable of generalizing.
> However there are plenty of disqualifiers that are more or less universally accepted (ie the negative)
Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI. There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them.
An AI controlled robot will be standing over the cooling corpse of the last human who will die certain that it wasn’t done by AGI.
First, you're moving the goalposts. Second, it's not actually true that any existing frontier AI can write a chess bot program that can beat a 1600 player ... not unless the program is derived from Stockfish or some other leading engine that has been in development for decades.
> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
These comments indicate a complete failure to understand the technology.
> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
This argument is fundamentally incompatible with all the breathless rhetoric about "AGI" coming from the providers' general direction.
The labs frequently apply their raw models to problems that do not make economic sense for their customers but that demonstrate the power and capability of their systems. These experiments can cost millions of dollars. That's not customer-shaped.
They're not going to give you access to that. It's not a product. The government might have an interest in this, but that's not something you'd be privileged to know about.
And when these labs do develop "AGI", they more than likely won't be selling it to end users. They've pretty much already said this.
>> the breathless rhetoric about "AGI" coming from the providers' general direction
So many commenters here see it as their ... duty? to argue against the most optimistic/unhinged (take your pick) arguments from "the other side" and then treat everybody who disagrees as a shill or an idiot.
Why is "being good at chess" a proxy for whatever AGI strawmen you want to argue against?
Maybe step back from your black-and-white ledge and think about discussing what's actually under discussion? For example, why or why not would an LLM be good at chess? Will they be good at chess? What technical limitations might preclude that?
carodgers · · focus · HN ↗
<a href="https://arxiv.org/html/2509.24239v4" rel="nofollow">https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
joefourier · · focus · HN ↗
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
sigmoid10 · · focus · HN ↗
<a href="https://chessbench-ai.github.io/#leaderboard" rel="nofollow">https://chessbench-ai.github.io/#leaderboard
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
minraws · · focus · HN ↗
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
echelon · · focus · HN ↗
You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
minraws · · focus · HN ↗
Delusion runs deep in HN circles.
I say that as someone heavily invested in AI startups and projects and as someone working in the field.
I think most people on HN should touch grass and find real human contact. Lmao
Incredible reasoning all around here.
echelon · · focus · HN ↗
People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.
minraws · · focus · HN ↗
We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own.
You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.
I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?
sdf32dsf · · focus · HN ↗
He definitely needs to touch grass.
echelon · · focus · HN ↗
Y'all seem to miss the point of this forum. Building and hacking and science and engineering.
I swear there's a whole lot of you who just like to look down instead of up. There's a whole universe up there.
bigstrat2003 · · focus · HN ↗
We know no such thing. LLMs are quite bad at generating code, worse than any capable human.
modulus1 · · focus · HN ↗
minraws · · focus · HN ↗
Is code omnipotent, I have been in software all my life and I would hard agree here.
Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it's the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.
For instance Maths is just code with different symbols and slightly less universally legible concepts.
AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.
But that's it, I am certain a bunch of companies will make a lot of money despite no AGI.
I think people either don't understand AGI or don't understand how real world works.
Until an LLM can bow it's head take responsibility for mistakes made and ensure they aren't repeated again with 100% confidence to the leadership it's inarguably a tool a rather questionable one at that.
simianwords · · focus · HN ↗
So.. like chess?
Anyway, do you have any prediction on what LLM's can or can't do in a few years?
Yizahi · · focus · HN ↗
diehunde · · focus · HN ↗
Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
hackinthebochs · · focus · HN ↗
Why should that matter?
janalsncm · · focus · HN ↗
So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
hackinthebochs · · focus · HN ↗
minraws · · focus · HN ↗
hackinthebochs · · focus · HN ↗
janalsncm · · focus · HN ↗
simianwords · · focus · HN ↗
>GPT-6 Astra xHigh: 0.06% rejected moves
frde_me · · focus · HN ↗
I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece.
Gregkion · · focus · HN ↗
And the stuff i'm using LLMs daily is just fake?
I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.
dosisking · · focus · HN ↗
It simply means that LLMs are smarter than you, but not smarter than the average person
zahlman · · focus · HN ↗
No, because we can, in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us.
> And the stuff i'm using LLMs daily is just fake?
This misses the point completely.
hackinthebochs · · focus · HN ↗
How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a fork or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
diehunde · · focus · HN ↗
hackinthebochs · · focus · HN ↗
zahlman · · focus · HN ↗
lelanthran · · focus · HN ↗
Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.
So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.
hackinthebochs · · focus · HN ↗
ncruces · · focus · HN ↗
lelanthran · · focus · HN ↗
Okay, lets go with that: it's the "shown the rules" bit that we are arguing about.
The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.
This does not point to generalisable and adaptable intelligence, such as we see in the average human.
hackinthebochs · · focus · HN ↗
See my comment here for more: <a href="https://news.ycombinator.com/item?id=49725306">https://news.ycombinator.com/item?id=49725306
HarHarVeryFunny · · focus · HN ↗
It only matters if you are claiming it to be general purpose.
If you admit that it's just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose.
The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can't do is very relevant.
lostmsu · · focus · HN ↗
recursive · · focus · HN ↗
lostmsu · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
lostmsu · · focus · HN ↗
zahlman · · focus · HN ↗
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
Sure they could, but that's irrelevant.
A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.
However, that's not how LLMs work. They don't memorize inputs - they predict them, based on disovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positons). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.
> Don't you know the legend about rice grains on a chess board?
Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
If not, then what are you talking about ?
If yes, then what is the relevance to an LLM playing chess ?
lostmsu · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
2) An LLM is not going to memorize vs generalize when there is no training pressure to do so. You might expect it to memorize book openings that occur over and over in the training data, but not some random non-celebrity game that occurs once in the Lichess dataset and is never again referred to.
> They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board?
If the wise man was a bit wiser, he'd have asked for his rice on a snakes & ladders board (100 squares, not 64) and would have had 2^36 more rice, which is equally irrelevant.
Gregkion · · focus · HN ↗
And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.
Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.
This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing
tsimionescu · · focus · HN ↗
That doesn't mean it has to be extraordinary at these things. But to be AGI, it has to have some level of competency when used on problems outside its training set. In particular, it the LLMs were to install a known chess engine and run that to get the moves when asked to play chess, that would qualify for more AGI-like behavior. But really, chess is such a simplistic game that they should be able to do decently well at it even without even needing that. At the very least, they should be able to consistently play without making illegal moves - something that many 7-year olds manage quite well.
rsfern · · focus · HN ↗
I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe
striking · · focus · HN ↗
willmarch · · focus · HN ↗
what · · focus · HN ↗
>If they cared to have it perform well in chess games, you'd see a different shape and behavior.
So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?
phoghed · · focus · HN ↗
fc417fc802 · · focus · HN ↗
phoghed · · focus · HN ↗
It can’t even count the R’s in strawberry
It can’t even add numbers
It can’t even solve a millennium puzzle
It’s not even a chess GM
It’s not even beyond human capability in Go
It can’t even drive a car
It can’t even self replicate
It can’t even build weapons
It doesn’t even have feelings
So how could someone conceivably convince everyone that some system is AGI when there are still tasks that some human or group of humans can do that the system cannot?
This will only happen, in my opinion, when the model/system can self-improve at a rate that scares people.
fc417fc802 · · focus · HN ↗
One, you're not addressing what I wrote above and two, yes, that's absolutely correct. Doing X doesn't qualify something as AGI. If you can't X you can't be AGI. The inverse doesn't hold though. In particular if you have to retain the model in order to X then it can't possibly be AGI since (being _general_) it would be capable of figuring X out on its own having never seen it before.
phoghed · · focus · HN ↗
fc417fc802 · · focus · HN ↗
phoghed · · focus · HN ↗
Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever.
Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not.
And for what it’s worth I just watched GitHub Copilot figure something out. So your definition is once again lacking.
fc417fc802 · · focus · HN ↗
However there are plenty of disqualifiers that are more or less universally accepted. In the above case it is literally by definition. Something cannot be termed general if it is incapable of generalizing.
phoghed · · focus · HN ↗
Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI. There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them.
An AI controlled robot will be standing over the cooling corpse of the last human who will die certain that it wasn’t done by AGI.
cindyllm · · focus · HN ↗
[dead]
jibal · · focus · HN ↗
> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
These comments indicate a complete failure to understand the technology.
I won't respond again.
zahlman · · focus · HN ↗
This argument is fundamentally incompatible with all the breathless rhetoric about "AGI" coming from the providers' general direction.
echelon · · focus · HN ↗
The labs frequently apply their raw models to problems that do not make economic sense for their customers but that demonstrate the power and capability of their systems. These experiments can cost millions of dollars. That's not customer-shaped.
They're not going to give you access to that. It's not a product. The government might have an interest in this, but that's not something you'd be privileged to know about.
And when these labs do develop "AGI", they more than likely won't be selling it to end users. They've pretty much already said this.
anthonyrstevens · · focus · HN ↗
So many commenters here see it as their ... duty? to argue against the most optimistic/unhinged (take your pick) arguments from "the other side" and then treat everybody who disagrees as a shill or an idiot.
Why is "being good at chess" a proxy for whatever AGI strawmen you want to argue against?
Maybe step back from your black-and-white ledge and think about discussing what's actually under discussion? For example, why or why not would an LLM be good at chess? Will they be good at chess? What technical limitations might preclude that?