Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
This just exposes that they don't even do the thing you said.
Not only is it still true that they can't do math directly, but not even indirectly.
They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.
Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.
That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.
> they found bits of code that are associated with "math" and the supplied arguments
How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.
>> they found bits of code that are associated with "math" and the supplied arguments
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to...
This doesn't make any sense at all. Was this supposed to be a gotcha? An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition. The pattern recognition is also particularly compressed into its most sparse and fundamental components, as this is key to generalization. This is not a sensible difference between human and LLM learning, we do the same thing.
I'll give you a recent example from my usage. Pi harness with extension for learning Chinese. When using it to feed drill questions to me and rate answers it would sometimes get lost in the sauce and start generating user aka me answer and then rate it and comment it. It's trivially wrong to the point that if a person would do that, they would be considered for some serious psych issues.
And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.
So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.
This is often (but not always, as it is often run with positive temperature which deliberately shifts the generation path) because of disconnected context and the issues around context compaction. Long context has always been critical to important work, but it remains a substantial challenge due to computational bottlenecks. It isn't a fundamental issue with the architecture, moreso the tricks to make them cheaper to use.
>>> How is this different from a human using an algorithm they have memorized ...
>> Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans ...
> This doesn't make any sense at all. Was this supposed to be a gotcha?
No, it was meant to be an explanation as to the difference between "memorization" and "understanding." In this context, people pick the algorithm they determine applicable and then the question of memorization is relevant.
> An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition.
Funny that you make this argument here, where when I wrote elsewhere in this thread:
[LLMs] are statistical token generators whose results are
dependent upon their training data set and involve a
degree of randomness.
Nothing more.
...
It is pattern recognition, a task in which ANNs excel.
To which you replied to the above with:
During conversation, we are statistical token generators
whose results are dependent upon our training set.
Seriously, write that definition out rigorously. It
encompasses virtually everything. It is totally
meaningless. So to say "nothing more" is effectively also a
tautology.
This argument was asinine in 2024. It is insane to be
saying these things in 2026. Where have you been?
...
It absolutely understands how to do math, by whatever
reasonable definition you want to provide to the word
"understand".
So which is it?
Are LLMs ANNs? Which themselves are pattern recognition algorithms (hint: they are)?
OR (setting aside the ad hominems you kindly provided)
Do LLMs possess "understanding" of concepts such as abstract mathematics (defined and interpreted by humans) and we, as simple humans, nothing more than statistical token generators as you assert?
It is both. I do not understand why you would assert that both cannot hold simultaneously. Pattern recognition becomes "understanding" once individually recognized concepts become sufficiently sparsified and compactified. Or, at least, that is to my knowledge the only mathematically valid definition of "understanding" one can produce at this scale (it is valid under Solomonoff induction). Philosophy is fine, but we need to have a consistent definition of what "understanding" means, or we will just talk past each other. I argue that for any proper definition you provide which humans satisfy, a strong LLM is very likely to satisfy that as well.
I also would not argue that humans are "simple token generators". That is not what I said. I said that just about everything can fall under the classification of "statistical token generators" at an abstract level, so it isn't a useful distinction. We are not talking about a Markov chain generator from the 90s, so if that is the frame of reference, I think we should all get that out of our heads.
Frankly, I don't think you understand what "statistically generating tokens" actually means. Write that definition out formally. Then compare that operation to what a human does, assuming no revisions. It is the same, and that is my point. If you believe that humans understand, then "statistically generating tokens" cannot be disjoint from understanding.
You have observed nothing more than that a human can turn a shaft the same as an electric motor, and that an mp3 player can say "hello" the same as a human.
That observation is my point. The definition of "statistically generating tokens" is too broad as to be meaningless in this context. So using it as a reason for lack of understanding is ridiculous.
I wrote a program that statistically generated tokens to consistently factor latge RSA numbers. Does this program actually factor numbers or is it just a statistical next token predictor?
> Pattern recognition becomes "understanding" once individually recognized concepts become sufficiently sparsified [sic] and compactified [sic].
Understanding is a state of mind. As such, it exists entirely within an individual and nowhere else.
For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.
> I argue that for any proper definition [of understanding] you provide which humans satisfy, a strong LLM is very likely to satisfy that as well.
This is demonstrably incorrect as detailed above. There is no "understanding" LLMs can satisfy as we know it, since to certify said "understanding", it requires interpretation by a person to "know" an LLM "understands."
> I also would not argue that humans are "simple token generators". That is not what I said.
That is the essence of what you wrote, unless you object to my use of "simple" instead of "statistical". In this context, I postulate this is a distinction without difference.
> I said that just about everything can fall under the classification of "statistical token generators" at an abstract level, so it isn't a useful distinction.
This only holds if one subscribes to statistical token generators being a/the fundamental underpinning of "everything". Here is a proof by contradiction:
If everything can be classified as a derivative of
statistical token generation, how does one explain
quantum physics?
>Understanding is a state of mind. As such, it exists entirely within an individual and nowhere else.
For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.
We cannot engage in an intellectual discussion if we do not agree on definitions. So far, your definitions of understanding seem to be whatever vibe you are going for in the statement and I would urge you to think about what a sensible mathematical definition of understanding is so that it can sensibly be assessed on neural networks. Otherwise, this isn't science, it's a debate about personal experience.
> Understanding is a state of mind
This is meaningless, it is a circular definition at best.
> It exists entirely within the individual and nowhere else
Then why are we talking about it? What is the point if it is something that can only be defined per individual?
> requires interpretation by a person to "know" an LLM "understands."
We are still not getting anywhere because you have not prescribed criteria to determine whether it understands. If it is a "know it when I see it" situation, that clearly isn't working. For example, if you say that you need to dig into its internals and figure out whether it is breaking things down appropriately, that doesn't work because you probably don't have the expertise to do that. The experts that do are telling you that it very likely understands because it pulls apart most concepts in the way we would expect.
I do object to the use of the word "simple". "Statistical" is so broad to be almost meaningless; it merely means that a prediction is being made in the presence of data which possibly contains some degree of uncertainty. "Simple" encompasses that which can be understood readily by a non-expert.
Quantum mechanics is statistical (this is literally the Born rule), but evolutions are not operating as stochastic processes in the sense of Kolmogorov. That is very different, and not relevant to our discussion.
> So far, your definitions of understanding seem to be whatever vibe you are going for in the statement and I would urge you to think about what a sensible mathematical definition of understanding is so that it can sensibly be assessed on neural networks.
Any reasonable definition of understanding is not dependent upon "whatever vibe you are going for", but instead must include at least an English dictionary definition of "understand" such as:
to grasp the meaning of[0]
And, for further clarification, "grasp" can be defined as:
to lay hold of with the mind[1]
Which makes an equivalent term-expanded definition of "understand" to be:
to lay hold of with the mind the meaning of
As such, there is no "sensible mathematical definition of understanding", unless you possess a complete mathematical model of the human mind.
>> Understanding is a state of mind
> This is meaningless, it is a circular definition at best.
See above to as to why there is meaning in what I wrote.
>> It exists entirely within the individual and nowhere else
> Then why are we talking about it? What is the point if it is something that can only be defined per individual?
I like to think analyzing fundamental premises, often implicit, explicitly can help to identify fallacious positions.
> We are still not getting anywhere because you have not prescribed criteria to determine whether [an LLM] understands.
My apologies for being opaque. Let me clarify:
LLMs do not "understand". People interpreting LLM output
are the only entities involved which can "understand",
because "understanding" exists strictly within each
person who possesses it.
If by "result" you mean the final code, then just asking someone else who understood the math to write it would also have achieved the same result.
On the other hand, if by "result" you mean that you gained knowledge or understanding of the code in a way where you could personally tailor its behavior to specific circumstances without asking for help, then it's not the same result at all.
I find a lot of the arguments that having LLMs write your code is no different from copy/pasting Stack Overflow answers to be specious. They blur the line between asking for help and asking for someone else (or something else) to do the work for you. What they ignore is that doing the work yourself has ancillary benefits and is a valuable end in its own right.
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
And how is _that_ different from making the human memorize a billion weights and do matrix calculations in their head, in order to generate tokens?
How is _that_ different from a hive of bees trained to do the same?
Everything is modelable by a tueung computer as fsr as we know. So yes, I don't see why these processes can't do the same, its on you to show why such a process implemented in any of these isn't a form of thought.
Hugsbox · · focus · HN ↗
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
beloch · · focus · HN ↗
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
VCFundedGenYer · · focus · HN ↗
walrus01 · · focus · HN ↗
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
reference: <a href="https://github.com/pbrod/karney" rel="nofollow">https://github.com/pbrod/karney
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
Brian_K_White · · focus · HN ↗
Not only is it still true that they can't do math directly, but not even indirectly.
They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.
Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.
That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.
It's nothing more than an sql query.
walrus01 · · focus · HN ↗
How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.
AdieuToLogic · · focus · HN ↗
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to...
Wait for it...
Understanding.
hodgehog11 · · focus · HN ↗
UpsideDownRide · · focus · HN ↗
And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.
So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.
hodgehog11 · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
>> Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans ...
> This doesn't make any sense at all. Was this supposed to be a gotcha?
No, it was meant to be an explanation as to the difference between "memorization" and "understanding." In this context, people pick the algorithm they determine applicable and then the question of memorization is relevant.
> An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition.
Funny that you make this argument here, where when I wrote elsewhere in this thread:
To which you replied to the above with: So which is it?Are LLMs ANNs? Which themselves are pattern recognition algorithms (hint: they are)?
OR (setting aside the ad hominems you kindly provided)
Do LLMs possess "understanding" of concepts such as abstract mathematics (defined and interpreted by humans) and we, as simple humans, nothing more than statistical token generators as you assert?
Because it cannot be both.
hodgehog11 · · focus · HN ↗
I also would not argue that humans are "simple token generators". That is not what I said. I said that just about everything can fall under the classification of "statistical token generators" at an abstract level, so it isn't a useful distinction. We are not talking about a Markov chain generator from the 90s, so if that is the frame of reference, I think we should all get that out of our heads.
tripzilch · · focus · HN ↗
Okay fine. I think we can agree to disagree on that.
hodgehog11 · · focus · HN ↗
Brian_K_White · · focus · HN ↗
You have observed nothing more than that a human can turn a shaft the same as an electric motor, and that an mp3 player can say "hello" the same as a human.
hodgehog11 · · focus · HN ↗
hardbass · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
Understanding is a state of mind. As such, it exists entirely within an individual and nowhere else.
For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.
> I argue that for any proper definition [of understanding] you provide which humans satisfy, a strong LLM is very likely to satisfy that as well.
This is demonstrably incorrect as detailed above. There is no "understanding" LLMs can satisfy as we know it, since to certify said "understanding", it requires interpretation by a person to "know" an LLM "understands."
> I also would not argue that humans are "simple token generators". That is not what I said.
That is the essence of what you wrote, unless you object to my use of "simple" instead of "statistical". In this context, I postulate this is a distinction without difference.
> I said that just about everything can fall under the classification of "statistical token generators" at an abstract level, so it isn't a useful distinction.
This only holds if one subscribes to statistical token generators being a/the fundamental underpinning of "everything". Here is a proof by contradiction:
hardbass · · focus · HN ↗
For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.
What? What are you even trying to say?
hodgehog11 · · focus · HN ↗
> Understanding is a state of mind
This is meaningless, it is a circular definition at best.
> It exists entirely within the individual and nowhere else
Then why are we talking about it? What is the point if it is something that can only be defined per individual?
> requires interpretation by a person to "know" an LLM "understands."
We are still not getting anywhere because you have not prescribed criteria to determine whether it understands. If it is a "know it when I see it" situation, that clearly isn't working. For example, if you say that you need to dig into its internals and figure out whether it is breaking things down appropriately, that doesn't work because you probably don't have the expertise to do that. The experts that do are telling you that it very likely understands because it pulls apart most concepts in the way we would expect.
I do object to the use of the word "simple". "Statistical" is so broad to be almost meaningless; it merely means that a prediction is being made in the presence of data which possibly contains some degree of uncertainty. "Simple" encompasses that which can be understood readily by a non-expert.
Quantum mechanics is statistical (this is literally the Born rule), but evolutions are not operating as stochastic processes in the sense of Kolmogorov. That is very different, and not relevant to our discussion.
AdieuToLogic · · focus · HN ↗
Any reasonable definition of understanding is not dependent upon "whatever vibe you are going for", but instead must include at least an English dictionary definition of "understand" such as:
And, for further clarification, "grasp" can be defined as: Which makes an equivalent term-expanded definition of "understand" to be: As such, there is no "sensible mathematical definition of understanding", unless you possess a complete mathematical model of the human mind.>> Understanding is a state of mind
> This is meaningless, it is a circular definition at best.
See above to as to why there is meaning in what I wrote.
>> It exists entirely within the individual and nowhere else
> Then why are we talking about it? What is the point if it is something that can only be defined per individual?
I like to think analyzing fundamental premises, often implicit, explicitly can help to identify fallacious positions.
> We are still not getting anywhere because you have not prescribed criteria to determine whether [an LLM] understands.
My apologies for being opaque. Let me clarify:
0 - <a href="https://www.merriam-webster.com/dictionary/understand" rel="nofollow">https://www.merriam-webster.com/dictionary/understand1 - <a href="https://www.merriam-webster.com/dictionary/grasp" rel="nofollow">https://www.merriam-webster.com/dictionary/grasp
noduerme · · focus · HN ↗
On the other hand, if by "result" you mean that you gained knowledge or understanding of the code in a way where you could personally tailor its behavior to specific circumstances without asking for help, then it's not the same result at all.
I find a lot of the arguments that having LLMs write your code is no different from copy/pasting Stack Overflow answers to be specious. They blur the line between asking for help and asking for someone else (or something else) to do the work for you. What they ignore is that doing the work yourself has ancillary benefits and is a valuable end in its own right.
tripzilch · · focus · HN ↗
And how is _that_ different from making the human memorize a billion weights and do matrix calculations in their head, in order to generate tokens?
How is _that_ different from a hive of bees trained to do the same?
Go ahead, argue these things are all the same ...
hardbass · · focus · HN ↗