Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
> Heck, just for fun I asked a reasonably smart LLM to ...
LLMs are neither smart nor stupid. They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.
> You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path ...
Again, LLMs do not "hallucinate." They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.
Nothing more.
See also anthropomorphism[0].
> More precisely it's that [LLMs] can't do the math internally but they're quite capable of producing the tool that does the math.
This still falls under the purvey of statistical token generation. To wit, given enough variations of:
bc -e '1 + 2'
bc -e '41 + 1'
...
LLMs can identify the addition expression in "What is 4 + 1?" and then emit a `'bc "4 + 1"'` command to produce a response. This is not "doing" or "understanding" math.
It is pattern recognition, a task in which ANNs[1] excel.
I'm not anthropomorphizing anything, I literally said that the training data for the formulas and equations is baked into it. It only "knows" things because a crawler and scraper acquired the information from an existing written source. In just about the same way that information is baked into a printed encyclopedia.
This is not even remotely accurate. "Baking information" like into a "printed encyclopedia" is memorization. It has been shown, time and time again, that LLMs do not merely memorize. It is not even possible for it to do so at scale. It can memorize some things, yes, but it is forced during the training procedure to bake general concepts into intermediate layers (this is why transfer learning works), analogous to compression. One can make several arguments that compression and intrinisic feature sparsity is the closest mathematical explanation to understanding that we have.
It is completely possible to ask an LLM a series of increasingly more esoteric and discrete questions until you find precisely what information did, or did not make it into the model. If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually.
>>> I'm not anthropomorphizing anything ...
Yes you are, regarding LLMs at least. Here's why:
just for fun I asked a reasonably smart LLM to ...
[be] capable of understanding if it's gone off on
a hallucinatory path ...
"Smart" in this context is a subjective value judgement.
"Hallucinations" are only experienced by living organisms.
You then went on to state:
> If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually.
Again, "hallucinating" is not something an algorithm can do. Also, determining factuality is again subjective based on the person assessing the information.
You ever heard of something called a metaphor, guy? I use the word "smart" as shorthand to describe something that scores highly in a number of coding and terminal use benchmarks (as compared to, let's say, a 30B size model from one and a half years ago which will score much worse), and "hallucinating" to mean "outputs plausible sounding gibberish that doesn't hold together consistently". Of course there's no actual hallucination going on.
> You ever heard of something called a metaphor, guy?
In this media (comments in HN threads), all I can do is interpret what people write. ;-)
> And "hallucinating" to mean "outputs plausible sounding gibberish that doesn't hold together consistently". Of course there's no actual hallucination going on.
This may very well be what you know to be true and I have no reason nor desire to assume otherwise. The problem is... Many people use the word "hallucinating" in this context literally and not metaphorically.
Since I do not know you, how am I to tell the difference?
Hugsbox · · focus · HN ↗
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
beloch · · focus · HN ↗
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
VCFundedGenYer · · focus · HN ↗
walrus01 · · focus · HN ↗
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
reference: <a href="https://github.com/pbrod/karney" rel="nofollow">https://github.com/pbrod/karney
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
AdieuToLogic · · focus · HN ↗
LLMs are neither smart nor stupid. They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.
> You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path ...
Again, LLMs do not "hallucinate." They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.
Nothing more.
See also anthropomorphism[0].
> More precisely it's that [LLMs] can't do the math internally but they're quite capable of producing the tool that does the math.
This still falls under the purvey of statistical token generation. To wit, given enough variations of:
LLMs can identify the addition expression in "What is 4 + 1?" and then emit a `'bc "4 + 1"'` command to produce a response. This is not "doing" or "understanding" math.It is pattern recognition, a task in which ANNs[1] excel.
0 - <a href="https://en.wikipedia.org/wiki/Anthropomorphism" rel="nofollow">https://en.wikipedia.org/wiki/Anthropomorphism
1 - <a href="https://en.wikipedia.org/wiki/Neural_network_(machine_learning)" rel="nofollow">https://en.wikipedia.org/wiki/Neural_network_(machine_learni...
walrus01 · · focus · HN ↗
hodgehog11 · · focus · HN ↗
walrus01 · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
Yes you are, regarding LLMs at least. Here's why:
"Smart" in this context is a subjective value judgement. "Hallucinations" are only experienced by living organisms.You then went on to state:
> If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually.
Again, "hallucinating" is not something an algorithm can do. Also, determining factuality is again subjective based on the person assessing the information.
walrus01 · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
In this media (comments in HN threads), all I can do is interpret what people write. ;-)
> And "hallucinating" to mean "outputs plausible sounding gibberish that doesn't hold together consistently". Of course there's no actual hallucination going on.
This may very well be what you know to be true and I have no reason nor desire to assume otherwise. The problem is... Many people use the word "hallucinating" in this context literally and not metaphorically.
Since I do not know you, how am I to tell the difference?