Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
This just exposes that they don't even do the thing you said.
Not only is it still true that they can't do math directly, but not even indirectly.
They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.
Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.
That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.
> they found bits of code that are associated with "math" and the supplied arguments
How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.
>> they found bits of code that are associated with "math" and the supplied arguments
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to...
This doesn't make any sense at all. Was this supposed to be a gotcha? An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition. The pattern recognition is also particularly compressed into its most sparse and fundamental components, as this is key to generalization. This is not a sensible difference between human and LLM learning, we do the same thing.
I'll give you a recent example from my usage. Pi harness with extension for learning Chinese. When using it to feed drill questions to me and rate answers it would sometimes get lost in the sauce and start generating user aka me answer and then rate it and comment it. It's trivially wrong to the point that if a person would do that, they would be considered for some serious psych issues.
And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.
So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.
This is often (but not always, as it is often run with positive temperature which deliberately shifts the generation path) because of disconnected context and the issues around context compaction. Long context has always been critical to important work, but it remains a substantial challenge due to computational bottlenecks. It isn't a fundamental issue with the architecture, moreso the tricks to make them cheaper to use.
Hugsbox · · focus · HN ↗
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
beloch · · focus · HN ↗
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
VCFundedGenYer · · focus · HN ↗
walrus01 · · focus · HN ↗
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
reference: <a href="https://github.com/pbrod/karney" rel="nofollow">https://github.com/pbrod/karney
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
Brian_K_White · · focus · HN ↗
Not only is it still true that they can't do math directly, but not even indirectly.
They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.
Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.
That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.
It's nothing more than an sql query.
walrus01 · · focus · HN ↗
How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.
AdieuToLogic · · focus · HN ↗
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to...
Wait for it...
Understanding.
hodgehog11 · · focus · HN ↗
UpsideDownRide · · focus · HN ↗
And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.
So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.
hodgehog11 · · focus · HN ↗