Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
I similarly noticed Gemini absolutely refuses to look at a url when I give it one and will instead just hallucinate based on what it thinks the url is. Here I am assuming Google will have the best web capabilities in its AI.
I searched for something, it told me that according to a YouTube video, the entire point of my search was wrong. I asked it for the source, I watched the video, it never made the claims Gemini hallucinated. I asked again and it claimed it scrubbed the video and found the point it made multiple times. I said those timecodes were wrong and it admitted it couldn't actually parse videos and just guessed.
Same experience with the interaction I described in my comment... after I finally got the correct answer, I asked where it got the faulty information from. It said it had just simply fabricated it. That's actively worse than just saying "I don't know", for something that's sold to us as an easy way to look up information.
These models are incapable of saying they don't know, because they have no concept of knowing. They simply predict the next word which is most likely.
The saddest part is when people take their experience with Google's idiotic AI implementation and assume that's how all LLMs work. Frontier-class models will, in fact, generally admit when they don't know something. That includes the one I run at home on my own graphics cards, but it seems that Google just doesn't GAF.
Your point about "predicting the next word" mostly means that your post was very easy to predict.
Maybe it's because they are trained on Internet comments, and the most rare thing to find on the Internet is someone admitting they don't know something.
But if they had been trained on comments saying "I don't know", they'd probably act the same as they do now but they'd treat "I don't know" as the answer.
This is such a 2024 take. Model will push back and tell you the knowledge gap they have, but people are just used to ask gooogle leading questions, which port badly to llm as it makes them defend the position instead of research data
Gemini is the worst gaslighter. It hallucinated a feature in an open source product. I said no that does not exist. It invented very realistic git commit messages. I said commit does not exist, so it said I was out of sync. I linked the repo and it said it might be a fork. Then it invented a mailing list thread where they supposedly removed the feature, that’s why I couldn’t find it.
> The scary (or funny) part is people use that for serious questions.
More scary than funny, IMO.
The US military almost took an action that could very plausibly have escalated into a hot war with China because people are already relying too heavily on these systems.
Despite the reporting, nobody in power seems sufficiently freaked out about this.
I don't understand how this isn't considered as active malice. Like purposefully outputting random stuff. Any other sort of computer system would get lot more flak than these are getting.
Throughout my career, I've almost always been close enough to the user that I hear about it quickly when something is wrong. It's a tough problem that so many Google engineers are typically so far removed from end users. Or maybe it's just that a small part of the company has long subsidized the rest of the employees to the point that it doesn't matter how good their work is because they'll get paid anyway.
The engineers are pressured to significantly reduce "dependencies" for projects. Anything that could become risk or create friction is dramatically less appetizing.
Simply because of how many people that *must* agree with your proposal. Getting all the relevant tech leads, some you have never heard of or ever spoken with, to agree on a proposal for your team's project is a nightmare.
So you keep it as simple and agreeable as possible. Given the circumstances, it makes sense as one of the engineers. It's fairly fine advice in general wherever you work, but it just haunts all the work you do at Google in particular. Nothing gets done otherwise.
If you have a dependency that can be dropped from an engineering perspective, that's the route the 9 leads reviewing your design doc will take:
"Let's iterate and start with just the basics (no user testing)", "let's get this working and user test in a later phase", "I think this problem is obvious enough we don't need to consult with users about it."
I worked in Ads Integrity at the time, and for one of my projects I was concerned how it would impact the manual reviewers. Then I learned I couldn't talk with them, only by proxy through another person if we really had to. And that proxy takes time, so...
Very insightful. I've worked in only one company of Google's size, but it was a vastly different industry. I personally never got to speak with a user on a massive internal application my team worked on for years. :)
Right. In a global financial services organization, a strategic analysis app used by thousands of internal users was pretty high touch. SME business analysts would sit with the users to make sure they liked it and had no trouble. Every user generates significant revenue. Worldwide B2C is a different ball game, and Google barely makes money off of each individual user.
> The engineers are pressured to significantly reduce "dependencies" for projects.
Is a major factor of this Google's tendency to kill entire products, so you obviously don't want to make your hot new thing dependent on some other internal thing that might be killed off?
Interestingly there is less opportunity for that to happen.
Most of the internal infrastructure is quite available and stable, and there are clear choices for almost all of your application-level and backend-level needs.
Guess, your argument reads so weird. It outputs PLAUSIBLE text. It doesn't actually have a connection to causally correct information, although plausiblness often coincidents with causally correctness.
I still believe this AI push out of nowhere is due to the current Government in power state side. Making everyone question themselves and each other and being uncertain about facts while being inundated with techbro fake news called hallucinations is a recipe for disaster for older populations that don't 'trust but verify' like most technology inclined people. This is all by design and we'll falling for it.
If you don't think the world is rife with criminal conspiracies you weren't paying attention in history class.
There's an enormous difference between outlandish claims about extraterrestrials or turning frog gays and believing there's a category of politicians trying to get wealthy off of their position. It's somewhat reasonable to hypothesize about criminal conspiracies in that 2nd scenario.
It takes either being scared to death to admit what's going on in the world, or an attention span of a guppy, to keep calling people conspiracy theorists these days.
Yup. I'm most familiar with Jill Lepore and Quinn Slobodian (and others in their respective orbits). They're both historians who've written extensively about Musk and Muskism (et al). Wild stuff.
Apparently the plan is to use AI slop, mediated thru social medias, to defeat the woke mind virus, perpetuated by the Anti-Christ, in order to safe guard humanity's future.
The more simple explanation is that the current government in power is over exposed in their AI investment. The normal conservative mainstream media was already doing a great job of propagandizing older populations.
When it says it can read videos you don’t trust it.
When it says it can’t read videos you think that’s an accurate introspection on its abilities, and not a statistically likely continuation of a conversation where one side seems to be reading videos and the other side says the read is inaccurate?
Claude also has very arbitrary and confusing rules about what web pages it allows itself to look at, and how much of the page it can read. Did you know for example you’ll get a much deeper analysis if you download a PDF yourself and upload it to Claude instead of giving it the url?
This is a key reason why I actually like Grok for factual queries based on web grounding. It’s fast and reliable. Maybe it’s ignoring robots.txt? Dunno. But it works well.
Literally every experience I've had with Gemini / Google Search AI answers followed this exact pattern, often repeated several times more if I remained persistent instead of just giving up.
Typical example:
"Where can I buy <thing I'm looking for that I can't find anywhere using normal search terms>?"
> You're looking for <related but different and widely available thing>. It is sold by <sites I never heard of>.
"No, that's different. I'm looking for <that thing but with the exact differences spelled out again>."
> Ah, my mistake. You're looking for <thing I described>. It is sold by <sites I never heard of but which don't actually sell it>.
"I've checked your links and none of those sites actually sell it, one doesn't even sell products and instead only offers manufacturing - but also not for what I asked you for."
> I'm sorry, my bad. Those sites don't sell what you are looking for. Instead you should check out <more sites I've never heard of>.
"Those sites sell the thing you initially thought I was asking about but not the thing I described."
> I'm sorry for the misunderstanding. You can find the thing you actually described at <yet more sites including some of the same>.
"No. None of these sites sell anything close to what I asked you for and two of them don't actually exist."
> Oh, sorry about that. You're completely right. The thing you asked me about isn't actually being sold by anyone. However you could buy <thing it first thought I meant and that wouldn't bring me any closer to solving my problem>.
The few times I've resigned myself to asking Gemini or ChatGPT something I couldn't find an answer to, my experience has been the same. 100% of the time. An LLM has never, not once, given me a correct answer or not lied to me. It's always been the same experience you describe. It goes in circles "Try X... Try Y... Try X?" Until I tell it to stop telling me X or Y, and then it goes, "LOL you can't do that at all, I was just wasting your time."
Most recently, I was considering moving away from iterm2 on MacOS, and I wanted to know if any other terminal emulator supported gestures for switching between tabs. So I asked Gemini, and it says, "Yes Ghostty supports gestures for switching between tabs."
"Ok I just installed Ghostty and I can't find anything about gestures."
"You need to add foo=bar to your conf file."
"I added foo=bar to my conf file and now it's saying the conf file is invalid."
"Sorry bro, remove foo=bar and add baz=boo to the conf file."
"It says baz=boo is invalid too."
"baz=boo isn't a real option. Remove that and add foo=bar to your conf file."
"You already told me to do that and I already told you that doesn't work."
"You shouldn't put foo=bar or baz=boo in the Ghostty conf file. Both are invalid. Ghostty doesn't support gestures. Have you considered iterm2?"
Gemini is utterly useless, but every other modern model would be capable of answering this correctly. If you ran a local coding agent, I wouldn’t be surprised if you could one-shot implement gesture support in Ghostty. It’d definitely configure BTT for you.
Here’s GPT 6’s answer to the prompt “What macOS terminal apps support gestures? Include a reference to the docs on how to enable/configure them.”:
iTerm2 supports configurable three-finger taps and swipes for switching tabs/panes, creating splits, pasting, etc. Set them up under Settings > Pointer > Bindings. Check for conflicting macOS trackpad assignments. [1]
The others are more limited: Ghostty supports macOS lookup/Quick Look gestures [2], while WezTerm lets you bind scroll events—for example, Ctrl+scroll to change font size. [3] Neither is equivalent to iTerm2’s gesture bindings.
For custom gestures without switching terminals, BetterTouchTool can map app-specific trackpad gestures to the terminal’s existing keyboard shortcuts. [4]
This whole tree just made me realize that people have wildly different prompting styles, and Gemini is probably too specialized for usage patterns of long-time Google Search users. The prompt my finger generated(before this comment was posted) was "are there any terminal emulator that supports gesture actions, on macOS, other than iterm2", and Gemini gave me Tabby, WezTerm, Kitty, and BetterTouchTool gestures.
This is a different experience to GP from query to result. I thought they've all fixed that issue of difference in tones affecting results. I guess it was never easily fixed.
Pretty sure it's the harness here, not Gemini per se. Plug Gemini into pi or any harness that has any competence pointing to web search and the hallucinations drop 75% immediately.
LLMs can actually be a good fit in this application if a better search engine feeds them a list of candidate websites that might sell the thing, and the LLM drives a web browser to see if any of the websites have it.
I have one google account where gemini constantly confidently hallucinates crap, and another where it doesn't (at least to the extent that other modern LLMs don't).
I take this to mean that my one account has been mistaken for a competitor and they're trying to poison its data. But who knows.
It's a reference to this old Dilbert cartoon: <a href="https://i.sstatic.net/Y3zE3.gif" rel="nofollow">https://i.sstatic.net/Y3zE3.gif
that's my 'who knows'-- It's a striking and seemingly reliable difference. I doubt it's just randomness but I'm not sure and with hosted AI You can never be sure.
Hugsbox · · focus · HN ↗
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
jswelker · · focus · HN ↗
Hugsbox · · focus · HN ↗
jswelker · · focus · HN ↗
tapoxi · · focus · HN ↗
What the fuck?
Hugsbox · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
CamperBob2 · · focus · HN ↗
Your point about "predicting the next word" mostly means that your post was very easy to predict.
ryandrake · · focus · HN ↗
mitxela · · focus · HN ↗
avereveard · · focus · HN ↗
jswelker · · focus · HN ↗
Me: "You bastard."
Gemini: "Fair callout. I should have been more up front that [has no idea what the fuck it is talking about]."
zerd · · focus · HN ↗
lukan · · focus · HN ↗
georgemcbay · · focus · HN ↗
More scary than funny, IMO.
The US military almost took an action that could very plausibly have escalated into a hot war with China because people are already relying too heavily on these systems.
Despite the reporting, nobody in power seems sufficiently freaked out about this.
<a href="https://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence-china-ship" rel="nofollow">https://www.cnn.com/2026/09/18/politics/us-military-ai-false...
queenkjuul · · focus · HN ↗
Ekaros · · focus · HN ↗
stephenhuey · · focus · HN ↗
lanyard-textile · · focus · HN ↗
The engineers are pressured to significantly reduce "dependencies" for projects. Anything that could become risk or create friction is dramatically less appetizing.
Simply because of how many people that *must* agree with your proposal. Getting all the relevant tech leads, some you have never heard of or ever spoken with, to agree on a proposal for your team's project is a nightmare.
So you keep it as simple and agreeable as possible. Given the circumstances, it makes sense as one of the engineers. It's fairly fine advice in general wherever you work, but it just haunts all the work you do at Google in particular. Nothing gets done otherwise.
If you have a dependency that can be dropped from an engineering perspective, that's the route the 9 leads reviewing your design doc will take:
"Let's iterate and start with just the basics (no user testing)", "let's get this working and user test in a later phase", "I think this problem is obvious enough we don't need to consult with users about it."
I worked in Ads Integrity at the time, and for one of my projects I was concerned how it would impact the manual reviewers. Then I learned I couldn't talk with them, only by proxy through another person if we really had to. And that proxy takes time, so...
stephenhuey · · focus · HN ↗
lanyard-textile · · focus · HN ↗
But without it... :)
stephenhuey · · focus · HN ↗
walrus01 · · focus · HN ↗
Is a major factor of this Google's tendency to kill entire products, so you obviously don't want to make your hot new thing dependent on some other internal thing that might be killed off?
<a href="https://killedbygoogle.com/" rel="nofollow">https://killedbygoogle.com/
lanyard-textile · · focus · HN ↗
Most of the internal infrastructure is quite available and stable, and there are clear choices for almost all of your application-level and backend-level needs.
1718627440 · · focus · HN ↗
darksim905 · · focus · HN ↗
arcanemachiner · · focus · HN ↗
dpc050505 · · focus · HN ↗
There's an enormous difference between outlandish claims about extraterrestrials or turning frog gays and believing there's a category of politicians trying to get wealthy off of their position. It's somewhat reasonable to hypothesize about criminal conspiracies in that 2nd scenario.
wartywhoa23 · · focus · HN ↗
specialist · · focus · HN ↗
Apparently the plan is to use AI slop, mediated thru social medias, to defeat the woke mind virus, perpetuated by the Anti-Christ, in order to safe guard humanity's future.
I wish I was making this up.
dfxm12 · · focus · HN ↗
ytcommentsectio · · focus · HN ↗
[dead]
jodrellblank · · focus · HN ↗
When it says it can’t read videos you think that’s an accurate introspection on its abilities, and not a statistically likely continuation of a conversation where one side seems to be reading videos and the other side says the read is inaccurate?
oofbey · · focus · HN ↗
This is a key reason why I actually like Grok for factual queries based on web grounding. It’s fast and reliable. Maybe it’s ignoring robots.txt? Dunno. But it works well.
hnbad · · focus · HN ↗
Typical example:
"Where can I buy <thing I'm looking for that I can't find anywhere using normal search terms>?"
> You're looking for <related but different and widely available thing>. It is sold by <sites I never heard of>.
"No, that's different. I'm looking for <that thing but with the exact differences spelled out again>."
> Ah, my mistake. You're looking for <thing I described>. It is sold by <sites I never heard of but which don't actually sell it>.
"I've checked your links and none of those sites actually sell it, one doesn't even sell products and instead only offers manufacturing - but also not for what I asked you for."
> I'm sorry, my bad. Those sites don't sell what you are looking for. Instead you should check out <more sites I've never heard of>.
"Those sites sell the thing you initially thought I was asking about but not the thing I described."
> I'm sorry for the misunderstanding. You can find the thing you actually described at <yet more sites including some of the same>.
"No. None of these sites sell anything close to what I asked you for and two of them don't actually exist."
> Oh, sorry about that. You're completely right. The thing you asked me about isn't actually being sold by anyone. However you could buy <thing it first thought I meant and that wouldn't bring me any closer to solving my problem>.
(ad nauseam)
coldfloor · · focus · HN ↗
Most recently, I was considering moving away from iterm2 on MacOS, and I wanted to know if any other terminal emulator supported gestures for switching between tabs. So I asked Gemini, and it says, "Yes Ghostty supports gestures for switching between tabs."
"Ok I just installed Ghostty and I can't find anything about gestures."
"You need to add foo=bar to your conf file."
"I added foo=bar to my conf file and now it's saying the conf file is invalid."
"Sorry bro, remove foo=bar and add baz=boo to the conf file."
"It says baz=boo is invalid too."
"baz=boo isn't a real option. Remove that and add foo=bar to your conf file."
"You already told me to do that and I already told you that doesn't work."
"You shouldn't put foo=bar or baz=boo in the Ghostty conf file. Both are invalid. Ghostty doesn't support gestures. Have you considered iterm2?"
fcarraldo · · focus · HN ↗
Here’s GPT 6’s answer to the prompt “What macOS terminal apps support gestures? Include a reference to the docs on how to enable/configure them.”:
iTerm2 supports configurable three-finger taps and swipes for switching tabs/panes, creating splits, pasting, etc. Set them up under Settings > Pointer > Bindings. Check for conflicting macOS trackpad assignments. [1]
The others are more limited: Ghostty supports macOS lookup/Quick Look gestures [2], while WezTerm lets you bind scroll events—for example, Ctrl+scroll to change font size. [3] Neither is equivalent to iTerm2’s gesture bindings.
For custom gestures without switching terminals, BetterTouchTool can map app-specific trackpad gestures to the terminal’s existing keyboard shortcuts. [4]
[1] <a href="https://iterm2.com/documentation-preferences-pointer.html" rel="nofollow">https://iterm2.com/documentation-preferences-pointer.html
[2] <a href="https://ghostty.org/docs/features#macos" rel="nofollow">https://ghostty.org/docs/features#macos
[3] <a href="https://wezterm.org/config/mouse.html" rel="nofollow">https://wezterm.org/config/mouse.html
[4] <a href="https://docs.folivora.ai/docs/trackpad-mouse/magic-mouse-trackpad/" rel="nofollow">https://docs.folivora.ai/docs/trackpad-mouse/magic-mouse-tra...
numpad0 · · focus · HN ↗
This is a different experience to GP from query to result. I thought they've all fixed that issue of difference in tones affecting results. I guess it was never easily fixed.
1: <a href="https://gist.github.com/numpad0/c40c16232288d544f7ea46521c644161" rel="nofollow">https://gist.github.com/numpad0/c40c16232288d544f7ea46521c64...
jswelker · · focus · HN ↗
selestify · · focus · HN ↗
dcrazy · · focus · HN ↗
nullc · · focus · HN ↗
I take this to mean that my one account has been mistaken for a competitor and they're trying to poison its data. But who knows.
xorcist · · focus · HN ↗
"That's the thing with randomness. You can never be sure."
wavewrangler · · focus · HN ↗
gjm11 · · focus · HN ↗
nullc · · focus · HN ↗
tempaccountabcd · · focus · HN ↗
[dead]
tzs · · focus · HN ↗
[dead]