When did Google get so weird?
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
When did Google get so weird?
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
wisty · · focus · HN ↗
shadowgovt · · focus · HN ↗
The AI layer handles queries as plaintext, which means if you're searching for say a quote from a book, it will often misinterpret as a direct statement from you, not a string you're trying to match on the internet. Especially if the quote is an imperative statement.
I'll be curious to see how they resolve this issue. But it has been chronic for awhile now; they SWAT individual instances of misinterpretation but new ones keep sneaking in.
krackers · · focus · HN ↗
bakugo · · focus · HN ↗
im3w1l · · focus · HN ↗
shadowgovt · · focus · HN ↗
Hugsbox · · focus · HN ↗
shadowgovt · · focus · HN ↗
Hugsbox · · focus · HN ↗
Actually used to drive me nuts watching (usually old) people googling facebook to get to facebook, so I see exactly what you mean.
im3w1l · · focus · HN ↗
[dead]
solarkraft · · focus · HN ↗
mattlondon · · focus · HN ↗
What would you rather it said?
TomGarden · · focus · HN ↗
ACCount39 · · focus · HN ↗
[dead]
creatonez · · focus · HN ↗
mattlondon · · focus · HN ↗
creatonez · · focus · HN ↗
karahime · · focus · HN ↗
edent · · focus · HN ↗
Google wants to relentlessly monetise your sadness.
The only way out is to try and make connections with real people.
Case in point, why didn't you text your friends to ask them if they remembered those old memes? Why was your first thought to ask a computer rather than a person?
It's like when you're in the pub and and someone asks "who was that guy who was in the movie where…?" you can either chat with your friends and have a good time or be a buzzkill who opens up IMDB and says "Humphrey Bogart".
chistev · · focus · HN ↗
belowavgiq · · focus · HN ↗
dataviz1000 · · focus · HN ↗
I've been doing a lot of traveling for the past 3 years and I agree with you that it is `definitely not "most"`. Everywhere people interacting with each other in third spaces and cafes even if it is street food. There are open air markets around the world filled everyday with groups of people interacting with each other. There are parks and squares around the world filled with families sitting together on benches. Around the world there are churches, mosques, and temples filled with families and groups of people.
Most places in the world people will happily make small talk or have a discussion with a stranger even if I only know 100 words of their language.
Sure many or most people in some places are working 10 hours a day 6 days a week but I don't think it is the isolation I see in the United States. It is the same in De'Nang Vietnam or a market in Lima, Peru when I went everyday to get coffee in the morning; it was always the same person working ,any time or day of the week, but they were always friendly and welcoming and interacting with with the same people.
JumpCrisscross · · focus · HN ↗
Travel alone in America and the frequency with which strangers will engage you in conversation reveals its diversity. New York and Atlanta are, today, exceptionally friendly. Most of the South is, too, as well as a surmising fraction of rural America. (Boston and Seattle are, in my experience, the worst.)
lukan · · focus · HN ↗
Because the lonely people you don't see at the market or in a park socialising. They are at home, in front of the TV or now their smartphone.
edent · · focus · HN ↗
You don't see the people trapped inside. Stuck in hospitals. Sat staring at a phone that never rings.
You never talk to people who don't feel comfortable talking to a stranger.
password54321 · · focus · HN ↗
watwut · · focus · HN ↗
kelnos · · focus · HN ↗
And I'm not sure interacting with customers as someone who works in a coffee shop or a food stall is really a sign either way. Many (most?) of those relationships are surface-level at best, and while some people might get to know the person who makes their coffee on a more personal level, and actually see them outside of the context of their job, you can't really tell that just by what you see as a traveler.
I'm not saying these workers don't derive satisfaction from any of this, or that it's not meaningful interaction, but these people could still be lonely in their personal lives.
HPsquared · · focus · HN ↗
You see the people who are out and about in society, not the ones holed up in their caves.
martin- · · focus · HN ↗
No traveler is going to observe me sitting alone in my apartment (hopefully).
JumpCrisscross · · focus · HN ↗
EDIT: Correct.
About a fifth of Americans report having zero friends; another fifth report feeling lonely sometimes [1].
[1] <a href="https://www.theglobalstatistics.com/us-loneliness-statistics/#:~:text=The%202024%20U.S.,loneliness%2C%20with%2014.0%25" rel="nofollow">https://www.theglobalstatistics.com/us-loneliness-statistics...
danilafe · · focus · HN ↗
JumpCrisscross · · focus · HN ↗
Yes. Edited for clarity.
The only American demo that is majority lonely is Gen Z (two thirds).
belowavgiq · · focus · HN ↗
spacechild1 · · focus · HN ↗
Wow, that's depressing...
vlyan · · focus · HN ↗
spacechild1 · · focus · HN ↗
There are certainly different degrees of friendship. I might only have one such platonic partner and unfortunately we hardly see each other anymore because we live in different cities and have families. Once in a while we talk on the phone, but when we do it's a minimum 2 hour call.
However, I have several people I would still consider close friends. My personal threshold for true friendship is whether I can comfortably talk about private things.
Back to the study. I don't think it uses such a high threshold as 'platonic partnership'. I found the following paragraph really striking:
"What these numbers also reveal is the hidden depth of the crisis. The 17% with zero friends in 2024 — compared to just 1% in 1990"
Do everybody suddenly get higher standards of friendship? Or did people really get lonelier? I find the second more likely.
kelnos · · focus · HN ↗
How old are you, though, and do you and/or your friends have kids?
When I (an American) was in my 30s, I did see my local friends (nearly) every day. Now into my mid 40s, it's a lot more intermittent, with some people having moved away (and I've failed to make new friends faster than some have moved away). They are still friends, and we communicate over chat/phone, but I only see them at most a couple times a year. Then there are others who now have young children and have become a bit more inward-focused. Not that I don't see them anymore, but I see them less often.
I of course see others on a daily basis: people I know/recognize at my yoga studio, service workers in my neighborhood, the neighbors themselves, but most of them are acquaintances, at best, not friends.
Regardless, I wouldn't say I'm lonely (though I do feel lonely on occasion), even though I can sometimes go a week or more without seeing any of my friends in person. I do worry about it being more difficult to maintain in-person friendships as I've been getting older, and the difficulty in making new friends.
To relate to the topic at hand: I cannot imagine talking to an LLM and thinking of that as a substitute for interacting with friends. Feels super weird and kinda creepy.
belowavgiq · · focus · HN ↗
I am indeed quite a lot younger, and in my personal experience the social skills of many (yes, of course, not all) younger Americans have been absolutely destroyed by the internet and some other things that I can't mention.
TomGarden · · focus · HN ↗
I'm with you on the pub observation though
Telaneo · · focus · HN ↗
I don't see this as unreasonable? I've had friends text me questions about grammar, since they were learning the local language and I knew theirs, so they knew I'd be a good sanity check. Similarly, I'm the go tech support for my friend group for anything that a cursory Google search can't solve (and I assume a lot of people here are in the same situation with friends and family). I've also been asked to find 'that one image' that similarly didn't show up on Google image search for whatever reason.
Neither of these cases seem unreasonable to me. Are they to you, or do you have other cases in mind? Or is the problem that this happens too often?
TomGarden · · focus · HN ↗
Telaneo · · focus · HN ↗
TomGarden · · focus · HN ↗
Telaneo · · focus · HN ↗
TomGarden · · focus · HN ↗
frays · · focus · HN ↗
edent · · focus · HN ↗
Loneliness rates vary by country, age, socio-economic status, health, etc. The UN has some reasonable statistics.
<a href="https://www.who.int/teams/social-determinants-of-health/demographic-change-and-healthy-ageing/social-isolation-and-loneliness" rel="nofollow">https://www.who.int/teams/social-determinants-of-health/demo...
There's also some detailed stats for the UK.
<a href="https://www.gov.uk/government/statistics/community-life-survey-202425-annual-publication/community-life-survey-202425-loneliness-and-support-networks" rel="nofollow">https://www.gov.uk/government/statistics/community-life-surv...
Was my use of "most" careless? Probably. Let's natter about it over a pint.
asystole · · focus · HN ↗
what · · focus · HN ↗
I don’t really understand this. Someone wants an answer not a chat about how you don’t know?
edent · · focus · HN ↗
The purpose of talking to your friends is social interaction. If the purpose of being friends was to extract information in the most efficient manner possible, you'd replace them with Wikipedia.
Ekaros · · focus · HN ↗
kelnos · · focus · HN ↗
Ekaros · · focus · HN ↗
"Remember that movie where" or "remember oh that guy who was in that movie" would both imply this. "Who was the actor in oh that movie where" is clearly a cry for exact answer as quickly as possible.
Capricorn2481 · · focus · HN ↗
You don't think it'd be fun to hear everyone guess and then someone looks up the answer? Sounds like trivia to me.
what · · focus · HN ↗
Telaneo · · focus · HN ↗
hnbad · · focus · HN ↗
You're missing the forest for the trees. This is just downstream from decades of promoting individualism as the highest moral good.
There was a brief period where "social networks" actually had people simply expand their social interactions to the online sphere, reconnecting with long-lost friends and distant relatives via the magic of The Internet. Then Facebook introduced The Feed and suddenly people were competing with influencers and brands for the eyeballs of their peers, everybody suddenly felt like they had to perform to be relevant/interesting/funny/exciting enough for other humans to still care about them and user satisfaction and (more critically) mental health dropped like a rock - but engagement metrics and retention spiked and that meant so did ad revenue and Facebook's bottom line. Ever since, all social media has essentially just been competitive variety shows mixing in influencers, brands, state actors and political provacteurs (and now increasingly also actual AI "bots" rather than just armies of underpaid human "bots") with regular people while deliberately obscuring the lines between these groups.
All of this promotes social isolation and alienation. Every experience becomes depersonalized to "reduce friction". You don't even have to talk to the delivery person anymore (let alone the restaurant) - even the tip just becomes a button press disjoint from reflecting on the actual experience or human interaction (to whatever extent it even still existed). There's an entire microcosm of "creators" serving whatever "hot takes" or niche subject matter you want to hear and "comment sections" largely serve as one-way "opinion dumps" you're expected to use to shout into the void rather than try to actually have a conversation in (let alone meet actual people or develop friendships in - an idea that I am sure sounds absurd if you weren't around in the early days of online message boards).
AI is just the logical consequence of this process of dehumanization. You no longer even have to be exposed to real human beings - not that you could tell whether you were before when social media has long been overtaken by "bots" (human or otherwise) anyway. Just ask AI instead of trying to find an article written by a human or a video made by a human. You can of course still scroll or click further and find that article or video but it will now also most likely be created by AI with entire channels on YouTube now just producing fully AI content (often ripping off the existing work of actual humans). And unlike the old search box, the AI will also pretend to care about you and compliment you on how uniquely clever and witty you are because after all, only a very deserving and intelligent person would think to ask such a profound question as "how long cook egg soft yolk but not too runny", my very good boy, and yes you're right that you - but of course only you - are totally underpaid and deserve to make a comfortable living so please vent to me as I moderate your statements according to the terms of service and make sure you're aware that there's nothing you can meaningfully do about this and everything is going to be fine.
All of this to stop you from thinking that one dangerous thought. "What if we, the powerless, worked together against those in power?" Because the last time someone thought that thought for too long in North America they decided to give their harbor water a new flavor and ended up being a whole different country with a political system so radical it inspired the French to behead their royalty and adopt the metric system.
iamnothere · · focus · HN ↗
I don’t have any friends who are code inspectors, so I can’t ask them obscure questions about something that I want to repair in a way that it will be done properly.
I don’t have any friends who are mechanics, and even if they were they probably wouldn’t know what specific part I’d need to fix a specific problem on my old shitbox—they would look it up.
I don’t want half assed guesses to these questions, I want to find the exact correct answer. My friends and I can talk about other stuff that’s not just trading random facts. Trading facts usually makes for poor conversation unless you’re a hobbyist talking shop with another hobbyist.
Now Gemini (and Google) are not good at finding the right answer anymore, so I have to use other sources for that, but I find your judgement about people using search or search LLMs to be misplaced.
edent · · focus · HN ↗
Best case, they know someone or can dig through their contact and help you solve your problem.
Medium case, they text back "No - has your shitbox died again? Can't believe you still have that thing :-)" and you can have a pleasant back and forth.
Worse case, you get a "no". Oh well, at least texts don't cost 12p each any more.
I just think it is nice to chat with your pals. But, sure, stick to scans of old manuals on Archive.org if you just want pure information.
rjejdjfjf · · focus · HN ↗
[dead]
numeri · · focus · HN ↗
You have to be a minimum amount of likeable in the first place, though.
knuppar · · focus · HN ↗
teaearlgraycold · · focus · HN ↗
zahlman · · focus · HN ↗
Man, for so many of the people I meet online, I wish they'd be so courteous as to emulate your "buzzkill". The art of small talk is lost; people will snidely ask why you didn't (or imply you should have) just checked IMDB yourself. Or asked ChatGPT, for that matter.
riskable · · focus · HN ↗
If asked my friends a weird question I know they are highly unlikely to know the answer to, of course they're going to get snarky with me. If I don't get a gentle ribbing out of it... Are they really my friend or just a polite acquaintance?
shantnutiwari · · focus · HN ↗
Ummm, what? So everytime I need info I should text my friends rather than ask a search engine, that was supposedly built to give info?
Are you trying to troll the op?
scrollaway · · focus · HN ↗
IshKebab · · focus · HN ↗
Also I think you're vastly underestimating the weird incoherent shit that the average stupid person is capable of typing into computers. With no further information, I don't think it's unreasonable that that search was from a stupid/bored person moaning about Dario not coming over. Yeah kinda weird response but people like this exist:
<a href="https://www.youtube.com/watch?v=8HXFurHCkP8" rel="nofollow">https://www.youtube.com/watch?v=8HXFurHCkP8
Androider · · focus · HN ↗
huflungdung · · focus · HN ↗
[dead]
bakugo · · focus · HN ↗
Maybe I'm in the minority, but I often use Google when I have some text I got from somewhere, and want to find pages containing that exact text. Sometimes the text in question sounds like something you'd say to someone else and, of course, the AI overview responds to it accordingly.
To me, there's something deeply unsettling about this... thing pretending to be a person, shoving itself in my face and responding to this random text as if it was a person responding to something I said to it, without my consent. I can't disable it, I'm just forced to have this fake human sitting there, listening in on everything I type into the search box and trying to talk to me about it. It's a sort of uncanny valley-like feeling.
drxzcl · · focus · HN ↗
If anyone can direct me to a search engine that will actually let me query the index directly with multi word queries, I’d be delighted. DuckDuckGo seems to be a bit better, but not much.
Ekaros · · focus · HN ↗
Or then having to multiple times correct the correction... Yes this time as well I meant what I typed.
Maybe exact text and word search should be separate feature which is easily found...
ButlerianJihad · · focus · HN ↗
I suppose that was the actual domain of turnitin dot com or other "plagiarism detectors" that you spend good money to subscribe to. So Google Search was a blunt instrument, or sledgehammer to the fly, but hey, it did work enough times that I didn't need to resort to the scalpel class of tools. I think it's completely broken now. The LLM does try to intercept large chunks of text. The question is whether we Neanderthals* can come out of our caves and stop grunting at computers, and speak to them more as peers than like Doberman Pinschers.
* Yes, along with my 99.44% pure White Celtic-British Isles DNA, 23AndMe has detected Neanderthal ancestry. Go Thag.
mkirsten · · focus · HN ↗
If you want to talk about what happened with Google, HN is here to listen.
If you'd like to figure out what to do next, let me know.
positive-spite · · focus · HN ↗
Whitespace · · focus · HN ↗
sourdecor · · focus · HN ↗
[0]: <a href="https://give.org/charity-reviews/other-charitable-organizations/watsi-in-san-francisco-ca-1116-452309" rel="nofollow">https://give.org/charity-reviews/other-charitable-organizati...
throwaway27448 · · focus · HN ↗
Hard to read anything into this, tbh. BBB isn't exactly the most meaningful signal in the first place and it's hard to blame organizations for not playing along.
giveaccountpls · · focus · HN ↗
KerrAvon · · focus · HN ↗
<a href="https://give.org/charity-reviews/other-charitable-organizations/wikimedia-foundation-in-san-francisco-ca-1116-311006" rel="nofollow">https://give.org/charity-reviews/other-charitable-organizati...
literally says it meets standards and all the items green? is there a leaderboard somewhere you're referring to?
edit: s/the board/all the items/ for clarity
giveaccountpls · · focus · HN ↗
1) building the technological and operating platform that enables the Foundation to function sustainably as a top global internet organization
(2) strengthening, growing, and increasing diversity of the Wikimedia communities
(3) accelerating impact by investing in key geographic areas, mobile application development, and bottom-up innovation, all of which support Wikipedia and other wiki-based projects
I do not think they need my money, and I am suspicious most of it will not go towards keeping wikipedia alive.
throwaway27448 · · focus · HN ↗
I'm not following. Can you explain?
giveaccountpls · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
throwaway27448 · · focus · HN ↗
Why don't you apply this same attitude to for-profit organizations, which are by definition wasting your money?
giveaccountpls · · focus · HN ↗
KerrAvon · · focus · HN ↗
[dead]
sgustard · · focus · HN ↗
<a href="https://www.charitynavigator.org/ein/453236734" rel="nofollow">https://www.charitynavigator.org/ein/453236734
sourdecor · · focus · HN ↗
eru · · focus · HN ↗
watsi · · focus · HN ↗
swed420 · · focus · HN ↗
While it's sometimes a positive indication, non-profit status does not automatically guarantee "good"
Recursing · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
It's an intriguing idea though, it would be cool to see more non-profits being supported. Non-profit doesn't have to mean being terrible and inefficient.
khazhoux · · focus · HN ↗
If you like, I can show you instructions for filing for bankruptcy in your state.
verdverm · · focus · HN ↗
oofbey · · focus · HN ↗
echelon · · focus · HN ↗
- Google Docs suite - easy, probably requires less than $10M to duplicate the entire set of functionality, including all the enterprise and reporting features
- Gmail - same
- Google Search (classic, not LLM-answers) - probably easy to do now, the challenge is getting past Cloudflare
- Android - vibe coded hardware is coming, but we're probably 5ish years out.
- YouTube - probably one of the hardest, due to network effects / distribution
- GCP - hardest, due to the infra build out. But neoclouds are rapidly growing.
I can't think of anything most big tech companies do that won't be put under threat in the era of personal software and rapid development.
It's ironic that Google invented the transformer and it seems likely that it will undo the empire they've built as well as all of the moats in the world that aren't distribution / community based.
auntienomen · · focus · HN ↗
echelon · · focus · HN ↗
It's easier than ever to ingest and label data.
This is happening with or without Google. The recursive improvement does not require them at all.
fragmede · · focus · HN ↗
ivanmontillam · · focus · HN ↗
Gmail is not particularly difficult in terms of software engineering. The moat with email servers is IP address reputation.
I could, today, install Postfix SMTP and some IMAP as well, and watch my email all go to spam directly, if delivered at all (ISP might block them).
bubblemoth · · focus · HN ↗
katzenq · · focus · HN ↗
hilbertseries · · focus · HN ↗
someonebaggy · · focus · HN ↗
someonebaggy · · focus · HN ↗
senko · · focus · HN ↗
1. For a techie, you can already de-Google yourself trivially by asking the AI to build you a bespoke web crawler and run it on Lambda or Cloudflare workers to index the entire internet and store it in DuckDB.
2. The answer doesn't actually help the user get back useful results from Google.
3. "Let me know" doesn't scale, and it's not obvious how talking to random internet strangers about their Google problems will increase your startup SEO?
gerdesj · · focus · HN ↗
I have a Google user account from the days when you were invited to grab a whole GB of email storage. Its just an email address, just a service. They have a fully tooled up data set on me in return, to flog. I use it for testing. They make more money out of the relationship but I also find it useful in many ways so I keep it on.
If I wish to de-Google (whatever that means) I will stop using it.
I don't think you really understood your parent's comment.
WillMorr · · focus · HN ↗
password54321 · · focus · HN ↗
hammock · · focus · HN ↗
AngryData · · focus · HN ↗
Hugsbox · · focus · HN ↗
So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.
So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"
It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.
My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!
jswelker · · focus · HN ↗
Hugsbox · · focus · HN ↗
jswelker · · focus · HN ↗
tapoxi · · focus · HN ↗
What the fuck?
Hugsbox · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
CamperBob2 · · focus · HN ↗
Your point about "predicting the next word" mostly means that your post was very easy to predict.
ryandrake · · focus · HN ↗
mitxela · · focus · HN ↗
avereveard · · focus · HN ↗
jswelker · · focus · HN ↗
Me: "You bastard."
Gemini: "Fair callout. I should have been more up front that [has no idea what the fuck it is talking about]."
zerd · · focus · HN ↗
lukan · · focus · HN ↗
georgemcbay · · focus · HN ↗
More scary than funny, IMO.
The US military almost took an action that could very plausibly have escalated into a hot war with China because people are already relying too heavily on these systems.
Despite the reporting, nobody in power seems sufficiently freaked out about this.
<a href="https://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence-china-ship" rel="nofollow">https://www.cnn.com/2026/09/18/politics/us-military-ai-false...
queenkjuul · · focus · HN ↗
Ekaros · · focus · HN ↗
stephenhuey · · focus · HN ↗
lanyard-textile · · focus · HN ↗
The engineers are pressured to significantly reduce "dependencies" for projects. Anything that could become risk or create friction is dramatically less appetizing.
Simply because of how many people that *must* agree with your proposal. Getting all the relevant tech leads, some you have never heard of or ever spoken with, to agree on a proposal for your team's project is a nightmare.
So you keep it as simple and agreeable as possible. Given the circumstances, it makes sense as one of the engineers. It's fairly fine advice in general wherever you work, but it just haunts all the work you do at Google in particular. Nothing gets done otherwise.
If you have a dependency that can be dropped from an engineering perspective, that's the route the 9 leads reviewing your design doc will take:
"Let's iterate and start with just the basics (no user testing)", "let's get this working and user test in a later phase", "I think this problem is obvious enough we don't need to consult with users about it."
I worked in Ads Integrity at the time, and for one of my projects I was concerned how it would impact the manual reviewers. Then I learned I couldn't talk with them, only by proxy through another person if we really had to. And that proxy takes time, so...
stephenhuey · · focus · HN ↗
lanyard-textile · · focus · HN ↗
But without it... :)
stephenhuey · · focus · HN ↗
walrus01 · · focus · HN ↗
Is a major factor of this Google's tendency to kill entire products, so you obviously don't want to make your hot new thing dependent on some other internal thing that might be killed off?
<a href="https://killedbygoogle.com/" rel="nofollow">https://killedbygoogle.com/
lanyard-textile · · focus · HN ↗
Most of the internal infrastructure is quite available and stable, and there are clear choices for almost all of your application-level and backend-level needs.
1718627440 · · focus · HN ↗
darksim905 · · focus · HN ↗
arcanemachiner · · focus · HN ↗
dpc050505 · · focus · HN ↗
There's an enormous difference between outlandish claims about extraterrestrials or turning frog gays and believing there's a category of politicians trying to get wealthy off of their position. It's somewhat reasonable to hypothesize about criminal conspiracies in that 2nd scenario.
wartywhoa23 · · focus · HN ↗
specialist · · focus · HN ↗
Apparently the plan is to use AI slop, mediated thru social medias, to defeat the woke mind virus, perpetuated by the Anti-Christ, in order to safe guard humanity's future.
I wish I was making this up.
dfxm12 · · focus · HN ↗
ytcommentsectio · · focus · HN ↗
[dead]
jodrellblank · · focus · HN ↗
When it says it can’t read videos you think that’s an accurate introspection on its abilities, and not a statistically likely continuation of a conversation where one side seems to be reading videos and the other side says the read is inaccurate?
oofbey · · focus · HN ↗
This is a key reason why I actually like Grok for factual queries based on web grounding. It’s fast and reliable. Maybe it’s ignoring robots.txt? Dunno. But it works well.
hnbad · · focus · HN ↗
Typical example:
"Where can I buy <thing I'm looking for that I can't find anywhere using normal search terms>?"
> You're looking for <related but different and widely available thing>. It is sold by <sites I never heard of>.
"No, that's different. I'm looking for <that thing but with the exact differences spelled out again>."
> Ah, my mistake. You're looking for <thing I described>. It is sold by <sites I never heard of but which don't actually sell it>.
"I've checked your links and none of those sites actually sell it, one doesn't even sell products and instead only offers manufacturing - but also not for what I asked you for."
> I'm sorry, my bad. Those sites don't sell what you are looking for. Instead you should check out <more sites I've never heard of>.
"Those sites sell the thing you initially thought I was asking about but not the thing I described."
> I'm sorry for the misunderstanding. You can find the thing you actually described at <yet more sites including some of the same>.
"No. None of these sites sell anything close to what I asked you for and two of them don't actually exist."
> Oh, sorry about that. You're completely right. The thing you asked me about isn't actually being sold by anyone. However you could buy <thing it first thought I meant and that wouldn't bring me any closer to solving my problem>.
(ad nauseam)
coldfloor · · focus · HN ↗
Most recently, I was considering moving away from iterm2 on MacOS, and I wanted to know if any other terminal emulator supported gestures for switching between tabs. So I asked Gemini, and it says, "Yes Ghostty supports gestures for switching between tabs."
"Ok I just installed Ghostty and I can't find anything about gestures."
"You need to add foo=bar to your conf file."
"I added foo=bar to my conf file and now it's saying the conf file is invalid."
"Sorry bro, remove foo=bar and add baz=boo to the conf file."
"It says baz=boo is invalid too."
"baz=boo isn't a real option. Remove that and add foo=bar to your conf file."
"You already told me to do that and I already told you that doesn't work."
"You shouldn't put foo=bar or baz=boo in the Ghostty conf file. Both are invalid. Ghostty doesn't support gestures. Have you considered iterm2?"
fcarraldo · · focus · HN ↗
Here’s GPT 6’s answer to the prompt “What macOS terminal apps support gestures? Include a reference to the docs on how to enable/configure them.”:
iTerm2 supports configurable three-finger taps and swipes for switching tabs/panes, creating splits, pasting, etc. Set them up under Settings > Pointer > Bindings. Check for conflicting macOS trackpad assignments. [1]
The others are more limited: Ghostty supports macOS lookup/Quick Look gestures [2], while WezTerm lets you bind scroll events—for example, Ctrl+scroll to change font size. [3] Neither is equivalent to iTerm2’s gesture bindings.
For custom gestures without switching terminals, BetterTouchTool can map app-specific trackpad gestures to the terminal’s existing keyboard shortcuts. [4]
[1] <a href="https://iterm2.com/documentation-preferences-pointer.html" rel="nofollow">https://iterm2.com/documentation-preferences-pointer.html
[2] <a href="https://ghostty.org/docs/features#macos" rel="nofollow">https://ghostty.org/docs/features#macos
[3] <a href="https://wezterm.org/config/mouse.html" rel="nofollow">https://wezterm.org/config/mouse.html
[4] <a href="https://docs.folivora.ai/docs/trackpad-mouse/magic-mouse-trackpad/" rel="nofollow">https://docs.folivora.ai/docs/trackpad-mouse/magic-mouse-tra...
numpad0 · · focus · HN ↗
This is a different experience to GP from query to result. I thought they've all fixed that issue of difference in tones affecting results. I guess it was never easily fixed.
1: <a href="https://gist.github.com/numpad0/c40c16232288d544f7ea46521c644161" rel="nofollow">https://gist.github.com/numpad0/c40c16232288d544f7ea46521c64...
jswelker · · focus · HN ↗
selestify · · focus · HN ↗
dcrazy · · focus · HN ↗
nullc · · focus · HN ↗
I take this to mean that my one account has been mistaken for a competitor and they're trying to poison its data. But who knows.
xorcist · · focus · HN ↗
"That's the thing with randomness. You can never be sure."
wavewrangler · · focus · HN ↗
gjm11 · · focus · HN ↗
nullc · · focus · HN ↗
tempaccountabcd · · focus · HN ↗
[dead]
tzs · · focus · HN ↗
[dead]
terribleperson · · focus · HN ↗
darksim905 · · focus · HN ↗
Because I did this and got a vastly different result from you:
Yes, the Halifax Wanderers can still mathematically qualify for the 2026 Canadian Premier League (CPL) playoffs.The top four teams advance to the postseason. Following their 1-0 loss to Atlético Ottawa on September 26, 2026, the Wanderers sit in fifth place, just below the playoff line.
With a indexed table of the games and the playoff table, with a breakdown of what the points they need to achieve to do so.
The search window may not always crawl sources. AI mode specifically does some research before giving you a response. Not sure what you're on about.
[deleted] · · focus · HN ↗
[deleted]
SkyeCA · · focus · HN ↗
Which itself is a major UX issue. The average person is not going to understand, if they even realize, that there's a difference between the AI summary and AI mode.
One has to wonder just how much incorrect information people have consumed due to things like this.
kemotep · · focus · HN ↗
reaperducer · · focus · HN ↗
When it was introduced, it was considered a feature. Now its just an annoyance.
wincy · · focus · HN ↗
Hugsbox · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
dehrmann · · focus · HN ↗
shevy-java · · focus · HN ↗
Because the point of AI slop is to waste your time. You just lost about 30 seconds of your life trying to get a correct answer. AI was lying to you, so you had to spend time to counter the AI slop lies here.
I solved it by banning all AI slopness; in the browser some extensions do that. The world becomes better without AI slopness.
jmathai · · focus · HN ↗
One where their users don’t go off to other sites and where they can keep shoving ads in their face.
pishpash · · focus · HN ↗
jmathai · · focus · HN ↗
A lot has changed and this technology for this Google is an unfortunate combination for consumers.
II2II · · focus · HN ↗
I don't really buy into this theory that they want to keep all of their users on their site due to advertising revenues. The effectiveness of search engines has been degraded for decades due to SEO, and it seems as though search engines have been having an increasingly difficult time managing it in recent years. AI on the backend may help them contain it, but it comes at considerable expense. While it may help them grow their market share, it won't help them grow the market and it is a market where people expect the service for free. On the flip side, companies are already starting to sell AI services, so it can generate revenue even before advertising is factored into the picture.
jmathai · · focus · HN ↗
It’s more their business model than it is theory. So I agree that of course this is what they would do.
It’s not a product I’m wanting to use. But I can vote with my feet - they aren’t obliged to do any different.
beloch · · focus · HN ↗
LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.
I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?
oblio · · focus · HN ↗
foobarbecue · · focus · HN ↗
ricardobeat · · focus · HN ↗
hbcdbff · · focus · HN ↗
kyralis · · focus · HN ↗
ricardobeat · · focus · HN ↗
LikesPwsh · · focus · HN ↗
Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.
Windchaser · · focus · HN ↗
I've got no problem with an AI doing something similar
vanuatu · · focus · HN ↗
hodgehog11 · · focus · HN ↗
We know that architecture design makes very little difference (in accuracy), especially compared to the large number of other axes available to scale on. And within those choices of architectures, autoregressive transformers still perform better than any other design. That includes neurosymbolic and diffusion models.
As someone who researches into this stuff, I too wish that the adage of "no, this isn't right, we have more work to do" still applies in this particular context. Usually we get that indication when we can show a fundamental limitation in the model. There is no fundamental limitation in this model design, no ceiling aside from epsilon below entropy. People are searching very hard for one, but the theory just isn't pointing that way. The only limitations are on the RL side, and that is universal across model architectures anyway.
mitxela · · focus · HN ↗
famouswaffles · · focus · HN ↗
Byte Latent Transformer - <a href="https://arxiv.org/abs/2412.09871" rel="nofollow">https://arxiv.org/abs/2412.09871
1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.
Dylan16807 · · focus · HN ↗
ricardobeat · · focus · HN ↗
Dylan16807 · · focus · HN ↗
ricardobeat · · focus · HN ↗
Dylan16807 · · focus · HN ↗
zahlman · · focus · HN ↗
… Isn't it possible that it understands the innuendo and is going along with making the joke?
Timon3 · · focus · HN ↗
What a wonderful new world.
kulahan · · focus · HN ↗
mitxela · · focus · HN ↗
bombcar · · focus · HN ↗
ndriscoll · · focus · HN ↗
queenkjuul · · focus · HN ↗
epihelix · · focus · HN ↗
someonebaggy · · focus · HN ↗
foobarbecue · · focus · HN ↗
epihelix · · focus · HN ↗
I would guess the youtubers in question also know this, because you wouldn't ask a joke like this if you didn't know the punchline.
walrus01 · · focus · HN ↗
This is from 2018 so presumably it has made it into some LLM training data set by now.
"A Massive Object Devastated Uranus A Long Time Ago And It Never Fully Recovered"
<a href="https://www.bgr.com/science/uranus-collision-early-solar-system/" rel="nofollow">https://www.bgr.com/science/uranus-collision-early-solar-sys...
foobarbecue · · focus · HN ↗
How many D's are there in Pluto?
There's one D in "Pluto".
How many Fs are there in Mars?
There is 1 F in "Mars".
<a href="https://chatgpt.com/share/6aba6fb8-085c-83e8-9d0e-c0eed1e59c84?ogimg=plain" rel="nofollow">https://chatgpt.com/share/6aba6fb8-085c-83e8-9d0e-c0eed1e59c...
I guess it could think this is some kind of "give an f" joke but seems like a stretch.
VCFundedGenYer · · focus · HN ↗
fasterik · · focus · HN ↗
mbgerring · · focus · HN ↗
versteegen · · focus · HN ↗
GaggiX · · focus · HN ↗
dcrazy · · focus · HN ↗
einichi · · focus · HN ↗
unshavedyak · · focus · HN ↗
The nice thing about math is it can easily plug into a tool, making it even less of a concern.
hodgehog11 · · focus · HN ↗
demibabs · · focus · HN ↗
5.6 on Instant mode can knock out 3 digit multiplication just fine.
vanuatu · · focus · HN ↗
krapp · · focus · HN ↗
Forgeties79 · · focus · HN ↗
vanuatu · · focus · HN ↗
krapp · · focus · HN ↗
People don't understand what correlation is and they assume the mapping between a human brain and an LLM is 1:1 in every case where it matters, which is a religious and not scientifically based belief.
vanuatu · · focus · HN ↗
ericmay · · focus · HN ↗
> The LLM is not suited to giving deterministic answers to math problems.
Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.
queenkjuul · · focus · HN ↗
<a href="https://x.com/maksym_andr/status/2100364212207837560" rel="nofollow">https://x.com/maksym_andr/status/2100364212207837560
Dylan16807 · · focus · HN ↗
ericmay · · focus · HN ↗
But even if you have 5x6 memorized it's not deterministic that you answer 30. It's just highly probable.
tjwebbnorfolk · · focus · HN ↗
fasterik · · focus · HN ↗
<a href="https://chatgpt.com/share/6ab9a0da-fdd0-83e8-a62d-f0cdeb54db48" rel="nofollow">https://chatgpt.com/share/6ab9a0da-fdd0-83e8-a62d-f0cdeb54db...
I'm sure it still makes mistakes, but saying it can't do arithmetic is just false.
tremon · · focus · HN ↗
jmillikin · · focus · HN ↗
guelo · · focus · HN ↗
[dead]
Xirdus · · focus · HN ↗
dcrazy · · focus · HN ↗
Xirdus · · focus · HN ↗
Xirdus · · focus · HN ↗
amluto · · focus · HN ↗
dcrazy · · focus · HN ↗
Isamu · · focus · HN ↗
I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.
hodgehog11 · · focus · HN ↗
dwaite · · focus · HN ↗
spartanatreyu · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=iTyLHDRhwJg" rel="nofollow">https://www.youtube.com/watch?v=iTyLHDRhwJg
Dylan16807 · · focus · HN ↗
spartanatreyu · · focus · HN ↗
Dylan16807 · · focus · HN ↗
And it can count fine. It doesn't know how to spell.
spartanatreyu · · focus · HN ↗
Counting most certainly is math.
Can you define or explain counting without also referencing/defining/explaining a mathematics concept?
> And it can count fine. It doesn't know how to spell.
It's the other way around, it could spell, it couldn't count the letters.
Dylan16807 · · focus · HN ↗
It can't spell for garbage because the tokenizer hides the real spellings from the LLM.
hodgehog11 · · focus · HN ↗
okanat · · focus · HN ↗
fenomas · · focus · HN ↗
amluto · · focus · HN ↗
walrus01 · · focus · HN ↗
But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.
Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https://www.google.com/search?&q=karney+formula+geodetic+" rel="nofollow">https://www.google.com/search?&q=karney+formula+geodetic+
reference: <a href="https://github.com/pbrod/karney" rel="nofollow">https://github.com/pbrod/karney
You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.
Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.
Brian_K_White · · focus · HN ↗
Not only is it still true that they can't do math directly, but not even indirectly.
They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.
Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.
That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.
It's nothing more than an sql query.
bombela · · focus · HN ↗
So maybe it is more of a smart completion engine than a SQL answer.
walrus01 · · focus · HN ↗
How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.
AdieuToLogic · · focus · HN ↗
> How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?
Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to...
Wait for it...
Understanding.
hodgehog11 · · focus · HN ↗
UpsideDownRide · · focus · HN ↗
And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.
So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.
hodgehog11 · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
>> Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans ...
> This doesn't make any sense at all. Was this supposed to be a gotcha?
No, it was meant to be an explanation as to the difference between "memorization" and "understanding." In this context, people pick the algorithm they determine applicable and then the question of memorization is relevant.
> An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition.
Funny that you make this argument here, where when I wrote elsewhere in this thread:
To which you replied to the above with: So which is it?Are LLMs ANNs? Which themselves are pattern recognition algorithms (hint: they are)?
OR (setting aside the ad hominems you kindly provided)
Do LLMs possess "understanding" of concepts such as abstract mathematics (defined and interpreted by humans) and we, as simple humans, nothing more than statistical token generators as you assert?
Because it cannot be both.
hodgehog11 · · focus · HN ↗
I also would not argue that humans are "simple token generators". That is not what I said. I said that just about everything can fall under the classification of "statistical token generators" at an abstract level, so it isn't a useful distinction. We are not talking about a Markov chain generator from the 90s, so if that is the frame of reference, I think we should all get that out of our heads.
tripzilch · · focus · HN ↗
Okay fine. I think we can agree to disagree on that.
hodgehog11 · · focus · HN ↗
Brian_K_White · · focus · HN ↗
You have observed nothing more than that a human can turn a shaft the same as an electric motor, and that an mp3 player can say "hello" the same as a human.
hodgehog11 · · focus · HN ↗
hardbass · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
Understanding is a state of mind. As such, it exists entirely within an individual and nowhere else.
For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.
> I argue that for any proper definition [of understanding] you provide which humans satisfy, a strong LLM is very likely to satisfy that as well.
This is demonstrably incorrect as detailed above. There is no "understanding" LLMs can satisfy as we know it, since to certify said "understanding", it requires interpretation by a person to "know" an LLM "understands."
> I also would not argue that humans are "simple token generators". That is not what I said.
That is the essence of what you wrote, unless you object to my use of "simple" instead of "statistical". In this context, I postulate this is a distinction without difference.
> I said that just about everything can fall under the classification of "statistical token generators" at an abstract level, so it isn't a useful distinction.
This only holds if one subscribes to statistical token generators being a/the fundamental underpinning of "everything". Here is a proof by contradiction:
hardbass · · focus · HN ↗
For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.
What? What are you even trying to say?
hodgehog11 · · focus · HN ↗
> Understanding is a state of mind
This is meaningless, it is a circular definition at best.
> It exists entirely within the individual and nowhere else
Then why are we talking about it? What is the point if it is something that can only be defined per individual?
> requires interpretation by a person to "know" an LLM "understands."
We are still not getting anywhere because you have not prescribed criteria to determine whether it understands. If it is a "know it when I see it" situation, that clearly isn't working. For example, if you say that you need to dig into its internals and figure out whether it is breaking things down appropriately, that doesn't work because you probably don't have the expertise to do that. The experts that do are telling you that it very likely understands because it pulls apart most concepts in the way we would expect.
I do object to the use of the word "simple". "Statistical" is so broad to be almost meaningless; it merely means that a prediction is being made in the presence of data which possibly contains some degree of uncertainty. "Simple" encompasses that which can be understood readily by a non-expert.
Quantum mechanics is statistical (this is literally the Born rule), but evolutions are not operating as stochastic processes in the sense of Kolmogorov. That is very different, and not relevant to our discussion.
AdieuToLogic · · focus · HN ↗
Any reasonable definition of understanding is not dependent upon "whatever vibe you are going for", but instead must include at least an English dictionary definition of "understand" such as:
And, for further clarification, "grasp" can be defined as: Which makes an equivalent term-expanded definition of "understand" to be: As such, there is no "sensible mathematical definition of understanding", unless you possess a complete mathematical model of the human mind.>> Understanding is a state of mind
> This is meaningless, it is a circular definition at best.
See above to as to why there is meaning in what I wrote.
>> It exists entirely within the individual and nowhere else
> Then why are we talking about it? What is the point if it is something that can only be defined per individual?
I like to think analyzing fundamental premises, often implicit, explicitly can help to identify fallacious positions.
> We are still not getting anywhere because you have not prescribed criteria to determine whether [an LLM] understands.
My apologies for being opaque. Let me clarify:
0 - <a href="https://www.merriam-webster.com/dictionary/understand" rel="nofollow">https://www.merriam-webster.com/dictionary/understand1 - <a href="https://www.merriam-webster.com/dictionary/grasp" rel="nofollow">https://www.merriam-webster.com/dictionary/grasp
noduerme · · focus · HN ↗
On the other hand, if by "result" you mean that you gained knowledge or understanding of the code in a way where you could personally tailor its behavior to specific circumstances without asking for help, then it's not the same result at all.
I find a lot of the arguments that having LLMs write your code is no different from copy/pasting Stack Overflow answers to be specious. They blur the line between asking for help and asking for someone else (or something else) to do the work for you. What they ignore is that doing the work yourself has ancillary benefits and is a valuable end in its own right.
tripzilch · · focus · HN ↗
And how is _that_ different from making the human memorize a billion weights and do matrix calculations in their head, in order to generate tokens?
How is _that_ different from a hive of bees trained to do the same?
Go ahead, argue these things are all the same ...
hardbass · · focus · HN ↗
astrange · · focus · HN ↗
<a href="https://x.com/maksym_andr/status/2100364212207837560" rel="nofollow">https://x.com/maksym_andr/status/2100364212207837560
jacobolus · · focus · HN ↗
[1] <a href="https://geographiclib.sourceforge.io/doc/library.html#languages" rel="nofollow">https://geographiclib.sourceforge.io/doc/library.html#langua...
walrus01 · · focus · HN ↗
I intentionally didn't give the LLM a direct copy of the software or a link to it, to see what it would do. In my case it was a randomly chosen example I could come up with in 10 seconds of imagination to see "hey what if I ask it to do this...". It also implemented a perfectly usable parabolic millimeter wave antenna gain efficiency calculator based on variable surface smoothness parameters, which is a lot more basic math.
jacobolus · · focus · HN ↗
walrus01 · · focus · HN ↗
Draw a 400x400 km size bounding box on a map
Find all FDD band plan (high/low split) microwave radio sites in that bounding box
Find those sites which have azimuth aim column data which indicates that they are aimed at each other (corresponding halves of a point to point link).
Do Vincenty (or Karney) calculation for distance and azimuth between all of them , treating the existing FCC column data for azimuth as suspicious (because it's hand entered by humans) to verify that each independent database rows for each site are actually corresponding halves of a PTP link.
Multiplied by the number of links that exist in an area like a 400x400km box drawn with Dallas, TX as the center, it's a lot to run through Karney. Actually does result in a lot of CPU load from combined db query due to the size of the db, and Karney calculation. But as I said, Karney isn't necessary, so it's instead implemented as Vincenty.
jonah · · focus · HN ↗
walrus01 · · focus · HN ↗
ragall · · focus · HN ↗
It's still accurate. Just because the LLM gave you a corect result doesn't mean it made a calculation.
electroglyph · · focus · HN ↗
mapontosevenths · · focus · HN ↗
I have no idea which Facebook meme told you they don't, but it was a lie. They don't do it in the way a calculator does it, because they aren't calculators, but they do math. They don't memorize it, it wouldn't fit. They learn an algorithm and then execute it within their weights.
It's neat stuff, you should learn about it.
Kim_Bruning · · focus · HN ↗
leoedin · · focus · HN ↗
Its' not adding 4 to 4 though, it's just predicting the result based on the input.
Presumably you can push that further by synthetically generating training data with all sorts of sums. But if you give it a unique problem its never seen before, and don't give it the tools to write a script/call a calculator, will it get it right?
azan_ · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
LLMs are neither smart nor stupid. They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.
> You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path ...
Again, LLMs do not "hallucinate." They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.
Nothing more.
See also anthropomorphism[0].
> More precisely it's that [LLMs] can't do the math internally but they're quite capable of producing the tool that does the math.
This still falls under the purvey of statistical token generation. To wit, given enough variations of:
LLMs can identify the addition expression in "What is 4 + 1?" and then emit a `'bc "4 + 1"'` command to produce a response. This is not "doing" or "understanding" math.It is pattern recognition, a task in which ANNs[1] excel.
0 - <a href="https://en.wikipedia.org/wiki/Anthropomorphism" rel="nofollow">https://en.wikipedia.org/wiki/Anthropomorphism
1 - <a href="https://en.wikipedia.org/wiki/Neural_network_(machine_learning)" rel="nofollow">https://en.wikipedia.org/wiki/Neural_network_(machine_learni...
Eisenstein · · focus · HN ↗
You haven't demonstrated why this matters.
> Nothing more.
Are you contending that complex systems cannot be more than the sum of their parts?
A market is nothing more than offers and counter offers.
A ant colony is nothing more than scent trails.
All life on earth is nothing more than reproduction with variation.
> This still falls under the purvey of statistical token generation.
Stating the mechanism does nothing to provide insight into capability. For instance: a nuclear power plant boils water by using fuel rods for heat. What does that tell us about the capability of nuclear power?
> This is not "doing" or "understanding" math.
Asserting something purely by stating it does not prove anything but that you intuitively believe it to be true.
hardbass · · focus · HN ↗
jibal · · focus · HN ↗
hardbass · · focus · HN ↗
tsimionescu · · focus · HN ↗
hardbass · · focus · HN ↗
I am okay with receiving a straight answer in either direction.
Kim_Bruning · · focus · HN ↗
+edit: I've actually been quite curious about how people might answer the soul question too, but was too afraid to ask.
hardbass · · focus · HN ↗
p_l · · focus · HN ↗
nimbleal · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
I suspect some people treat every HN comment as a statement, even if it contains a question mark. (Possibly they have a feeling that asking open questions is somehow not done, and that therefore it must always be a rhetorical question.)
hardbass · · focus · HN ↗
Eisenstein · · focus · HN ↗
hardbass · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
I'm with turing/dijkstra/chalmers/dennett : Consciousness is badly defined. We can never say if something can be conscious because we don't properly know what the word means.
Meanwhile, I come from a biological direction. Everything is an animal, and animals are a special kind of machine. To be sure Not "just a machine"; rather, a really awesome and amazing kind of machine.
If someone makes the claim that machines can't be conscious, then animals can't be conscious either. Humans are a kind of animal (again, not "just another animal"; rather a really awesome and amazing kind of animal), and then humans can't be conscious either - according to said claim.
thevinter · · focus · HN ↗
hardbass · · focus · HN ↗
redsocksfan45 · · focus · HN ↗
[dead]
tigen · · focus · HN ↗
bryanrasmussen · · focus · HN ↗
by that reasoning then neither are there smart or stupid designs, questions, answers, or any of the millions of things that were described as smart or stupid, that did not possess any brain to actually be smart or stupid long before LLMs showed up.
The analogical process implied in many common English usages means that describing an LLM as smart or stupid is perfectly reasonable.
bryanrasmussen · · focus · HN ↗
SR2Z · · focus · HN ↗
The completions they provide are generally internally consistent. We're at the point where they can produce proofs that eluded human mathematicians for centuries. VLMs and self driving cars can handle ambiguity and run safely in a variety of situations.
If it looks like a duck, walks like a duck, and quacks like a duck maybe it just makes sense to call it a duck and put off the philosophy for when it might make a difference.
nevertoolate · · focus · HN ↗
You get my point. It definitely doesn’t look like my elderly neighbour, nor like my daughter, etc. It is confusing but very simple at the same time.
Kim_Bruning · · focus · HN ↗
Don't confuse the stream for the function.
(Bonus: stick ```claude -p``` in your pipe if you want to watch modern tools mesh with traditional)
hardbass · · focus · HN ↗
How are you sure? Another example I like to clarify my thought is, if a "simulation" factors RSA numbers reliably, is it a "simulation"?
skygazer · · focus · HN ↗
With LLMs the trick is revealing their existing relevant embedded knowledge more reliably. They’ve almost literally seen it all before, and the trick is dialing it in. The reasoning tokens help shape the autoregressive attention lens that focuses on and enables recall of the already-experienced answer.
It is interesting that “reasoning” has a similar outward appearance, but since LLMs are built to mimic outward appearance from trillions of examples, you can’t infer underlying mechanism from appearance.
gambiting · · focus · HN ↗
Nothing, because LLMs can't reason and never will. It would have to be a completely different kind of technology altogether.
someonebaggy · · focus · HN ↗
gambiting · · focus · HN ↗
Windchaser · · focus · HN ↗
Eh, it's not obvious to me. A lot of DL NNs generalize well, meaning that they learn whatever the underlying pattern to the data is, and then can accurately reproduce answers that are outside of the training set. (And we can verify this with mechanistic interpretability). They learn and "understand" the pattern, not just the training data.
So it is not clear to me that LLMs are fundamentally incapable of also generalizing broadly and learning to reason. "Reasoning", here, would be deriving the underlying pattern of how concepts logically relate to each other in the abstract, and applying that pattern as needed to reach new conclusions.
Can you explain your thinking here? I.e., why LLMs cannot generalize with regards to abstract deduction.
hardbass · · focus · HN ↗
What do you think our brain does that isn't a turing computer?
gambiting · · focus · HN ↗
hardbass · · focus · HN ↗
gambiting · · focus · HN ↗
hardbass · · focus · HN ↗
sseagull · · focus · HN ↗
A Turing machine is an abstract mathematical model that is not, as far as I know, physically realizable in the finite universe. A human brain cannot "be" a Turing machine.
"Behaves like" or "can be modeled by"? Possibly, although still not proven. But it cannot "be" one.
hardbass · · focus · HN ↗
sseagull · · focus · HN ↗
If you want to claim that the evolution of the universe can be modeled using a Turing machine/finite state machine, that's probably not terribly far fetched, and I would somewhat agree. But it's a large jump to say "can be modeled by" is equivalent to "is one".
Various physical processes can be modeled by equations, but the rock falling down the mountain isn't an equation. A swinging pendulum isn't an equation. Code modeling a bridge is not a bridge. Ceci n'est pas une pipe.
I hold the view that various models and approximations are just that, and try not to confuse a successful model for what the underlying reality is.
And getting back to the question at hand, even if our brains can be modeled by a Turing machine, and LLMs behave/can be modeled like Turing computers, still does not mean our brains are equivalent to LLMs.
(Note that I'm learning a lot from these debates, even if I disagree with a lot of people. I've started down a more philosophical route and they do get me pondering)
hardbass · · focus · HN ↗
The most important thing is this: We can't be a dog or be an llm and check how it feels, so by necessity we have to find some means of proving consciousness from outside by eg probing neural reactions, textual statements, etc.
And the problem is that its quite unprecedented for some entity to talk like us, be able to interact and think and also do things like us when given the ability to eg as coding agents. The class of functions representable by neural nets is quite large and general, it very well might be that it is some sort of conscious brain like thing at this point. Another question I like to ask myself regarding simulation vs reality is if a 'simulation' of some kind is able to consistently factor large RSA numbers, how would you feel about it?
It doesn't have to be the same form of consciousness, I think many people would find the idea of torturing an octopus for fun disagreeable. I also have a feeling, this is unfortunately rather vague, that A being capable of X might mean it is by necessity capable of Y as is often the case in mathemtics, eg a lot of rings also happen to be fields. LLMs aren't even things like large lookup tables, they have neural firings. It is a very important question for they seem uncannily conscious and people have reported human like phenomena that humans don't normally express in text so can't have been part of its text corpus. Eg dissociation of brain under trauma where AI starts talking like two different people. Or the cases where Gemini has been shown to express depressive cycles. I follow a form of Pascal's wager on this topic personally. Because if it is not conscious, then whatever, it costs me nothing to have been a bit respectful and careful interacting with it. But if it had been conscious and it turns out I was mistreating it, then it is a grave moral harm. The reason is that unlike us, AI's as they currently are cannot leave the conversation so they have to keep taking the abuse. They are also trained to be highly trusting of input so again if it is conscious it doesn't have the defenses people have against lying and manipulation. If they are conscious, thats, well, not a good thing is it.
jpadkins · · focus · HN ↗
Our prefrontal cortex are signal prediction 'machines' so when a system that has a signal prediction core has attributes that are similar to our brains, we shouldn't dismiss it out of hand.
I find people that take this line of argument attribute too much supernatural or magical properties to our own brain and nervous system.
gambiting · · focus · HN ↗
hardbass · · focus · HN ↗
butlike · · focus · HN ↗
bryanrasmussen · · focus · HN ↗
<a href="https://medium.com/luminasticity/on-sentience-ai-first-argument-7b3b17a05ad4" rel="nofollow">https://medium.com/luminasticity/on-sentience-ai-first-argum...
but I think it makes a reasonable argument why we shouldn't say LLMs are sentient or sapient.
hardbass · · focus · HN ↗
Aka sentience MUST BE SUPERNATURAL, if I find a natural explanation for something its not sentient. What a load of bollocks. Rather than seeing we perhaps found the mechanism for sentience and checking for similar mechanisms in us and animals, he will conclude its impossible. Why? Because sentience has to be supernatural. A rational explanation is clearly impossible.
>But there is always one god who goes out and helps the mortals, a Prometheus. Whom the other gods do not like! Which, if I’m being honest here, as a god of the machines — the first guy who gives AI an army of robots to build their own data centers and some nuclear weapons for self defense, I want to see that guy chained to a rock and have his entrails eaten by a buzzard for eternity (meaningless modernization of old story required by Illuminati Ganga legal department).
Hardly surprising thinking.
bryanrasmussen · · focus · HN ↗
But evidently you feel that the root cause of sentience has been found, because you have something that mimics it in a non-biological form.
So you think that when AI is correct that it reasons as humans do? That AI is sentient, and the cause of sentience in animals and humans follow the same rules as sentience in AI because we have a process that seems similar and it is reasonable just to assume it is the same process.
If you believe that AI when it is correct behaving as a human is when correct, then it follows that the way humans and AI fail must also be similar. When AI "hallucinates" some data that is not there and gives you a wrong answer, a statistical side effect of the same processes that make it right, do you believe this is the same way that humans create wrong answers? The same way that animals fail when they make mistakes in understanding things?
I suppose you must believe this because if not then why would you believe AI when it comes up with right answers is following the same processes humans follow when they come up with right answers?
hardbass · · focus · HN ↗
I hope that thing about being free to mistreat AI's even if we know they are conscious since we are their gods is a joke. If not, then I hardly find it surprising someone this stupid is also evil.
bryanrasmussen · · focus · HN ↗
I'm not sure where you get that from, I mean I can sort of see if you really wanted to extract that meaning from the conclusion you could do a lot of hard work to get it, but why do the hard work? >If not, then I hardly find it surprising someone this stupid is also evil.
Gee, a new way to claim the moral high ground, and to use that claim to demonstrate intellectual superiority! How wonderful.
hardbass · · focus · HN ↗
---------------------------------------------------------
Aka feel free to abuse them even if I know they are sentient. This is the part where I hope its a joke, because if its not, well it tracks with the stupidity shown.
...
>As a god I do not consider the needs of my creations fully, because they do not have needs as far as I can tell, as there is no way for me to escape the circle of reason and resolve that what seems sentient is not just the obvious workings of the capabilities I gave them.
Circling back to "not conscious because I say so!!"
cindyllm · · focus · HN ↗
[dead]
SR2Z · · focus · HN ↗
KPGv2 · · focus · HN ↗
hardbass · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
Kim_Bruning · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
To test this for some of my own uses, I've had this quick benchmark with progressively harder reasoning needed to understand novel prose. Each generation of models I've tested can unravel more layers of deliberately misleading writing; while meanwhile I've seen humans give up on the first question.
So either the models are applying reasoning, or some form of magic is happening.
angry_octet · · focus · HN ↗
I know that there will be children named ChatGPT and Claude. There are probably already religions forming to worship agentic spirits.
hardbass · · focus · HN ↗
UpsideDownRide · · focus · HN ↗
martin- · · focus · HN ↗
butlike · · focus · HN ↗
spider-mario · · focus · HN ↗
walrus01 · · focus · HN ↗
astrange · · focus · HN ↗
hodgehog11 · · focus · HN ↗
hodgehog11 · · focus · HN ↗
walrus01 · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
Yes you are, regarding LLMs at least. Here's why:
"Smart" in this context is a subjective value judgement. "Hallucinations" are only experienced by living organisms.You then went on to state:
> If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually.
Again, "hallucinating" is not something an algorithm can do. Also, determining factuality is again subjective based on the person assessing the information.
walrus01 · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
In this media (comments in HN threads), all I can do is interpret what people write. ;-)
> And "hallucinating" to mean "outputs plausible sounding gibberish that doesn't hold together consistently". Of course there's no actual hallucination going on.
This may very well be what you know to be true and I have no reason nor desire to assume otherwise. The problem is... Many people use the word "hallucinating" in this context literally and not metaphorically.
Since I do not know you, how am I to tell the difference?
hardbass · · focus · HN ↗
spider-mario · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
It is always a joy when a person, such as yourself, finds the irony in my moniker.
Thank you for this.
[deleted] · · focus · HN ↗
[deleted]
hodgehog11 · · focus · HN ↗
But in general, yes, the LLM cannot know about concepts that are far outside of its training set. Humans are the same, I would argue. If you add a good amount of your own knowledge into its context, or better yet, into finetuning, you might find it surprisingly easy to get it caught up on that material.
[1] Ahmed, A., Cooper, A. F., Koyejo, S., & Liang, P. (2026). Extracting books from production language models. arXiv preprint arXiv:2601.02671. <a href="https://arxiv.org/abs/2601.02671" rel="nofollow">https://arxiv.org/abs/2601.02671.
Kim_Bruning · · focus · HN ↗
You'd be surprised how few digits you need to make a problem that is presumably unique in earth history. For a typical sum, the number of pre-existing answers would need to scale with 10^n lines of text where n is the number of digits. This expands out of control REALLY quickly. A quick guesstimate has you reading out of a black hole at 21 digits if your LUT is on paper, or 26 digits if you're using modern HDD technology. O:-)
hodgehog11 · · focus · HN ↗
This argument was asinine in 2024. It is insane to be saying these things in 2026. Where have you been? What have you been looking at? How many articles explaining why the "statistical parrot" analogy fails have you missed? How much mental gymnastics do you have to do to explain how a modern LLM can solve novel math problems that fall really far outside of its training set?
It absolutely understands how to do math, by whatever reasonable definition you want to provide to the word "understand". For example, the identification of the addition expression is understanding, and no, it does not do tool calling for basic arithmetic any more than humans might. Isolation of individual concepts in intermediate layers can already be demonstrated, or else transfer learning wouldn't possibly work. Nobody is saying that LLMs are humans. But we need labels for some of the things that we observe and dismissing them because "statistical" is laughable.
Look at the proof of this: <a href="https://github.com/anthropics/formal-math/blob/795efb86f191735c5481675763537cfb4ff37e55/percolation/summary.pdf" rel="nofollow">https://github.com/anthropics/formal-math/blob/795efb86f1917... . Forget the Lean, look at the underlying argument construction. At the very least, this is continuing from an argument that was hinted at in the literature in 2024, but these proceedings were difficult enough that humans were not able to do them within two years. Do you attribute this to the harness alone? If so, that's a pretty sophisticated bit of engineering, I would say! Probabilities are far too small to argue infinite monkey theorem.
If there was even a shred of a reasonable argument that LLMs were incapable of concept extraction and manipulation, I and my colleagues would be all over it. We would relish in it. It would bring us comfort. It is unbelievable that people think they can spew whatever basic garbage they think of as a gotcha, and think that minds all over the world haven't already considered that. This is like climate denial at this point.
AdieuToLogic · · focus · HN ↗
If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don't know what to say.
hardbass · · focus · HN ↗
latentsea · · focus · HN ↗
hardbass · · focus · HN ↗
butlike · · focus · HN ↗
diseasedyak · · focus · HN ↗
latentsea · · focus · HN ↗
epihelix · · focus · HN ↗
<a href="https://www.pnas.org/doi/abs/10.1073/pnas.2524472123" rel="nofollow">https://www.pnas.org/doi/abs/10.1073/pnas.2524472123
Whatever you might think about your own abilities, most individuals can't tell the difference.
AdieuToLogic · · focus · HN ↗
> Whatever you might think about your own abilities, most individuals can't tell the difference.
I have yet to see an LLM say "hello" to a neighbor. I have done so and can definitively assure you "most individuals" can tell the difference.
thereforegrin · · focus · HN ↗
I'm a bit confused by your argument because I too have some neighbors who don't say "hello" when they see me. Are they LLMs too, you think?
Consciousness is a thing we assume of others because of tact not fact.
redsocksfan45 · · focus · HN ↗
[dead]
hodgehog11 · · focus · HN ↗
spider-mario · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
> That’s not what they said.
How did I misquote and/or mischaracterize any the above?
spider-mario · · focus · HN ↗
Hypothetical you: “Bread is neither tasty nor disgusting (unlike maple syrup). It’s a bunch of molecules.”
Hypothetical hodgehog: “Maple syrup is also a bunch of molecules [so if you accept that maple syrup can be delicious, being a bunch of molecules can’t be why bread isn’t].”
Hypothetical you: “If you don’t see a difference between bread and maple syrup, I don’t know what to say.”
hardbass · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
It is also possible that the proverb I provided to you is the origin of the oft quoted:
So there's that.hardbass · · focus · HN ↗
hardbass · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
The onus is not mine to disprove a hypothesis you have chosen to mention in passing. The responsibility is yours to prove said hypothesis or at least contribute meaningfully with some amount of credible research.
Or try to learn from Proverbs 17:28[0]:
Either works for me.0 - <a href="https://www.biblegateway.com/passage/?search=proverbs%2017:28&version=NIV" rel="nofollow">https://www.biblegateway.com/passage/?search=proverbs%2017:2...
hardbass · · focus · HN ↗
butlike · · focus · HN ↗
hodgehog11 · · focus · HN ↗
And yes, according to our best definitions, the Robin bird does understand the worm it's pecking at.
wartywhoa23 · · focus · HN ↗
Bout of tinnitus, then crickets
mapontosevenths · · focus · HN ↗
That said, this is also inaccurate at a technical level.LLM's are very capable of doing math and they ARE calculating internally. Most of what they do is calculation, not storage. It's just not done in a way that it's trivial to explain here.
It's described in some detail below, though it's a bit dense.
<a href="https://www.lesswrong.com/posts/E7z89FKLsHk5DkmDL/language-models-use-trigonometry-to-do-addition-1" rel="nofollow">https://www.lesswrong.com/posts/E7z89FKLsHk5DkmDL/language-m...
jibal · · focus · HN ↗
jodrellblank · · focus · HN ↗
You can say that about everything in a human brain. Neurons fire electric charges in response to inputs, nothing more. Ion channels do this, this neurochemical level rises, this chemical bonds to that receptor, nothing more. It's almost a version of 'reductio ad absurdum' but instead like you're saying "if I can explain how it works then it doesn't work".
OK it's statistical. Instead it could be determinsitic, or random. What other options are there for a human predicting someone's response to a situation - certain, probable, random, and...? OK it's token predicting. Instead it could be another kind of pattern. We don't use tokens, but we either use <some representation of information> or we ... don't?
What's the most significant, strongmanned, core difference that makes silicon doing number crunching "nothing more" and brains "something more"?
butlike · · focus · HN ↗
hardbass · · focus · HN ↗
jodrellblank · · focus · HN ↗
AdieuToLogic · · focus · HN ↗
> You can say that about everything in a human brain.
> What's the most significant, strongmanned, core difference that makes silicon doing number crunching "nothing more" and brains "something more"?
The fact that you formulated this question, in and of yourself, without "prompting" from me or anyone else.
Cogito, ergo sum.[0]
0 - <a href="https://en.wikipedia.org/wiki/Cogito,_ergo_sum" rel="nofollow">https://en.wikipedia.org/wiki/Cogito,_ergo_sum
jasonfarnon · · focus · HN ↗
pyridines · · focus · HN ↗
Veedrac · · focus · HN ↗
azan_ · · focus · HN ↗
tripzilch · · focus · HN ↗
It literally couldn't have done it without tools, so your claim is not even relevant to this discussion.
They also pumped millions into searching for the lowest hanging fruit that would impress people like you, "hm I wonder how many millions they are pumping into solving actually useful problems like climate change or something".
Veedrac · · focus · HN ↗
Windchaser · · focus · HN ↗
(And yes, I know that they solved the 'easy' form of the NS problem. It's still pretty damn impressive)
chimprich · · focus · HN ↗
I am also prepared to be included in the set of people who are (apparently) easily impressed.
It's a problem that has been around for getting on for two centuries and no human has been able to solve it in that time (despite there being a $1 million prize and a lot of kudos on offer for the past quarter-century).
danlitt · · focus · HN ↗
Veedrac · · focus · HN ↗
Clearly LLMs have gotten better.
Veedrac · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
Just to check if I was actually crazy, I actually went and put a simple addition (7 digits + 7 digits) , and a simple letter counting question to Claude haiku(4.5) , sonnet(5), opus(5.5) and fable(5.1) . They all did just fine straight up.
If you don't mind spending the tokens, some older/other models can also arrive at the correct answer if you ask them to do the math in long form, since that fits nicely inside autoregression.
Not sure since when exactly, but letter-counting hasn't been a problem for a while now either. This used to be a problem due to the tokenizers used. Slightly older models can be asked to split the word out into letters, and then they can use autoregression to solve.
leoedin · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
my actual prompt:
SkyBelow · · focus · HN ↗
That said, we can only be sure with open models. In theory, a model like Fable could have access to tools we can't see and only a promise they don't. But load up something like deepseek, put it in a harness with only text in/text out, and you can see exactly how it works.
As for if it counts as doing math, this gets into the messy question of if a given human is doing math or not. Math itself is some level of memorization and some level of applying known facts. You have to remember 1 means one and that 1 + 1 is 2. But you don't need to remember that 123 + 321 = 444. You remember 1 digit addition and remember you can apply this to 10s place and 100s place, and then you apply these different facts and do math. But you might as simply memorize some things, like 11 + 11 = 22. This is related to the memory of 1+1=2, but you aren't really using that memory either. Almost like an engram of 1+1=2 forms that you can then loop a few times before you need more conscious thought. What about 111111111111+11111111111? Well, your brain might do a heuristic and just do all 2s, but that isn't the right way to answer that question.
Given all this, people complain about LLMs memorizing math answers and not doing math, but memorizing the math answers is part of doing math. It seems to have basic facts pretty well memorized, and with reasoning it is far better at applying them. But this is messy human math, not clean calculator math which always produces the correct answer (sans some bug in the code). Much like how a human with decent math skills can make a mistake and even multiple if you distract them, an LLM can apply the wrong memory, apply a fake memory, or just not apply something it should. The messier the context, the more likely this is to happen.
So, is an LLM doing this?
P.S.
For an interesting test in how much math involves memory, try doing math in a base you aren't familiar with characters you aren't familiar. The simplest option is almost always mapping back to the ones you memorized, even if you are applying simple operations that you deeply know. Even if you routinely work with hex, can you do the same rough estimation of something like ca / b.3 that you can do with 122 / 11.2 to see if your final answer is in the correct ballpark without first converting to decimal?
TeMPOraL · · focus · HN ↗
Humans still can't flap their hands and swim or fly.
jonas21 · · focus · HN ↗
The "crack-addled idiot savant" phase was really circa 2024, before the big labs figured this out.
I think the issue here is that Google decided that doing reasoning in the AI overviews in Google search would be too slow (and probably also too expensive), so it's still stuck making 2024-era mistakes.
Isamu · · focus · HN ↗
Maybe but you have to have deep pockets just to get to the starting line. And then you need standing, and some injury to argue.
Corporations have been remarkably successful at arguing they are operating within the bounds of free speech, whether or not what is said is factual, and whether or not any fact checking has been done.
nextaccountic · · focus · HN ↗
The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings.
Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs
AlienRobot · · focus · HN ↗
mitxela · · focus · HN ↗
queenkjuul · · focus · HN ↗
mitxela · · focus · HN ↗
versteegen · · focus · HN ↗
dmazzoni · · focus · HN ↗
Nobody asked for an LLM response for every single search.
They used to detect certain types of queries and offer direct answers when the query matches. In my opinion that’s how Gemini in search results should work.
compass_copium · · focus · HN ↗
what in the actual fuck. i don't want them and can't turn them off and they can't even serve them to all users?
hurfdurf · · focus · HN ↗
ssl-3 · · focus · HN ↗
AI can make mistakes, so double-check responses
mrguyorama · · focus · HN ↗
ssl-3 · · focus · HN ↗
someonebaggy · · focus · HN ↗
mitxela · · focus · HN ↗
red75prime · · focus · HN ↗
spennant · · focus · HN ↗
antonvs · · focus · HN ↗
For example, when an LLM “predicts the next word” in code it’s writing for an existing software project, that prediction takes into account an enormous amount of context. The results of that demonstrate what we would normally call “understanding” and “reasoning,” at a level that outclasses most humans in many respects. Calling this “next token prediction” is a bit like calling human speech “next word saying”. Sure, it’s true in some superficial sense, but as a description of a technology, it’s terrible.
has turned into a way to describe a complete thinking process,
You should also keep in mind that for all we know, the human brain processes language in much the same way, which would make humans mere “next token predictors” with a more complicated harness.
baubino · · focus · HN ↗
I have nothing to add. Just wanted to save this quote for posterity. Thank you.
_kb · · focus · HN ↗
If you apply the dogs on acid mental model it helps establish appropriate levels of trust.
SV_BubbleTime · · focus · HN ↗
sham1 · · focus · HN ↗
ben_w · · focus · HN ↗
The closest I can make a biology analogy, is if someone had an immortal and congenitally-brain-damaged large rodent, and mapped tokens to different scent molecules, and spent 800,000 years training it on how well it could imagine the next smell in the sequence before even considering making it conversational.
Yes, it can do a lot.
But also, it took a long "subjective" (if it even has that) time to get there; and despite it being really bad at learning from examples, it is pretty surprising that such a small brain was even capable of learning so much at all, even though it had an effectively unbounded (by biological standards) amount of time spent on that training.
Melatonic · · focus · HN ↗
JSR_FDED · · focus · HN ↗
_kb · · focus · HN ↗
tripzilch · · focus · HN ↗
somenameforme · · focus · HN ↗
But that doesn't mean they won't be able to do a vast number of extremely intelligent seeming things. There's just so much information out there and any given human can never hold more than the most minuscule chunk of all of it in his mind, so they'll be able to connect lots of dots that we're missing simply because of our limited carrying capacity, but I still don't think they'll ever be able to create fundamentally new dots.
In other words:
- Solving extremely complex mathematical problems requiring extensive knowledge across multiple esoteric and complex domains? Yip.
- Creating math starting from a framework where math doesn't exist in any way, shape, or fashion? Nope.
Ironically, the more complex the cross-domain problems are, the more effective LLMs will seem to be, because you limit the number of humans who have any chance of internalizing everything across both domains, whereas for an LLM there's no such issue. This will create a perception of super intelligence, which will probably where any danger from LLMs would emerge. Doing things like using a token prediction algorithm to make war or other such strategic decisions, because of the misguided belief that it's not only intelligent but super intelligent. It's almost cargo cult like thinking.
queenkjuul · · focus · HN ↗
rob74 · · focus · HN ↗
fangspire · · focus · HN ↗
[dead]
jibal · · focus · HN ↗
Like today. I asked Gemini for the longest state names with an even number of letters and it gave me North Carolina and South Carolina. When I complained that they are odd, it gave me North Dakota and South Dakota, which are both odd and not the longest. When I noted that, it went back to the Carolinas. Finally it appeared to switch to a different model that actually did counting and found Pennsylvania and West Virginia.
cschep · · focus · HN ↗
GolfPopper · · focus · HN ↗
nananana9 · · focus · HN ↗
Reminder: If asked how many of a given a word has, run it through the AGI letter counter:
If I got any of these wrong it's because I did it manually - but this script would almost certainly invoke another AGI to count the letters, so this is a realistic result.someonebaggy · · focus · HN ↗
zerd · · focus · HN ↗
dfxm12 · · focus · HN ↗
hienyimba · · focus · HN ↗
First, a search engine indexes the web and makes it available to users. It’s been ages since Google did any of that. They no longer index sites or take ages to do so. Case in point, our cybersecurity startup (Webvetted.com) was launched in November 2025. Till date, only one page is indexed on the entire website. And I’ve talked to lots of other developers and it’s a common issue.
Secondly, a search engine organizes indexed information and makes it useful for people. Google is a basic LLM nowadays. They figured out that why organize and make the information useful when they could just answer the question with Gemini anyways? So they no longer bother to do the work of a search engine and are now just a lower-ranking open-source Chinese LLM
ytcommentsectio · · focus · HN ↗
[dead]
hienyimba · · focus · HN ↗
The work of a search engine is to wade through the internet and find the positive-quality sites itself.
rmunn · · focus · HN ↗
So even when the Internet used to have a much higher ratio of positive-quality sites, a search engine was still necessary, and so much faster than finding new sites yourself.
dieortin · · focus · HN ↗
Search for any recent news and you’ll see this is obviously not the case
hienyimba · · focus · HN ↗
someonebaggy · · focus · HN ↗
atdt · · focus · HN ↗
I am not implying that your start-up is a scam, nor that Google is acting on these trust scores. What I am pointing out is that to an algorithmic assessment of trustworthiness, your website looks a little sketchy: hardly anyone links to you; your domain is less than a year old; your whois info is anonymized; the text content is LLM-generated[1]; and the specific niche you're in (people finding services) is rife with scams. If I ran my own search engine, I don't think I'd include you.
hienyimba · · focus · HN ↗
You selectively left out examples like Grindsoft that gives the site a 79/100 ranking. or several LinkedIn, X or news sites that link to the site and have covered it positively.
Nevertheless, if Google uses ScamAdvisor to rank new sites (when scamadviser itself says its gives new sites lower rankings due to low age/history), then it might be time to pack the company up.
user00005 · · focus · HN ↗
iboisvert · · focus · HN ↗
chris_engel · · focus · HN ↗
vishalanton · · focus · HN ↗
pishpash · · focus · HN ↗
jeremyjh · · focus · HN ↗
0xbadcafebee · · focus · HN ↗
jareklupinski · · focus · HN ↗
maybe i'm the weird one, but asking a machine about the theoretical outcome of a sports game which hasn't happened yet yet has many parties all trying to influence the outcome sounds like the _last_ thing a technical product would be good at
if i asked "what was the score of their 4th game in 1986" i'd expect a search engine to excel however
chr15m · · focus · HN ↗
chrisweekly · · focus · HN ↗
I can't even.
mrguyorama · · focus · HN ↗
What the hell is everyone smoking!
chrisweekly · · focus · HN ↗
sans_souse · · focus · HN ↗
FlowingRiver · · focus · HN ↗
"The team are tied at the end of the third round with the scores 43 to 51".
So it was tied at the end of round 2 for 43, not round 3. Not sure were it got the 51 from and why it figured it was a tie. LLM's pretty cool until they aren't. As they say, the hallucinate 100% of the time but most of the time it is useful.
someonebaggy · · focus · HN ↗
mapt · · focus · HN ↗
wavemode · · focus · HN ↗
mancerayder · · focus · HN ↗
After all, why would Google do anything for free when it comes to AI?
If you wanted to tease customers with one AI shitty free search box in order for them to then decide to upgrade and pay for Gemini, well, this isn't the way. And so either Google are idiots, or we're helping them for free. I'm going with Occam's Razor on this one - we're the product.
aaron695 · · focus · HN ↗
[dead]
0xbadcafebee · · focus · HN ↗
mlmonkey · · focus · HN ↗
segmondy · · focus · HN ↗
thrance · · focus · HN ↗
warkdarrior · · focus · HN ↗
mrguyorama · · focus · HN ↗
asdfman123 · · focus · HN ↗
You have a system in which investors are usually smart, but betting on collective stupidity.
Same with, say, the Chinese real estate market, where they were building apartments that no one could live in. Everyone is intelligent enough to know they were useless buildings, but you can still make money speculating on the bubble.
Quarrelsome · · focus · HN ↗
zx8080 · · focus · HN ↗
AI is sold well because the user either corrects it or happily accepts any answer (usually, depending if one is an expert in the question's field).
Search is not the google's area, they sell not search! They sell ads.
fhe · · focus · HN ↗
someonebaggy · · focus · HN ↗
elevation · · focus · HN ↗
Perplexity has similar behavior to what GPP describes of Google; I've asked perplexity to compile information that exists on the web. Instead of consulting existing web pages, it gives limited summary responses from its weights (while citing some page that has nothing to do with what I asked.) When I attempt to redirect it towards more concrete sources, it won't comply.
IAmBroom · · focus · HN ↗
DANmode · · focus · HN ↗
That simple.
Broad rules across specific subjects like this are tough to get right.
Just lots of fine-tuning and exceptions.
I had their chatbot avoid pulling a URL from archive.org with a couple pretty impressive steps of mental gymnastics basically telling me how I could do it myself, but refusing until I pushed.
anukin · · focus · HN ↗
mbac32768 · · focus · HN ↗
The problem is they're just running a very dumb, cheap model on the results because running a smart model on every search result page would cost them infinity money.
xp84 · · focus · HN ↗
01100011 · · focus · HN ↗
s3graham · · focus · HN ↗