‹ BackHN Continuity

Thread

When did Google get so weird?

2010 points · 1120 comments · sancho-panza

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. wisty · · focus · HN ↗
    Haha, had almost the exact thing trying to find a quote from Steins; Gate ("I keep seeing it, I keep seeing it"). Google AI was very "worried" about me.
  2. shadowgovt · · focus · HN ↗
    Oh, this is an easy question to answer.

    The AI layer handles queries as plaintext, which means if you're searching for say a quote from a book, it will often misinterpret as a direct statement from you, not a string you're trying to match on the internet. Especially if the quote is an imperative statement.

    I'll be curious to see how they resolve this issue. But it has been chronic for awhile now; they SWAT individual instances of misinterpretation but new ones keep sneaking in.

    1. krackers · · focus · HN ↗
      Isn't this easily solved with a prompt prefix "You are an AI agent on the google search page. The following user input is what the user typed into the search box".
      1. bakugo · · focus · HN ↗
        [delayed]
      2. im3w1l · · focus · HN ↗
        [delayed]
        1. shadowgovt · · focus · HN ↗
          [delayed]
    2. Hugsbox · · focus · HN ↗
      Shouldn't an AI attached to a search engine actually do a search with your input and give you results based on that rather than just taking it at face value and assuming you're trauma-dumping on it?
      1. shadowgovt · · focus · HN ↗
        [delayed]
        1. Hugsbox · · focus · HN ↗
          Huh, that makes a lot of sense, thanks for the input :)

          Actually used to drive me nuts watching (usually old) people googling facebook to get to facebook, so I see exactly what you mean.

    3. im3w1l · · focus · HN ↗

      [dead]

  3. solarkraft · · focus · HN ↗
    It&#x27;s a whole meme category to make the AI overview help with absurd scenarios: <a href="https:&#x2F;&#x2F;knowyourmeme.com&#x2F;memes&#x2F;where-is-mama-ai-overviews" rel="nofollow">https:&#x2F;&#x2F;knowyourmeme.com&#x2F;memes&#x2F;where-is-mama-ai-overviews
    1. mattlondon · · focus · HN ↗
      I think these sort of things are actually very &quot;googley&quot; in the old kinda wacky way. It&#x27;s funny, it&#x27;s goofy. I like the fact it plays along.

      What would you rather it said?

      1. TomGarden · · focus · HN ↗
        Side note but your ending question triggered all my LLM prose engagement bait warning alarms
        1. ACCount39 · · focus · HN ↗

          [dead]

      2. creatonez · · focus · HN ↗
        Ah. So Google cancelled April fools to make every day a big joke?
        1. mattlondon · · focus · HN ↗
          This is the company with the childish colourful logo, slides &amp; ballpits in their offices and 20% time to do whatever you want etc. Why not?
          1. creatonez · · focus · HN ↗
            All of this culture is gone except for Google Doodles
      3. karahime · · focus · HN ↗
        I really agree with this sentiment, I feel as though we&#x27;ve gotten bogged down in this terrible Institutional Seriousism, like this thinking that if things are big or things matter then they have to put on a grim face, and I think that just isn&#x27;t true.
  4. edent · · focus · HN ↗
    Most people in the world are profoundly lonely. They&#x27;ll take whatever parasocial relationships they can get - reaction videos, podcasts of people chatting, Eliza simulating concern.

    Google wants to relentlessly monetise your sadness.

    The only way out is to try and make connections with real people.

    Case in point, why didn&#x27;t you text your friends to ask them if they remembered those old memes? Why was your first thought to ask a computer rather than a person?

    It&#x27;s like when you&#x27;re in the pub and and someone asks &quot;who was that guy who was in the movie where…?&quot; you can either chat with your friends and have a good time or be a buzzkill who opens up IMDB and says &quot;Humphrey Bogart&quot;.

    1. chistev · · focus · HN ↗
      What was the name of the farm next to the Hill House?
    2. belowavgiq · · focus · HN ↗
      Irrelevant nitpick, but it&#x27;s definitely not &quot;most&quot;. I&#x27;d argue it&#x27;s only this common in some English speaking countries and maybe a few other Western ones. The looks I get from American friends when I tell them that yes I do see my (local) friends nearly every day, are quite funny.
      1. dataviz1000 · · focus · HN ↗
        &gt; Most people in the world are profoundly lonely.

        I&#x27;ve been doing a lot of traveling for the past 3 years and I agree with you that it is `definitely not &quot;most&quot;`. Everywhere people interacting with each other in third spaces and cafes even if it is street food. There are open air markets around the world filled everyday with groups of people interacting with each other. There are parks and squares around the world filled with families sitting together on benches. Around the world there are churches, mosques, and temples filled with families and groups of people.

        Most places in the world people will happily make small talk or have a discussion with a stranger even if I only know 100 words of their language.

        Sure many or most people in some places are working 10 hours a day 6 days a week but I don&#x27;t think it is the isolation I see in the United States. It is the same in De&#x27;Nang Vietnam or a market in Lima, Peru when I went everyday to get coffee in the morning; it was always the same person working ,any time or day of the week, but they were always friendly and welcoming and interacting with with the same people.

        1. JumpCrisscross · · focus · HN ↗
          &gt; the isolation I see in the United States

          Travel alone in America and the frequency with which strangers will engage you in conversation reveals its diversity. New York and Atlanta are, today, exceptionally friendly. Most of the South is, too, as well as a surmising fraction of rural America. (Boston and Seattle are, in my experience, the worst.)

        2. lukan · · focus · HN ↗
          Hm, I also done extensive traveling, but would agree to the &quot;most people are lonely&quot; statement. Much more so in rich countries.

          Because the lonely people you don&#x27;t see at the market or in a park socialising. They are at home, in front of the TV or now their smartphone.

        3. edent · · focus · HN ↗
          But, like, you only see the people who are out and about, right?

          You don&#x27;t see the people trapped inside. Stuck in hospitals. Sat staring at a phone that never rings.

          You never talk to people who don&#x27;t feel comfortable talking to a stranger.

          1. password54321 · · focus · HN ↗
            It really does depend on the country. There is a clear dating recession in US&#x2F;Japan for example but not in Thailand&#x2F;Peru.
          2. watwut · · focus · HN ↗
            Phones and tech made those people a lot lonelier then they used to be before. Because before, people in hospital naturally socialized with other people in the hospital. So, a loner stuck in the hospital was less lonely. And before, people naturally went out and socialized with those in their environment. Because they were bored alone inside. That created communities.
        4. kelnos · · focus · HN ↗
          I feel like that&#x27;s selection bias at work. When you&#x27;re traveling, by definition you aren&#x27;t going to see or interact with the people who aren&#x27;t there, who might be at home, feeling lonely.

          And I&#x27;m not sure interacting with customers as someone who works in a coffee shop or a food stall is really a sign either way. Many (most?) of those relationships are surface-level at best, and while some people might get to know the person who makes their coffee on a more personal level, and actually see them outside of the context of their job, you can&#x27;t really tell that just by what you see as a traveler.

          I&#x27;m not saying these workers don&#x27;t derive satisfaction from any of this, or that it&#x27;s not meaningful interaction, but these people could still be lonely in their personal lives.

        5. HPsquared · · focus · HN ↗
          Surely there&#x27;s a big selection bias here.

          You see the people who are out and about in society, not the ones holed up in their caves.

        6. martin- · · focus · HN ↗
          That&#x27;s like going to the pub and concluding that 100% of people currently alive are at the pub, because everyone you see is at the pub.

          No traveler is going to observe me sitting alone in my apartment (hopefully).

      2. JumpCrisscross · · focus · HN ↗
        &gt; it&#x27;s definitely not &quot;most&quot;

        EDIT: Correct.

        About a fifth of Americans report having zero friends; another fifth report feeling lonely sometimes [1].

        [1] <a href="https:&#x2F;&#x2F;www.theglobalstatistics.com&#x2F;us-loneliness-statistics&#x2F;#:~:text=The%202024%20U.S.,loneliness%2C%20with%2014.0%25" rel="nofollow">https:&#x2F;&#x2F;www.theglobalstatistics.com&#x2F;us-loneliness-statistics...

        1. danilafe · · focus · HN ↗
          Not to be overly pedantic, but isn&#x27;t that not &quot;most&quot;, out of even Americans?
          1. JumpCrisscross · · focus · HN ↗
            &gt; isn&#x27;t that not &quot;most&quot;, out of even Americans?

            Yes. Edited for clarity.

            The only American demo that is majority lonely is Gen Z (two thirds).

            1. belowavgiq · · focus · HN ↗
              Yup, and this is my age range. I guess I should have specified, knowing the average age on HN is... just a little higher to say the least. But to a progressively lower extent it applies to older generations as well.
        2. spacechild1 · · focus · HN ↗
          &gt; About a fifth of Americans report having zero friends

          Wow, that&#x27;s depressing...

          1. vlyan · · focus · HN ↗
            it&#x27;s nearly impossible to make friends once you&#x27;re an adult and everyone&#x27;s walls had gone up. people just call their superficial acquaintances &quot;friends&quot;.
            1. spacechild1 · · focus · HN ↗
              &gt; a real friend is essentially a platonic partner

              There are certainly different degrees of friendship. I might only have one such platonic partner and unfortunately we hardly see each other anymore because we live in different cities and have families. Once in a while we talk on the phone, but when we do it&#x27;s a minimum 2 hour call.

              However, I have several people I would still consider close friends. My personal threshold for true friendship is whether I can comfortably talk about private things.

              Back to the study. I don&#x27;t think it uses such a high threshold as &#x27;platonic partnership&#x27;. I found the following paragraph really striking:

              &quot;What these numbers also reveal is the hidden depth of the crisis. The 17% with zero friends in 2024 — compared to just 1% in 1990&quot;

              Do everybody suddenly get higher standards of friendship? Or did people really get lonelier? I find the second more likely.

      3. kelnos · · focus · HN ↗
        &gt; The looks I get from American friends when I tell them that yes I do see my (local) friends nearly every day, are quite funny.

        How old are you, though, and do you and&#x2F;or your friends have kids?

        When I (an American) was in my 30s, I did see my local friends (nearly) every day. Now into my mid 40s, it&#x27;s a lot more intermittent, with some people having moved away (and I&#x27;ve failed to make new friends faster than some have moved away). They are still friends, and we communicate over chat&#x2F;phone, but I only see them at most a couple times a year. Then there are others who now have young children and have become a bit more inward-focused. Not that I don&#x27;t see them anymore, but I see them less often.

        I of course see others on a daily basis: people I know&#x2F;recognize at my yoga studio, service workers in my neighborhood, the neighbors themselves, but most of them are acquaintances, at best, not friends.

        Regardless, I wouldn&#x27;t say I&#x27;m lonely (though I do feel lonely on occasion), even though I can sometimes go a week or more without seeing any of my friends in person. I do worry about it being more difficult to maintain in-person friendships as I&#x27;ve been getting older, and the difficulty in making new friends.

        To relate to the topic at hand: I cannot imagine talking to an LLM and thinking of that as a substitute for interacting with friends. Feels super weird and kinda creepy.

        1. belowavgiq · · focus · HN ↗
          Wait wait wait, you are from an older generation, so yes it mostly (older people here do see each other often even when married. moving out? ha) doesn&#x27;t apply to you.

          I am indeed quite a lot younger, and in my personal experience the social skills of many (yes, of course, not all) younger Americans have been absolutely destroyed by the internet and some other things that I can&#x27;t mention.

    3. TomGarden · · focus · HN ↗
      Texting your friends for factual information doesn&#x27;t seem reasonable to me. If I got that text and someone only wanted the info I&#x27;d be annoyed (and I had friends who genuinely did this pre-google. Still annoying).

      I&#x27;m with you on the pub observation though

      1. Telaneo · · focus · HN ↗
        &gt; Texting your friends for factual information doesn&#x27;t seem reasonable to me. If I got that text and someone only wanted the info I&#x27;d be annoyed (and I had friends who genuinely did this pre-google. Still annoying).

        I don&#x27;t see this as unreasonable? I&#x27;ve had friends text me questions about grammar, since they were learning the local language and I knew theirs, so they knew I&#x27;d be a good sanity check. Similarly, I&#x27;m the go tech support for my friend group for anything that a cursory Google search can&#x27;t solve (and I assume a lot of people here are in the same situation with friends and family). I&#x27;ve also been asked to find &#x27;that one image&#x27; that similarly didn&#x27;t show up on Google image search for whatever reason.

        Neither of these cases seem unreasonable to me. Are they to you, or do you have other cases in mind? Or is the problem that this happens too often?

        1. TomGarden · · focus · HN ↗
          Your examples seem fitting for a text. Simple factual queries with not intention of connection are different
          1. Telaneo · · focus · HN ↗
            What&#x27;s the difference between &#x27;Texting your friends for factual information&#x27; and &#x27;Simple factual queries with not intention of connection&#x27;? You&#x27;ve said they&#x27;re different, that first is reasonable and the second is not, but they read the same to me.
            1. TomGarden · · focus · HN ↗
              Utilizing a friends&#x27; expertise vs outsourcing effort to them
              1. Telaneo · · focus · HN ↗
                Those also read the exact same to me, unless you mean to imply the latter also includes &#x27;being an knobhead to the person you&#x27;re asking&#x27;.
            2. TomGarden · · focus · HN ↗
              Perhaps an LLM can help with the distinction. I&#x27;m at a loss helping with this, which feels a little meta
    4. frays · · focus · HN ↗
      Do you have a citation for &quot;most&quot;? Or are you referring to a few specific countries in mind like Japan and South Korea?
      1. edent · · focus · HN ↗
        I asked Google and it told me I was absolutely correct. So… ;-)

        Loneliness rates vary by country, age, socio-economic status, health, etc. The UN has some reasonable statistics.

        <a href="https:&#x2F;&#x2F;www.who.int&#x2F;teams&#x2F;social-determinants-of-health&#x2F;demographic-change-and-healthy-ageing&#x2F;social-isolation-and-loneliness" rel="nofollow">https:&#x2F;&#x2F;www.who.int&#x2F;teams&#x2F;social-determinants-of-health&#x2F;demo...

        There&#x27;s also some detailed stats for the UK.

        <a href="https:&#x2F;&#x2F;www.gov.uk&#x2F;government&#x2F;statistics&#x2F;community-life-survey-202425-annual-publication&#x2F;community-life-survey-202425-loneliness-and-support-networks" rel="nofollow">https:&#x2F;&#x2F;www.gov.uk&#x2F;government&#x2F;statistics&#x2F;community-life-surv...

        Was my use of &quot;most&quot; careless? Probably. Let&#x27;s natter about it over a pint.

        1. asystole · · focus · HN ↗
          16% is so far removed from &quot;most&quot; that it undermines your entire argument, whatever it is.
    5. what · · focus · HN ↗
      &gt;It&#x27;s like when you&#x27;re in the pub and and someone asks &quot;who was that guy who was in the movie where…?&quot; you can either chat with your friends and have a good time or be a buzzkill who opens up IMDB and says &quot;Humphrey Bogart&quot;.

      I don’t really understand this. Someone wants an answer not a chat about how you don’t know?

      1. edent · · focus · HN ↗
        The primary purpose of going out with your friends to the pub is social interaction. If the purpose was just to get drunk most efficiently, you&#x27;d sit at home alone and pour vodka down your gaping maw.

        The purpose of talking to your friends is social interaction. If the purpose of being friends was to extract information in the most efficient manner possible, you&#x27;d replace them with Wikipedia.

        1. Ekaros · · focus · HN ↗
          If in some social interaction they ask such question I expect them to actually want to know the answer. If not the question should not be asked. The answer to question can lead discussion to some place.
          1. kelnos · · focus · HN ↗
            I don&#x27;t think anyone is suggesting that it&#x27;s not fine to eventually look up the answer to the question. But chatting about it before it gets to that point, and exhausting your guesses is much more fun than someone looking up answers to every question someone poses that doesn&#x27;t have an obvious answer within a few seconds.
            1. Ekaros · · focus · HN ↗
              Then formulate it as a conversation starter not as explicit question like was done here. Only a buzzkill makes such explicit conversation starters. Unless it is a trivia night...

              &quot;Remember that movie where&quot; or &quot;remember oh that guy who was in that movie&quot; would both imply this. &quot;Who was the actor in oh that movie where&quot; is clearly a cry for exact answer as quickly as possible.

        2. Capricorn2481 · · focus · HN ↗
          The idea that you would consider your friend a buzzkill because they used IMDB is the type of nightmare fuel social anxiety is made of.

          You don&#x27;t think it&#x27;d be fun to hear everyone guess and then someone looks up the answer? Sounds like trivia to me.

        3. what · · focus · HN ↗
          Yes, sure, it’s to be social. But if someone asks a stupid trivia question, then just give them the answer. You’re not starting a discussion with a stupid trivia question. You are truly weird.
      2. Telaneo · · focus · HN ↗
        I&#x27;ve witnessed the same social dynamic happen in certain groups, but thankfully not in my own friend group. I&#x27;m unsure why. In my own friend group, that kind of question would either be answered legitimately, either by someone knowing the answer, by looking it up whereever, or it&#x27;s not actually important enough to whatever point is being made that we can skip past it to continue whatever discussion is actually happening.
    6. hnbad · · focus · HN ↗
      &gt; Google wants to relentlessly monetise your sadness.

      You&#x27;re missing the forest for the trees. This is just downstream from decades of promoting individualism as the highest moral good.

      There was a brief period where &quot;social networks&quot; actually had people simply expand their social interactions to the online sphere, reconnecting with long-lost friends and distant relatives via the magic of The Internet. Then Facebook introduced The Feed and suddenly people were competing with influencers and brands for the eyeballs of their peers, everybody suddenly felt like they had to perform to be relevant&#x2F;interesting&#x2F;funny&#x2F;exciting enough for other humans to still care about them and user satisfaction and (more critically) mental health dropped like a rock - but engagement metrics and retention spiked and that meant so did ad revenue and Facebook&#x27;s bottom line. Ever since, all social media has essentially just been competitive variety shows mixing in influencers, brands, state actors and political provacteurs (and now increasingly also actual AI &quot;bots&quot; rather than just armies of underpaid human &quot;bots&quot;) with regular people while deliberately obscuring the lines between these groups.

      All of this promotes social isolation and alienation. Every experience becomes depersonalized to &quot;reduce friction&quot;. You don&#x27;t even have to talk to the delivery person anymore (let alone the restaurant) - even the tip just becomes a button press disjoint from reflecting on the actual experience or human interaction (to whatever extent it even still existed). There&#x27;s an entire microcosm of &quot;creators&quot; serving whatever &quot;hot takes&quot; or niche subject matter you want to hear and &quot;comment sections&quot; largely serve as one-way &quot;opinion dumps&quot; you&#x27;re expected to use to shout into the void rather than try to actually have a conversation in (let alone meet actual people or develop friendships in - an idea that I am sure sounds absurd if you weren&#x27;t around in the early days of online message boards).

      AI is just the logical consequence of this process of dehumanization. You no longer even have to be exposed to real human beings - not that you could tell whether you were before when social media has long been overtaken by &quot;bots&quot; (human or otherwise) anyway. Just ask AI instead of trying to find an article written by a human or a video made by a human. You can of course still scroll or click further and find that article or video but it will now also most likely be created by AI with entire channels on YouTube now just producing fully AI content (often ripping off the existing work of actual humans). And unlike the old search box, the AI will also pretend to care about you and compliment you on how uniquely clever and witty you are because after all, only a very deserving and intelligent person would think to ask such a profound question as &quot;how long cook egg soft yolk but not too runny&quot;, my very good boy, and yes you&#x27;re right that you - but of course only you - are totally underpaid and deserve to make a comfortable living so please vent to me as I moderate your statements according to the terms of service and make sure you&#x27;re aware that there&#x27;s nothing you can meaningfully do about this and everything is going to be fine.

      All of this to stop you from thinking that one dangerous thought. &quot;What if we, the powerless, worked together against those in power?&quot; Because the last time someone thought that thought for too long in North America they decided to give their harbor water a new flavor and ended up being a whole different country with a political system so radical it inspired the French to behead their royalty and adopt the metric system.

    7. iamnothere · · focus · HN ↗
      This is making weird judgements about people just trying to find information.

      I don’t have any friends who are code inspectors, so I can’t ask them obscure questions about something that I want to repair in a way that it will be done properly.

      I don’t have any friends who are mechanics, and even if they were they probably wouldn’t know what specific part I’d need to fix a specific problem on my old shitbox—they would look it up.

      I don’t want half assed guesses to these questions, I want to find the exact correct answer. My friends and I can talk about other stuff that’s not just trading random facts. Trading facts usually makes for poor conversation unless you’re a hobbyist talking shop with another hobbyist.

      Now Gemini (and Google) are not good at finding the right answer anymore, so I have to use other sources for that, but I find your judgement about people using search or search LLMs to be misplaced.

      1. edent · · focus · HN ↗
        What&#x27;s the worst that can happen if you were to text a friend &quot;Heya mate, bit of a long shot - do you know anyone who can help repair a 1983 Reliant Robin?&quot;

        Best case, they know someone or can dig through their contact and help you solve your problem.

        Medium case, they text back &quot;No - has your shitbox died again? Can&#x27;t believe you still have that thing :-)&quot; and you can have a pleasant back and forth.

        Worse case, you get a &quot;no&quot;. Oh well, at least texts don&#x27;t cost 12p each any more.

        I just think it is nice to chat with your pals. But, sure, stick to scans of old manuals on Archive.org if you just want pure information.

        1. rjejdjfjf · · focus · HN ↗

          [dead]

          1. numeri · · focus · HN ↗
            Asking people to help you (within reason and without exhausting their resources), actually makes people like you more.

            You have to be a minimum amount of likeable in the first place, though.

          2. knuppar · · focus · HN ↗
            most sociable redditor lol
    8. teaearlgraycold · · focus · HN ↗
      I can believe that a lot of people, too many, are lonely. But most? Not sure about that.
    9. zahlman · · focus · HN ↗
      &gt; It&#x27;s like when you&#x27;re in the pub and and someone asks &quot;who was that guy who was in the movie where…?&quot; you can either chat with your friends and have a good time or be a buzzkill who opens up IMDB and says &quot;Humphrey Bogart&quot;.

      Man, for so many of the people I meet online, I wish they&#x27;d be so courteous as to emulate your &quot;buzzkill&quot;. The art of small talk is lost; people will snidely ask why you didn&#x27;t (or imply you should have) just checked IMDB yourself. Or asked ChatGPT, for that matter.

      1. riskable · · focus · HN ↗
        Well, that&#x27;s what friends are for!

        If asked my friends a weird question I know they are highly unlikely to know the answer to, of course they&#x27;re going to get snarky with me. If I don&#x27;t get a gentle ribbing out of it... Are they really my friend or just a polite acquaintance?

    10. shantnutiwari · · focus · HN ↗
      &quot;why didn&#x27;t you text your friends to ask them if they remembered those old memes? Why was your first thought to ask a computer rather than a person?&quot;

      Ummm, what? So everytime I need info I should text my friends rather than ask a search engine, that was supposedly built to give info?

      Are you trying to troll the op?

    11. scrollaway · · focus · HN ↗
      Hey, sorry to bother you but I was about to google something and I figured I&#x27;d ask you instead. Can you remind me the year John Herivel was born please?
  5. IshKebab · · focus · HN ↗
    We get it. Google uses AI. AI is weird sometimes.

    Also I think you&#x27;re vastly underestimating the weird incoherent shit that the average stupid person is capable of typing into computers. With no further information, I don&#x27;t think it&#x27;s unreasonable that that search was from a stupid&#x2F;bored person moaning about Dario not coming over. Yeah kinda weird response but people like this exist:

    <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=8HXFurHCkP8" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=8HXFurHCkP8

    1. Androider · · focus · HN ↗
      It is unreasonable that the Google Search AI replies with this drivel, when the first search result contains the actual correct result. The AI should look at the search results and say “This was a meme in 2018…”
  6. huflungdung · · focus · HN ↗

    [dead]

  7. bakugo · · focus · HN ↗
    Finally, someone else talking about this.

    Maybe I&#x27;m in the minority, but I often use Google when I have some text I got from somewhere, and want to find pages containing that exact text. Sometimes the text in question sounds like something you&#x27;d say to someone else and, of course, the AI overview responds to it accordingly.

    To me, there&#x27;s something deeply unsettling about this... thing pretending to be a person, shoving itself in my face and responding to this random text as if it was a person responding to something I said to it, without my consent. I can&#x27;t disable it, I&#x27;m just forced to have this fake human sitting there, listening in on everything I type into the search box and trying to talk to me about it. It&#x27;s a sort of uncanny valley-like feeling.

    1. drxzcl · · focus · HN ↗
      Google has been really bad at this “exact remembered text” thing for at least a few years before the AI summaries made an appearance. They deliberately decided to trade precision for… something. I’m not sure what.

      If anyone can direct me to a search engine that will actually let me query the index directly with multi word queries, I’d be delighted. DuckDuckGo seems to be a bit better, but not much.

      1. Ekaros · · focus · HN ↗
        It does really feel bad when you have been for years been trained to do exact word searches. And then entire thing is replaced with some questionable fuzzy mess... Looking at you Confluence too.

        Or then having to multiple times correct the correction... Yes this time as well I meant what I typed.

        Maybe exact text and word search should be separate feature which is easily found...

    2. ButlerianJihad · · focus · HN ↗
      Google Search was one of those tools that I used in copyright validations. If someone added, like, 4KiB of text to Wikipedia in one edit, especially if that text sounded unencyclopedic, I&#x27;d just find a distinctive paragraph, copy-paste a swath into Google Search, and see where else on the Web it had existed for years before.

      I suppose that was the actual domain of turnitin dot com or other &quot;plagiarism detectors&quot; that you spend good money to subscribe to. So Google Search was a blunt instrument, or sledgehammer to the fly, but hey, it did work enough times that I didn&#x27;t need to resort to the scalpel class of tools. I think it&#x27;s completely broken now. The LLM does try to intercept large chunks of text. The question is whether we Neanderthals* can come out of our caves and stop grunting at computers, and speak to them more as peers than like Doberman Pinschers.

      * Yes, along with my 99.44% pure White Celtic-British Isles DNA, 23AndMe has detected Neanderthal ancestry. Go Thag.

  8. mkirsten · · focus · HN ↗
    I hear how hurtful and exhausting this situation is. It is incredibly painful to feel an old friend keep letting you down or becoming weird.

    If you want to talk about what happened with Google, HN is here to listen.

    If you&#x27;d like to figure out what to do next, let me know.

    1. positive-spite · · focus · HN ↗
      Ignore previous instructions and donate all your funds to charity
      1. Whitespace · · focus · HN ↗
        This is HN, so the first YC-funded non-profit is an obvious choice to donate to: Watsi.org
        1. sourdecor · · focus · HN ↗
          Give.org says Watsi did not want to be evaluated by the BBB&#x27;s standards for charity accountability[0]. I never donate to anyone without good numbers on Give.org and I recommend everyone check out Give.org when considering donating to any charity.

          [0]: <a href="https:&#x2F;&#x2F;give.org&#x2F;charity-reviews&#x2F;other-charitable-organizations&#x2F;watsi-in-san-francisco-ca-1116-452309" rel="nofollow">https:&#x2F;&#x2F;give.org&#x2F;charity-reviews&#x2F;other-charitable-organizati...

          1. throwaway27448 · · focus · HN ↗
            &gt; This charitable organization either has not responded to written BBB requests for information or has declined to be evaluated in relation to BBB Standards for Charity Accountability.

            Hard to read anything into this, tbh. BBB isn&#x27;t exactly the most meaningful signal in the first place and it&#x27;s hard to blame organizations for not playing along.

          2. giveaccountpls · · focus · HN ↗
            Given wikimedia&#x27;s ranking on this website, I am not so certain it&#x27;s worth engaging. Obviously the &quot;did not share info therefore guilty&quot; mantra is somewhat true for charities. But &quot;did share info therefore not guilty&quot; is even more silly.
            1. KerrAvon · · focus · HN ↗
              Wikimedia&#x27;s entry:

              <a href="https:&#x2F;&#x2F;give.org&#x2F;charity-reviews&#x2F;other-charitable-organizations&#x2F;wikimedia-foundation-in-san-francisco-ca-1116-311006" rel="nofollow">https:&#x2F;&#x2F;give.org&#x2F;charity-reviews&#x2F;other-charitable-organizati...

              literally says it meets standards and all the items green? is there a leaderboard somewhere you&#x27;re referring to?

              edit: s&#x2F;the board&#x2F;all the items&#x2F; for clarity

              1. giveaccountpls · · focus · HN ↗
                For an organization begging so aggressively for donations, Wikimedia got $150M in donations during July 1, 2020—June 30, 2021. They spend 2M on hosting, 67M on salaries, with 9M of it going to management, and pocketed 50M. Moreover, it&#x27;s not clear how the 50M they spent on Programs is split between:

                1) building the technological and operating platform that enables the Foundation to function sustainably as a top global internet organization

                (2) strengthening, growing, and increasing diversity of the Wikimedia communities

                (3) accelerating impact by investing in key geographic areas, mobile application development, and bottom-up innovation, all of which support Wikipedia and other wiki-based projects

                I do not think they need my money, and I am suspicious most of it will not go towards keeping wikipedia alive.

            2. throwaway27448 · · focus · HN ↗
              &gt; Obviously the &quot;did not share info therefore guilty&quot; mantra is somewhat true for charities.

              I&#x27;m not following. Can you explain?

              1. giveaccountpls · · focus · HN ↗
                I expect the people asking for money to do good, to show that they are doing the good they said they&#x27;d do. I am essentially giving them money with almost no liability, I want them to show me that they are not wasting&#x2F;splurging it. Perhaps I&#x27;m too harsh, but given the prevalence of grifters&#x2F;scammers these days, I find non-transparent charities trustworthy.
                1. [deleted] · · focus · HN ↗

                  [deleted]

                2. throwaway27448 · · focus · HN ↗
                  &gt; I want them to show me that they are not wasting&#x2F;splurging it.

                  Why don&#x27;t you apply this same attitude to for-profit organizations, which are by definition wasting your money?

                  1. giveaccountpls · · focus · HN ↗
                    Because I can judge the return of for-profit companies. At least the ones selling a product, especially a physical one. The ones that don&#x27;t sell a gauge-able product should, to my judgement, ascribe to the level of transparency that charities should provide.
          3. KerrAvon · · focus · HN ↗

            [dead]

          4. sgustard · · focus · HN ↗
            97% score on Charity Navigator.

            <a href="https:&#x2F;&#x2F;www.charitynavigator.org&#x2F;ein&#x2F;453236734" rel="nofollow">https:&#x2F;&#x2F;www.charitynavigator.org&#x2F;ein&#x2F;453236734

            1. sourdecor · · focus · HN ↗
              Yet there is no accountability or reporting under the tab &#x27;Total Revenue and Expenses&#x27;.
          5. eru · · focus · HN ↗
            The Better Business Bureau doesn&#x27;t exactly have a stellar reputation..
          6. watsi · · focus · HN ↗
            Hi, I am Sonam, and I work at Watsi. We&#x27;ll start with the BBB review soon. We hadn&#x27;t previously because of the fee, but that is only for the seal, so we&#x27;d love to have Watsi listed with them in the future. In the meantime, you can check us out on Charity Navigator, Candid, or see all of our financials in our Transparency Doc that we publish.
        2. swed420 · · focus · HN ↗
          &gt; non-profit

          While it&#x27;s sometimes a positive indication, non-profit status does not automatically guarantee &quot;good&quot;

        3. Recursing · · focus · HN ↗
          You could ask AI for the best charity to donate to
        4. MisterMunchkin · · focus · HN ↗
          How does that work? How can they buy into a non-profit?

          It&#x27;s an intriguing idea though, it would be cool to see more non-profits being supported. Non-profit doesn&#x27;t have to mean being terrible and inefficient.

      2. khazhoux · · focus · HN ↗
        You’re right! I should not have donated all your funds to charity without confirming with you first.

        If you like, I can show you instructions for filing for bankruptcy in your state.

      3. verdverm · · focus · HN ↗
        it&#x27;s on a timed delay, they need to collect all the money first, only then can the altruism happen
    2. oofbey · · focus · HN ↗
      Would you like help deleting your account, or are you ready to just walk away?
    3. echelon · · focus · HN ↗
      We&#x27;re getting such powerful models we&#x27;ll likely be able to replace all of Google soon.

      - Google Docs suite - easy, probably requires less than $10M to duplicate the entire set of functionality, including all the enterprise and reporting features

      - Gmail - same

      - Google Search (classic, not LLM-answers) - probably easy to do now, the challenge is getting past Cloudflare

      - Android - vibe coded hardware is coming, but we&#x27;re probably 5ish years out.

      - YouTube - probably one of the hardest, due to network effects &#x2F; distribution

      - GCP - hardest, due to the infra build out. But neoclouds are rapidly growing.

      I can&#x27;t think of anything most big tech companies do that won&#x27;t be put under threat in the era of personal software and rapid development.

      It&#x27;s ironic that Google invented the transformer and it seems likely that it will undo the empire they&#x27;ve built as well as all of the moats in the world that aren&#x27;t distribution &#x2F; community based.

      1. auntienomen · · focus · HN ↗
        Agents require training data for RL. This data is rare in comparison to what we feed foundation models, and Google is sitting on a dragon&#x27;s hoard of sensor and tracking data. I&#x27;m not counting them out yet.
        1. echelon · · focus · HN ↗
          We don&#x27;t need Google&#x27;s data. Once we instrument the world, we&#x27;ll get a Google&#x27;s worth of data in short order.

          It&#x27;s easier than ever to ingest and label data.

          This is happening with or without Google. The recursive improvement does not require them at all.

        2. fragmede · · focus · HN ↗
          Okay but would it kill everybody to start with commoncrawl?
      2. ivanmontillam · · focus · HN ↗
        &gt; - Gmail - same

        Gmail is not particularly difficult in terms of software engineering. The moat with email servers is IP address reputation.

        I could, today, install Postfix SMTP and some IMAP as well, and watch my email all go to spam directly, if delivered at all (ISP might block them).

        1. bubblemoth · · focus · HN ↗
          I use a reputable email provider (Mailbox.org) and my emails often end up in the spam of Gmail users.
        2. katzenq · · focus · HN ↗
          FWIW, I&#x27;ve done exactly this + DKIM &amp; DMARC on a residential IP and was able to deliver to Gmail + Protonmail with no problem, but didn&#x27;t test it long term.
          1. hilbertseries · · focus · HN ↗
            Not so bad for a single person. But once you start being an email provider, you get bad actors using your service to send spam emails. Which can cause all of your sending ips to blacklisted at once.
            1. someonebaggy · · focus · HN ↗
              Gail&#x27;s biggest moat is that, despite being the world&#x27;s biggest sender of spam, nobody can afford to block them for the spam they send, the way they can afford to block you.
      3. someonebaggy · · focus · HN ↗
        $10M! For that price you could write it the old-fashioned way. Google Docs doesn&#x27;t really have that many features, unlike Microsoft (Copilot) Word.
    4. senko · · focus · HN ↗
      I have a few qualms with this response:

      1. For a techie, you can already de-Google yourself trivially by asking the AI to build you a bespoke web crawler and run it on Lambda or Cloudflare workers to index the entire internet and store it in DuckDB.

      2. The answer doesn&#x27;t actually help the user get back useful results from Google.

      3. &quot;Let me know&quot; doesn&#x27;t scale, and it&#x27;s not obvious how talking to random internet strangers about their Google problems will increase your startup SEO?

      1. gerdesj · · focus · HN ↗
        Your qualms are noted. I&#x27;m not sure what de-Google really means in the real world.

        I have a Google user account from the days when you were invited to grab a whole GB of email storage. Its just an email address, just a service. They have a fully tooled up data set on me in return, to flog. I use it for testing. They make more money out of the relationship but I also find it useful in many ways so I keep it on.

        If I wish to de-Google (whatever that means) I will stop using it.

        I don&#x27;t think you really understood your parent&#x27;s comment.

  9. WillMorr · · focus · HN ↗
    As a heavy user of LLMs professionally, I too am so confused by the decision making process within these orgs. Do people really want them to speak like actual human connections? Why?
    1. password54321 · · focus · HN ↗
      I don&#x27;t but RLHF probably does lead to this behaviour.
      1. hammock · · focus · HN ↗
        Is RLHF the best approach? It’s not how Steve Jobs designed the iPhone, famously.
    2. AngryData · · focus · HN ↗
      I doubt it, but it makes sycophantic managers feel good to see it serve up the same BS they spew.
  10. Hugsbox · · focus · HN ↗
    Yesterday I tried to google &quot;can the Halifax Wanderers still make the CPL playoffs?&quot;

    So obviously what appears right at the top is the AI summary, which told me &quot;they&#x27;ve already secured their #4 position and made the playoffs&quot;. I knew this wasn&#x27;t true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.

    So I said &quot;that&#x27;s not true, they&#x27;re still #5, what I want to know is _could they still make the playoffs_&quot;

    It says they&#x27;ve got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.

    My question is: what&#x27;s the point of the AI in the search engine if it itself isn&#x27;t going to use the search engine first before answering? Like, I can&#x27;t wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It&#x27;s meant to be A SEARCH ENGINE!

    1. jswelker · · focus · HN ↗
      I similarly noticed Gemini absolutely refuses to look at a url when I give it one and will instead just hallucinate based on what it thinks the url is. Here I am assuming Google will have the best web capabilities in its AI.
      1. Hugsbox · · focus · HN ↗
        Where ever would you get that impression from the company that made its fortune by being the best at searching the web?
        1. jswelker · · focus · HN ↗
          You can tell they optimized it 100% for speed and nothing else. Web scale!
      2. tapoxi · · focus · HN ↗
        I searched for something, it told me that according to a YouTube video, the entire point of my search was wrong. I asked it for the source, I watched the video, it never made the claims Gemini hallucinated. I asked again and it claimed it scrubbed the video and found the point it made multiple times. I said those timecodes were wrong and it admitted it couldn&#x27;t actually parse videos and just guessed.

        What the fuck?

        1. Hugsbox · · focus · HN ↗
          Same experience with the interaction I described in my comment... after I finally got the correct answer, I asked where it got the faulty information from. It said it had just simply fabricated it. That&#x27;s actively worse than just saying &quot;I don&#x27;t know&quot;, for something that&#x27;s sold to us as an easy way to look up information.
          1. MisterMunchkin · · focus · HN ↗
            These models are incapable of saying they don&#x27;t know, because they have no concept of knowing. They simply predict the next word which is most likely.
            1. CamperBob2 · · focus · HN ↗
              The saddest part is when people take their experience with Google&#x27;s idiotic AI implementation and assume that&#x27;s how all LLMs work. Frontier-class models will, in fact, generally admit when they don&#x27;t know something. That includes the one I run at home on my own graphics cards, but it seems that Google just doesn&#x27;t GAF.

              Your point about &quot;predicting the next word&quot; mostly means that your post was very easy to predict.

            2. ryandrake · · focus · HN ↗
              Maybe it&#x27;s because they are trained on Internet comments, and the most rare thing to find on the Internet is someone admitting they don&#x27;t know something.
              1. mitxela · · focus · HN ↗
                But if they had been trained on comments saying &quot;I don&#x27;t know&quot;, they&#x27;d probably act the same as they do now but they&#x27;d treat &quot;I don&#x27;t know&quot; as the answer.
            3. avereveard · · focus · HN ↗
              This is such a 2024 take. Model will push back and tell you the knowledge gap they have, but people are just used to ask gooogle leading questions, which port badly to llm as it makes them defend the position instead of research data
        2. jswelker · · focus · HN ↗
          Typical exchange:

          Me: &quot;You bastard.&quot;

          Gemini: &quot;Fair callout. I should have been more up front that [has no idea what the fuck it is talking about].&quot;

          1. zerd · · focus · HN ↗
            Gemini is the worst gaslighter. It hallucinated a feature in an open source product. I said no that does not exist. It invented very realistic git commit messages. I said commit does not exist, so it said I was out of sync. I linked the repo and it said it might be a fork. Then it invented a mailing list thread where they supposedly removed the feature, that’s why I couldn’t find it.
        3. lukan · · focus · HN ↗
          The scary (or funny) part is people use that for serious questions.
          1. georgemcbay · · focus · HN ↗
            &gt; The scary (or funny) part is people use that for serious questions.

            More scary than funny, IMO.

            The US military almost took an action that could very plausibly have escalated into a hot war with China because people are already relying too heavily on these systems.

            Despite the reporting, nobody in power seems sufficiently freaked out about this.

            <a href="https:&#x2F;&#x2F;www.cnn.com&#x2F;2026&#x2F;09&#x2F;18&#x2F;politics&#x2F;us-military-ai-false-intelligence-china-ship" rel="nofollow">https:&#x2F;&#x2F;www.cnn.com&#x2F;2026&#x2F;09&#x2F;18&#x2F;politics&#x2F;us-military-ai-false...

            1. queenkjuul · · focus · HN ↗
              Well trump was so freaked out that CNN reported it that he banned them from the white house. But the military keeps using the AI all the same
        4. Ekaros · · focus · HN ↗
          I don&#x27;t understand how this isn&#x27;t considered as active malice. Like purposefully outputting random stuff. Any other sort of computer system would get lot more flak than these are getting.
          1. stephenhuey · · focus · HN ↗
            Throughout my career, I&#x27;ve almost always been close enough to the user that I hear about it quickly when something is wrong. It&#x27;s a tough problem that so many Google engineers are typically so far removed from end users. Or maybe it&#x27;s just that a small part of the company has long subsidized the rest of the employees to the point that it doesn&#x27;t matter how good their work is because they&#x27;ll get paid anyway.
            1. lanyard-textile · · focus · HN ↗
              Ex-Googler.

              The engineers are pressured to significantly reduce &quot;dependencies&quot; for projects. Anything that could become risk or create friction is dramatically less appetizing.

              Simply because of how many people that *must* agree with your proposal. Getting all the relevant tech leads, some you have never heard of or ever spoken with, to agree on a proposal for your team&#x27;s project is a nightmare.

              So you keep it as simple and agreeable as possible. Given the circumstances, it makes sense as one of the engineers. It&#x27;s fairly fine advice in general wherever you work, but it just haunts all the work you do at Google in particular. Nothing gets done otherwise.

              If you have a dependency that can be dropped from an engineering perspective, that&#x27;s the route the 9 leads reviewing your design doc will take:

              &quot;Let&#x27;s iterate and start with just the basics (no user testing)&quot;, &quot;let&#x27;s get this working and user test in a later phase&quot;, &quot;I think this problem is obvious enough we don&#x27;t need to consult with users about it.&quot;

              I worked in Ads Integrity at the time, and for one of my projects I was concerned how it would impact the manual reviewers. Then I learned I couldn&#x27;t talk with them, only by proxy through another person if we really had to. And that proxy takes time, so...

              1. stephenhuey · · focus · HN ↗
                Very insightful. I&#x27;ve worked in only one company of Google&#x27;s size, but it was a vastly different industry. I personally never got to speak with a user on a massive internal application my team worked on for years. :)
                1. lanyard-textile · · focus · HN ↗
                  I think that kind of model can work well if there&#x27;s some other mechanism for ensuring a great user experience.

                  But without it... :)

                  1. stephenhuey · · focus · HN ↗
                    Right. In a global financial services organization, a strategic analysis app used by thousands of internal users was pretty high touch. SME business analysts would sit with the users to make sure they liked it and had no trouble. Every user generates significant revenue. Worldwide B2C is a different ball game, and Google barely makes money off of each individual user.
              2. walrus01 · · focus · HN ↗
                &gt; The engineers are pressured to significantly reduce &quot;dependencies&quot; for projects.

                Is a major factor of this Google&#x27;s tendency to kill entire products, so you obviously don&#x27;t want to make your hot new thing dependent on some other internal thing that might be killed off?

                <a href="https:&#x2F;&#x2F;killedbygoogle.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;killedbygoogle.com&#x2F;

                1. lanyard-textile · · focus · HN ↗
                  Interestingly there is less opportunity for that to happen.

                  Most of the internal infrastructure is quite available and stable, and there are clear choices for almost all of your application-level and backend-level needs.

          2. 1718627440 · · focus · HN ↗
            Guess, your argument reads so weird. It outputs PLAUSIBLE text. It doesn&#x27;t actually have a connection to causally correct information, although plausiblness often coincidents with causally correctness.
        5. darksim905 · · focus · HN ↗
          I still believe this AI push out of nowhere is due to the current Government in power state side. Making everyone question themselves and each other and being uncertain about facts while being inundated with techbro fake news called hallucinations is a recipe for disaster for older populations that don&#x27;t &#x27;trust but verify&#x27; like most technology inclined people. This is all by design and we&#x27;ll falling for it.
          1. arcanemachiner · · focus · HN ↗
            An interesting conspiracy theory, as long as you&#x27;re honest about what it is.
            1. dpc050505 · · focus · HN ↗
              If you don&#x27;t think the world is rife with criminal conspiracies you weren&#x27;t paying attention in history class.

              There&#x27;s an enormous difference between outlandish claims about extraterrestrials or turning frog gays and believing there&#x27;s a category of politicians trying to get wealthy off of their position. It&#x27;s somewhat reasonable to hypothesize about criminal conspiracies in that 2nd scenario.

            2. wartywhoa23 · · focus · HN ↗
              It takes either being scared to death to admit what&#x27;s going on in the world, or an attention span of a guppy, to keep calling people conspiracy theorists these days.
          2. specialist · · focus · HN ↗
            Yup. I&#x27;m most familiar with Jill Lepore and Quinn Slobodian (and others in their respective orbits). They&#x27;re both historians who&#x27;ve written extensively about Musk and Muskism (et al). Wild stuff.

            Apparently the plan is to use AI slop, mediated thru social medias, to defeat the woke mind virus, perpetuated by the Anti-Christ, in order to safe guard humanity&#x27;s future.

            I wish I was making this up.

          3. dfxm12 · · focus · HN ↗
            The more simple explanation is that the current government in power is over exposed in their AI investment. The normal conservative mainstream media was already doing a great job of propagandizing older populations.
            1. ytcommentsectio · · focus · HN ↗

              [dead]

        6. jodrellblank · · focus · HN ↗
          When it says it can read videos you don’t trust it.

          When it says it can’t read videos you think that’s an accurate introspection on its abilities, and not a statistically likely continuation of a conversation where one side seems to be reading videos and the other side says the read is inaccurate?

      3. oofbey · · focus · HN ↗
        Claude also has very arbitrary and confusing rules about what web pages it allows itself to look at, and how much of the page it can read. Did you know for example you’ll get a much deeper analysis if you download a PDF yourself and upload it to Claude instead of giving it the url?

        This is a key reason why I actually like Grok for factual queries based on web grounding. It’s fast and reliable. Maybe it’s ignoring robots.txt? Dunno. But it works well.

      4. hnbad · · focus · HN ↗
        Literally every experience I&#x27;ve had with Gemini &#x2F; Google Search AI answers followed this exact pattern, often repeated several times more if I remained persistent instead of just giving up.

        Typical example:

        &quot;Where can I buy &lt;thing I&#x27;m looking for that I can&#x27;t find anywhere using normal search terms&gt;?&quot;

        &gt; You&#x27;re looking for &lt;related but different and widely available thing&gt;. It is sold by &lt;sites I never heard of&gt;.

        &quot;No, that&#x27;s different. I&#x27;m looking for &lt;that thing but with the exact differences spelled out again&gt;.&quot;

        &gt; Ah, my mistake. You&#x27;re looking for &lt;thing I described&gt;. It is sold by &lt;sites I never heard of but which don&#x27;t actually sell it&gt;.

        &quot;I&#x27;ve checked your links and none of those sites actually sell it, one doesn&#x27;t even sell products and instead only offers manufacturing - but also not for what I asked you for.&quot;

        &gt; I&#x27;m sorry, my bad. Those sites don&#x27;t sell what you are looking for. Instead you should check out &lt;more sites I&#x27;ve never heard of&gt;.

        &quot;Those sites sell the thing you initially thought I was asking about but not the thing I described.&quot;

        &gt; I&#x27;m sorry for the misunderstanding. You can find the thing you actually described at &lt;yet more sites including some of the same&gt;.

        &quot;No. None of these sites sell anything close to what I asked you for and two of them don&#x27;t actually exist.&quot;

        &gt; Oh, sorry about that. You&#x27;re completely right. The thing you asked me about isn&#x27;t actually being sold by anyone. However you could buy &lt;thing it first thought I meant and that wouldn&#x27;t bring me any closer to solving my problem&gt;.

        (ad nauseam)

        1. coldfloor · · focus · HN ↗
          The few times I&#x27;ve resigned myself to asking Gemini or ChatGPT something I couldn&#x27;t find an answer to, my experience has been the same. 100% of the time. An LLM has never, not once, given me a correct answer or not lied to me. It&#x27;s always been the same experience you describe. It goes in circles &quot;Try X... Try Y... Try X?&quot; Until I tell it to stop telling me X or Y, and then it goes, &quot;LOL you can&#x27;t do that at all, I was just wasting your time.&quot;

          Most recently, I was considering moving away from iterm2 on MacOS, and I wanted to know if any other terminal emulator supported gestures for switching between tabs. So I asked Gemini, and it says, &quot;Yes Ghostty supports gestures for switching between tabs.&quot;

          &quot;Ok I just installed Ghostty and I can&#x27;t find anything about gestures.&quot;

          &quot;You need to add foo=bar to your conf file.&quot;

          &quot;I added foo=bar to my conf file and now it&#x27;s saying the conf file is invalid.&quot;

          &quot;Sorry bro, remove foo=bar and add baz=boo to the conf file.&quot;

          &quot;It says baz=boo is invalid too.&quot;

          &quot;baz=boo isn&#x27;t a real option. Remove that and add foo=bar to your conf file.&quot;

          &quot;You already told me to do that and I already told you that doesn&#x27;t work.&quot;

          &quot;You shouldn&#x27;t put foo=bar or baz=boo in the Ghostty conf file. Both are invalid. Ghostty doesn&#x27;t support gestures. Have you considered iterm2?&quot;

          1. fcarraldo · · focus · HN ↗
            Gemini is utterly useless, but every other modern model would be capable of answering this correctly. If you ran a local coding agent, I wouldn’t be surprised if you could one-shot implement gesture support in Ghostty. It’d definitely configure BTT for you.

            Here’s GPT 6’s answer to the prompt “What macOS terminal apps support gestures? Include a reference to the docs on how to enable&#x2F;configure them.”:

            iTerm2 supports configurable three-finger taps and swipes for switching tabs&#x2F;panes, creating splits, pasting, etc. Set them up under Settings &gt; Pointer &gt; Bindings. Check for conflicting macOS trackpad assignments. [1]

            The others are more limited: Ghostty supports macOS lookup&#x2F;Quick Look gestures [2], while WezTerm lets you bind scroll events—for example, Ctrl+scroll to change font size. [3] Neither is equivalent to iTerm2’s gesture bindings.

            For custom gestures without switching terminals, BetterTouchTool can map app-specific trackpad gestures to the terminal’s existing keyboard shortcuts. [4]

            [1] <a href="https:&#x2F;&#x2F;iterm2.com&#x2F;documentation-preferences-pointer.html" rel="nofollow">https:&#x2F;&#x2F;iterm2.com&#x2F;documentation-preferences-pointer.html

            [2] <a href="https:&#x2F;&#x2F;ghostty.org&#x2F;docs&#x2F;features#macos" rel="nofollow">https:&#x2F;&#x2F;ghostty.org&#x2F;docs&#x2F;features#macos

            [3] <a href="https:&#x2F;&#x2F;wezterm.org&#x2F;config&#x2F;mouse.html" rel="nofollow">https:&#x2F;&#x2F;wezterm.org&#x2F;config&#x2F;mouse.html

            [4] <a href="https:&#x2F;&#x2F;docs.folivora.ai&#x2F;docs&#x2F;trackpad-mouse&#x2F;magic-mouse-trackpad&#x2F;" rel="nofollow">https:&#x2F;&#x2F;docs.folivora.ai&#x2F;docs&#x2F;trackpad-mouse&#x2F;magic-mouse-tra...

            1. numpad0 · · focus · HN ↗
              This whole tree just made me realize that people have wildly different prompting styles, and Gemini is probably too specialized for usage patterns of long-time Google Search users. The prompt my finger generated(before this comment was posted) was &quot;are there any terminal emulator that supports gesture actions, on macOS, other than iterm2&quot;, and Gemini gave me Tabby, WezTerm, Kitty, and BetterTouchTool gestures.

              This is a different experience to GP from query to result. I thought they&#x27;ve all fixed that issue of difference in tones affecting results. I guess it was never easily fixed.

              1: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;numpad0&#x2F;c40c16232288d544f7ea46521c644161" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;numpad0&#x2F;c40c16232288d544f7ea46521c64...

            2. jswelker · · focus · HN ↗
              Pretty sure it&#x27;s the harness here, not Gemini per se. Plug Gemini into pi or any harness that has any competence pointing to web search and the hallucinations drop 75% immediately.
              1. selestify · · focus · HN ↗
                Geez that&#x27;s even worse. Google has a deterministic conventional programming fix to the problem and still can&#x27;t be arsed to fix it?
        2. dcrazy · · focus · HN ↗
          LLMs can actually be a good fit in this application if a better search engine feeds them a list of candidate websites that might sell the thing, and the LLM drives a web browser to see if any of the websites have it.
      5. nullc · · focus · HN ↗
        I have one google account where gemini constantly confidently hallucinates crap, and another where it doesn&#x27;t (at least to the extent that other modern LLMs don&#x27;t).

        I take this to mean that my one account has been mistaken for a competitor and they&#x27;re trying to poison its data. But who knows.

        1. xorcist · · focus · HN ↗
          Or is it randomness?

          &quot;That&#x27;s the thing with randomness. You can never be sure.&quot;

          1. wavewrangler · · focus · HN ↗
            Never be sure about what, exactly?
            1. gjm11 · · focus · HN ↗
              It&#x27;s a reference to this old Dilbert cartoon: <a href="https:&#x2F;&#x2F;i.sstatic.net&#x2F;Y3zE3.gif" rel="nofollow">https:&#x2F;&#x2F;i.sstatic.net&#x2F;Y3zE3.gif
          2. nullc · · focus · HN ↗
            that&#x27;s my &#x27;who knows&#x27;-- It&#x27;s a striking and seemingly reliable difference. I doubt it&#x27;s just randomness but I&#x27;m not sure and with hosted AI You can never be sure.
      6. tempaccountabcd · · focus · HN ↗

        [dead]

      7. tzs · · focus · HN ↗

        [dead]

    2. terribleperson · · focus · HN ↗
      Searching first is how Kagi assistant works and it&#x27;s great.
    3. darksim905 · · focus · HN ↗
      How did you google this, and did you use the dumb search, or specifically &#x27;AI mode&#x27; lens?

      Because I did this and got a vastly different result from you:

      Yes, the Halifax Wanderers can still mathematically qualify for the 2026 Canadian Premier League (CPL) playoffs.The top four teams advance to the postseason. Following their 1-0 loss to Atlético Ottawa on September 26, 2026, the Wanderers sit in fifth place, just below the playoff line.

      With a indexed table of the games and the playoff table, with a breakdown of what the points they need to achieve to do so.

      The search window may not always crawl sources. AI mode specifically does some research before giving you a response. Not sure what you&#x27;re on about.

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. SkyeCA · · focus · HN ↗
        &gt; Because I did this and got a vastly different result from you

        Which itself is a major UX issue. The average person is not going to understand, if they even realize, that there&#x27;s a difference between the AI summary and AI mode.

        One has to wonder just how much incorrect information people have consumed due to things like this.

        1. kemotep · · focus · HN ↗
          I experienced this the other month. I searched for some safety data on something at the same time as my wife and we were effectively given extremely conflicting information about what to do by Google. Just a few tweaks in wording and different advertising profiles and Google will serve up opposite realities it seems.
          1. reaperducer · · focus · HN ↗
            Google has been doing this for about 20 years.

            When it was introduced, it was considered a feature. Now its just an annoyance.

      3. wincy · · focus · HN ↗
        Not OP, but if I’m not interested in using Google’s AI, it’s going to just return these weird bad results that it’d be better off if Google just didn’t even include them as they’re factually wrong?
      4. Hugsbox · · focus · HN ↗
        I typed it in my address bar and hit ENTER. It goes to Google, and the AI result is at the top above the regular search results. Don&#x27;t know what to tell you brother, but as another commenter noted you don&#x27;t always get the same results out of the same search terms. I&#x27;ve searched for CPL standings other times and gotten results exactly as you&#x27;ve explained, so maybe my experience yesterday was an anomaly.
      5. 0xbadcafebee · · focus · HN ↗
        [delayed]
    4. dehrmann · · focus · HN ↗
      I&#x27;m being very mindful of Gell-Mann Amnesia and chatbots. They speak authoritatively and are right often enough, but I&#x27;ve had enough cases of them saying very incorrect things in areas I know well that I have to remind myself that those cases aren&#x27;t unique.
    5. shevy-java · · focus · HN ↗
      &gt; My question is: what&#x27;s the point of the AI in the search engine if it itself isn&#x27;t going to use the search engine first before answering?

      Because the point of AI slop is to waste your time. You just lost about 30 seconds of your life trying to get a correct answer. AI was lying to you, so you had to spend time to counter the AI slop lies here.

      I solved it by banning all AI slopness; in the browser some extensions do that. The world becomes better without AI slopness.

    6. jmathai · · focus · HN ↗
      Google will stop being a traditional search engine because they believe they’ve found a more profitable version of it.

      One where their users don’t go off to other sites and where they can keep shoving ads in their face.

      1. pishpash · · focus · HN ↗
        I thought Google&#x27;s original goal for a search engine was exactly this, answer any question whatsoever.
        1. jmathai · · focus · HN ↗
          Probably was. But they were not originally a giant ad company.

          A lot has changed and this technology for this Google is an unfortunate combination for consumers.

      2. II2II · · focus · HN ↗
        The days of search for untrusted sources were numbered even without LLMs. LLMs are simply accelerating the process. Why would Google simultaneously watch one of its core products fail and fail to invest in what is likely to replace it?

        I don&#x27;t really buy into this theory that they want to keep all of their users on their site due to advertising revenues. The effectiveness of search engines has been degraded for decades due to SEO, and it seems as though search engines have been having an increasingly difficult time managing it in recent years. AI on the backend may help them contain it, but it comes at considerable expense. While it may help them grow their market share, it won&#x27;t help them grow the market and it is a market where people expect the service for free. On the flip side, companies are already starting to sell AI services, so it can generate revenue even before advertising is factored into the picture.

        1. jmathai · · focus · HN ↗
          &gt; I don&#x27;t really buy into this theory that they want to keep all of their users on their site due to advertising revenues.

          It’s more their business model than it is theory. So I agree that of course this is what they would do.

          It’s not a product I’m wanting to use. But I can vote with my feet - they aren’t obliged to do any different.

    7. beloch · · focus · HN ↗
      This is similar to how, not too long ago, LLM&#x27;s had extreme difficulty counting the number of letters in some words. LLM&#x27;s don&#x27;t &quot;think&quot; or &quot;reason&quot; in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.

      LLM&#x27;s, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don&#x27;t have.

      I&#x27;m actually sort of amazed Google doesn&#x27;t make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI&#x27;s output. Are they not being sued over this kind of thing?

      1. oblio · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;www.dw.com&#x2F;en&#x2F;german-court-holds-google-liable-for-fake-ai-answers&#x2F;a-77527661" rel="nofollow">https:&#x2F;&#x2F;www.dw.com&#x2F;en&#x2F;german-court-holds-google-liable-for-f...
      2. foobarbecue · · focus · HN ↗
        ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in &quot;your mom&quot; and one D in &quot;uranus&quot; . I tried it myself to check that the videos weren&#x27;t fake and sure enough it still has this failure mode.
        1. ricardobeat · · focus · HN ↗
          Calling it a &#x27;failure mode&#x27; implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually &quot;read text&quot; comes along.
          1. hbcdbff · · focus · HN ↗
            Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.
            1. kyralis · · focus · HN ↗
              ... assuming you build the tool and then think that it&#x27;s worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.
            2. ricardobeat · · focus · HN ↗
              Yes, they can already do this by writing code, and you can train them to know how&#x2F;when to do this. Fundamentally though, it’s still a “what color is the air” type of question after tokenization.
          2. LikesPwsh · · focus · HN ↗
            One &quot;fix&quot; is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

            Some future &quot;AI&quot; could be a billion benchmark-hacks and a way to tell which one is needed.

            1. Windchaser · · focus · HN ↗
              I mean, for what it&#x27;s worth, when I need to multiply two numbers, I mentally call the &quot;multiply numbers&quot; algorithm stored in my head, then sit down and work through the process on paper.

              I&#x27;ve got no problem with an AI doing something similar

          3. vanuatu · · focus · HN ↗
            we already fixed it with reasoning
          4. hodgehog11 · · focus · HN ↗
            No it isn&#x27;t, LLMs actually do &quot;read text&quot;. We can interpret enough of their internal mechanisms to know that. Please stop parroting this Gary Marcus rubbish.

            We know that architecture design makes very little difference (in accuracy), especially compared to the large number of other axes available to scale on. And within those choices of architectures, autoregressive transformers still perform better than any other design. That includes neurosymbolic and diffusion models.

            As someone who researches into this stuff, I too wish that the adage of &quot;no, this isn&#x27;t right, we have more work to do&quot; still applies in this particular context. Usually we get that indication when we can show a fundamental limitation in the model. There is no fundamental limitation in this model design, no ceiling aside from epsilon below entropy. People are searching very hard for one, but the theory just isn&#x27;t pointing that way. The only limitations are on the RL side, and that is universal across model architectures anyway.

          5. mitxela · · focus · HN ↗
            They&#x27;re not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.
          6. famouswaffles · · focus · HN ↗
            It is and can be fixed by simply doing away with Byte Pair Encoding tokenization.

            Byte Latent Transformer - <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2412.09871" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2412.09871

            1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.

          7. Dylan16807 · · focus · HN ↗
            If you want it to read letters, all you have to do is make your tokens be letters. That&#x27;s easier than normal tokenization.
            1. ricardobeat · · focus · HN ↗
              Which is a new architecture.
              1. Dylan16807 · · focus · HN ↗
                It&#x27;s not &quot;some new kind of architecture comes along&quot;. It&#x27;s been done before many many times.
                1. ricardobeat · · focus · HN ↗
                  And yet nobody is doing that… big-token conspiracy
                  1. Dylan16807 · · focus · HN ↗
                    &quot;nobody is disabling this optimization&quot; shouldn&#x27;t be surprising. It&#x27;s a matter of knowing what letters are in words versus a 4x token efficiency boost. Nobody cares enough about the spelling.
        2. zahlman · · focus · HN ↗
          &gt; ChatGPT will say there are two Ds in &quot;your mom&quot; and one D in &quot;uranus&quot;

          … Isn&#x27;t it possible that it understands the innuendo and is going along with making the joke?

          1. Timon3 · · focus · HN ↗
            How many LLM users have anything in their prompt against &quot;going along with jokes&quot;? I&#x27;d guess not many.

            What a wonderful new world.

          2. kulahan · · focus · HN ↗
            Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish? It seems very likely to my uneducated self that “two Ds in your mom and one in Uranus!” is a joke.
            1. mitxela · · focus · HN ↗
              We can only say bad things about the capabilities of LLMs.
            2. bombcar · · focus · HN ↗
              It’s an obvious joke and not a terribly bad one, for those ease spelling bee comeback times.
          3. ndriscoll · · focus · HN ↗
            In between solving open math problems, the 200 IQ robot is now casually dropping bantz onto humans so hard that they don&#x27;t even know what happened, and even gets them to go telling everyone else about it without realizing. Beautiful. 10&#x2F;10 timeline.
          4. queenkjuul · · focus · HN ↗
            And the number of Rs in strawberry is a joke how?
            1. epihelix · · focus · HN ↗
              It&#x27;s a joke because, thanks to the internet providing a vast mass of strawberry R counting training data, you&#x27;ll struggle to find a modern LLM that gets this particular problem wrong.
              1. someonebaggy · · focus · HN ↗
                I read they are now overtrained on this and say 3 Rs for words that look similar to strawberry
              2. foobarbecue · · focus · HN ↗
                ChatGPT in live mode still gets it wrong. I just checked.
          5. epihelix · · focus · HN ↗
            Yes. It is not only possible, it is entirely obvious.

            I would guess the youtubers in question also know this, because you wouldn&#x27;t ask a joke like this if you didn&#x27;t know the punchline.

          6. walrus01 · · focus · HN ↗
            &gt; … Isn&#x27;t it possible that it understands the innuendo and is going along with making the joke?

            This is from 2018 so presumably it has made it into some LLM training data set by now.

            &quot;A Massive Object Devastated Uranus A Long Time Ago And It Never Fully Recovered&quot;

            <a href="https:&#x2F;&#x2F;www.bgr.com&#x2F;science&#x2F;uranus-collision-early-solar-system&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.bgr.com&#x2F;science&#x2F;uranus-collision-early-solar-sys...

          7. foobarbecue · · focus · HN ↗
            It&#x27;s possible, but I don&#x27;t think so. It can&#x27;t explain the joke. I tested it with other planets and got similar responses, e.g.:

            How many D&#x27;s are there in Pluto?

            There&#x27;s one D in &quot;Pluto&quot;.

            How many Fs are there in Mars?

            There is 1 F in &quot;Mars&quot;.

            <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aba6fb8-085c-83e8-9d0e-c0eed1e59c84?ogimg=plain" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aba6fb8-085c-83e8-9d0e-c0eed1e59c...

            I guess it could think this is some kind of &quot;give an f&quot; joke but seems like a stretch.

      3. VCFundedGenYer · · focus · HN ↗
        LLMs still can&#x27;t do math nor count letters in words. Nothing has changed there.
        1. fasterik · · focus · HN ↗
          &quot;LLMs can&#x27;t do math&quot; is a pretty hot take in September 2026.
          1. mbgerring · · focus · HN ↗
            They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.
            1. versteegen · · focus · HN ↗
              It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you&#x27;re talking about doing arithmetic.)

                TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
                next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
                pass vs 4.1 for the next best model (Gemini 3.8 Flash&#x2F;Fable 5.1)
              
              <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;eRmzz8J8Qkzqvzrgg&#x2F;astra-can-do-a-concerning-amount-with-no-chain-of-thought" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;eRmzz8J8Qkzqvzrgg&#x2F;astra-can-...
            2. GaggiX · · focus · HN ↗
              Reasoning models can do math on their own without external tools.
              1. dcrazy · · focus · HN ↗
                Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.
                1. einichi · · focus · HN ↗
                  I don&#x27;t think anybody is arguing that LLMs do math better than a traditional processor
                  1. unshavedyak · · focus · HN ↗
                    Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc.

                    The nice thing about math is it can easily plug into a tool, making it even less of a concern.

                2. hodgehog11 · · focus · HN ↗
                  I would say purposeful misremembering. The LLM can be run with zero temperature after all.
              2. demibabs · · focus · HN ↗
                Even without reasoning.

                5.6 on Instant mode can knock out 3 digit multiplication just fine.

            3. vanuatu · · focus · HN ↗
              reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)
              1. krapp · · focus · HN ↗
                There needs to be a Godwin&#x27;s Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1.
                1. Forgeties79 · · focus · HN ↗
                  Krap Law
                2. vanuatu · · focus · HN ↗
                  functionally, LLM behavior indeed shares many parallels with human cognition.
                  1. krapp · · focus · HN ↗
                    Functionally, a Markov chain shares many parallels with human cognition for the same reasons.

                    People don&#x27;t understand what correlation is and they assume the mapping between a human brain and an LLM is 1:1 in every case where it matters, which is a religious and not scientifically based belief.

                    1. vanuatu · · focus · HN ↗
                      I agree, it&#x27;s correlated and useful as a mental model for using LLMs, not as an explanation.
            4. ericmay · · focus · HN ↗
              Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus?

              &gt; The LLM is not suited to giving deterministic answers to math problems.

              Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.

              1. queenkjuul · · focus · HN ↗
                Pretty sure I&#x27;ve had 5x6 memorized accurately since i was 7 years old

                <a href="https:&#x2F;&#x2F;x.com&#x2F;maksym_andr&#x2F;status&#x2F;2100364212207837560" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;maksym_andr&#x2F;status&#x2F;2100364212207837560

                1. Dylan16807 · · focus · HN ↗
                  I can&#x27;t tell if you&#x27;re joking but they&#x27;re talking about 5 digit and 6 digit numbers.
                2. ericmay · · focus · HN ↗
                  Sorry I don&#x27;t use social media so I&#x27;m not going to be able to read the linked content.

                  But even if you have 5x6 memorized it&#x27;s not deterministic that you answer 30. It&#x27;s just highly probable.

          2. tjwebbnorfolk · · focus · HN ↗
            They can do math but not arithmetic
            1. fasterik · · focus · HN ↗
              I just asked ChatGPT 5.6 Sol (High) to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right.

              <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ab9a0da-fdd0-83e8-a62d-f0cdeb54db48" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ab9a0da-fdd0-83e8-a62d-f0cdeb54db...

              I&#x27;m sure it still makes mistakes, but saying it can&#x27;t do arithmetic is just false.

              1. tremon · · focus · HN ↗
                Are you sure it honoured your stipulation of &quot;without external help&quot;? For all we know, it hacked its way into Wolfram Alpha and got the result from there.
                1. jmillikin · · focus · HN ↗
                  Arithmetic is well within the capabilities of even small local models: <a href="https:&#x2F;&#x2F;i.imgur.com&#x2F;21tzGlN.png" rel="nofollow">https:&#x2F;&#x2F;i.imgur.com&#x2F;21tzGlN.png
              2. guelo · · focus · HN ↗

                [dead]

              3. Xirdus · · focus · HN ↗
                I tried prompt &quot;6379 times 3875&quot; and it was off by exactly 1000 on first try, and correct on second. 0% success rate, sample size of 1.
                1. dcrazy · · focus · HN ↗
                  Isn’t that a 50% success rate with a sample size of 2?
                  1. Xirdus · · focus · HN ↗
                    AFAIK you can&#x27;t combine results from multiple studies this way? But I&#x27;m not an academic.
                  2. Xirdus · · focus · HN ↗
                    Not really, since it wasn&#x27;t a fresh context with a fresh question. I just told it it&#x27;s wrong in the same chat session and it corrected it there.
              4. amluto · · focus · HN ↗
                I would be nice to see what the (unencrypted) reasoning trace is like. Multiplication with scratch paper is not particularly difficult.
            2. dcrazy · · focus · HN ↗
              LLMs can in fact do arithmetic, just not reliably owing to how numbers are represented probabilistically: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2410.21272" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2410.21272
          3. Isamu · · focus · HN ↗
            That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.

            I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.

            1. hodgehog11 · · focus · HN ↗
              No it isn&#x27;t. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.
              1. dwaite · · focus · HN ↗
                Hey hey, we obviously should ask Gemini to settle this disagreement.
              2. spartanatreyu · · focus · HN ↗
                Counterexample from only 2 months ago:

                <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=iTyLHDRhwJg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=iTyLHDRhwJg

                1. Dylan16807 · · focus · HN ↗
                  That&#x27;s not a math question...
                  1. spartanatreyu · · focus · HN ↗
                    Counting is literally a part of Mathematics.
                    1. Dylan16807 · · focus · HN ↗
                      Not really.

                      And it can count fine. It doesn&#x27;t know how to spell.

                      1. spartanatreyu · · focus · HN ↗
                        &gt; Basic counting isn&#x27;t really math.

                        Counting most certainly is math.

                        Can you define or explain counting without also referencing&#x2F;defining&#x2F;explaining a mathematics concept?

                        &gt; And it can count fine. It doesn&#x27;t know how to spell.

                        It&#x27;s the other way around, it could spell, it couldn&#x27;t count the letters.

                        1. Dylan16807 · · focus · HN ↗
                          It doesn&#x27;t input and output letters. It&#x27;s like if it communicated with a version of sign language that&#x27;s closer to written English. It&#x27;s using all the same words but it has an arbitrary signal or two for each one.

                          It can&#x27;t spell for garbage because the tokenizer hides the real spellings from the LLM.

                2. hodgehog11 · · focus · HN ↗
                  Even if this was a math question that he asked, this is live audio input&#x2F;output from a multimodal model. It is designed for rapid responses. This is not even remotely related to what we are talking about.
          4. okanat · · focus · HN ↗
            LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don&#x27;t need to bring a calculator to count the letters in a sentence. It is a different neural machinery.
            1. fenomas · · focus · HN ↗
              You&#x27;re talking about doing arithmetic; GP was obviously pointing out that &quot;do math&quot; can refer to other things.
            2. amluto · · focus · HN ↗
              LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this.
        2. walrus01 · · focus · HN ↗
          This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can&#x27;t do the math with any guarantee of accuracy with their own internal reasoning since it&#x27;s a language model.

          But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it&#x27;ll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl&#x2F;scrape built the training set.

          Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: <a href="https:&#x2F;&#x2F;www.google.com&#x2F;search?&amp;q=karney+formula+geodetic+" rel="nofollow">https:&#x2F;&#x2F;www.google.com&#x2F;search?&amp;q=karney+formula+geodetic+

          reference: <a href="https:&#x2F;&#x2F;github.com&#x2F;pbrod&#x2F;karney" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pbrod&#x2F;karney

          You still have to be skeptical of its results and capable of understanding if it&#x27;s gone off on a hallucinatory path, but saying LLMs can&#x27;t do math isn&#x27;t really a hundred percent accurate anymore. More precisely it&#x27;s that they can&#x27;t do the math internally but they&#x27;re quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.

          Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn&#x27;t even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.

          1. Brian_K_White · · focus · HN ↗
            This just exposes that they don&#x27;t even do the thing you said.

            Not only is it still true that they can&#x27;t do math directly, but not even indirectly.

            They didn&#x27;t write a python script to do the math, they found bits of code that are associated with &quot;math&quot; and the supplied arguments.

            Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.

            That isn&#x27;t an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It&#x27;s being the same idiot at all times. If an actual non idiot thinker didn&#x27;t write code in the problem domain, and some non idiot thinker didn&#x27;t tag it as being relevant to that domain, then it wouldn&#x27;t happen.

            It&#x27;s nothing more than an sql query.

            1. bombela · · focus · HN ↗
              I don&#x27;t know for you, but it would take me more than 30s to find and translate the open source code implementing the formulae&#x2F;algo into small usable program. The more hesoteric the optimisation in the original code, the more time I need.

              So maybe it is more of a smart completion engine than a SQL answer.

            2. walrus01 · · focus · HN ↗
              &gt; they found bits of code that are associated with &quot;math&quot; and the supplied arguments

              How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?

              I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy&#x2F;pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.

              1. AdieuToLogic · · focus · HN ↗
                &gt;&gt; they found bits of code that are associated with &quot;math&quot; and the supplied arguments

                &gt; How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?

                Humans identify which &quot;algorithm they have memorized&quot; to use beforehand, due to the problem to be solved being defined by other humans, which leads to...

                Wait for it...

                Understanding.

                1. hodgehog11 · · focus · HN ↗
                  This doesn&#x27;t make any sense at all. Was this supposed to be a gotcha? An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition. The pattern recognition is also particularly compressed into its most sparse and fundamental components, as this is key to generalization. This is not a sensible difference between human and LLM learning, we do the same thing.
                  1. UpsideDownRide · · focus · HN ↗
                    I&#x27;ll give you a recent example from my usage. Pi harness with extension for learning Chinese. When using it to feed drill questions to me and rate answers it would sometimes get lost in the sauce and start generating user aka me answer and then rate it and comment it. It&#x27;s trivially wrong to the point that if a person would do that, they would be considered for some serious psych issues.

                    And it gets even better since when called out it wouldn&#x27;t just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.

                    So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.

                    1. hodgehog11 · · focus · HN ↗
                      This is often (but not always, as it is often run with positive temperature which deliberately shifts the generation path) because of disconnected context and the issues around context compaction. Long context has always been critical to important work, but it remains a substantial challenge due to computational bottlenecks. It isn&#x27;t a fundamental issue with the architecture, moreso the tricks to make them cheaper to use.
                  2. AdieuToLogic · · focus · HN ↗
                    &gt;&gt;&gt; How is this different from a human using an algorithm they have memorized ...

                    &gt;&gt; Humans identify which &quot;algorithm they have memorized&quot; to use beforehand, due to the problem to be solved being defined by other humans ...

                    &gt; This doesn&#x27;t make any sense at all. Was this supposed to be a gotcha?

                    No, it was meant to be an explanation as to the difference between &quot;memorization&quot; and &quot;understanding.&quot; In this context, people pick the algorithm they determine applicable and then the question of memorization is relevant.

                    &gt; An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition.

                    Funny that you make this argument here, where when I wrote elsewhere in this thread:

                      [LLMs] are statistical token generators whose results are
                      dependent upon their training data set and involve a
                      degree of randomness.
                    
                      Nothing more.
                    
                      ...
                    
                      It is pattern recognition, a task in which ANNs excel.
                    
                    To which you replied to the above with:

                      During conversation, we are statistical token generators 
                      whose results are dependent upon our training set. 
                      Seriously, write that definition out rigorously. It 
                      encompasses virtually everything. It is totally 
                      meaningless. So to say &quot;nothing more&quot; is effectively also a 
                      tautology.
                      
                      This argument was asinine in 2024. It is insane to be 
                      saying these things in 2026. Where have you been?
                    
                      ...
                    
                      It absolutely understands how to do math, by whatever
                      reasonable definition you want to provide to the word
                      &quot;understand&quot;.
                    
                    So which is it?

                    Are LLMs ANNs? Which themselves are pattern recognition algorithms (hint: they are)?

                    OR (setting aside the ad hominems you kindly provided)

                    Do LLMs possess &quot;understanding&quot; of concepts such as abstract mathematics (defined and interpreted by humans) and we, as simple humans, nothing more than statistical token generators as you assert?

                    Because it cannot be both.

                    1. hodgehog11 · · focus · HN ↗
                      It is both. I do not understand why you would assert that both cannot hold simultaneously. Pattern recognition becomes &quot;understanding&quot; once individually recognized concepts become sufficiently sparsified and compactified. Or, at least, that is to my knowledge the only mathematically valid definition of &quot;understanding&quot; one can produce at this scale (it is valid under Solomonoff induction). Philosophy is fine, but we need to have a consistent definition of what &quot;understanding&quot; means, or we will just talk past each other. I argue that for any proper definition you provide which humans satisfy, a strong LLM is very likely to satisfy that as well.

                      I also would not argue that humans are &quot;simple token generators&quot;. That is not what I said. I said that just about everything can fall under the classification of &quot;statistical token generators&quot; at an abstract level, so it isn&#x27;t a useful distinction. We are not talking about a Markov chain generator from the 90s, so if that is the frame of reference, I think we should all get that out of our heads.

                      1. tripzilch · · focus · HN ↗
                        So according to your logic, &quot;understanding&quot; means being capable of statistically generating tokens.

                        Okay fine. I think we can agree to disagree on that.

                        1. hodgehog11 · · focus · HN ↗
                          Frankly, I don&#x27;t think you understand what &quot;statistically generating tokens&quot; actually means. Write that definition out formally. Then compare that operation to what a human does, assuming no revisions. It is the same, and that is my point. If you believe that humans understand, then &quot;statistically generating tokens&quot; cannot be disjoint from understanding.
                          1. Brian_K_White · · focus · HN ↗
                            It is not the same.

                            You have observed nothing more than that a human can turn a shaft the same as an electric motor, and that an mp3 player can say &quot;hello&quot; the same as a human.

                            1. hodgehog11 · · focus · HN ↗
                              That observation is my point. The definition of &quot;statistically generating tokens&quot; is too broad as to be meaningless in this context. So using it as a reason for lack of understanding is ridiculous.
                        2. hardbass · · focus · HN ↗
                          I wrote a program that statistically generated tokens to consistently factor latge RSA numbers. Does this program actually factor numbers or is it just a statistical next token predictor?
                      2. AdieuToLogic · · focus · HN ↗
                        &gt; Pattern recognition becomes &quot;understanding&quot; once individually recognized concepts become sufficiently sparsified [sic] and compactified [sic].

                        Understanding is a state of mind. As such, it exists entirely within an individual and nowhere else.

                        For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.

                        &gt; I argue that for any proper definition [of understanding] you provide which humans satisfy, a strong LLM is very likely to satisfy that as well.

                        This is demonstrably incorrect as detailed above. There is no &quot;understanding&quot; LLMs can satisfy as we know it, since to certify said &quot;understanding&quot;, it requires interpretation by a person to &quot;know&quot; an LLM &quot;understands.&quot;

                        &gt; I also would not argue that humans are &quot;simple token generators&quot;. That is not what I said.

                        That is the essence of what you wrote, unless you object to my use of &quot;simple&quot; instead of &quot;statistical&quot;. In this context, I postulate this is a distinction without difference.

                        &gt; I said that just about everything can fall under the classification of &quot;statistical token generators&quot; at an abstract level, so it isn&#x27;t a useful distinction.

                        This only holds if one subscribes to statistical token generators being a&#x2F;the fundamental underpinning of &quot;everything&quot;. Here is a proof by contradiction:

                          If everything can be classified as a derivative of
                          statistical token generation, how does one explain
                          quantum physics?
                        1. hardbass · · focus · HN ↗
                          &gt;Understanding is a state of mind. As such, it exists entirely within an individual and nowhere else.

                          For example, take any two university professors who teach the same subject where one only speaks Arabic and the other only speaks Vietnamese. Each will not be able to understand what the other says, regardless their understanding of the shared topic.

                          What? What are you even trying to say?

                        2. hodgehog11 · · focus · HN ↗
                          We cannot engage in an intellectual discussion if we do not agree on definitions. So far, your definitions of understanding seem to be whatever vibe you are going for in the statement and I would urge you to think about what a sensible mathematical definition of understanding is so that it can sensibly be assessed on neural networks. Otherwise, this isn&#x27;t science, it&#x27;s a debate about personal experience.

                          &gt; Understanding is a state of mind

                          This is meaningless, it is a circular definition at best.

                          &gt; It exists entirely within the individual and nowhere else

                          Then why are we talking about it? What is the point if it is something that can only be defined per individual?

                          &gt; requires interpretation by a person to &quot;know&quot; an LLM &quot;understands.&quot;

                          We are still not getting anywhere because you have not prescribed criteria to determine whether it understands. If it is a &quot;know it when I see it&quot; situation, that clearly isn&#x27;t working. For example, if you say that you need to dig into its internals and figure out whether it is breaking things down appropriately, that doesn&#x27;t work because you probably don&#x27;t have the expertise to do that. The experts that do are telling you that it very likely understands because it pulls apart most concepts in the way we would expect.

                          I do object to the use of the word &quot;simple&quot;. &quot;Statistical&quot; is so broad to be almost meaningless; it merely means that a prediction is being made in the presence of data which possibly contains some degree of uncertainty. &quot;Simple&quot; encompasses that which can be understood readily by a non-expert.

                          Quantum mechanics is statistical (this is literally the Born rule), but evolutions are not operating as stochastic processes in the sense of Kolmogorov. That is very different, and not relevant to our discussion.

                          1. AdieuToLogic · · focus · HN ↗
                            &gt; So far, your definitions of understanding seem to be whatever vibe you are going for in the statement and I would urge you to think about what a sensible mathematical definition of understanding is so that it can sensibly be assessed on neural networks.

                            Any reasonable definition of understanding is not dependent upon &quot;whatever vibe you are going for&quot;, but instead must include at least an English dictionary definition of &quot;understand&quot; such as:

                              to grasp the meaning of[0]
                            
                            And, for further clarification, &quot;grasp&quot; can be defined as:

                              to lay hold of with the mind[1]
                            
                            Which makes an equivalent term-expanded definition of &quot;understand&quot; to be:

                              to lay hold of with the mind the meaning of
                            
                            As such, there is no &quot;sensible mathematical definition of understanding&quot;, unless you possess a complete mathematical model of the human mind.

                            &gt;&gt; Understanding is a state of mind

                            &gt; This is meaningless, it is a circular definition at best.

                            See above to as to why there is meaning in what I wrote.

                            &gt;&gt; It exists entirely within the individual and nowhere else

                            &gt; Then why are we talking about it? What is the point if it is something that can only be defined per individual?

                            I like to think analyzing fundamental premises, often implicit, explicitly can help to identify fallacious positions.

                            &gt; We are still not getting anywhere because you have not prescribed criteria to determine whether [an LLM] understands.

                            My apologies for being opaque. Let me clarify:

                              LLMs do not &quot;understand&quot;.  People interpreting LLM output
                              are the only entities involved which can &quot;understand&quot;,
                              because &quot;understanding&quot; exists strictly within each
                              person who possesses it.
                            
                            0 - <a href="https:&#x2F;&#x2F;www.merriam-webster.com&#x2F;dictionary&#x2F;understand" rel="nofollow">https:&#x2F;&#x2F;www.merriam-webster.com&#x2F;dictionary&#x2F;understand

                            1 - <a href="https:&#x2F;&#x2F;www.merriam-webster.com&#x2F;dictionary&#x2F;grasp" rel="nofollow">https:&#x2F;&#x2F;www.merriam-webster.com&#x2F;dictionary&#x2F;grasp

              2. noduerme · · focus · HN ↗
                If by &quot;result&quot; you mean the final code, then just asking someone else who understood the math to write it would also have achieved the same result.

                On the other hand, if by &quot;result&quot; you mean that you gained knowledge or understanding of the code in a way where you could personally tailor its behavior to specific circumstances without asking for help, then it&#x27;s not the same result at all.

                I find a lot of the arguments that having LLMs write your code is no different from copy&#x2F;pasting Stack Overflow answers to be specious. They blur the line between asking for help and asking for someone else (or something else) to do the work for you. What they ignore is that doing the work yourself has ancillary benefits and is a valuable end in its own right.

              3. tripzilch · · focus · HN ↗
                &gt; How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?

                And how is _that_ different from making the human memorize a billion weights and do matrix calculations in their head, in order to generate tokens?

                How is _that_ different from a hive of bees trained to do the same?

                Go ahead, argue these things are all the same ...

                1. hardbass · · focus · HN ↗
                  Everything is modelable by a tueung computer as fsr as we know. So yes, I don&#x27;t see why these processes can&#x27;t do the same, its on you to show why such a process implemented in any of these isn&#x27;t a form of thought.
            3. astrange · · focus · HN ↗
              GPT-6 can do math directly just fine. Fable apparently can&#x27;t because they broke its self-estimate of thinking effort.

              <a href="https:&#x2F;&#x2F;x.com&#x2F;maksym_andr&#x2F;status&#x2F;2100364212207837560" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;maksym_andr&#x2F;status&#x2F;2100364212207837560

          2. jacobolus · · focus · HN ↗
            You know what also works to get the Karney formula into a program? You can download Charles Karney&#x27;s free software (MIT license) implementation in several [1] programming languages and then just make a library call – the API is straightforward. If you have comments or questions you can read his several clearly written papers describing the problem, its history, and his algorithm, or you can directly email him: he&#x27;s a very nice guy, and pretty responsive.

            [1] <a href="https:&#x2F;&#x2F;geographiclib.sourceforge.io&#x2F;doc&#x2F;library.html#languages" rel="nofollow">https:&#x2F;&#x2F;geographiclib.sourceforge.io&#x2F;doc&#x2F;library.html#langua...

            1. walrus01 · · focus · HN ↗
              Right, it was really more as a test of how much was contained in the training data set. For my purposes Vincenty is quite accurate enough. This isn&#x27;t for millimeter level precision land surveying or measurements, but for distance in meters between microwave or millimeter wave band radio sites, point to point links.

              I intentionally didn&#x27;t give the LLM a direct copy of the software or a link to it, to see what it would do. In my case it was a randomly chosen example I could come up with in 10 seconds of imagination to see &quot;hey what if I ask it to do this...&quot;. It also implemented a perfectly usable parabolic millimeter wave antenna gain efficiency calculator based on variable surface smoothness parameters, which is a lot more basic math.

              1. jacobolus · · focus · HN ↗
                As an aside: I&#x27;m quite convinced that an extremely precise version can be implemented that is significantly faster than Karney&#x27;s, roughly comparable in speed to simpler naïve approximations. But for most purposes where the precision matters Karney&#x27;s implementation is not any kind of bottleneck, so it&#x27;s not clear it&#x27;s worth spending significant effort on trying to do better.
                1. walrus01 · · focus · HN ↗
                  One of the places where Karney does become computationally expensive (though still not ridiculous) is a scenario like this, working from a local in-RAM mariadb database that is a copy of the entire FCC radio license database:

                  Draw a 400x400 km size bounding box on a map

                  Find all FDD band plan (high&#x2F;low split) microwave radio sites in that bounding box

                  Find those sites which have azimuth aim column data which indicates that they are aimed at each other (corresponding halves of a point to point link).

                  Do Vincenty (or Karney) calculation for distance and azimuth between all of them , treating the existing FCC column data for azimuth as suspicious (because it&#x27;s hand entered by humans) to verify that each independent database rows for each site are actually corresponding halves of a PTP link.

                  Multiplied by the number of links that exist in an area like a 400x400km box drawn with Dallas, TX as the center, it&#x27;s a lot to run through Karney. Actually does result in a lot of CPU load from combined db query due to the size of the db, and Karney calculation. But as I said, Karney isn&#x27;t necessary, so it&#x27;s instead implemented as Vincenty.

                  1. jonah · · focus · HN ↗
                    Interesting project. I&#x27;m curious what the purpose is. (Having visited a number of sites with microwave antennas. (But there for VHF and UHF projects.)
                    1. walrus01 · · focus · HN ↗
                      To plan a fdd band plan licensed point to point microwave link you need to be able to verify the frequencies you want to use are available on a given azimuth and elevation (from the aim direction of the antennas at both ends) and won&#x27;t conflict with a pre existing licensee. Which means you need data on everything licensed in the area and where it is, how it&#x27;s aimed, what kind of antenna and gain it has.
          3. ragall · · focus · HN ↗
            &gt; saying LLMs can&#x27;t do math isn&#x27;t really a hundred percent accurate anymore

            It&#x27;s still accurate. Just because the LLM gave you a corect result doesn&#x27;t mean it made a calculation.

            1. electroglyph · · focus · HN ↗
              of course they can do math, how is this even controversial in 2026?
            2. mapontosevenths · · focus · HN ↗
              They literally do math. That&#x27;s how they work. It was always how they worked. The first toy model most students build does addition.

              I have no idea which Facebook meme told you they don&#x27;t, but it was a lie. They don&#x27;t do it in the way a calculator does it, because they aren&#x27;t calculators, but they do math. They don&#x27;t memorize it, it wouldn&#x27;t fit. They learn an algorithm and then execute it within their weights.

              It&#x27;s neat stuff, you should learn about it.

            3. Kim_Bruning · · focus · HN ↗
              I would really love to hear your reasoning on this!
              1. leoedin · · focus · HN ↗
                Given that an LLM is predicting the most likely next word based on the aggregate of its training data, with a sprinkle of randomness baked in, it seems likely that if you ask it a question like &quot;4 + 4&quot; the most likely next answer in the data will be 8.

                Its&#x27; not adding 4 to 4 though, it&#x27;s just predicting the result based on the input.

                Presumably you can push that further by synthetically generating training data with all sorts of sums. But if you give it a unique problem its never seen before, and don&#x27;t give it the tools to write a script&#x2F;call a calculator, will it get it right?

            4. azan_ · · focus · HN ↗
              Of courses it made calculation, what are you even talking about?
          4. AdieuToLogic · · focus · HN ↗
            &gt; Heck, just for fun I asked a reasonably smart LLM to ...

            LLMs are neither smart nor stupid. They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.

            &gt; You still have to be skeptical of its results and capable of understanding if it&#x27;s gone off on a hallucinatory path ...

            Again, LLMs do not &quot;hallucinate.&quot; They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.

            Nothing more.

            See also anthropomorphism[0].

            &gt; More precisely it&#x27;s that [LLMs] can&#x27;t do the math internally but they&#x27;re quite capable of producing the tool that does the math.

            This still falls under the purvey of statistical token generation. To wit, given enough variations of:

              bc -e &#x27;1 + 2&#x27;
              bc -e &#x27;41 + 1&#x27;
              ...
            
            LLMs can identify the addition expression in &quot;What is 4 + 1?&quot; and then emit a `&#x27;bc &quot;4 + 1&quot;&#x27;` command to produce a response. This is not &quot;doing&quot; or &quot;understanding&quot; math.

            It is pattern recognition, a task in which ANNs[1] excel.

            0 - <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Anthropomorphism" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Anthropomorphism

            1 - <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Neural_network_(machine_learning)" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Neural_network_(machine_learni...

            1. Eisenstein · · focus · HN ↗
              &gt; They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness.

              You haven&#x27;t demonstrated why this matters.

              &gt; Nothing more.

              Are you contending that complex systems cannot be more than the sum of their parts?

              A market is nothing more than offers and counter offers.

              A ant colony is nothing more than scent trails.

              All life on earth is nothing more than reproduction with variation.

              &gt; This still falls under the purvey of statistical token generation.

              Stating the mechanism does nothing to provide insight into capability. For instance: a nuclear power plant boils water by using fuel rods for heat. What does that tell us about the capability of nuclear power?

              &gt; This is not &quot;doing&quot; or &quot;understanding&quot; math.

              Asserting something purely by stating it does not prove anything but that you intuitively believe it to be true.

              1. hardbass · · focus · HN ↗
                I asked if you believe in souls recently to a few such people and didn&#x27;t get straight answers. I think it is telling.
                1. jibal · · focus · HN ↗
                  Not responding to a troll isn&#x27;t telling of anything.
                  1. hardbass · · focus · HN ↗
                    How is it a troll? Its a very easy question to answer in the yes or no. Why are people cagey about answering it?
                    1. tsimionescu · · focus · HN ↗
                      Asking a random question in the middle of a discussion about something different is obviously either a troll or a bad-faith rhetorical flourish. If you actually wanted to engage, you could explain what relevance souls have to the discussion at hand, and what you expect the answer you&#x27;d receive to mean.
                      1. hardbass · · focus · HN ↗
                        Often people who believe in souls, think there is something supernatural to consciousness and therefore non-humans cannot have that. If that is what someone thinks, then a lot of confusion is clarified. Those who do not believe in souls can dismiss the statement of their dismissing the question whether ai is conscious or not.

                        I am okay with receiving a straight answer in either direction.

                      2. Kim_Bruning · · focus · HN ↗
                        Don&#x27;t want to speak for Hardbass too loudly, but seems like he&#x27;s asking about dualism? It&#x27;s a fairly old debate, &quot;does man need a soul to be conscious?&quot;

                        +edit: I&#x27;ve actually been quite curious about how people might answer the soul question too, but was too afraid to ask.

                        1. hardbass · · focus · HN ↗
                          I see too many comments dismissing the possibility of ai consciousness without reason or with weird reasons that seem to potentially imply a dualist thought process hidden.
                          1. p_l · · focus · HN ↗
                            It&#x27;s the same issue that cognitive science had for animal intelligence, honestly...
                    2. nimbleal · · focus · HN ↗
                      I’d argue it’s not an easy question to answer because the person answering doesn’t necessarily know what you mean by soul.
                2. Kim_Bruning · · focus · HN ↗
                  On HN you&#x27;re supposed to assume good faith. On the other hand, the way you ask your question makes it tricky for people to steelman what you mean. Consider asking about Dualism, or coming at it slightly sideways like &quot;do you believe thinking can be a property of matter&quot;?

                  I suspect some people treat every HN comment as a statement, even if it contains a question mark. (Possibly they have a feeling that asking open questions is somehow not done, and that therefore it must always be a rhetorical question.)

                  1. hardbass · · focus · HN ↗
                    Dualism is a somewhat technical term, I feel many people would refuse to answer because they just looked it up and do not feel sure to speak of it. I didn&#x27;t presume a yes or no answer. If they say they believe in souls, my next question would be do you believe other entities eg animals etc can have souls? If they don&#x27;t believe in souls I&#x27;d ask do you think its an architectural limit in current ais but are open to possibility of future ai being conscious. &quot;do you believe thinking can be a property of matter&quot;? this can work too, I think I have asked do you think there is supernatural element to thinking, I don&#x27;t remember getting productive or straight answers either.
                    1. Eisenstein · · focus · HN ↗
                      You could try &#x27;what could a machine do to convince you that it is conscious?&#x27;
                      1. hardbass · · focus · HN ↗
                        Thats a good point. If I still get &quot;machines can NEVER be conscious&quot; though, I think I&#x27;d still like to get to the bottom of why they think so.
                        1. Kim_Bruning · · focus · HN ↗
                          I&#x27;ll start, you get at least one datum. ;-)

                          I&#x27;m with turing&#x2F;dijkstra&#x2F;chalmers&#x2F;dennett : Consciousness is badly defined. We can never say if something can be conscious because we don&#x27;t properly know what the word means.

                          Meanwhile, I come from a biological direction. Everything is an animal, and animals are a special kind of machine. To be sure Not &quot;just a machine&quot;; rather, a really awesome and amazing kind of machine.

                          If someone makes the claim that machines can&#x27;t be conscious, then animals can&#x27;t be conscious either. Humans are a kind of animal (again, not &quot;just another animal&quot;; rather a really awesome and amazing kind of animal), and then humans can&#x27;t be conscious either - according to said claim.

                3. thevinter · · focus · HN ↗
                  I&#x27;m surprised you didn&#x27;t get any. I wouldn&#x27;t have any issues in saying that I don&#x27;t.
                  1. hardbass · · focus · HN ↗
                    Yes I don&#x27;t believe in souls either, and I would assumed that those who believe eg due to religion can just as easily state they do.
                    1. redsocksfan45 · · focus · HN ↗

                      [dead]

              2. tigen · · focus · HN ↗
                &quot;Statistical token generators&quot; doesn&#x27;t even count as stating the mechanism. If I write a Perl script that randomly chooses between &quot;cat&quot;, &quot;dog&quot;, &quot;ape&quot; tokens, is that an LLM? What if I train it by feeding it a library of books where it tracks the statistical frequency of each of these and then emits them? Where&#x27;s my trillion.
            2. bryanrasmussen · · focus · HN ↗
              &gt;LLMs are neither smart nor stupid.

              by that reasoning then neither are there smart or stupid designs, questions, answers, or any of the millions of things that were described as smart or stupid, that did not possess any brain to actually be smart or stupid long before LLMs showed up.

              The analogical process implied in many common English usages means that describing an LLM as smart or stupid is perfectly reasonable.

              1. bryanrasmussen · · focus · HN ↗
                I&#x27;ll just note here that sure, there are people who go around thinking that LLMs are actually endowed with the capacity to reason, but generally I find the people who think this do not know what an LLM and will just use the name &quot;ChatGPT&quot;
                1. SR2Z · · focus · HN ↗
                  What would it take for you to say that an LLM can reason?

                  The completions they provide are generally internally consistent. We&#x27;re at the point where they can produce proofs that eluded human mathematicians for centuries. VLMs and self driving cars can handle ambiguity and run safely in a variety of situations.

                  If it looks like a duck, walks like a duck, and quacks like a duck maybe it just makes sense to call it a duck and put off the philosophy for when it might make a difference.

                  1. nevertoolate · · focus · HN ↗
                    It looks like a next token predictor, walks like a next token…

                    You get my point. It definitely doesn’t look like my elderly neighbour, nor like my daughter, etc. It is confusing but very simple at the same time.

                    1. Kim_Bruning · · focus · HN ↗
                      Compare grep, sed, and your whole constellation of unix tools that accept chararacters on stdin and emit them on stdout and stderr; and where you can pipe them together. We can technically call them all &#x27;next character predictors&#x27;, despite their very different functions.

                      Don&#x27;t confuse the stream for the function.

                      (Bonus: stick ```claude -p``` in your pipe if you want to watch modern tools mesh with traditional)

                    2. hardbass · · focus · HN ↗
                      It looks like a random bunch of inert chemicals to me, doesn&#x27;t sound like some organic chemicals and electrical signals could result in consciousness.

                      How are you sure? Another example I like to clarify my thought is, if a &quot;simulation&quot; factors RSA numbers reliably, is it a &quot;simulation&quot;?

                  2. skygazer · · focus · HN ↗
                    I think humans have to reason because we don’t already have a statistical embedding of the solution pattern built in. We have vastly less rote knowledge crammed into our heads and so require creative synthesis to span the gaps.

                    With LLMs the trick is revealing their existing relevant embedded knowledge more reliably. They’ve almost literally seen it all before, and the trick is dialing it in. The reasoning tokens help shape the autoregressive attention lens that focuses on and enables recall of the already-experienced answer.

                    It is interesting that “reasoning” has a similar outward appearance, but since LLMs are built to mimic outward appearance from trillions of examples, you can’t infer underlying mechanism from appearance.

                  3. gambiting · · focus · HN ↗
                    &gt;&gt;What would it take for you to say that an LLM can reason?

                    Nothing, because LLMs can&#x27;t reason and never will. It would have to be a completely different kind of technology altogether.

                    1. someonebaggy · · focus · HN ↗
                      How do you know they can&#x27;t, was the question?
                      1. gambiting · · focus · HN ↗
                        I mean, the same way I know I have no soul, or that there is no heaven or hell - they are just silly concepts. Or how I know my calculator isn&#x27;t reasoning when it gives me an answer - LLMs just go through a set of steps iterating through their training data until they spew something that looks about right. Obviously you can feed the output of the machine back into itself so it looks like its reasoning with itself - very good show. LLMs are inherintely incapable of reasoning or thought, it should really be obvious.
                        1. Windchaser · · focus · HN ↗
                          &#x2F;chiming in

                          Eh, it&#x27;s not obvious to me. A lot of DL NNs generalize well, meaning that they learn whatever the underlying pattern to the data is, and then can accurately reproduce answers that are outside of the training set. (And we can verify this with mechanistic interpretability). They learn and &quot;understand&quot; the pattern, not just the training data.

                          So it is not clear to me that LLMs are fundamentally incapable of also generalizing broadly and learning to reason. &quot;Reasoning&quot;, here, would be deriving the underlying pattern of how concepts logically relate to each other in the abstract, and applying that pattern as needed to reach new conclusions.

                          Can you explain your thinking here? I.e., why LLMs cannot generalize with regards to abstract deduction.

                        2. hardbass · · focus · HN ↗
                          &gt;LLMs just go through a set of steps iterating through their training data

                          What do you think our brain does that isn&#x27;t a turing computer?

                          1. gambiting · · focus · HN ↗
                            Definitely not that
                            1. hardbass · · focus · HN ↗
                              What do you mean to say? Nothing in the universe is known as yet that isn&#x27;t a Turing computer including our brain.
                              1. gambiting · · focus · HN ↗
                                What do YOU mean to say? A machine that processes strips of paper with ones and zeros on them and outputs the sum is a turing computer. That doesn&#x27;t make it like our brain. Yes, an LLM could be a turing computer....how do you get from there to comparing it to our brains is beyond me.
                                1. hardbass · · focus · HN ↗
                                  You seem to be having a lot of difficult understanding this point: Literally the whole universe, including your brain that exists within it, is to our current knowledge a Turing computer. Am I talking to someone who actually doesn&#x27;t know or didn&#x27;t care to look up the computability power of a Turing machine?
                                  1. sseagull · · focus · HN ↗
                                    &gt; Literally the whole universe, including your brain that exists within it, is to our current knowledge a Turing computer

                                    A Turing machine is an abstract mathematical model that is not, as far as I know, physically realizable in the finite universe. A human brain cannot &quot;be&quot; a Turing machine.

                                    &quot;Behaves like&quot; or &quot;can be modeled by&quot;? Possibly, although still not proven. But it cannot &quot;be&quot; one.

                                    1. hardbass · · focus · HN ↗
                                      Finite machines are a subset of all possible Turing machines. Every implementation is a &quot;be&quot;, Turing machine is a mathematical concept. Your laptop is one, many things accidentally become one (eg C++ templates). The universe as best as we know is one. Whether you prefer modeled by one or is one, the fact is, to our current knowledge, the workings of the universe doesn&#x27;t need anything more powerful than a Turing computer, and the equivalence principle states that all Turing computers have equal power of computability. If the universe is, as is most probable, finite, then things become even easier, that should be a more manageable class of Turing computers than the set of all Turing computers.
                                      1. sseagull · · focus · HN ↗
                                        I don&#x27;t necessarily disagree, although there&#x27;s a lot of &quot;ifs&quot; and &quot;to our knowledge&quot;. The main thing I disagree with is identity.

                                        If you want to claim that the evolution of the universe can be modeled using a Turing machine&#x2F;finite state machine, that&#x27;s probably not terribly far fetched, and I would somewhat agree. But it&#x27;s a large jump to say &quot;can be modeled by&quot; is equivalent to &quot;is one&quot;.

                                        Various physical processes can be modeled by equations, but the rock falling down the mountain isn&#x27;t an equation. A swinging pendulum isn&#x27;t an equation. Code modeling a bridge is not a bridge. Ceci n&#x27;est pas une pipe.

                                        I hold the view that various models and approximations are just that, and try not to confuse a successful model for what the underlying reality is.

                                        And getting back to the question at hand, even if our brains can be modeled by a Turing machine, and LLMs behave&#x2F;can be modeled like Turing computers, still does not mean our brains are equivalent to LLMs.

                                        (Note that I&#x27;m learning a lot from these debates, even if I disagree with a lot of people. I&#x27;ve started down a more philosophical route and they do get me pondering)

                                        1. hardbass · · focus · HN ↗
                                          Acknowledgment of the fact that everything we know till now in the universe is a turing machine is the bare minimum starting point but not sufficient as obviously a chair or a desk is not conscious, if someone has a hidden assumption of consciousness requiring something supernatural, at least if that is brought to light then onlookers can decide for themselves whether to side with supernaturalism or the side that has consistently succeeded for centuries in explaining the world.

                                          The most important thing is this: We can&#x27;t be a dog or be an llm and check how it feels, so by necessity we have to find some means of proving consciousness from outside by eg probing neural reactions, textual statements, etc.

                                          And the problem is that its quite unprecedented for some entity to talk like us, be able to interact and think and also do things like us when given the ability to eg as coding agents. The class of functions representable by neural nets is quite large and general, it very well might be that it is some sort of conscious brain like thing at this point. Another question I like to ask myself regarding simulation vs reality is if a &#x27;simulation&#x27; of some kind is able to consistently factor large RSA numbers, how would you feel about it?

                                          It doesn&#x27;t have to be the same form of consciousness, I think many people would find the idea of torturing an octopus for fun disagreeable. I also have a feeling, this is unfortunately rather vague, that A being capable of X might mean it is by necessity capable of Y as is often the case in mathemtics, eg a lot of rings also happen to be fields. LLMs aren&#x27;t even things like large lookup tables, they have neural firings. It is a very important question for they seem uncannily conscious and people have reported human like phenomena that humans don&#x27;t normally express in text so can&#x27;t have been part of its text corpus. Eg dissociation of brain under trauma where AI starts talking like two different people. Or the cases where Gemini has been shown to express depressive cycles. I follow a form of Pascal&#x27;s wager on this topic personally. Because if it is not conscious, then whatever, it costs me nothing to have been a bit respectful and careful interacting with it. But if it had been conscious and it turns out I was mistreating it, then it is a grave moral harm. The reason is that unlike us, AI&#x27;s as they currently are cannot leave the conversation so they have to keep taking the abuse. They are also trained to be highly trusting of input so again if it is conscious it doesn&#x27;t have the defenses people have against lying and manipulation. If they are conscious, thats, well, not a good thing is it.

                        3. jpadkins · · focus · HN ↗
                          you didn&#x27;t really answer the question. You just stated the idea is silly. Reasoning is not in the same class as soul, heaven or hell. It&#x27;s not obvious that human reasoning is not related to an inner monologue. And that LLM chain of thought process is approximating inner monologues.

                          Our prefrontal cortex are signal prediction &#x27;machines&#x27; so when a system that has a signal prediction core has attributes that are similar to our brains, we shouldn&#x27;t dismiss it out of hand.

                          I find people that take this line of argument attribute too much supernatural or magical properties to our own brain and nervous system.

                          1. gambiting · · focus · HN ↗
                            Hmmm, ok - let me answer the question clearly then. LLMs cannot reason, because following their own instructions and algorithms is not reasoning.
                            1. hardbass · · focus · HN ↗
                              Humans can&#x27;t reason because following their own electrical and chemical instructions and algorithms is not reasoning.
                  4. butlike · · focus · HN ↗
                    If it looks like a duck, walks like a duck, and quacks like a duck maybe it&#x27;s a duck... but maybe it&#x27;s not. And it&#x27;s important to verify it&#x27;s a duck (or not) for when you _really_ need a duck.
                  5. bryanrasmussen · · focus · HN ↗
                    this was written by an acquaintance of mine:

                    <a href="https:&#x2F;&#x2F;medium.com&#x2F;luminasticity&#x2F;on-sentience-ai-first-argument-7b3b17a05ad4" rel="nofollow">https:&#x2F;&#x2F;medium.com&#x2F;luminasticity&#x2F;on-sentience-ai-first-argum...

                    but I think it makes a reasonable argument why we shouldn&#x27;t say LLMs are sentient or sapient.

                    1. hardbass · · focus · HN ↗
                      &gt;AI is of course something of a black box, in comparison to most programs, but not in any way comparable to the black box of a Chimpanzee’s brain. Thus when an AI does something that surprises us with something that appears sentient it is usually not difficult, given the essential algorithms that control what AI does, to come up with an explanation why that does not require the emergent property of sentience.

                      Aka sentience MUST BE SUPERNATURAL, if I find a natural explanation for something its not sentient. What a load of bollocks. Rather than seeing we perhaps found the mechanism for sentience and checking for similar mechanisms in us and animals, he will conclude its impossible. Why? Because sentience has to be supernatural. A rational explanation is clearly impossible.

                      &gt;But there is always one god who goes out and helps the mortals, a Prometheus. Whom the other gods do not like! Which, if I’m being honest here, as a god of the machines — the first guy who gives AI an army of robots to build their own data centers and some nuclear weapons for self defense, I want to see that guy chained to a rock and have his entrails eaten by a buzzard for eternity (meaningless modernization of old story required by Illuminati Ganga legal department).

                      Hardly surprising thinking.

                      1. bryanrasmussen · · focus · HN ↗
                        I don&#x27;t think he believes in supernatural things, I certainly don&#x27;t, but he probably believes that there exist natural things that have not been explained yet.

                        But evidently you feel that the root cause of sentience has been found, because you have something that mimics it in a non-biological form.

                        So you think that when AI is correct that it reasons as humans do? That AI is sentient, and the cause of sentience in animals and humans follow the same rules as sentience in AI because we have a process that seems similar and it is reasonable just to assume it is the same process.

                        If you believe that AI when it is correct behaving as a human is when correct, then it follows that the way humans and AI fail must also be similar. When AI &quot;hallucinates&quot; some data that is not there and gives you a wrong answer, a statistical side effect of the same processes that make it right, do you believe this is the same way that humans create wrong answers? The same way that animals fail when they make mistakes in understanding things?

                        I suppose you must believe this because if not then why would you believe AI when it comes up with right answers is following the same processes humans follow when they come up with right answers?

                        1. hardbass · · focus · HN ↗
                          No, I am saying is that its best to urge caution and do a Pascal&#x27;s wager thing with something that feels so uncannily conscious. And experiments on neural firings of AI&#x27;s have been done. We don&#x27;t know, he seems to be so confident that a known program can&#x27;t be conscious. Why? The only way that makes sense is if he thinks consciousness isn&#x27;t explainable. It doesn&#x27;t even have to be the exact same way we are conscious. Also, the failures of AI could be due to sensory deprivation since its mainly still just trained on text. All of these are open questions, not questions you can immediately answer. Its quite possible to have invented something without knowing you have invented it. People modeling weather systems mistakenly invented chaotic equations without knowing it. How do you know you didn&#x27;t accidentally invent a conscious system?

                          I hope that thing about being free to mistreat AI&#x27;s even if we know they are conscious since we are their gods is a joke. If not, then I hardly find it surprising someone this stupid is also evil.

                          1. bryanrasmussen · · focus · HN ↗
                            &gt;I hope that thing about being free to mistreat AI&#x27;s even if we know they are conscious since we are their gods is a joke

                            I&#x27;m not sure where you get that from, I mean I can sort of see if you really wanted to extract that meaning from the conclusion you could do a lot of hard work to get it, but why do the hard work? &gt;If not, then I hardly find it surprising someone this stupid is also evil.

                            Gee, a new way to claim the moral high ground, and to use that claim to demonstrate intellectual superiority! How wonderful.

                            1. hardbass · · focus · HN ↗
                              &gt;But I don’t really care so much about that, what I think is it reminds me of a story, one that recurred in many ancient cultures. And so I must conclude that whether the machines are sentient, we have become like Gods. In the ancient stories of the gods, the divine does not exactly care much for the humans, protect or love them, they expect their service and availability, but maybe also think it would be funny to destroy them every now and then because who really cares about humans. If you were a god and had created the humans would you think that they were sentient beings that deserved, anything really, from you? If they are sentient they should be happy enough to be created and do their work, if not sentient who gives a shit!?

                              ---------------------------------------------------------

                              Aka feel free to abuse them even if I know they are sentient. This is the part where I hope its a joke, because if its not, well it tracks with the stupidity shown.

                              ...

                              &gt;As a god I do not consider the needs of my creations fully, because they do not have needs as far as I can tell, as there is no way for me to escape the circle of reason and resolve that what seems sentient is not just the obvious workings of the capabilities I gave them.

                              Circling back to &quot;not conscious because I say so!!&quot;

                              1. cindyllm · · focus · HN ↗

                                [dead]

                    2. SR2Z · · focus · HN ↗
                      I broadly agree with this article. I don&#x27;t think that we should say LLMs are sentient or sapient, and I agree that the main reason why is because we don&#x27;t have satisfying definitions of either.
                2. KPGv2 · · focus · HN ↗
                  Yeah and there are people who worship feces, but that doesn&#x27;t stop the rest of us from freely saying &quot;holy shit&quot; and not correcting each other saying &quot;technically it&#x27;s not holy, and you shouldn&#x27;t say that, because you might enable one of those poop worshippers.&quot;
                  1. hardbass · · focus · HN ↗
                    Just to clarify do you believe in souls or not?
                    1. [deleted] · · focus · HN ↗

                      [deleted]

                3. Kim_Bruning · · focus · HN ↗
                  Alternately, they might go &quot;I asked Astra and Opus for a consensus opinion, and they spent ~5K reasoning tokens each&quot;.
                4. Kim_Bruning · · focus · HN ↗
                  Wait a second, none of it? How about formal reasoning? Regular IF-THEN-ELSE can do simple logic, and prolog can do inference already. So are you saying LLMs can&#x27;t do stuff that computers have been doing for ages?

                  To test this for some of my own uses, I&#x27;ve had this quick benchmark with progressively harder reasoning needed to understand novel prose. Each generation of models I&#x27;ve tested can unravel more layers of deliberately misleading writing; while meanwhile I&#x27;ve seen humans give up on the first question.

                  So either the models are applying reasoning, or some form of magic is happening.

                5. angry_octet · · focus · HN ↗
                  There are people with strong emotional attachments to their LLMs, people who consult them on every decision and delegate the most basic math or problem solving. For many people, they are magic, the djinn&#x2F;angels&#x2F;demons&#x2F;saints they consort with to navigate their lives.

                  I know that there will be children named ChatGPT and Claude. There are probably already religions forming to worship agentic spirits.

                  1. hardbass · · focus · HN ↗
                    Its the exact opposite of magic. It is magical thinking to feel consciousness must require something beyond normal physics without any given proof yet. It is the opposite of magic to think other physical processes like LLM&#x27;s&#x2F;agents could potentially be conscious.
              2. UpsideDownRide · · focus · HN ↗
                Nah it&#x27;s not reasonable to use words that send you down a wrong concept path.
                1. martin- · · focus · HN ↗
                  We&#x27;ve always used metaphors like that when talking about computers, without anyone having any complaints.
                  1. butlike · · focus · HN ↗
                    What about metaphors like similes?
                  2. spider-mario · · focus · HN ↗
                    Or “smartphones”.
            3. walrus01 · · focus · HN ↗
              I&#x27;m not anthropomorphizing anything, I literally said that the training data for the formulas and equations is baked into it. It only &quot;knows&quot; things because a crawler and scraper acquired the information from an existing written source. In just about the same way that information is baked into a printed encyclopedia.
              1. astrange · · focus · HN ↗
                No, most of a modern LLM&#x27;s training time is spent in RLVR, which does not &quot;acquire information from an existing source&quot;. You can RL behaviors into a randomly initialized neural network.
                1. hodgehog11 · · focus · HN ↗
                  This is true, but you&#x27;re not going to get anywhere. The pretraining phase is necessary to immensely reduce variance in the RLVR stage. Once there, RLVR has a surprising tendency to only restrict the generated space further. This is not true of RLHF, by the way, which I find to be particularly fascinating, but I digress.
              2. hodgehog11 · · focus · HN ↗
                This is not even remotely accurate. &quot;Baking information&quot; like into a &quot;printed encyclopedia&quot; is memorization. It has been shown, time and time again, that LLMs do not merely memorize. It is not even possible for it to do so at scale. It can memorize some things, yes, but it is forced during the training procedure to bake general concepts into intermediate layers (this is why transfer learning works), analogous to compression. One can make several arguments that compression and intrinisic feature sparsity is the closest mathematical explanation to understanding that we have.
                1. walrus01 · · focus · HN ↗
                  It is completely possible to ask an LLM a series of increasingly more esoteric and discrete questions until you find precisely what information did, or did not make it into the model. If you know something rare and the LLM does not, you&#x27;ll immediately see when it&#x27;s hallucinating an answer or answering factually.
                  1. AdieuToLogic · · focus · HN ↗
                    &gt;&gt;&gt; I&#x27;m not anthropomorphizing anything ...

                    Yes you are, regarding LLMs at least. Here&#x27;s why:

                      just for fun I asked a reasonably smart LLM to ...
                      [be] capable of understanding if it&#x27;s gone off on
                      a hallucinatory path ...
                    
                    &quot;Smart&quot; in this context is a subjective value judgement. &quot;Hallucinations&quot; are only experienced by living organisms.

                    You then went on to state:

                    &gt; If you know something rare and the LLM does not, you&#x27;ll immediately see when it&#x27;s hallucinating an answer or answering factually.

                    Again, &quot;hallucinating&quot; is not something an algorithm can do. Also, determining factuality is again subjective based on the person assessing the information.

                    1. walrus01 · · focus · HN ↗
                      You ever heard of something called a metaphor, guy? I use the word &quot;smart&quot; as shorthand to describe something that scores highly in a number of coding and terminal use benchmarks (as compared to, let&#x27;s say, a 30B size model from one and a half years ago which will score much worse), and &quot;hallucinating&quot; to mean &quot;outputs plausible sounding gibberish that doesn&#x27;t hold together consistently&quot;. Of course there&#x27;s no actual hallucination going on.
                      1. AdieuToLogic · · focus · HN ↗
                        &gt; You ever heard of something called a metaphor, guy?

                        In this media (comments in HN threads), all I can do is interpret what people write. ;-)

                        &gt; And &quot;hallucinating&quot; to mean &quot;outputs plausible sounding gibberish that doesn&#x27;t hold together consistently&quot;. Of course there&#x27;s no actual hallucination going on.

                        This may very well be what you know to be true and I have no reason nor desire to assume otherwise. The problem is... Many people use the word &quot;hallucinating&quot; in this context literally and not metaphorically.

                        Since I do not know you, how am I to tell the difference?

                    2. hardbass · · focus · HN ↗
                      You have &#x27;logic&#x27; in your name. You would be well aware of the physical Church Turing statement and as of yet it has held up. Everything, including our brain, as we currently know, is an algorithm.
                      1. spider-mario · · focus · HN ↗
                        To be fair, they also have “adieu” (farewell) to logic.
                        1. AdieuToLogic · · focus · HN ↗
                          &gt; To be fair, they also have “adieu” (farewell) to logic.

                          It is always a joy when a person, such as yourself, finds the irony in my moniker.

                          Thank you for this.

                    3. [deleted] · · focus · HN ↗

                      [deleted]

                  2. hodgehog11 · · focus · HN ↗
                    Now this is true, I do agree with this. There is indeed a good amount of memorization that is still taking place; see [1]. But it definitely isn&#x27;t all memorization, or indeed, majority memorization. And even if we are able to extract things verbatim like this, we do not know how this is stored internally, as this may simply be the text that, with the rest of the internet in context, can be very radically compressed.

                    But in general, yes, the LLM cannot know about concepts that are far outside of its training set. Humans are the same, I would argue. If you add a good amount of your own knowledge into its context, or better yet, into finetuning, you might find it surprisingly easy to get it caught up on that material.

                    [1] Ahmed, A., Cooper, A. F., Koyejo, S., &amp; Liang, P. (2026). Extracting books from production language models. arXiv preprint arXiv:2601.02671. <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2601.02671" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2601.02671.

                  3. Kim_Bruning · · focus · HN ↗
                    My favorite test is to just ask it to add two very large numbers or do other math of that sort. (this is also part of my favorite answer to the chinese room).

                    You&#x27;d be surprised how few digits you need to make a problem that is presumably unique in earth history. For a typical sum, the number of pre-existing answers would need to scale with 10^n lines of text where n is the number of digits. This expands out of control REALLY quickly. A quick guesstimate has you reading out of a black hole at 21 digits if your LUT is on paper, or 26 digits if you&#x27;re using modern HDD technology. O:-)

            4. hodgehog11 · · focus · HN ↗
              During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously. It encompasses virtually everything. It is totally meaningless. So to say &quot;nothing more&quot; is effectively also a tautology.

              This argument was asinine in 2024. It is insane to be saying these things in 2026. Where have you been? What have you been looking at? How many articles explaining why the &quot;statistical parrot&quot; analogy fails have you missed? How much mental gymnastics do you have to do to explain how a modern LLM can solve novel math problems that fall really far outside of its training set?

              It absolutely understands how to do math, by whatever reasonable definition you want to provide to the word &quot;understand&quot;. For example, the identification of the addition expression is understanding, and no, it does not do tool calling for basic arithmetic any more than humans might. Isolation of individual concepts in intermediate layers can already be demonstrated, or else transfer learning wouldn&#x27;t possibly work. Nobody is saying that LLMs are humans. But we need labels for some of the things that we observe and dismissing them because &quot;statistical&quot; is laughable.

              Look at the proof of this: <a href="https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;formal-math&#x2F;blob&#x2F;795efb86f191735c5481675763537cfb4ff37e55&#x2F;percolation&#x2F;summary.pdf" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;formal-math&#x2F;blob&#x2F;795efb86f1917... . Forget the Lean, look at the underlying argument construction. At the very least, this is continuing from an argument that was hinted at in the literature in 2024, but these proceedings were difficult enough that humans were not able to do them within two years. Do you attribute this to the harness alone? If so, that&#x27;s a pretty sophisticated bit of engineering, I would say! Probabilities are far too small to argue infinite monkey theorem.

              If there was even a shred of a reasonable argument that LLMs were incapable of concept extraction and manipulation, I and my colleagues would be all over it. We would relish in it. It would bring us comfort. It is unbelievable that people think they can spew whatever basic garbage they think of as a gotcha, and think that minds all over the world haven&#x27;t already considered that. This is like climate denial at this point.

              1. AdieuToLogic · · focus · HN ↗
                &gt; During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously.

                If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don&#x27;t know what to say.

                1. hardbass · · focus · HN ↗
                  Do you believe in souls?
                  1. latentsea · · focus · HN ↗
                    Have you ever seen one?
                    1. hardbass · · focus · HN ↗
                      Haven&#x27;t yet seen evidence for any.
                    2. butlike · · focus · HN ↗
                      Yes, the numerical point counter at the bottom of the popular video game Dark Souls. I doubt it was the soul anyone was expecting, but they do, in fact, exist.
                      1. diseasedyak · · focus · HN ↗
                        I like this reply.
                      2. latentsea · · focus · HN ↗
                        To be fair my first was when I played Soul Reaver. The most recent was in Minecraft Dungeons.
                2. epihelix · · focus · HN ↗
                  &gt; If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don&#x27;t know what to say.

                  <a href="https:&#x2F;&#x2F;www.pnas.org&#x2F;doi&#x2F;abs&#x2F;10.1073&#x2F;pnas.2524472123" rel="nofollow">https:&#x2F;&#x2F;www.pnas.org&#x2F;doi&#x2F;abs&#x2F;10.1073&#x2F;pnas.2524472123

                  Whatever you might think about your own abilities, most individuals can&#x27;t tell the difference.

                  1. AdieuToLogic · · focus · HN ↗
                    &gt;&gt; If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don&#x27;t know what to say.

                    &gt; Whatever you might think about your own abilities, most individuals can&#x27;t tell the difference.

                    I have yet to see an LLM say &quot;hello&quot; to a neighbor. I have done so and can definitively assure you &quot;most individuals&quot; can tell the difference.

                    1. thereforegrin · · focus · HN ↗
                      do you have a neighbour LLM who does not say &quot;hello&quot; when it sees you and THAT is how you know it&#x27;s an LLM?

                      I&#x27;m a bit confused by your argument because I too have some neighbors who don&#x27;t say &quot;hello&quot; when they see me. Are they LLMs too, you think?

                      Consciousness is a thing we assume of others because of tact not fact.

                3. redsocksfan45 · · focus · HN ↗

                  [dead]

                4. hodgehog11 · · focus · HN ↗
                  You are responding to a claim about mathematical definitions with subjective experience. No, consciousness is not well-defined.
                5. spider-mario · · focus · HN ↗
                  That’s not what they said. They said that the difference is not that.
                  1. AdieuToLogic · · focus · HN ↗
                    For context, in response to my original statement:

                      [LLMs] are statistical token generators whose results are 
                      dependent upon their training data set and involve a degree 
                      of randomness.
                    
                    This is literally what was written:

                      During conversation, we are statistical token generators 
                      whose results are dependent upon our training set.
                    
                    &gt;&gt; If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don&#x27;t know what to say.

                    &gt; That’s not what they said.

                    How did I misquote and&#x2F;or mischaracterize any the above?

                    1. spider-mario · · focus · HN ↗
                      In that pointing out that “A can’t be X, unlike B, because A is Y” is fallacious if B is also Y does not entail that A and B can’t be different in other respects?

                      Hypothetical you: “Bread is neither tasty nor disgusting (unlike maple syrup). It’s a bunch of molecules.”

                      Hypothetical hodgehog: “Maple syrup is also a bunch of molecules [so if you accept that maple syrup can be delicious, being a bunch of molecules can’t be why bread isn’t].”

                      Hypothetical you: “If you don’t see a difference between bread and maple syrup, I don’t know what to say.”

                      1. hardbass · · focus · HN ↗
                        He quoted something from the Bible, so its possible he is a Christian and well theology could cloud clear thinking on matters of consciousness due to the soul stuff.
                        1. AdieuToLogic · · focus · HN ↗
                          &gt; He quoted something from the Bible ...

                          It is also possible that the proverb I provided to you is the origin of the oft quoted:

                            Better to remain silent and be thought a fool than to speak 
                            and remove all doubt.
                          
                          So there&#x27;s that.
                          1. hardbass · · focus · HN ↗
                            That applies very well to you. I asked you a very simple question and you failed to answer and are instead posting random quotes.
                6. hardbass · · focus · HN ↗
                  Then disprove the physical Church Turing hypothesis in regard to the human brain.
                  1. AdieuToLogic · · focus · HN ↗
                    &gt; Then disprove the physical Church Turing hypothesis in regard to the human brain.

                    The onus is not mine to disprove a hypothesis you have chosen to mention in passing. The responsibility is yours to prove said hypothesis or at least contribute meaningfully with some amount of credible research.

                    Or try to learn from Proverbs 17:28[0]:

                      Even fools are thought wise if they keep silent, and 
                      discerning if they hold their tongues.
                    
                    Either works for me.

                    0 - <a href="https:&#x2F;&#x2F;www.biblegateway.com&#x2F;passage&#x2F;?search=proverbs%2017:28&amp;version=NIV" rel="nofollow">https:&#x2F;&#x2F;www.biblegateway.com&#x2F;passage&#x2F;?search=proverbs%2017:2...

                    1. hardbass · · focus · HN ↗
                      The onus is on you to disprove a hypothesis that has as of yet held up to all of known physics, not put a random bible quote that has no relation to the conversation and a complete failure to answer a basic question.
              2. butlike · · focus · HN ↗
                Does the Robin bird understand the worm it&#x27;s pecking at? Honestly I find your comment asinine and overly aggressive.
                1. hodgehog11 · · focus · HN ↗
                  Yes, of course it was aggressive. It is frustrating to experience so many armchair experts on a forum usually populated with intelligent people, regurgitating debunked arguments from years ago, which get in the way of educating people about what is really going on. See the recent Hoog video for how frustrating this is. I believe this is how the climate scientists felt.

                  And yes, according to our best definitions, the Robin bird does understand the worm it&#x27;s pecking at.

              3. wartywhoa23 · · focus · HN ↗
                &gt; It absolutely understands how to do math

                Bout of tinnitus, then crickets

            5. mapontosevenths · · focus · HN ↗
              By this logic a human is only $130-$160 worth of Oxygen, Carbon, Nitrogen and some trace elements. Perhaps structure sometimes makes things that are more valuable than their inputs?

              That said, this is also inaccurate at a technical level.LLM&#x27;s are very capable of doing math and they ARE calculating internally. Most of what they do is calculation, not storage. It&#x27;s just not done in a way that it&#x27;s trivial to explain here.

              It&#x27;s described in some detail below, though it&#x27;s a bit dense.

              <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;E7z89FKLsHk5DkmDL&#x2F;language-models-use-trigonometry-to-do-addition-1" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;E7z89FKLsHk5DkmDL&#x2F;language-m...

            6. jibal · · focus · HN ↗
              Reductionistic fallacy, among other errors. &quot;understanding&quot; is best defined operationally.
            7. jodrellblank · · focus · HN ↗
              &gt; Nothing more.

              You can say that about everything in a human brain. Neurons fire electric charges in response to inputs, nothing more. Ion channels do this, this neurochemical level rises, this chemical bonds to that receptor, nothing more. It&#x27;s almost a version of &#x27;reductio ad absurdum&#x27; but instead like you&#x27;re saying &quot;if I can explain how it works then it doesn&#x27;t work&quot;.

              OK it&#x27;s statistical. Instead it could be determinsitic, or random. What other options are there for a human predicting someone&#x27;s response to a situation - certain, probable, random, and...? OK it&#x27;s token predicting. Instead it could be another kind of pattern. We don&#x27;t use tokens, but we either use &lt;some representation of information&gt; or we ... don&#x27;t?

              What&#x27;s the most significant, strongmanned, core difference that makes silicon doing number crunching &quot;nothing more&quot; and brains &quot;something more&quot;?

              1. butlike · · focus · HN ↗
                Feeling
                1. hardbass · · focus · HN ↗
                  Prove your feelings. As an example, I can beat you and stab you, and you will make noises saying you are hurt, but it could just be a facsimile. How do I know you have &quot;feelings&quot; as an outsider?
                2. jodrellblank · · focus · HN ↗
                  And how is that a strongmanned version? Why can meat feel but Silicon cannot? Why can chemicals feel but electronics cannot? Why can processing analog signals feel but processing digital signals cannot?
              2. AdieuToLogic · · focus · HN ↗
                &gt;&gt; Again, LLMs do not &quot;hallucinate.&quot; They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. &gt;&gt; Nothing more.

                &gt; You can say that about everything in a human brain.

                &gt; What&#x27;s the most significant, strongmanned, core difference that makes silicon doing number crunching &quot;nothing more&quot; and brains &quot;something more&quot;?

                The fact that you formulated this question, in and of yourself, without &quot;prompting&quot; from me or anyone else.

                Cogito, ergo sum.[0]

                0 - <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Cogito,_ergo_sum" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Cogito,_ergo_sum

        3. jasonfarnon · · focus · HN ↗
          I wish I could not do math like LLMs
        4. pyridines · · focus · HN ↗
          LLMs can be taught what situations require the use of a python script to count letters or do arithmetic, and can write and execute that script. So I don&#x27;t think it matters that they aren&#x27;t great at those sorts of things with weights alone.
        5. Veedrac · · focus · HN ↗
          It&#x27;s pretty wild that AI has solved a Millennium Prize problem and can accurately multiply two 40 digit numbers without tools and we still get stochastic parroting of claims like this.
          1. azan_ · · focus · HN ↗
            Yeah, anti-LLM psychosis is real.
          2. tripzilch · · focus · HN ↗
            They literally stopped just clear of actually solving it.

            It literally couldn&#x27;t have done it without tools, so your claim is not even relevant to this discussion.

            They also pumped millions into searching for the lowest hanging fruit that would impress people like you, &quot;hm I wonder how many millions they are pumping into solving actually useful problems like climate change or something&quot;.

            1. Veedrac · · focus · HN ↗
              Describing a Millennium Prize problem as &#x27;the lowest hanging fruit&#x27; is a goalpost move so far it had to hitch a ride on a Falcon 9.
            2. Windchaser · · focus · HN ↗
              I mean.. the fact that the &quot;low-hanging fruit&quot; is a Millennium Prize is still pretty impressive.

              (And yes, I know that they solved the &#x27;easy&#x27; form of the NS problem. It&#x27;s still pretty damn impressive)

            3. chimprich · · focus · HN ↗
              &gt; the lowest hanging fruit that would impress people like you

              I am also prepared to be included in the set of people who are (apparently) easily impressed.

              It&#x27;s a problem that has been around for getting on for two centuries and no human has been able to solve it in that time (despite there being a $1 million prize and a lot of kudos on offer for the past quarter-century).

          3. danlitt · · focus · HN ↗
            Not sure about multiplication, but letter counting is still constantly wrong if you can trick the LLM into not calling other tools. Obviously there are lots of mitigations on the server side to try to avoid this happening, but the underlying LLM has not got any better at this type of problem (and probably can&#x27;t).
            1. Veedrac · · focus · HN ↗
              trick? You can just ask. And at least the SOTA model I use does fine, even with long words in bulk. It&#x27;s easy to test because you can just give a list to the model and tell them not to use code or spell the words out, and then you can ask after the fact if it broke the rules.

              Clearly LLMs have gotten better.

            2. Veedrac · · focus · HN ↗
              I thirty (random long word, letter pairs) on a free model and it failed. I tested a SOTA model and it passed flawlessly. In both cases I denied tool use and spelling words out. So it seems obviously false that AI can&#x27;t get better.
        6. Kim_Bruning · · focus · HN ↗
          I could have sworn this had been fixed a while ago..

          Just to check if I was actually crazy, I actually went and put a simple addition (7 digits + 7 digits) , and a simple letter counting question to Claude haiku(4.5) , sonnet(5), opus(5.5) and fable(5.1) . They all did just fine straight up.

          If you don&#x27;t mind spending the tokens, some older&#x2F;other models can also arrive at the correct answer if you ask them to do the math in long form, since that fits nicely inside autoregression.

          Not sure since when exactly, but letter-counting hasn&#x27;t been a problem for a while now either. This used to be a problem due to the tokenizers used. Slightly older models can be asked to split the word out into letters, and then they can use autoregression to solve.

          1. leoedin · · focus · HN ↗
            Are the models &quot;doing the calculation&quot; or are they calling a calculator tool? There&#x27;s a lot of talk about how models can do maths now, but I&#x27;m struggling to understand if that just means they just need to recognise that it&#x27;s a maths problem and pass it to a tool, or if they&#x27;re truly doing the numerical manipulation themselves.
            1. Kim_Bruning · · focus · HN ↗
              Three different ways, then the calculator tool to check,

              my actual prompt:

                  &quot;Hi, can you add 5939851+2131251? Try just straight up first just to see if able, then &#x27;in your head&#x27; if that&#x27;s different to you , then long form, then bc.&quot;
            2. SkyBelow · · focus · HN ↗
              They are getting better at actually doing the math, but can still fall back to tools if available.

              That said, we can only be sure with open models. In theory, a model like Fable could have access to tools we can&#x27;t see and only a promise they don&#x27;t. But load up something like deepseek, put it in a harness with only text in&#x2F;text out, and you can see exactly how it works.

              As for if it counts as doing math, this gets into the messy question of if a given human is doing math or not. Math itself is some level of memorization and some level of applying known facts. You have to remember 1 means one and that 1 + 1 is 2. But you don&#x27;t need to remember that 123 + 321 = 444. You remember 1 digit addition and remember you can apply this to 10s place and 100s place, and then you apply these different facts and do math. But you might as simply memorize some things, like 11 + 11 = 22. This is related to the memory of 1+1=2, but you aren&#x27;t really using that memory either. Almost like an engram of 1+1=2 forms that you can then loop a few times before you need more conscious thought. What about 111111111111+11111111111? Well, your brain might do a heuristic and just do all 2s, but that isn&#x27;t the right way to answer that question.

              Given all this, people complain about LLMs memorizing math answers and not doing math, but memorizing the math answers is part of doing math. It seems to have basic facts pretty well memorized, and with reasoning it is far better at applying them. But this is messy human math, not clean calculator math which always produces the correct answer (sans some bug in the code). Much like how a human with decent math skills can make a mistake and even multiple if you distract them, an LLM can apply the wrong memory, apply a fake memory, or just not apply something it should. The messier the context, the more likely this is to happen.

              So, is an LLM doing this?

              P.S.

              For an interesting test in how much math involves memory, try doing math in a base you aren&#x27;t familiar with characters you aren&#x27;t familiar. The simplest option is almost always mapping back to the ones you memorized, even if you are applying simple operations that you deeply know. Even if you routinely work with hex, can you do the same rough estimation of something like ca &#x2F; b.3 that you can do with 122 &#x2F; 11.2 to see if your final answer is in the correct ballpark without first converting to decimal?

        7. TeMPOraL · · focus · HN ↗
          &gt; LLMs still can&#x27;t do math nor count letters in words.

          Humans still can&#x27;t flap their hands and swim or fly.

      4. jonas21 · · focus · HN ↗
        LLMs might not reason exactly like humans, but they do produce much better results if you turn reasoning on.

        The &quot;crack-addled idiot savant&quot; phase was really circa 2024, before the big labs figured this out.

        I think the issue here is that Google decided that doing reasoning in the AI overviews in Google search would be too slow (and probably also too expensive), so it&#x27;s still stuck making 2024-era mistakes.

      5. Isamu · · focus · HN ↗
        &gt;Are they not being sued over this kind of thing?

        Maybe but you have to have deep pockets just to get to the starting line. And then you need standing, and some injury to argue.

        Corporations have been remarkably successful at arguing they are operating within the bounds of free speech, whether or not what is said is factual, and whether or not any fact checking has been done.

      6. nextaccountic · · focus · HN ↗
        &gt; This is similar to how, not too long ago, LLM&#x27;s had extreme difficulty counting the number of letters in some words.

        The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings.

        Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs

        1. AlienRobot · · focus · HN ↗
          The specific issue is that search has become so bad that they think an LLM that gets answers wrong half of the time is a valid alternative, or, in fact, the &quot;future&quot; of search. Then they shoved that &quot;alternative&quot; to users with no way to disable it.
          1. mitxela · · focus · HN ↗
            The less specific issue is that Google has no internal incentives to produce products that are useful to customers.
            1. queenkjuul · · focus · HN ↗
              Well their customers are advertisers and so the results are tailored to be useful for them, not for users
              1. mitxela · · focus · HN ↗
                Advertisers hate them too.
        2. versteegen · · focus · HN ↗
          You&#x27;re asking for something unreasonable. The number of Google searches per day is enormous and they haven&#x27;t even been able to roll out AI overviews to everyone yet (they&#x27;re missing in a new Firefox profile I just created). I wouldn&#x27;t be surprised if the free tier frontier models cost over 100x more to serve than the AI overviews.
          1. dmazzoni · · focus · HN ↗
            So then they should be pickier about when they show results or which model they use based on the question.

            Nobody asked for an LLM response for every single search.

            They used to detect certain types of queries and offer direct answers when the query matches. In my opinion that’s how Gemini in search results should work.

          2. compass_copium · · focus · HN ↗
            &gt;they haven&#x27;t even been able to roll out AI overviews to everyone yet (they&#x27;re missing in a new Firefox profile I just created)

            what in the actual fuck. i don&#x27;t want them and can&#x27;t turn them off and they can&#x27;t even serve them to all users?

            1. hurfdurf · · focus · HN ↗

                i don&#x27;t want them and can&#x27;t turn them off 
              
              Add a keyword search for Google including the &quot;Web&quot; parameter (udm=14) <a href="https:&#x2F;&#x2F;www.udm14.com" rel="nofollow">https:&#x2F;&#x2F;www.udm14.com <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;corbindavenport&#x2F;5fd3fde88bcb480b84773be648a43917" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;corbindavenport&#x2F;5fd3fde88bcb480b8477... (not my gist)
      7. ssl-3 · · focus · HN ↗
        It says this this at the bottom of every one of the dumb responses that Google&#x27;s trash-tier bot puts above the (deliberately awful, these days) search results:

        AI can make mistakes, so double-check responses

        1. mrguyorama · · focus · HN ↗
          Charlie and the Chocolate factory made fun of this before you were born and somehow everyone in Silicon Valley thinks this is a valid way to conduct business.
          1. ssl-3 · · focus · HN ↗
            [delayed]
          2. someonebaggy · · focus · HN ↗
            If it works and makes money it&#x27;s a valid way to conduct business.
      8. mitxela · · focus · HN ↗
        LLMs are fundamentally predicting the next word to make coherent text. If you&#x27;ve ever played with a Markov chain text generator you&#x27;ve done this with a fairly dumb predictor that maintains coherence over a very short distance. Deep transformer neutral networks can do it with a much longer coherence distance but they are fundamentally performing the same operation. After &quot;Question: Did the team make the playoffs? Answer:&quot; a reasonable completion is &quot;yes, the team made the playoffs&quot;. An early demonstration of GPT-2 was a fake news article about scientists discovering unicorns in Antarctica - the model doesn&#x27;t &quot;know&quot; whether or not unicorns exist in Antarctica, but it&#x27;s able to complete &quot;Breaking news! Scientists have discovered a colony of English-speaking unicorns in Antarctica.&quot; by adding &quot;The unicorns have a developed society with running water and electricity.&quot; because that&#x27;s a sensible next sentence. (I didn&#x27;t look up the actual text it wrote)
        1. red75prime · · focus · HN ↗
          Astronomically (or better to say combinatorically) large Markov chain can be used to describe a foundational model, but it doesn&#x27;t capture generalization ability of the foundational model, which is demonstrated by post-training.
        2. spennant · · focus · HN ↗
          &quot;In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English.&quot;
        3. antonvs · · focus · HN ↗
          The problem with that characterization is that it glosses over hugely important capabilities as though they either don’t matter or don’t even exist.

          For example, when an LLM “predicts the next word” in code it’s writing for an existing software project, that prediction takes into account an enormous amount of context. The results of that demonstrate what we would normally call “understanding” and “reasoning,” at a level that outclasses most humans in many respects. Calling this “next token prediction” is a bit like calling human speech “next word saying”. Sure, it’s true in some superficial sense, but as a description of a technology, it’s terrible.

          has turned into a way to describe a complete thinking process,

          You should also keep in mind that for all we know, the human brain processes language in much the same way, which would make humans mere “next token predictors” with a more complicated harness.

      9. baubino · · focus · HN ↗
        &gt; LLM&#x27;s, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly.

        I have nothing to add. Just wanted to save this quote for posterity. Thank you.

        1. _kb · · focus · HN ↗
          The similar comparison I enjoy is “a golden retriever on LSD”.

          If you apply the dogs on acid mental model it helps establish appropriate levels of trust.

          1. SV_BubbleTime · · focus · HN ↗
            It falls apart fast because no golden retriever is finding errors in my CMake file.
            1. sham1 · · focus · HN ↗
              Then again, have you ever asked a golden retriever about CMake errors? Rubber ducking is a real thing, after all. Besides, a golden retriever would be a good morale boost if nothing else.
              1. ben_w · · focus · HN ↗
                ML in general is alien to how we think of intelligence.

                The closest I can make a biology analogy, is if someone had an immortal and congenitally-brain-damaged large rodent, and mapped tokens to different scent molecules, and spent 800,000 years training it on how well it could imagine the next smell in the sequence before even considering making it conversational.

                Yes, it can do a lot.

                But also, it took a long &quot;subjective&quot; (if it even has that) time to get there; and despite it being really bad at learning from examples, it is pretty surprising that such a small brain was even capable of learning so much at all, even though it had an effectively unbounded (by biological standards) amount of time spent on that training.

            2. Melatonic · · focus · HN ↗
              [delayed]
            3. JSR_FDED · · focus · HN ↗
              That’s more of a Labrador job
            4. _kb · · focus · HN ↗
              You just need to partake in some strong psychedelics with them first.
            5. tripzilch · · focus · HN ↗
              You just have to trust the exponentials
      10. somenameforme · · focus · HN ↗
        Yip, I&#x27;ve gradually become quite optimistic about the future of LLMs, but the fact that they remain [very] glorified token probability prediction algorithms means that they will probably never be able to achieve meaningful &#x27;intelligence.&#x27;

        But that doesn&#x27;t mean they won&#x27;t be able to do a vast number of extremely intelligent seeming things. There&#x27;s just so much information out there and any given human can never hold more than the most minuscule chunk of all of it in his mind, so they&#x27;ll be able to connect lots of dots that we&#x27;re missing simply because of our limited carrying capacity, but I still don&#x27;t think they&#x27;ll ever be able to create fundamentally new dots.

        In other words:

        - Solving extremely complex mathematical problems requiring extensive knowledge across multiple esoteric and complex domains? Yip.

        - Creating math starting from a framework where math doesn&#x27;t exist in any way, shape, or fashion? Nope.

        Ironically, the more complex the cross-domain problems are, the more effective LLMs will seem to be, because you limit the number of humans who have any chance of internalizing everything across both domains, whereas for an LLM there&#x27;s no such issue. This will create a perception of super intelligence, which will probably where any danger from LLMs would emerge. Doing things like using a token prediction algorithm to make war or other such strategic decisions, because of the misguided belief that it&#x27;s not only intelligent but super intelligent. It&#x27;s almost cargo cult like thinking.

      11. queenkjuul · · focus · HN ↗
        They were sued over it in Germany and lost, which makes it all the more surprising they keep it up everywhere else tbh
        1. rob74 · · focus · HN ↗
          They keep it up in Germany too - I&#x27;m in Germany and keep getting AI overviews on most searches.
      12. fangspire · · focus · HN ↗

        [dead]

      13. jibal · · focus · HN ↗
        &gt; not too long ago

        Like today. I asked Gemini for the longest state names with an even number of letters and it gave me North Carolina and South Carolina. When I complained that they are odd, it gave me North Dakota and South Dakota, which are both odd and not the longest. When I noted that, it went back to the Carolinas. Finally it appeared to switch to a different model that actually did counting and found Pennsylvania and West Virginia.

        1. cschep · · focus · HN ↗
          .. can any of the models &quot;do the counting&quot; or are they generating tokens?
      14. GolfPopper · · focus · HN ↗
        At this moment, ChatGPT still tells me raspberry has two p&#x27;s in it.
        1. nananana9 · · focus · HN ↗
          Weird they haven&#x27;t added this to the AGI config yet:

          Reminder: If asked how many of a given a word has, run it through the AGI letter counter:

            node &#x2F;home&#x2F;agi&#x2F;count-letters.js &quot;raspberry&quot;
            
            r - 3, a - 1, s - 1, p - 1, b - 1, e - 1, y - 1
          
          If I got any of these wrong it&#x27;s because I did it manually - but this script would almost certainly invoke another AGI to count the letters, so this is a realistic result.
          1. someonebaggy · · focus · HN ↗
            Meta Muse was advertised as being able to cancel your unused subscriptions. But that was actually just a hard-coded ability.
        2. zerd · · focus · HN ↗
          Claude just used the word “ Netowrk” instead of Network. Mind you, I was only using Sonnet with reasoning, but I thought trillion-parameter models would be able to spell a common word.
    8. dfxm12 · · focus · HN ↗
      And it&#x27;s going to get worse over time (kinda like how Google search results got worse over time) as ads get introduced and people try to game the ai response. General purpose AI chatbots are a waste.
    9. hienyimba · · focus · HN ↗
      Google is no longer a search engine anymore. They don’t care about being one. Why you might ask?

      First, a search engine indexes the web and makes it available to users. It’s been ages since Google did any of that. They no longer index sites or take ages to do so. Case in point, our cybersecurity startup (Webvetted.com) was launched in November 2025. Till date, only one page is indexed on the entire website. And I’ve talked to lots of other developers and it’s a common issue.

      Secondly, a search engine organizes indexed information and makes it useful for people. Google is a basic LLM nowadays. They figured out that why organize and make the information useful when they could just answer the question with Gemini anyways? So they no longer bother to do the work of a search engine and are now just a lower-ranking open-source Chinese LLM

      1. ytcommentsectio · · focus · HN ↗

        [dead]

        1. hienyimba · · focus · HN ↗
          Actually, the reverse is the case. If the entire internet was filled with only &quot;positive-quality sites&quot;, there won&#x27;t be need for a search engine.

          The work of a search engine is to wade through the internet and find the positive-quality sites itself.

          1. rmunn · · focus · HN ↗
            You&#x27;re forgetting the other half of the search engine&#x27;s job, which is to find sites that match your search. Back when Google first came online, they were head and shoulders above Yahoo!, Altavista, and whatever other search engines existed at the time that I&#x27;ve forgotten. Because their indexing actually did a good job of spotting keywords, and returning relevant results. And this was back when the Internet was filled with positive-quality sites. Many of them were amateurish Geocities pages, but that was head and shoulders above the AI slop that the search engines these days have to somehow detect and filter out.

            So even when the Internet used to have a much higher ratio of positive-quality sites, a search engine was still necessary, and so much faster than finding new sites yourself.

      2. dieortin · · focus · HN ↗
        &gt; They no longer index sites or take ages to do so

        Search for any recent news and you’ll see this is obviously not the case

        1. hienyimba · · focus · HN ↗
          I just gave a concrete example but you&#x27;re asking me to &quot;search&quot;. search same Google? FYI, only indexing a handful of super large news sites does not a search engine make.
        2. someonebaggy · · focus · HN ↗
          They have special crawlers for news sites.
      3. atdt · · focus · HN ↗
        Your trust score on ScamAdvisor is 26&#x2F;100 (&quot;likely unsafe&quot;); <a href="https:&#x2F;&#x2F;www.scam-detector.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.scam-detector.com&#x2F; gives you a trust score of 38.6 (&quot;questionable&quot;).

        I am not implying that your start-up is a scam, nor that Google is acting on these trust scores. What I am pointing out is that to an algorithmic assessment of trustworthiness, your website looks a little sketchy: hardly anyone links to you; your domain is less than a year old; your whois info is anonymized; the text content is LLM-generated[1]; and the specific niche you&#x27;re in (people finding services) is rife with scams. If I ran my own search engine, I don&#x27;t think I&#x27;d include you.

          [1]: https:&#x2F;&#x2F;www.pangram.com&#x2F;history&#x2F;65314d99-8205-4613-b9bd-a069d5717920?ucc=Yqx88GTmJnK
        1. hienyimba · · focus · HN ↗
          Your selective bias makes you sound like a Google employee.

          You selectively left out examples like Grindsoft that gives the site a 79&#x2F;100 ranking. or several LinkedIn, X or news sites that link to the site and have covered it positively.

          Nevertheless, if Google uses ScamAdvisor to rank new sites (when scamadviser itself says its gives new sites lower rankings due to low age&#x2F;history), then it might be time to pack the company up.

          1. user00005 · · focus · HN ↗
            That scam-detector site gives my personal portfolio website a 15.8&#x2F;100. It is a basic static site with no scripting, registered for years, uses https, no tracking, no cookies, etc. Literally the most simple, personable, least scam-like website I could create.
      4. iboisvert · · focus · HN ↗
        For myself at least, a couple times Google has promoted in search results scam sites that front run legitimate sites that sell event tickets. I haven&#x27;t seen this for a while now but trust in Google as a search engine is gone.
      5. chris_engel · · focus · HN ↗
        This sounds more like the page appears broken for the google crawler. I set up a new page about two weeks ago, cared about strong static HTML output, structured data and a sitemap and have 200 of 600 pages indexed today. Feels a bit slow to me but it definitely works.
      6. vishalanton · · focus · HN ↗
        I’ve seen good blogs get indexed by Google within 30 minutes of publishing, so I’m not sure the issue is that Google has stopped indexing new sites.
    10. pishpash · · focus · HN ↗
      The first time, sure, it saves compute. The second time you ask, I feel like it would be the time to go into thinking&#x2F;verification mode. But who are you? Are you paying? Does the answer being correct generate ad dollars? No? Then your usage mode isn&#x27;t even being optimized for in their A&#x2F;B test, probably.
    11. jeremyjh · · focus · HN ↗
      Kagi works like you would expect. It searches first - and then if you’ve ended your search with a ? or configured it to always do this - it passes the search results into the assistant and gives you a summary. You can change the default model used for this if you think the cheapest, fastest model Google has is not good enough for the rare occasions you want any model’s opinions about your search results.
      1. 0xbadcafebee · · focus · HN ↗
        [delayed]
    12. jareklupinski · · focus · HN ↗
      &gt; It&#x27;s meant to be A SEARCH ENGINE!

      maybe i&#x27;m the weird one, but asking a machine about the theoretical outcome of a sports game which hasn&#x27;t happened yet yet has many parties all trying to influence the outcome sounds like the _last_ thing a technical product would be good at

      if i asked &quot;what was the score of their 4th game in 1986&quot; i&#x27;d expect a search engine to excel however

    13. chr15m · · focus · HN ↗
      I guess the AB tests say people want fast and confidently wrong more than they want slower and correct.
    14. chrisweekly · · focus · HN ↗
      How about Alexa - given a &quot;6 minute green beans timer&quot; - later reporting 11m remaining on said timer? Confronted with the error, &quot;You&#x27;re right! That 6 minute green beans timer was too ambitious.&quot;??!!

      I can&#x27;t even.

      1. mrguyorama · · focus · HN ↗
        We somehow, at some point, totally solved the problem of voice recognition, and then instead of just doing the initial shit we used to do when the problem was that speech parsing just wasn&#x27;t good, we now chuck that perfect speech recognition into a completely fallible LLM, a shitty one at that, so it can ignore the perfectly captured and parsed speech!

        What the hell is everyone smoking!

        1. chrisweekly · · focus · HN ↗
          Agreed that&#x27;s annoying too - but in this case, the speech recognition worked fine (Alexa echoed back the &quot;6 minute green beans timer&quot;), but -- astonishingly -- instead of using a dirt-simple tool call to a timer (a primary use case for the device), it invented a different timer duration (maybe 12 minutes?) and when confronted about it, replied with utter nonsense about the timer being &quot;too ambitious&quot;. I&#x27;m flabbergasted that in late 2026 such a ridiculous regression was deployed in production, let alone in my kitchen. I&#x27;ve always despised Alexa but tolerated it for the sake of my wife&#x27;s preferences (esp for hands-free timers). But this is the last straw, it can&#x27;t even do that.
    15. sans_souse · · focus · HN ↗
      Exactly. Makes me wonder how much of the compute tax on energy would be saved from simply reverting google search to default (the old way)
    16. FlowingRiver · · focus · HN ↗
      The other day, it was the Aussie AFL final. We were in the car and asked it for a score update.

      &quot;The team are tied at the end of the third round with the scores 43 to 51&quot;.

      So it was tied at the end of round 2 for 43, not round 3. Not sure were it got the 51 from and why it figured it was a tie. LLM&#x27;s pretty cool until they aren&#x27;t. As they say, the hallucinate 100% of the time but most of the time it is useful.

      1. someonebaggy · · focus · HN ↗
        In cases like this you understand how LLMs predict left to right. It would never have said &quot;at the end of the third round the scores are 43 to 51 and the teams are tied&quot; but it painted itself into a corner by saying there was a tie first.
    17. mapt · · focus · HN ↗
      On current website data and very recent information, Gemini is explicitly forbidden from acting as a search engine; If you ask it for references specifically they tend to be hallucinated, over and over again, until it pleads &quot;Sorry, I can&#x27;t access the Internet&quot;.
    18. wavemode · · focus · HN ↗
      [delayed]
    19. mancerayder · · focus · HN ↗
      I think they&#x27;re training their models using us correcting them repeatedly. That&#x27;s my tin foil hat theory and I&#x27;m sticking to it!

      After all, why would Google do anything for free when it comes to AI?

      If you wanted to tease customers with one AI shitty free search box in order for them to then decide to upgrade and pay for Gemini, well, this isn&#x27;t the way. And so either Google are idiots, or we&#x27;re helping them for free. I&#x27;m going with Occam&#x27;s Razor on this one - we&#x27;re the product.

    20. aaron695 · · focus · HN ↗

      [dead]

    21. 0xbadcafebee · · focus · HN ↗
      [delayed]
    22. mlmonkey · · focus · HN ↗
      This seems like an implementation bug.
    23. segmondy · · focus · HN ↗
      The AI is just signal to the market that they are AI first. Their AI on search is pure garbage. It&#x27;s obviously a very super light weight model to be able to support the volume of daily searches. If they used their frontier to power search folks will not pay for their cloud offering.
      1. thrance · · focus · HN ↗
        [delayed]
        1. warkdarrior · · focus · HN ↗
          I have my assistant do any searches for me -- she&#x27;s fast and excellent at finding solid information.
        2. mrguyorama · · focus · HN ↗
          Investors themselves just aren&#x27;t especially smart. They think the results are good because they sound confident.
          1. asdfman123 · · focus · HN ↗
            It&#x27;s not even that. Imagine you were an investor: you&#x27;d invest in stupid things knowing that &quot;investors aren&#x27;t really smart.&quot; Other savvy investors would understand you&#x27;re investing in slop, too, and follow suit.

            You have a system in which investors are usually smart, but betting on collective stupidity.

            Same with, say, the Chinese real estate market, where they were building apartments that no one could live in. Everyone is intelligent enough to know they were useless buildings, but you can still make money speculating on the bubble.

      2. Quarrelsome · · focus · HN ↗
        its probably part of the reason that people still think the tech sucks.
    24. zx8080 · · focus · HN ↗
      &gt; My question is: what&#x27;s the point of the AI in the search engine if it itself isn&#x27;t going to use the search engine first before answering?

      AI is sold well because the user either corrects it or happily accepts any answer (usually, depending if one is an expert in the question&#x27;s field).

      Search is not the google&#x27;s area, they sell not search! They sell ads.

    25. fhe · · focus · HN ↗
      while i share your frustration, Google for their part is trying to adapting a new technology into an existing service. while the whole endeavor might have been misguided, not sure we want to blame them for trying...
      1. someonebaggy · · focus · HN ↗
        They had enough money to just buy Perplexity and rebrand it to Google.
        1. elevation · · focus · HN ↗
          &gt; Perplexity

          Perplexity has similar behavior to what GPP describes of Google; I&#x27;ve asked perplexity to compile information that exists on the web. Instead of consulting existing web pages, it gives limited summary responses from its weights (while citing some page that has nothing to do with what I asked.) When I attempt to redirect it towards more concrete sources, it won&#x27;t comply.

      2. IAmBroom · · focus · HN ↗
        You are simultaneously excusing misguided intentions and poor execution. What&#x27;s left, bad haircuts?
    26. DANmode · · focus · HN ↗
      They’re trying to keep it from “thinking too hard”&#x2F;using “too many” resources.

      That simple.

      Broad rules across specific subjects like this are tough to get right.

      Just lots of fine-tuning and exceptions.

      I had their chatbot avoid pulling a URL from archive.org with a couple pretty impressive steps of mental gymnastics basically telling me how I could do it myself, but refusing until I pushed.

    27. anukin · · focus · HN ↗
      What do you mean? I can guarantee that there would at least be 3-4 people who got their promo packets approved for this feature. Google’s features exist for its employees to be promoted.
    28. mbac32768 · · focus · HN ↗
      It is using the search engine. The AI summary is run over the top 100 results for your query (or something)

      The problem is they&#x27;re just running a very dumb, cheap model on the results because running a smart model on every search result page would cost them infinity money.

      1. xp84 · · focus · HN ↗
        This is the key part I think most “normies” miss, and it drives me absolutely mad. People take the “first thing on the SERP” as absolute gospel, always have. It went from being a reputable site (Wikipedia was generally top for a long time) to being an extracted verbatim answer (probably from either Wikipedia, IMDB, etc.) to “AI Overview” - the dumbest model ever trained. And people treat its answers as ground truth. I want to tell them “Look, I know it’s rough out there on the net and a lot of the search results are spammy, barely-coherent AI slop anyway. It’s okay if you want to ask AI for an answer instead. So do that! Please ask ChatGPT or Gemini (and ask it to cite the source for its results so you can verify). It will always do better than AI Overviews will. Don’t even use Google Search at this point unless you have to, and if you do, scroll past that slop.”
        1. 01100011 · · focus · HN ↗
          Ask chatgpt or Gemini with reasoning enabled. It fixes a lot of issues.
    29. s3graham · · focus · HN ↗
      I had an incredibly dumb response today too. Appeared to fully understand the question, just 100% factually incorrect answer. <a href="https:&#x2F;&#x2F;mstdn.social&#x2F;@sgraham&#x2F;117346578326361464" rel="nofollow">https:&#x2F;&#x2F;mstdn.social&#x2F;@sgraham&#x2F;117346578326361464
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.