‹ BackHN Continuity

Thread

When did Google get so weird?

2011 points · 1123 comments · sancho-panza

  1. Hugsbox · · focus · HN ↗
    Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"

    So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.

    So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"

    It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.

    My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!

    1. beloch · · focus · HN ↗
      This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.

      LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.

      I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?

      1. VCFundedGenYer · · focus · HN ↗
        LLMs still can't do math nor count letters in words. Nothing has changed there.
        1. fasterik · · focus · HN ↗
          "LLMs can't do math" is a pretty hot take in September 2026.
          1. mbgerring · · focus · HN ↗
            They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.
            1. versteegen · · focus · HN ↗
              It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)

                TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
                next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
                pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
              
              <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;eRmzz8J8Qkzqvzrgg&#x2F;astra-can-do-a-concerning-amount-with-no-chain-of-thought" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;eRmzz8J8Qkzqvzrgg&#x2F;astra-can-...
            2. GaggiX · · focus · HN ↗
              Reasoning models can do math on their own without external tools.
              1. dcrazy · · focus · HN ↗
                Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.
                1. einichi · · focus · HN ↗
                  I don&#x27;t think anybody is arguing that LLMs do math better than a traditional processor
                  1. unshavedyak · · focus · HN ↗
                    Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc.

                    The nice thing about math is it can easily plug into a tool, making it even less of a concern.

                2. hodgehog11 · · focus · HN ↗
                  I would say purposeful misremembering. The LLM can be run with zero temperature after all.
              2. demibabs · · focus · HN ↗
                Even without reasoning.

                5.6 on Instant mode can knock out 3 digit multiplication just fine.

            3. vanuatu · · focus · HN ↗
              reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)
              1. krapp · · focus · HN ↗
                There needs to be a Godwin&#x27;s Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1.
                1. Forgeties79 · · focus · HN ↗
                  Krap Law
                2. vanuatu · · focus · HN ↗
                  functionally, LLM behavior indeed shares many parallels with human cognition.
                  1. krapp · · focus · HN ↗
                    Functionally, a Markov chain shares many parallels with human cognition for the same reasons.

                    People don&#x27;t understand what correlation is and they assume the mapping between a human brain and an LLM is 1:1 in every case where it matters, which is a religious and not scientifically based belief.

                    1. vanuatu · · focus · HN ↗
                      I agree, it&#x27;s correlated and useful as a mental model for using LLMs, not as an explanation.
            4. ericmay · · focus · HN ↗
              Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus?

              &gt; The LLM is not suited to giving deterministic answers to math problems.

              Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.

              1. queenkjuul · · focus · HN ↗
                Pretty sure I&#x27;ve had 5x6 memorized accurately since i was 7 years old

                <a href="https:&#x2F;&#x2F;x.com&#x2F;maksym_andr&#x2F;status&#x2F;2100364212207837560" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;maksym_andr&#x2F;status&#x2F;2100364212207837560

                1. Dylan16807 · · focus · HN ↗
                  I can&#x27;t tell if you&#x27;re joking but they&#x27;re talking about 5 digit and 6 digit numbers.
                2. ericmay · · focus · HN ↗
                  Sorry I don&#x27;t use social media so I&#x27;m not going to be able to read the linked content.

                  But even if you have 5x6 memorized it&#x27;s not deterministic that you answer 30. It&#x27;s just highly probable.

          2. tjwebbnorfolk · · focus · HN ↗
            They can do math but not arithmetic
            1. fasterik · · focus · HN ↗
              I just asked ChatGPT 5.6 Sol (High) to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right.

              <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ab9a0da-fdd0-83e8-a62d-f0cdeb54db48" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6ab9a0da-fdd0-83e8-a62d-f0cdeb54db...

              I&#x27;m sure it still makes mistakes, but saying it can&#x27;t do arithmetic is just false.

              1. tremon · · focus · HN ↗
                Are you sure it honoured your stipulation of &quot;without external help&quot;? For all we know, it hacked its way into Wolfram Alpha and got the result from there.
                1. jmillikin · · focus · HN ↗
                  Arithmetic is well within the capabilities of even small local models: <a href="https:&#x2F;&#x2F;i.imgur.com&#x2F;21tzGlN.png" rel="nofollow">https:&#x2F;&#x2F;i.imgur.com&#x2F;21tzGlN.png
              2. guelo · · focus · HN ↗

                [dead]

              3. Xirdus · · focus · HN ↗
                I tried prompt &quot;6379 times 3875&quot; and it was off by exactly 1000 on first try, and correct on second. 0% success rate, sample size of 1.
                1. dcrazy · · focus · HN ↗
                  Isn’t that a 50% success rate with a sample size of 2?
                  1. Xirdus · · focus · HN ↗
                    AFAIK you can&#x27;t combine results from multiple studies this way? But I&#x27;m not an academic.
                  2. Xirdus · · focus · HN ↗
                    Not really, since it wasn&#x27;t a fresh context with a fresh question. I just told it it&#x27;s wrong in the same chat session and it corrected it there.
              4. amluto · · focus · HN ↗
                I would be nice to see what the (unencrypted) reasoning trace is like. Multiplication with scratch paper is not particularly difficult.
            2. dcrazy · · focus · HN ↗
              LLMs can in fact do arithmetic, just not reliably owing to how numbers are represented probabilistically: <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2410.21272" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2410.21272
          3. Isamu · · focus · HN ↗
            That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.

            I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.

            1. hodgehog11 · · focus · HN ↗
              No it isn&#x27;t. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.
              1. dwaite · · focus · HN ↗
                Hey hey, we obviously should ask Gemini to settle this disagreement.
              2. spartanatreyu · · focus · HN ↗
                Counterexample from only 2 months ago:

                <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=iTyLHDRhwJg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=iTyLHDRhwJg

                1. Dylan16807 · · focus · HN ↗
                  That&#x27;s not a math question...
                  1. spartanatreyu · · focus · HN ↗
                    Counting is literally a part of Mathematics.
                    1. Dylan16807 · · focus · HN ↗
                      Not really.

                      And it can count fine. It doesn&#x27;t know how to spell.

                      1. spartanatreyu · · focus · HN ↗
                        &gt; Basic counting isn&#x27;t really math.

                        Counting most certainly is math.

                        Can you define or explain counting without also referencing&#x2F;defining&#x2F;explaining a mathematics concept?

                        &gt; And it can count fine. It doesn&#x27;t know how to spell.

                        It&#x27;s the other way around, it could spell, it couldn&#x27;t count the letters.

                        1. Dylan16807 · · focus · HN ↗
                          It doesn&#x27;t input and output letters. It&#x27;s like if it communicated with a version of sign language that&#x27;s closer to written English. It&#x27;s using all the same words but it has an arbitrary signal or two for each one.

                          It can&#x27;t spell for garbage because the tokenizer hides the real spellings from the LLM.

                2. hodgehog11 · · focus · HN ↗
                  Even if this was a math question that he asked, this is live audio input&#x2F;output from a multimodal model. It is designed for rapid responses. This is not even remotely related to what we are talking about.
          4. okanat · · focus · HN ↗
            LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don&#x27;t need to bring a calculator to count the letters in a sentence. It is a different neural machinery.
            1. fenomas · · focus · HN ↗
              You&#x27;re talking about doing arithmetic; GP was obviously pointing out that &quot;do math&quot; can refer to other things.
            2. amluto · · focus · HN ↗
              LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.