‹ BackHN Continuity

Thread

When did Google get so weird?

2011 points · 1123 comments · sancho-panza

  1. Hugsbox · · focus · HN ↗
    Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"

    So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.

    So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"

    It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.

    My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!

    1. beloch · · focus · HN ↗
      This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.

      LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.

      I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?

      1. VCFundedGenYer · · focus · HN ↗
        LLMs still can't do math nor count letters in words. Nothing has changed there.
        1. Kim_Bruning · · focus · HN ↗
          I could have sworn this had been fixed a while ago..

          Just to check if I was actually crazy, I actually went and put a simple addition (7 digits + 7 digits) , and a simple letter counting question to Claude haiku(4.5) , sonnet(5), opus(5.5) and fable(5.1) . They all did just fine straight up.

          If you don't mind spending the tokens, some older/other models can also arrive at the correct answer if you ask them to do the math in long form, since that fits nicely inside autoregression.

          Not sure since when exactly, but letter-counting hasn't been a problem for a while now either. This used to be a problem due to the tokenizers used. Slightly older models can be asked to split the word out into letters, and then they can use autoregression to solve.

          1. leoedin · · focus · HN ↗
            Are the models "doing the calculation" or are they calling a calculator tool? There's a lot of talk about how models can do maths now, but I'm struggling to understand if that just means they just need to recognise that it's a maths problem and pass it to a tool, or if they're truly doing the numerical manipulation themselves.
            1. SkyBelow · · focus · HN ↗
              They are getting better at actually doing the math, but can still fall back to tools if available.

              That said, we can only be sure with open models. In theory, a model like Fable could have access to tools we can't see and only a promise they don't. But load up something like deepseek, put it in a harness with only text in/text out, and you can see exactly how it works.

              As for if it counts as doing math, this gets into the messy question of if a given human is doing math or not. Math itself is some level of memorization and some level of applying known facts. You have to remember 1 means one and that 1 + 1 is 2. But you don't need to remember that 123 + 321 = 444. You remember 1 digit addition and remember you can apply this to 10s place and 100s place, and then you apply these different facts and do math. But you might as simply memorize some things, like 11 + 11 = 22. This is related to the memory of 1+1=2, but you aren't really using that memory either. Almost like an engram of 1+1=2 forms that you can then loop a few times before you need more conscious thought. What about 111111111111+11111111111? Well, your brain might do a heuristic and just do all 2s, but that isn't the right way to answer that question.

              Given all this, people complain about LLMs memorizing math answers and not doing math, but memorizing the math answers is part of doing math. It seems to have basic facts pretty well memorized, and with reasoning it is far better at applying them. But this is messy human math, not clean calculator math which always produces the correct answer (sans some bug in the code). Much like how a human with decent math skills can make a mistake and even multiple if you distract them, an LLM can apply the wrong memory, apply a fake memory, or just not apply something it should. The messier the context, the more likely this is to happen.

              So, is an LLM doing this?

              P.S.

              For an interesting test in how much math involves memory, try doing math in a base you aren't familiar with characters you aren't familiar. The simplest option is almost always mapping back to the ones you memorized, even if you are applying simple operations that you deeply know. Even if you routinely work with hex, can you do the same rough estimation of something like ca / b.3 that you can do with 122 / 11.2 to see if your final answer is in the correct ballpark without first converting to decimal?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.