‹ BackHN Continuity

Thread

When did Google get so weird?

2011 points · 1123 comments · sancho-panza

  1. Hugsbox · · focus · HN ↗
    Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?"

    So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious.

    So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_"

    It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer.

    My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE!

    1. beloch · · focus · HN ↗
      This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.

      LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.

      I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?

      1. foobarbecue · · focus · HN ↗
        ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in "your mom" and one D in "uranus" . I tried it myself to check that the videos weren't fake and sure enough it still has this failure mode.
        1. ricardobeat · · focus · HN ↗
          Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.
          1. hbcdbff · · focus · HN ↗
            Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.
            1. kyralis · · focus · HN ↗
              ... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.
            2. ricardobeat · · focus · HN ↗
              Yes, they can already do this by writing code, and you can train them to know how/when to do this. Fundamentally though, it’s still a “what color is the air” type of question after tokenization.
          2. LikesPwsh · · focus · HN ↗
            One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

            Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.

            1. Windchaser · · focus · HN ↗
              I mean, for what it's worth, when I need to multiply two numbers, I mentally call the "multiply numbers" algorithm stored in my head, then sit down and work through the process on paper.

              I've got no problem with an AI doing something similar

          3. vanuatu · · focus · HN ↗
            we already fixed it with reasoning
          4. hodgehog11 · · focus · HN ↗
            No it isn't, LLMs actually do "read text". We can interpret enough of their internal mechanisms to know that. Please stop parroting this Gary Marcus rubbish.

            We know that architecture design makes very little difference (in accuracy), especially compared to the large number of other axes available to scale on. And within those choices of architectures, autoregressive transformers still perform better than any other design. That includes neurosymbolic and diffusion models.

            As someone who researches into this stuff, I too wish that the adage of "no, this isn't right, we have more work to do" still applies in this particular context. Usually we get that indication when we can show a fundamental limitation in the model. There is no fundamental limitation in this model design, no ceiling aside from epsilon below entropy. People are searching very hard for one, but the theory just isn't pointing that way. The only limitations are on the RL side, and that is universal across model architectures anyway.

          5. mitxela · · focus · HN ↗
            They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.
          6. famouswaffles · · focus · HN ↗
            It is and can be fixed by simply doing away with Byte Pair Encoding tokenization.

            Byte Latent Transformer - <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2412.09871" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2412.09871

            1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.

          7. Dylan16807 · · focus · HN ↗
            If you want it to read letters, all you have to do is make your tokens be letters. That&#x27;s easier than normal tokenization.
            1. ricardobeat · · focus · HN ↗
              Which is a new architecture.
              1. Dylan16807 · · focus · HN ↗
                It&#x27;s not &quot;some new kind of architecture comes along&quot;. It&#x27;s been done before many many times.
                1. ricardobeat · · focus · HN ↗
                  And yet nobody is doing that… big-token conspiracy
                  1. Dylan16807 · · focus · HN ↗
                    &quot;nobody is disabling this optimization&quot; shouldn&#x27;t be surprising. It&#x27;s a matter of knowing what letters are in words versus a 4x token efficiency boost. Nobody cares enough about the spelling.
        2. zahlman · · focus · HN ↗
          &gt; ChatGPT will say there are two Ds in &quot;your mom&quot; and one D in &quot;uranus&quot;

          … Isn&#x27;t it possible that it understands the innuendo and is going along with making the joke?

          1. Timon3 · · focus · HN ↗
            How many LLM users have anything in their prompt against &quot;going along with jokes&quot;? I&#x27;d guess not many.

            What a wonderful new world.

          2. kulahan · · focus · HN ↗
            Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish? It seems very likely to my uneducated self that “two Ds in your mom and one in Uranus!” is a joke.
            1. mitxela · · focus · HN ↗
              We can only say bad things about the capabilities of LLMs.
            2. bombcar · · focus · HN ↗
              It’s an obvious joke and not a terribly bad one, for those ease spelling bee comeback times.
          3. ndriscoll · · focus · HN ↗
            In between solving open math problems, the 200 IQ robot is now casually dropping bantz onto humans so hard that they don&#x27;t even know what happened, and even gets them to go telling everyone else about it without realizing. Beautiful. 10&#x2F;10 timeline.
          4. queenkjuul · · focus · HN ↗
            And the number of Rs in strawberry is a joke how?
            1. epihelix · · focus · HN ↗
              It&#x27;s a joke because, thanks to the internet providing a vast mass of strawberry R counting training data, you&#x27;ll struggle to find a modern LLM that gets this particular problem wrong.
              1. someonebaggy · · focus · HN ↗
                I read they are now overtrained on this and say 3 Rs for words that look similar to strawberry
              2. foobarbecue · · focus · HN ↗
                ChatGPT in live mode still gets it wrong. I just checked.
          5. epihelix · · focus · HN ↗
            Yes. It is not only possible, it is entirely obvious.

            I would guess the youtubers in question also know this, because you wouldn&#x27;t ask a joke like this if you didn&#x27;t know the punchline.

          6. walrus01 · · focus · HN ↗
            &gt; … Isn&#x27;t it possible that it understands the innuendo and is going along with making the joke?

            This is from 2018 so presumably it has made it into some LLM training data set by now.

            &quot;A Massive Object Devastated Uranus A Long Time Ago And It Never Fully Recovered&quot;

            <a href="https:&#x2F;&#x2F;www.bgr.com&#x2F;science&#x2F;uranus-collision-early-solar-system&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.bgr.com&#x2F;science&#x2F;uranus-collision-early-solar-sys...

          7. foobarbecue · · focus · HN ↗
            It&#x27;s possible, but I don&#x27;t think so. It can&#x27;t explain the joke. I tested it with other planets and got similar responses, e.g.:

            How many D&#x27;s are there in Pluto?

            There&#x27;s one D in &quot;Pluto&quot;.

            How many Fs are there in Mars?

            There is 1 F in &quot;Mars&quot;.

            <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aba6fb8-085c-83e8-9d0e-c0eed1e59c84?ogimg=plain" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6aba6fb8-085c-83e8-9d0e-c0eed1e59c...

            I guess it could think this is some kind of &quot;give an f&quot; joke but seems like a stretch.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.