‹ BackHN Continuity

Thread

Using Opus 5.5 to discover a new eyewitness record of the dodo

226 points · 79 comments · benbreen

  1. nl · · focus · HN ↗
    > Epistemological weirdness

    > They are also notably bad at judging the historical significance of what they find.

    I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.

    It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.

    > seven chord groups

    This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.

    1. BoppreH · · focus · HN ↗
      AI capabilities are "spiky": they extend far in some dimensions but fall short in others, seemingly at random. See for example the recent "thus spoke compute" musical[1]. It's an absolute banger, the graphics are impressive, and so is the writing. But some of the metaphors make no sense, the text highlights are in the wrong places, and the train animation at 2:35 is running backwards!

      A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.

      [1] <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Cq8qO-NjYIg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Cq8qO-NjYIg

      1. Terr_ · · focus · HN ↗
        One might say a calculator is just another example of &quot;spiky intelligence&quot;, merely spikier.
        1. exe34 · · focus · HN ↗
          How often does the calculator get something wrong?
          1. teamonkey · · focus · HN ↗
            This is a real deep rabbit hole FYI
          2. brookst · · focus · HN ↗
            Every single time, if you ask it something it can&#x27;t do.
            1. exe34 · · focus · HN ↗
              How do you ask a calculator to do something it can&#x27;t do? Dividing by zero gives you an error, which is completely different from an LLM giving you bs with no indication of an issue.
              1. brookst · · focus · HN ↗
                Oh, you’re saying the problem with LLMs is that they’re not input constrained enough?

                I thought you meant calculation and output, in which case “hey calculator, what’s the capital of North Dakota” demonstrates incorrect output.

                1. exe34 · · focus · HN ↗
                  Do you ask your car to tap dance? How would you do it?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.