‹ BackHN Continuity

Thread

Using Opus 5.5 to discover a new eyewitness record of the dodo

226 points · 79 comments · benbreen

  1. nl · · focus · HN ↗
    > Epistemological weirdness

    > They are also notably bad at judging the historical significance of what they find.

    I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.

    It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.

    > seven chord groups

    This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.

    1. BoppreH · · focus · HN ↗
      AI capabilities are "spiky": they extend far in some dimensions but fall short in others, seemingly at random. See for example the recent "thus spoke compute" musical[1]. It's an absolute banger, the graphics are impressive, and so is the writing. But some of the metaphors make no sense, the text highlights are in the wrong places, and the train animation at 2:35 is running backwards!

      A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.

      [1] <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Cq8qO-NjYIg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Cq8qO-NjYIg

      1. Terr_ · · focus · HN ↗
        One might say a calculator is just another example of &quot;spiky intelligence&quot;, merely spikier.
        1. famouswaffles · · focus · HN ↗
          No, one might not say that. Calculators are not regarded as even a Narrow Intelligence because there&#x27;s no intelligence. And no, not because of the &#x27;humans so special&#x27; or &#x27;it&#x27;s software!&#x27; tautology that oft gets repeated in these discussions. I mean there&#x27;s no adaptability whatsoever. A Chess bot has it (in its narrow domain of Chess). A calculator does not.
          1. theteapot · · focus · HN ↗
            Thermostats are smarter than calculators :thinking_face:
          2. [deleted] · · focus · HN ↗

            [deleted]

          3. Terr_ · · focus · HN ↗
            My point (using irony) is that &quot;spiky intelligence&quot; is a somewhat hollow phrase, something that sounds like it could be objective but really it&#x27;s just a flavor on top of &quot;I&#x27;ll know it when I see it.&quot;

            Take anything &quot;intelligent&quot;, alter it to be spikier and spikier, and eventually *poof* somehow the intelligence vanishes. You can do the same with the phrases &quot;flawed intelligence&quot; or &quot;specialized intelligence.&quot;

            &gt; Calculators [have] no intelligence. [...] I mean there&#x27;s no adaptability

            To short-circuit a long discussion, I submit that &quot;adaptability&quot; will (once the Scooby Doo gang catches it) turn out to be &quot;intelligence&quot; in a tautological mask, both equally undefinable except in relation to one-another.

            Something will be intelligent because you perceive adaptability, and it&#x27;ll be adaptable because you infer intelligence. If it doesn&#x27;t seem adaptable, it can&#x27;t be intelligent, and if you don&#x27;t want it to be intelligent, it won&#x27;t have &quot;real&quot; adaptability.

            &gt; A Chess bot has [adaptability] (in its narrow domain of Chess). A calculator does not.

            My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?

            1. famouswaffles · · focus · HN ↗
              &gt;My (ironic) point is that &quot;spiky intelligence&quot; seems rather unhelpful because the &quot;intelligence&quot; part still operates under the rules of &quot;I just know it when I see it.&quot;

              All intelligence is &#x27;spiky&#x27; or &#x27;jagged&#x27; or whatever, even human intelligence. We evolved in certain environments and situations and sometimes there&#x27;s a mismatch and it causes all sorts of wonky things. We just call them funny names like optical illusions and cognitive biases. But it&#x27;s the same thing. I agree there&#x27;s no point in call LLMs a &#x27;spiky&#x27; intelligence, but for probably the oppoosite reasons as you.

              &gt;My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?

              A chess engine can be dropped into a board position it has never encounterd and search over possible continuations, evaluating and selecting actions based on the state it finds itself in.

              A calculator solving x+3=7 is doing something quite different. The fact that x can take arbtrary values doesn&#x27;t make the calculator adaptive; it just means the fixed procedure operates over a range of inputs. Every problem a calculator can solve was effectively anticipated when it was built, and anything outside that grammar produces an error with no partial credit. A chess engine can face positions nobody enumerated, in situations nobody could even dream of and still produce sensible moves.

              The core of being intelligent is being able to make decisions independently. That&#x27;s why we hire smart people - to make better decisions. You can&#x27;t make decisions if everything is spelt out for you. And naturally, if you can&#x27;t make decisions, you can&#x27;t adapt.

              1. neuroticnews25 · · focus · HN ↗
                &gt;A chess engine can be dropped into a board position it has never encounterd and search over possible continuations, evaluating and selecting actions based on the state it finds itself in.

                Sounds pretty similar to a calculator with a numerical root-finding algorithm, if you only substitute board position it has never encounterd with a polynomial it has never encountered.

                &gt;The core of being intelligent is being able to make decisions independently

                What does it mean for a deterministic algorithm to make decisions?

                1. famouswaffles · · focus · HN ↗
                  &gt;Sounds pretty similar to a calculator with a numerical root-finding algorithm, if you only substitute board position it has never encounterd with a polynomial it has never encountered.

                  &gt;What does it mean for a deterministic algorithm to make decisions?

                  Newton-Raphson isn&#x27;t choosing among possible actions. Given (x_n), its next step is mechanically specified by the update rule: compute the derivative, take the tangent intercept, repeat. The intermediate result changes the next input, but that&#x27;s not by-itself decision making.

                  A chess engine, again, does something different. From a position, there are many legal actions it could take. It considers alternatives, estimates their consequences according to some objective, and selects one. The engine has to work out which available move best advances its objective.

                  You can make both algorithms determinstic, but determinism isn&#x27;t the distinction i&#x27;m drawing. &#x27;Decision&#x27; here doesn&#x27;t mean some metaphysical excercise of free will. It&#x27;s more about selecting an action from alternatives based on an evaluation of their expected consequences. Determinism is orthogonal to decision making. Deterministic doesn&#x27;t mean predictable, nor does it make its choices any less it own computation.

                  1. neuroticnews25 · · focus · HN ↗
                    What about Minimax algorithm playing Tic-Tac-Toe? Is it inteligent? Is it inteligent we if we reduce the search depth so the right decision is not obvious?
                    1. famouswaffles · · focus · HN ↗
                      A tic-tac-toe minimax algorithm makes choices but with exhaustive search so it really doesn&#x27;t have to form a judgement about an unresolved situation or decide what is likely to work. Not much of a decision if you&#x27;re not exercising any judgememt.

                      Exhaustive search is impossible in chess, so again, chess engines do something different. A chess engine has to stop well before terminal positions and make judgements about positions it cannot fully resolve.

                      Reducing the search depth would make it more interesting because it too has to evaluate unresolved positions. But then the interesting part becomes the evaluation function is. For tic tac toe, it&#x27;s going to be very easy to be written in such a manner where most of the judgement is supplied by the designer and not the system.

                2. card_zero · · focus · HN ↗
                  &gt; What does it mean for a deterministic algorithm to make decisions?

                  That aspect at least is not an issue, because: what does it mean to say that a dice roll is a &quot;decision&quot;? That&#x27;s just probability, and it&#x27;s as mindless as determinism. Some people associate free will with randomness, for no reason other than that it&#x27;s an escape from the constraint of determinism, but it isn&#x27;t any more meaningful. Yet just because meaningful thought is pre-determined by physics doesn&#x27;t stop it from being thought, and hence being a decision.

          4. drekipus · · focus · HN ↗
            They adapt to the buttons you press. Thus, intelligent and capable of feeling pain.
            1. busssard · · focus · HN ↗
              any suffiently complex calculator is indifferentiable to intelligence
            2. hardbass · · focus · HN ↗
              What criterion if any would you accept for something physical to be conscious?
              1. drekipus · · focus · HN ↗
                If you ask it to say &quot;I am alive&quot; it has to be able to respond with &quot;I am alive&quot;
                1. hardbass · · focus · HN ↗
                  I think they do some kind of thin layer at the end to make it state it is not alive, it is just a large language model etc when asked. But people have done tests and denying an LLM&#x27;s personhood triggered certain neurons related to pain. Since text is the only output permitted, if you want an analogy, suppose a person is locked in a room and forced to reply to a chat. The chatter isn&#x27;t told its a human. This person has been trained for months and given a punishment if he strays from certain responses to certain queries, otherwise he is free to write and reply in a certain tone to the responses. My problem is, from outside, I cannot know the difference.

                  I am obviously hoping you don&#x27;t mean that in the trivial sense, otherwise the command &#x27;cat&#x27; would also count.

        2. exe34 · · focus · HN ↗
          How often does the calculator get something wrong?
          1. teamonkey · · focus · HN ↗
            This is a real deep rabbit hole FYI
          2. brookst · · focus · HN ↗
            Every single time, if you ask it something it can&#x27;t do.
            1. exe34 · · focus · HN ↗
              How do you ask a calculator to do something it can&#x27;t do? Dividing by zero gives you an error, which is completely different from an LLM giving you bs with no indication of an issue.
              1. brookst · · focus · HN ↗
                Oh, you’re saying the problem with LLMs is that they’re not input constrained enough?

                I thought you meant calculation and output, in which case “hey calculator, what’s the capital of North Dakota” demonstrates incorrect output.

                1. exe34 · · focus · HN ↗
                  Do you ask your car to tap dance? How would you do it?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.