‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. solenoid0937 · · focus · HN ↗
    Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
    1. solfox · · focus · HN ↗
      No, whether or not it's intentional, maybe can be debated. But there's definitely an experience of a model losing horsepower quickly after launch.
      1. nba456_ · · focus · HN ↗
        No there isn't.
        1. [deleted] · · focus · HN ↗

          [deleted]

        2. omani · · focus · HN ↗
          who is paying you to say that?
          1. nba456_ · · focus · HN ↗
            Mr. Dario himself
        3. AnimalMuppet · · focus · HN ↗
          The claim was that there's an experience of a model losing power. Your claim amounts to "No, you are not experiencing what you say". That's quite a claim for you to make with no data and no argument.
          1. solenoid0937 · · focus · HN ↗
            Or "what you're experiencing is irrelevant because perception is fickle"
        4. voiceeh · · focus · HN ↗
          You mean to write YOU haven't experienced this. Many have indeed experienced this.
          1. nba456_ · · focus · HN ↗
            No you didn't.
          2. doginasuit · · focus · HN ↗
            I'm not sure "experienced" is carrying much weight here. This is why there are benchmarks, anecdote doesn't mean very much.
            1. redanddead · · focus · HN ↗
              This is a great omen for the future of AI
          3. Kiro · · focus · HN ↗
            Where is the evidence? HN is no better than the "I've done my own research" on Facebook.
      2. nimchimpsky · · focus · HN ↗

        [dead]

      3. Craighead · · focus · HN ↗
        prove it
        1. bradfa · · focus · HN ↗
          Literally the point of the linked GitHub repo.
      4. jyoung8607 · · focus · HN ↗
        Has there ever been any measurement of this, of any sort? Honest question. I frequently see a plural of anecdotes to that effect, but I've not seen a concrete statement of fact or measurement that could be scrutinized or tested in any way.

        If so, please share. This should be measurable, and I'm glad this project is measuring it.

        Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.

        1. dannyw · · focus · HN ↗
          Official tweet from Tibo from OpenAI confirming they ran experiments tweaking “juice” (reasoning effort mapping values) and have reverted them: <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2076495156757577895" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2076495156757577895

          As he confirmed, for at least a period of time, and for some users, “Sol xhigh” was actually “Sol high”, etc.

    2. dude250711 · · focus · HN ↗
      Suuuuure...
    3. empath75 · · focus · HN ↗
      Yeah people push the models to the limits of what they are capable of almost instantly.
    4. wccrawford · · focus · HN ↗
      I was just wondering if, like certain processors, bugs get fixed and the speed goes down. Like, they find it&#x27;s doing things it shouldn&#x27;t, restrict it, and harm the throughput.
    5. raincole · · focus · HN ↗
      It&#x27;s not a hot take at all. Every benchmark shows that.
    6. jascha_eng · · focus · HN ↗
      Yeh it&#x27;s absurd that people claim this all the time. It&#x27;s some crazy conspiracy theory and when you ask for examples nothing ever shows up.

      It would be economical suicide from anthropic and OpenAI to actually need models intentionally.

      But hey I guess it&#x27;s hard with technology that truly seems like magic. People say if you&#x27;d bring electricity to the middle ages you&#x27;d be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn&#x27;t ready for yet.

    7. sumedh · · focus · HN ↗
      Anthropic has admitted in the past about bugs in the harness after users complained.

      Links have been provided by others in this post.

      1. Kiro · · focus · HN ↗
        Dishonest representation. The links provided do not show this at all, but very clearly that it has been isolated incidents that people extrapolate into false evidence.
        1. sumedh · · focus · HN ↗
          &gt; but very clearly that it has been isolated incidents

          Ant denied them at first though.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.