‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. solenoid0937 · · focus · HN ↗
    Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
    1. solfox · · focus · HN ↗
      No, whether or not it's intentional, maybe can be debated. But there's definitely an experience of a model losing horsepower quickly after launch.
      1. nba456_ · · focus · HN ↗
        No there isn't.
        1. [deleted] · · focus · HN ↗

          [deleted]

        2. omani · · focus · HN ↗
          who is paying you to say that?
          1. nba456_ · · focus · HN ↗
            Mr. Dario himself
        3. AnimalMuppet · · focus · HN ↗
          The claim was that there's an experience of a model losing power. Your claim amounts to "No, you are not experiencing what you say". That's quite a claim for you to make with no data and no argument.
          1. solenoid0937 · · focus · HN ↗
            Or "what you're experiencing is irrelevant because perception is fickle"
        4. voiceeh · · focus · HN ↗
          You mean to write YOU haven't experienced this. Many have indeed experienced this.
          1. nba456_ · · focus · HN ↗
            No you didn't.
          2. doginasuit · · focus · HN ↗
            I'm not sure "experienced" is carrying much weight here. This is why there are benchmarks, anecdote doesn't mean very much.
            1. redanddead · · focus · HN ↗
              This is a great omen for the future of AI
          3. Kiro · · focus · HN ↗
            Where is the evidence? HN is no better than the "I've done my own research" on Facebook.
      2. nimchimpsky · · focus · HN ↗

        [dead]

      3. Craighead · · focus · HN ↗
        prove it
        1. bradfa · · focus · HN ↗
          Literally the point of the linked GitHub repo.
      4. jyoung8607 · · focus · HN ↗
        Has there ever been any measurement of this, of any sort? Honest question. I frequently see a plural of anecdotes to that effect, but I've not seen a concrete statement of fact or measurement that could be scrutinized or tested in any way.

        If so, please share. This should be measurable, and I'm glad this project is measuring it.

        Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.

        1. dannyw · · focus · HN ↗
          Official tweet from Tibo from OpenAI confirming they ran experiments tweaking “juice” (reasoning effort mapping values) and have reverted them: <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2076495156757577895" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2076495156757577895

          As he confirmed, for at least a period of time, and for some users, “Sol xhigh” was actually “Sol high”, etc.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.