‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. Razengan · · focus · HN ↗
      Theory (Conjecture? Hypothesis?): What we notice as &quot;model nerfing&quot; is the company diverting compute to training&#x2F;running new unreleased models..

      Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of &quot;Astra&quot; more than a month before it was officially announced

      1. Centigonal · · focus · HN ↗
        wouldn&#x27;t less compute result in slower inference, rather than worse performance?
        1. poizan42 · · focus · HN ↗
          My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.