‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. user3939382 · · focus · HN ↗
      Anthropic A&#x2F;Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly&#x2F;5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
      1. apitman · · focus · HN ↗
        Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
        1. Aeolun · · focus · HN ↗
          Tokens used &#x2F; percentage change is a pretty obvious metric. They give you both, but they don’t do the math for you.
          1. apitman · · focus · HN ↗
            It&#x27;s obvious unless you have multiple requests from different models in flight at the same time, and the sum total usage comes out to less than a single percentage.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.