‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. user3939382 · · focus · HN ↗
      Anthropic A&#x2F;Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly&#x2F;5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
      1. apitman · · focus · HN ↗
        Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
        1. ffsm8 · · focus · HN ↗
          Claude code supposedly has otel you can set via env. I haven&#x27;t set it up, so I&#x27;m just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost

          It&#x27;s meant for their test env I think, so is not documented to my knowledge

          1. adastra22 · · focus · HN ↗
            Has otel? What is that?
            1. reubenmorais · · focus · HN ↗
              OpenTelemetry
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.