‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. user3939382 · · focus · HN ↗
      Anthropic A&#x2F;Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly&#x2F;5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
      1. jacquesm · · focus · HN ↗
        How did pissing off your customers ever become a business model?

        I can&#x27;t imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don&#x27;t control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don&#x27;t care because eventually it worked. It&#x27;s almost an ideal vehicle to scam people.

        Imagine the power company being able to decide how much you consume and at which price point.

        1. cavoirom · · focus · HN ↗
          Their fate is coming. Until the open-source models will be usable in machine with 256GB memory, they are done. Their behavior is unacceptable (Anthropic) recently but it won&#x27;t last long.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.