They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Often times people think of "nerfs" as my first prompt (which was greenfield - no or little code existed) used 5% of my plan usage. And then 2 weeks later (as the agent is busy reading hundreds of .rs and .ts files it previously generated) the user complains the usage is going down 30% for a single prompt instead of 5%. Attributing this to a "NERF" makes little sense because it's the same model.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
throwitaway222 · · focus · HN ↗