They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
These guys are on twitter angry about the rate limit decrease and allegedly cancelled all their OpenAI accounts. Wonder how they'll maintain this.
I am quite convinced that the whole nerfing phenomenon is 90% AI psychosis. I have the word muted on X.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
stbenjam · · focus · HN ↗
I am quite convinced that the whole nerfing phenomenon is 90% AI psychosis. I have the word muted on X.
cyberes · · focus · HN ↗
Barbing · · focus · HN ↗
stbenjam · · focus · HN ↗