They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Razengan · · focus · HN ↗
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
jackmott42 · · focus · HN ↗
fuck
[deleted] · · focus · HN ↗
[deleted]