They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
The tracker for Codex resonates with me. Its thick as pig shit the last few days: <a href="https://marginlab.ai/trackers/codex/" rel="nofollow">https://marginlab.ai/trackers/codex/
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
scrollop · · focus · HN ↗
<a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/
sscaryterry · · focus · HN ↗
phoghed · · focus · HN ↗
lxgr · · focus · HN ↗
This alone makes the benchmark unsound.
shawabawa3 · · focus · HN ↗