They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this "conspiracy" theory is perfectly plausible to lots of people
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Rapzid · · focus · HN ↗
Of course it's almost entirely unsubstantiated BS.
fbrncci · · focus · HN ↗
Rapzid · · focus · HN ↗
stackghost · · focus · HN ↗
jackmott42 · · focus · HN ↗
[dead]
stackghost · · focus · HN ↗
> in short, yall dumb, shut up.
no u
Rapzid · · focus · HN ↗
Everyone wants to be a software engineer, until it's time to do software engineering shit.
You know, like scientific method shit we learned in 5th/6th grade.
It's the great bro science incursion.
stackghost · · focus · HN ↗
The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.