They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
It's at least highly implausible. Why would they engage in pretty uncontroversially illegal deception/fraud if they have so many other legal ways of gaming benchmarks, selling more tokens etc. available to them?
It's like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.
They wouldn't be breaking any laws at all. It'd be simply tuning their output as they see fit. And that's if you could actually prove it was even intentional, when in reality they have endless plausible deniability of 'oh it was just a technical glitch we've since corrected.'
Banks have far greater transparency and legal requirements. LLM companies are just delivering a black box that they have complete control over. And given the regular 'How's Claude doing this session?' stuff, it's almost certain that they're A-B testing various tweaks on a per session basis.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Rapzid · · focus · HN ↗
Of course it's almost entirely unsubstantiated BS.
fbrncci · · focus · HN ↗
Rapzid · · focus · HN ↗
somenameforme · · focus · HN ↗
lxgr · · focus · HN ↗
It's like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.
somenameforme · · focus · HN ↗
Banks have far greater transparency and legal requirements. LLM companies are just delivering a black box that they have complete control over. And given the regular 'How's Claude doing this session?' stuff, it's almost certain that they're A-B testing various tweaks on a per session basis.