They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it's released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
If you're going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you're not debating in good faith. Let alone when we're speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we're speaking of companies who are all racing to IPO where short term results matter more than just about anything.
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
> an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Rapzid · · focus · HN ↗
Of course it's almost entirely unsubstantiated BS.
fbrncci · · focus · HN ↗
Rapzid · · focus · HN ↗
somenameforme · · focus · HN ↗
Rapzid · · focus · HN ↗
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
somenameforme · · focus · HN ↗
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
Rapzid · · focus · HN ↗
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.
So where is the evidence?!
airstrike · · focus · HN ↗
Not too mention these companies could easily offer one product to enteprises and another to everyone else
Model nerfing is real
Rapzid · · focus · HN ↗
I was so certain it wasn't, based on the complete lack of evidence.
But then you said it's real. NVM, I don't need evidence! Somebody said it's real!
This place has fallen off.
grim_io · · focus · HN ↗