They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
It's cheaper to run a quantization of a model, but its quality is reduced.
For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.
The game is to balance between the savings of quantization and making the model dumb enough the people notice
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
zerop · · focus · HN ↗
StableAlkyne · · focus · HN ↗
For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.
The game is to balance between the savings of quantization and making the model dumb enough the people notice
sscaryterry · · focus · HN ↗
ForHackernews · · focus · HN ↗
<a href="https://www.fool.com/investing/2026/07/25/spacexs-performance-looks-almost-identical-to-past/" rel="nofollow">https://www.fool.com/investing/2026/07/25/spacexs-performanc...