They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
This reminds me of the fact that true random does not feel random to users due to the clumpiness that the average person does not anticipate existing in true random.
e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"
How to Shuffle Songs? - <a href="https://web.archive.org/web/20220215030739/https://engineering.atspotify.com/2014/02/how-to-shuffle-songs/" rel="nofollow">https://web.archive.org/web/20220215030739/https://engineeri... ( <a href="https://news.ycombinator.com/item?id=38330877">https://news.ycombinator.com/item?id=38330877 78 points, 65 comments)
Took a little bit of digging to find it - I remembered the graphic at the top and found a blog post that copied it and linked to the blog post, but the blog post isn't there anymore... so web archive.
The current version of the blog post is from 2025 - <a href="https://engineering.atspotify.com/2025/11/shuffle-making-random-feel-more-human" rel="nofollow">https://engineering.atspotify.com/2025/11/shuffle-making-ran... (which didn't get any traction on HN)
I've been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc. Just could not find it. Streak continued to another roll, another roll... still couldn't find it. Then the streak just broke. Apparently just random being random.
I find it useful to consider something like shaken rice. If you take a 10x10 grid and shake 100 grains of rice on it, you'll find that some cells contain no rice while others contain as many as 5 or 6. Run the experiment enough and you'll converge on each cell getting one grain per run, but any individual sample will likely be clumpy and the state in which each only has one will occur infrequently.
Also, consider that in flipping 10 coins, you'll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you'll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...
Widening the range from "rolling exactly 15" to "rolls 15 or 16" or "rolls between 14-17" makes the strings even more likely as you're doubling the success rate from "only 9 15s" to the "any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s" space.
To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:
For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:
General Variables:
N = total number of rolls/trials
k = target streak length
p = probability of getting the target outcome (e.g., 1/20 for a specific roll on d20 or 1/10 for two specific results)
q = probability of getting any other outcome (1 - p)
Expected runs of AT LEAST length k:
E(runs >= k) = p^k * (1 + (N - k) * q)
Personally, I find that 'sticky' dice always provide a nice narrative device, at least in narrative games. A character who's player can't seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.
> I've been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc
Been through that this week as well with 100% success on 40% odds over multiple iterations on my game.. I tend to not dig into random but just rather ensure it works 'as closely to intended' as possible..
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
nsarrazin · · focus · HN ↗
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
goodmythical · · focus · HN ↗
e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"
shagie · · focus · HN ↗
How to Shuffle Songs? - <a href="https://web.archive.org/web/20220215030739/https://engineering.atspotify.com/2014/02/how-to-shuffle-songs/" rel="nofollow">https://web.archive.org/web/20220215030739/https://engineeri... ( <a href="https://news.ycombinator.com/item?id=38330877">https://news.ycombinator.com/item?id=38330877 78 points, 65 comments)
Took a little bit of digging to find it - I remembered the graphic at the top and found a blog post that copied it and linked to the blog post, but the blog post isn't there anymore... so web archive.
The current version of the blog post is from 2025 - <a href="https://engineering.atspotify.com/2025/11/shuffle-making-random-feel-more-human" rel="nofollow">https://engineering.atspotify.com/2025/11/shuffle-making-ran... (which didn't get any traction on HN)
svachalek · · focus · HN ↗
goodmythical · · focus · HN ↗
Also, consider that in flipping 10 coins, you'll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you'll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...
Widening the range from "rolling exactly 15" to "rolls 15 or 16" or "rolls between 14-17" makes the strings even more likely as you're doubling the success rate from "only 9 15s" to the "any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s" space.
To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:
For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:
General Variables: N = total number of rolls/trials k = target streak length p = probability of getting the target outcome (e.g., 1/20 for a specific roll on d20 or 1/10 for two specific results) q = probability of getting any other outcome (1 - p)
Expected runs of AT LEAST length k: E(runs >= k) = p^k * (1 + (N - k) * q)
Expected runs of EXACT length k: E(exact k) = p^k * q * (2 + (N - k - 1) * q)
Personally, I find that 'sticky' dice always provide a nice narrative device, at least in narrative games. A character who's player can't seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.
x______________ · · focus · HN ↗
Been through that this week as well with 100% success on 40% odds over multiple iterations on my game.. I tend to not dig into random but just rather ensure it works 'as closely to intended' as possible..