"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.
I made a graphic to explain why people feel like the models get nerfed:
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
johnfn · · focus · HN ↗
I made a graphic to explain why people feel like the models get nerfed:
<a href="https://x.com/thesilenceturns/status/2103551351825543610" rel="nofollow">https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
prodigycorp · · focus · HN ↗
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
simonw · · focus · HN ↗
Where?
prodigycorp · · focus · HN ↗
consumer451 · · focus · HN ↗
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
[0] <a href="https://status.claude.com/incidents/72f99lh1cj2c" rel="nofollow">https://status.claude.com/incidents/72f99lh1cj2c
[1] <a href="https://anthropic.com/engineering/a-postmortem-of-three-recent-issues" rel="nofollow">https://anthropic.com/engineering/a-postmortem-of-three-rece...
[2] <a href="https://texxr.com/handle/claudedevs" rel="nofollow">https://texxr.com/handle/claudedevs
source: <a href="https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2" rel="nofollow">https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2
what · · focus · HN ↗
There are no bugs, just happy little accidents.
consumer451 · · focus · HN ↗
Spooky23 · · focus · HN ↗
That just means they don’t reduce model quality for those reasons.
They didn’t mention other reason, for example, “Make more money”.