"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.
I made a graphic to explain why people feel like the models get nerfed:
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
I have not been doing increasingly complex things since Opus 4.6 when models got really good.
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
Open AI admits to such here: Open AI aims to have a stable API and admits to meddling with effort levels and such for subscriptions -<a href="https://news.ycombinator.com/item?id=49804316#49809266">https://news.ycombinator.com/item?id=49804316#49809266
It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc.
> It’s not a crazy conspiracy that the same model can be stupider
Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.
> I have not been doing increasingly complex things since Opus 4.6 when models got really good.
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
What sorts of things, if you can say? Is it a similar sized/complexity codebase? Most projects do become larger and/or more complex over time. And most people's standards do creep up as they learn.
Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.
My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.
How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?
There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs, e.g.: <a href="https://www.anthropic.com/engineering/april-23-postmortem" rel="nofollow">https://www.anthropic.com/engineering/april-23-postmortem
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
I mean there is a direct link two comments down from here from 30 minutes before your comment: <a href="https://news.ycombinator.com/item?id=49902477">https://news.ycombinator.com/item?id=49902477
I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
There was this:
<a href="https://www.reddit.com/r/Anthropic/comments/1sl5wfh/the_degradation_of_claude_opus_46_people_are/" rel="nofollow">https://www.reddit.com/r/Anthropic/comments/1sl5wfh/the_degr...
does it really matter wherever its the harness or the model for the vibe coder user claude code?
fwiw, i think almost all regressions are down to a/b testing in the harness by anthropic, but it is objectively indistinguishable beyond "the coding agent ceases to be usable" and i'm back to traditional coding for a few hours until its back to normal again
It absolutely matters because something like Claude Code has no guarantee that there won't be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...
Yes. The agent allows you to switch models, so you could sidestep a bug in one model by temporarily using other models.
The /r/antigravity SubReddit is full of users who very much notice bugs with the tool/agent. We should be thankful that Claude Code is pretty stable by comparison.
Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.
I am more skeptical about the compute provider claim - do you have any evidence of that?
I've seen that looping and glitching behaviour in over-quantized (~Q4 or lower) local models like Llama and Qwen 3.x. Thus, it is likely that they quantize the model after release to save on compute costs (while giving a favourable result at launch). That quantization can result in changes to the model's behaviour (you are changing the weights) that could be interpreted as nerfing.
It's very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it's a just matter of a single GPU runtime layer bug/update to break inference outputs. The model weights don't necessarily change/get quantized.
I suspect they play with their quants and perform weight sensitive tensor/parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.
I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.
Note: I know what quantization is so don't hold back.
I would guess that if they do use such methods, it'd be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there's excess capacity
like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions
During peak hours requests queue and inference slows. During off-peak they can move systems over to training.
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
It's all just conjecture, your hypothesis about moving systems equally so.
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
Trimming parameter size that can be reduced while surviving regression evals. They have so much data they know exactly where to shave the models. Most people will never see it in their work loads. It won’t affect core benches because that is part of the regression evaluation.
That’s right, and it’s been like this ever since we stopped programming in assembly language. Programmers’ brains used to grow manly and strong on a strict diet of manual memory management and custom stack frame handling. Once we transitioned to soft, weak modern languages like C it’s been all downhill.
Admittedly, I didn't click your link, however, based on what you've stated, there is some inaccuracy. All these big companies take your requests and the context, and route it based on the content, cost, etc.
What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.
Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.
There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.
I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.
That's a good observation, though I'd say here that two things could be true at the same time. But, I do personally believe that most of the reported nerfing is the case of your chart + latent evidence-less complaining. Honeymoon phases are real.
You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?
Where are these careful measurements? Are you 100% sure they don't change the harness between runs and have a large enough sample size to be statistically significant?
I'm not sure I'm arguing what you think I'm arguing. I'm not arguing against measurements. I'm saying that it's relevant to recall that honeymoon phases exist in the general human behavior, whether it applies here or not.
if that's the theory people won't keep using 4.6. Personally I've felt the nerf for 4.8, when 5.0 is (near) launching. And my theory of a model being nerfed several days / weeks after launching has to do with the number of users. At launch there won't be too many users so the computing power per user is huge. As time goes, users and agent has been adjusted to newer model, the computing power per person gets reduced
It's probably -only- for subscriptions. The frontier labs seem to use the subscription models on a dial to serve and prioritize the API users better, because they get more profit there.
If you click through to [1] that seems like a clear downwards trend (beyond the usual noise) about two weeks before the release of Opus 4.7, Opus 4.8, and Opus 5.5. Opus 5 is the only launch that looks clean without the previous model being nerfed beforehand
Sure, maybe it isn't the model getting nerved but the harness getting updates that make it better with the new model but substantially worse with the old (at that point still current) model.
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
Yes, but then the model wasn't nerfed, the harness/overall product just had a plain old regression.
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.
Off topic for sure, but why do people insist on using X/Twitter in this day and age?
The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.
I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.
There is an absolutely massive tech community on Twitter and its by far the place to get real time updates on tech news (yes - better than HN). It isn't all a far right cess pit and that is easy to avoid by just using the following tab
Nerfing is certainly real and I don't see how you could argue it isn't.
A/B testing alone would result in a performance nerf for one group.
8-bit quantized models will barely show degradation on benchmarks. The performance is reliably at 99% of the non-quantized model. 4-bit quantization retains somewhere around 95-98% performance on benchmarks. But if you've ever used a 4-bit model, it feels lobotomized.
And just consider what a compny serving these models would do if they were at capacity. Would they stop serving the model altogether? Of course they wouldn't...
Denying that models experience purposeful degradation is gaslighting.
johnfn · · focus · HN ↗
I made a graphic to explain why people feel like the models get nerfed:
<a href="https://x.com/thesilenceturns/status/2103551351825543610" rel="nofollow">https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
486sx33 · · focus · HN ↗
[dead]
hbn · · focus · HN ↗
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
Computer0 · · focus · HN ↗
johnfn · · focus · HN ↗
> It’s not a crazy conspiracy that the same model can be stupider
Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.
frde_me · · focus · HN ↗
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
usef- · · focus · HN ↗
Grimblewald · · focus · HN ↗
My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.
How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?
gobdovan · · focus · HN ↗
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
johnfn · · focus · HN ↗
prodigycorp · · focus · HN ↗
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
simonw · · focus · HN ↗
Where?
prodigycorp · · focus · HN ↗
consumer451 · · focus · HN ↗
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
[0] <a href="https://status.claude.com/incidents/72f99lh1cj2c" rel="nofollow">https://status.claude.com/incidents/72f99lh1cj2c
[1] <a href="https://anthropic.com/engineering/a-postmortem-of-three-recent-issues" rel="nofollow">https://anthropic.com/engineering/a-postmortem-of-three-rece...
[2] <a href="https://texxr.com/handle/claudedevs" rel="nofollow">https://texxr.com/handle/claudedevs
source: <a href="https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2" rel="nofollow">https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2
what · · focus · HN ↗
There are no bugs, just happy little accidents.
consumer451 · · focus · HN ↗
Spooky23 · · focus · HN ↗
That just means they don’t reduce model quality for those reasons.
They didn’t mention other reason, for example, “Make more money”.
p-e-w · · focus · HN ↗
erinnh · · focus · HN ↗
p-e-w · · focus · HN ↗
Christ this forum has become intellectually dishonest.
jibalt · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
jibalt · · focus · HN ↗
> There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs
erinnh · · focus · HN ↗
So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.
prodigycorp · · focus · HN ↗
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
winwang · · focus · HN ↗
computerex · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
QwenGlazer9000 · · focus · HN ↗
Unintentional tbf.
weird-eye-issue · · focus · HN ↗
ffsm8 · · focus · HN ↗
fwiw, i think almost all regressions are down to a/b testing in the harness by anthropic, but it is objectively indistinguishable beyond "the coding agent ceases to be usable" and i'm back to traditional coding for a few hours until its back to normal again
weird-eye-issue · · focus · HN ↗
It absolutely matters because something like Claude Code has no guarantee that there won't be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...
thephyber · · focus · HN ↗
The /r/antigravity SubReddit is full of users who very much notice bugs with the tool/agent. We should be thankful that Claude Code is pretty stable by comparison.
weedfroglozenge · · focus · HN ↗
Rapzid · · focus · HN ↗
swader999 · · focus · HN ↗
johnfn · · focus · HN ↗
I am more skeptical about the compute provider claim - do you have any evidence of that?
r_lee · · focus · HN ↗
and there's sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence
rhdunn · · focus · HN ↗
vikramkr · · focus · HN ↗
HawtAds · · focus · HN ↗
bitexploder · · focus · HN ↗
Rapzid · · focus · HN ↗
Note: I know what quantization is so don't hold back.
r_lee · · focus · HN ↗
like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions
dannyw · · focus · HN ↗
At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
Rapzid · · focus · HN ↗
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
sampullman · · focus · HN ↗
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
Rapzid · · focus · HN ↗
Nerfing conspiracy doesn't need to be proven false. Where is the evidence it's true?
Dylan16807 · · focus · HN ↗
They just want some evidence. It should be pretty easy to measure, shouldn't it?
sampullman · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
ashdksnndck · · focus · HN ↗
bitexploder · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
[dead]
mwigdahl · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
physicallyIllfr · · focus · HN ↗
[dead]
kdkdjcjejxowjdj · · focus · HN ↗
[dead]
physicallyIllfr · · focus · HN ↗
[dead]
cheevly · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
[dead]
[deleted] · · focus · HN ↗
[deleted]
cindyllm · · focus · HN ↗
[dead]
eek2121 · · focus · HN ↗
What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.
Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.
There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.
I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.
winwang · · focus · HN ↗
khalic · · focus · HN ↗
You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?
lxgr · · focus · HN ↗
khalic · · focus · HN ↗
lxgr · · focus · HN ↗
khalic · · focus · HN ↗
lxgr · · focus · HN ↗
winwang · · focus · HN ↗
fendy3002 · · focus · HN ↗
topspin · · focus · HN ↗
That sentence... This conversation is indistinguishable from a billion conversations had around multi-player online gaming.
[deleted] · · focus · HN ↗
[deleted]
itemize123 · · focus · HN ↗
nullbio · · focus · HN ↗
scrollop · · focus · HN ↗
<a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/
This site has been documenting it for a while
cbg0 · · focus · HN ↗
wongarsu · · focus · HN ↗
<a href="https://marginlab.ai/trackers/claude-code-historical-performance/" rel="nofollow">https://marginlab.ai/trackers/claude-code-historical-perform...
lxgr · · focus · HN ↗
> We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).
Changing the harness can have a big impact on performance even when leaving the model completely unchanged.
wongarsu · · focus · HN ↗
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
lxgr · · focus · HN ↗
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.
ricardobeat · · focus · HN ↗
sspiff · · focus · HN ↗
The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.
I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.
stratos123 · · focus · HN ↗
s08148692 · · focus · HN ↗
bigmadshoe · · focus · HN ↗
rybosworld · · focus · HN ↗
A/B testing alone would result in a performance nerf for one group.
8-bit quantized models will barely show degradation on benchmarks. The performance is reliably at 99% of the non-quantized model. 4-bit quantization retains somewhere around 95-98% performance on benchmarks. But if you've ever used a 4-bit model, it feels lobotomized.
And just consider what a compny serving these models would do if they were at capacity. Would they stop serving the model altogether? Of course they wouldn't...
Denying that models experience purposeful degradation is gaslighting.
applicative · · focus · HN ↗