"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.
I made a graphic to explain why people feel like the models get nerfed:
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
I mean there is a direct link two comments down from here from 30 minutes before your comment: <a href="https://news.ycombinator.com/item?id=49902477">https://news.ycombinator.com/item?id=49902477
I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
There was this:
<a href="https://www.reddit.com/r/Anthropic/comments/1sl5wfh/the_degradation_of_claude_opus_46_people_are/" rel="nofollow">https://www.reddit.com/r/Anthropic/comments/1sl5wfh/the_degr...
johnfn · · focus · HN ↗
I made a graphic to explain why people feel like the models get nerfed:
<a href="https://x.com/thesilenceturns/status/2103551351825543610" rel="nofollow">https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
prodigycorp · · focus · HN ↗
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
simonw · · focus · HN ↗
Where?
prodigycorp · · focus · HN ↗
p-e-w · · focus · HN ↗
erinnh · · focus · HN ↗
p-e-w · · focus · HN ↗
Christ this forum has become intellectually dishonest.
jibalt · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
jibalt · · focus · HN ↗
> There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs
erinnh · · focus · HN ↗
So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.
prodigycorp · · focus · HN ↗
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
winwang · · focus · HN ↗
computerex · · focus · HN ↗