They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
Regarding lazy mode, I recall ChatGPT sometimes almost refusing to do a web search despite me asking explicitly for it, instead replying with speculation about what the search results likely would tell us. If I pretend to be angry that it didn’t search the web it would however do it. Haven’t noticed this in a while either.
It's like doing the calculations yourself (with whatever tools you want) and knowing what you're doing, steering the process yourself, rather than begging someone to give you the result so you can then show it off like you did the work. I can understand and explain what my calculator is doing and I'm not offloading decisions to it. Using a language model to shit out projects that you can't fully understand the structure of yourself is insane. Ceding control of any design decisions to a language model is insane. They're fine as second-order autocomplete and semantic search engines. Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can't build humane systems.
> Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can't build humane systems.
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
Friend of mine works for a corp that is one of the top spenders on Claude models. He complained about these nerfs during peak demand.
Their Anthropic contact changed something and it did not happen since.
Anecdote - I work in a time zone offset from continental US. The performance of anthropic models would noticeably drop, around the time US work day started. It was so bad around 4.x time that multiple colleagues re-arranged their schedule to have least overlap with US work day. Admittedly it's been better recently.
This is what pissed me off the most. Make it slower, rate limit it, move the credits to other time slots, idk.. but returning BAD results? That's the worst approach you could take.
I do remember times in the past where when I was up super early (4am EST) I would get super high quality results, then in mid afternoon EST it seemed to be degraded.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
rplnt · · focus · HN ↗
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
transcriptase · · focus · HN ↗
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
setopt · · focus · HN ↗
natpalmer1776 · · focus · HN ↗
lukan · · focus · HN ↗
You have to pretend to be angry in such situations?
katzenq · · focus · HN ↗
AbsurdCensor · · focus · HN ↗
katzenq · · focus · HN ↗
sejje · · focus · HN ↗
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
Humans should use more llms.
smurf9852 · · focus · HN ↗
silversmith · · focus · HN ↗
Barbing · · focus · HN ↗
Anthropic cut a deal with SpaceXAI in May - $1.25b/mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.
<a href="https://www.anthropic.com/news/higher-limits-spacex" rel="nofollow">https://www.anthropic.com/news/higher-limits-spacex
rplnt · · focus · HN ↗
This is what pissed me off the most. Make it slower, rate limit it, move the credits to other time slots, idk.. but returning BAD results? That's the worst approach you could take.
jclardy · · focus · HN ↗