They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
Regarding lazy mode, I recall ChatGPT sometimes almost refusing to do a web search despite me asking explicitly for it, instead replying with speculation about what the search results likely would tell us. If I pretend to be angry that it didn’t search the web it would however do it. Haven’t noticed this in a while either.
It's like doing the calculations yourself (with whatever tools you want) and knowing what you're doing, steering the process yourself, rather than begging someone to give you the result so you can then show it off like you did the work. I can understand and explain what my calculator is doing and I'm not offloading decisions to it. Using a language model to shit out projects that you can't fully understand the structure of yourself is insane. Ceding control of any design decisions to a language model is insane. They're fine as second-order autocomplete and semantic search engines. Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can't build humane systems.
> Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can't build humane systems.
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
rplnt · · focus · HN ↗
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
transcriptase · · focus · HN ↗
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
setopt · · focus · HN ↗
natpalmer1776 · · focus · HN ↗
lukan · · focus · HN ↗
You have to pretend to be angry in such situations?
katzenq · · focus · HN ↗
AbsurdCensor · · focus · HN ↗
katzenq · · focus · HN ↗
sejje · · focus · HN ↗
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
Humans should use more llms.