> It could also mean nothing happened and people are pattern-matching on noise.
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
> It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.
> It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
What "stack" do you have in mind here?
An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. <a href="https://developers.openai.com/api/docs/models/gpt-4o" rel="nofollow">https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and <a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="nofollow">https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.
> An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
The API documentation linked in my comment seems to say that (with two exceptions*) when we ask for a specific models, we get that specific model.
The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.
> Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another.
Good points.
> And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
The behaviour I'm seeing from the companies these days, the A/B is what I'm saying is not showing up like this, they present user A/B options openly, and the impression I have is this is to train n+1 models; the other stuff (but I say with low certainty) appears to be done in a more headline-grabbing manner, "model taken offline due to ${news}"? Short update cycles seem to allow that.
But the prompts you're probably right, I wasn't giving that enough consideration.
* the two exceptions being automated safety downgrade for dangerous topics, and "Auto-switch to Thinking" as a used-specified option in ChatGPT
Serving LLMs at Anthropic scale is very, very different. It’s not SGLang or vLLM.
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
That is an extreme misrepresentation of what really happens. Running a relatively small LLM locally (serving a single user) is very different from doing it at scale, notwithstanding for a model 100x in size or more.
Frontier models are huge. Astra could be 10 trillion parameters or more. That will probably need > 10 TB of VRAM (HBM3 if you don’t want to wait forever for a response) and need a mini-cluster just to run one instance/copy.
And because token generation is a sequential process, all the code to orchestrate tensor parallelism (spreading a single request across multiple GPUs) is non-trivial. Each new token depends on the previous ones.
Combine that with KV caching optimizations, loading/unloading from cheaper cache storage, session management, load balancing, parallel sessions, and daily software updates plus testing optimizations to kernels; it’s a lot of moving parts.
We haven’t even discussed coordinating stuff across different data centres, failover mechanism, and what not.
I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".
> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
So, not the sort of thing referred to.
> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.
sheepscreek · · focus · HN ↗
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
zahlman · · focus · HN ↗
I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.
ben_w · · focus · HN ↗
What "stack" do you have in mind here?
An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. <a href="https://developers.openai.com/api/docs/models/gpt-4o" rel="nofollow">https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and <a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="nofollow">https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.
TeMPOraL · · focus · HN ↗
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
ben_w · · focus · HN ↗
The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.
> Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another.
Good points.
> And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
The behaviour I'm seeing from the companies these days, the A/B is what I'm saying is not showing up like this, they present user A/B options openly, and the impression I have is this is to train n+1 models; the other stuff (but I say with low certainty) appears to be done in a more headline-grabbing manner, "model taken offline due to ${news}"? Short update cycles seem to allow that.
But the prompts you're probably right, I wasn't giving that enough consideration.
* the two exceptions being automated safety downgrade for dangerous topics, and "Auto-switch to Thinking" as a used-specified option in ChatGPT
dannyw · · focus · HN ↗
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
jeffybefffy519 · · focus · HN ↗
sheepscreek · · focus · HN ↗
Frontier models are huge. Astra could be 10 trillion parameters or more. That will probably need > 10 TB of VRAM (HBM3 if you don’t want to wait forever for a response) and need a mini-cluster just to run one instance/copy.
And because token generation is a sequential process, all the code to orchestrate tensor parallelism (spreading a single request across multiple GPUs) is non-trivial. Each new token depends on the previous ones.
Combine that with KV caching optimizations, loading/unloading from cheaper cache storage, session management, load balancing, parallel sessions, and daily software updates plus testing optimizations to kernels; it’s a lot of moving parts.
We haven’t even discussed coordinating stuff across different data centres, failover mechanism, and what not.
jibalt · · focus · HN ↗
I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".
> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
So, not the sort of thing referred to.
> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.