‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. sheepscreek · · focus · HN ↗
    > It could also mean nothing happened and people are pattern-matching on noise.

    It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

    The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

    What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.

    But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.

    1. ben_w · · focus · HN ↗
      > It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

      What "stack" do you have in mind here?

      An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.

      Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the &quot;snapshots&quot; section in their recent and old models: e.g. <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;models&#x2F;gpt-4o" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;models&#x2F;gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and <a href="https:&#x2F;&#x2F;platform.claude.com&#x2F;docs&#x2F;en&#x2F;about-claude&#x2F;model-deprecations" rel="nofollow">https:&#x2F;&#x2F;platform.claude.com&#x2F;docs&#x2F;en&#x2F;about-claude&#x2F;model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.

      1. sheepscreek · · focus · HN ↗
        That is an extreme misrepresentation of what really happens. Running a relatively small LLM locally (serving a single user) is very different from doing it at scale, notwithstanding for a model 100x in size or more.

        Frontier models are huge. Astra could be 10 trillion parameters or more. That will probably need &gt; 10 TB of VRAM (HBM3 if you don’t want to wait forever for a response) and need a mini-cluster just to run one instance&#x2F;copy.

        And because token generation is a sequential process, all the code to orchestrate tensor parallelism (spreading a single request across multiple GPUs) is non-trivial. Each new token depends on the previous ones.

        Combine that with KV caching optimizations, loading&#x2F;unloading from cheaper cache storage, session management, load balancing, parallel sessions, and daily software updates plus testing optimizations to kernels; it’s a lot of moving parts.

        We haven’t even discussed coordinating stuff across different data centres, failover mechanism, and what not.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.