‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. sheepscreek · · focus · HN ↗
    > It could also mean nothing happened and people are pattern-matching on noise.

    It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

    The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

    What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.

    But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.

    1. ben_w · · focus · HN ↗
      > It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

      What "stack" do you have in mind here?

      An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.

      Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the &quot;snapshots&quot; section in their recent and old models: e.g. <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;models&#x2F;gpt-4o" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;models&#x2F;gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and <a href="https:&#x2F;&#x2F;platform.claude.com&#x2F;docs&#x2F;en&#x2F;about-claude&#x2F;model-deprecations" rel="nofollow">https:&#x2F;&#x2F;platform.claude.com&#x2F;docs&#x2F;en&#x2F;about-claude&#x2F;model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.

      1. TeMPOraL · · focus · HN ↗
        &gt; An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I&#x27;d be surprised if there&#x27;s a way for their efforts to show up directly in a &quot;has this model been nerfed?&quot; sense.

        It most definitely is not, hasn&#x27;t been for a while now.

        I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API&#x2F;UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.

        Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn&#x27;t touch load balancing, A&#x2F;B testing, shunting token burners (&quot;Hi chat, how are you?&quot;), protecting user from themselves (refusing to answer &quot;bad&quot; queries, stopping &quot;bad&quot; responses), protecting user from third parties (&quot;prompt injection&quot; mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.

        There&#x27;s a lot of things to tune there, and just as many reasons to do it.

        1. ben_w · · focus · HN ↗
          The API documentation linked in my comment seems to say that (with two exceptions*) when we ask for a specific models, we get that specific model.

          The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.

          &gt; Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another.

          Good points.

          &gt; And that doesn&#x27;t touch load balancing, A&#x2F;B testing, shunting token burners (&quot;Hi chat, how are you?&quot;), protecting user from themselves (refusing to answer &quot;bad&quot; queries, stopping &quot;bad&quot; responses), protecting user from third parties (&quot;prompt injection&quot; mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.

          The behaviour I&#x27;m seeing from the companies these days, the A&#x2F;B is what I&#x27;m saying is not showing up like this, they present user A&#x2F;B options openly, and the impression I have is this is to train n+1 models; the other stuff (but I say with low certainty) appears to be done in a more headline-grabbing manner, &quot;model taken offline due to ${news}&quot;? Short update cycles seem to allow that.

          But the prompts you&#x27;re probably right, I wasn&#x27;t giving that enough consideration.

          * the two exceptions being automated safety downgrade for dangerous topics, and &quot;Auto-switch to Thinking&quot; as a used-specified option in ChatGPT

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.