‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. prodigycorp · · focus · HN ↗
      Incorrect.

      Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

      Your chart is wrong.

      1. simonw · · focus · HN ↗
        &gt; Anthropic has admitted to nerfing in the past

        Where?

        1. prodigycorp · · focus · HN ↗
          Man this was last year and some Claude subreddit drama that I can’t furnish offf the top of my head but maybe one of the historians remember it.
          1. consumer451 · · focus · HN ↗
            I asked a historian:

            Two postmortems, neither quite &quot;admitted to nerfing&quot;:

            Sept 2025, infra bugs: &quot;A small percentage of Claude Sonnet 4 requests experienced degraded output quality&quot; [0], alongside &quot;We never reduce model quality due to demand, time of day, or server load.&quot; [1]

            April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, &quot;The models themselves didn&#x27;t regress, and the Claude API was not affected.&quot; [2]

            So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.

            [0] <a href="https:&#x2F;&#x2F;status.claude.com&#x2F;incidents&#x2F;72f99lh1cj2c" rel="nofollow">https:&#x2F;&#x2F;status.claude.com&#x2F;incidents&#x2F;72f99lh1cj2c

            [1] <a href="https:&#x2F;&#x2F;anthropic.com&#x2F;engineering&#x2F;a-postmortem-of-three-recent-issues" rel="nofollow">https:&#x2F;&#x2F;anthropic.com&#x2F;engineering&#x2F;a-postmortem-of-three-rece...

            [2] <a href="https:&#x2F;&#x2F;texxr.com&#x2F;handle&#x2F;claudedevs" rel="nofollow">https:&#x2F;&#x2F;texxr.com&#x2F;handle&#x2F;claudedevs

            source: <a href="https:&#x2F;&#x2F;claude.ai&#x2F;share&#x2F;4435bbcf-d6df-44a0-b1db-f08a11858bc2" rel="nofollow">https:&#x2F;&#x2F;claude.ai&#x2F;share&#x2F;4435bbcf-d6df-44a0-b1db-f08a11858bc2

            1. what · · focus · HN ↗
              &gt; bugs

              There are no bugs, just happy little accidents.

              1. consumer451 · · focus · HN ↗
                u&#x2F;bcherny does sound a bit like Bob Ross now that you mention it.
            2. Spooky23 · · focus · HN ↗
              &gt; &quot;We never reduce model quality due to demand, time of day, or server load.&quot;

              That just means they don’t reduce model quality for those reasons.

              They didn’t mention other reason, for example, “Make more money”.

          2. p-e-w · · focus · HN ↗
            Noone can seem to remember anything with certainty when asked to actually substantiate these claims.
            1. erinnh · · focus · HN ↗
              I mean there is a direct link two comments down from here from 30 minutes before your comment: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49902477">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49902477
              1. p-e-w · · focus · HN ↗
                That’s NOT Anthropic admitting to “nerfing” their model as claimed above (which implies intent), that’s a regression which they quickly fixed.

                Christ this forum has become intellectually dishonest.

                1. jibalt · · focus · HN ↗
                  There has always been plenty of intellectual dishonesty here, but remember Hanlon&#x27;s Razor.
                2. [deleted] · · focus · HN ↗

                  [deleted]

              2. jibalt · · focus · HN ↗
                Er, perhaps take a look after the edit:

                &gt; There are recorded cases of real regressions, but they&#x27;re better characterised as incidents, not nerfs

                1. erinnh · · focus · HN ↗
                  I dont really characterize a nerf as always by intent.

                  So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.

            2. prodigycorp · · focus · HN ↗
              I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.

              It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.

              1. winwang · · focus · HN ↗
                You can just have your agent find the evidence, review it, copypaste it.
            3. computerex · · focus · HN ↗
              There was this: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;Anthropic&#x2F;comments&#x2F;1sl5wfh&#x2F;the_degradation_of_claude_opus_46_people_are&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;Anthropic&#x2F;comments&#x2F;1sl5wfh&#x2F;the_degr...
          3. [deleted] · · focus · HN ↗

            [deleted]

        2. QwenGlazer9000 · · focus · HN ↗
          Earlier in march&#x2F;aprile, there was a regression in Claude code.

          Unintentional tbf.

          1. weird-eye-issue · · focus · HN ↗
            Unrelated to the models
            1. ffsm8 · · focus · HN ↗
              does it really matter wherever its the harness or the model for the vibe coder user claude code?

              fwiw, i think almost all regressions are down to a&#x2F;b testing in the harness by anthropic, but it is objectively indistinguishable beyond &quot;the coding agent ceases to be usable&quot; and i&#x27;m back to traditional coding for a few hours until its back to normal again

              1. weird-eye-issue · · focus · HN ↗
                ...

                It absolutely matters because something like Claude Code has no guarantee that there won&#x27;t be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...

              2. thephyber · · focus · HN ↗
                Yes. The agent allows you to switch models, so you could sidestep a bug in one model by temporarily using other models.

                The &#x2F;r&#x2F;antigravity SubReddit is full of users who very much notice bugs with the tool&#x2F;agent. We should be thankful that Claude Code is pretty stable by comparison.

        3. weedfroglozenge · · focus · HN ↗
          Honestly you need to start making pelicans on day of release, and following up a month later. Only way we are going to catch them
          1. Rapzid · · focus · HN ↗
            This is the way. Pelicans in a coal mine.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.