‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. scrollop · · focus · HN ↗
      It&#x27;s a real thing

      <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;

      This site has been documenting it for a while

      1. cbg0 · · focus · HN ↗
        You&#x27;ve posted a link that doesn&#x27;t support your statement.
        1. wongarsu · · focus · HN ↗
          If you click through to [1] that seems like a clear downwards trend (beyond the usual noise) about two weeks before the release of Opus 4.7, Opus 4.8, and Opus 5.5. Opus 5 is the only launch that looks clean without the previous model being nerfed beforehand

          <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code-historical-performance&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code-historical-perform...

          1. lxgr · · focus · HN ↗
            From that site:

            &gt; We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).

            Changing the harness can have a big impact on performance even when leaving the model completely unchanged.

            1. wongarsu · · focus · HN ↗
              Sure, maybe it isn&#x27;t the model getting nerved but the harness getting updates that make it better with the new model but substantially worse with the old (at that point still current) model.

              The test doesn&#x27;t differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release

              1. lxgr · · focus · HN ↗
                Yes, but then the model wasn&#x27;t nerfed, the harness&#x2F;overall product just had a plain old regression.

                This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.