‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. prodigycorp · · focus · HN ↗
      Incorrect.

      Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

      Your chart is wrong.

      1. johnfn · · focus · HN ↗
        Sorry, you are correct - I modified my original post. I get frustrated every time there&#x27;s a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it&#x27;s technically not 100% due to a few edge cases.

        I am more skeptical about the compute provider claim - do you have any evidence of that?

        1. r_lee · · focus · HN ↗
          I noticed that a few weeks back 5.6 sol would regularly glitch out and start speeding random words or loop and then the next day it&#x27;d be fine

          and there&#x27;s sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence

          1. rhdunn · · focus · HN ↗
            I&#x27;ve seen that looping and glitching behaviour in over-quantized (~Q4 or lower) local models like Llama and Qwen 3.x. Thus, it is likely that they quantize the model after release to save on compute costs (while giving a favourable result at launch). That quantization can result in changes to the model&#x27;s behaviour (you are changing the weights) that could be interpreted as nerfing.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.