‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. hbn · · focus · HN ↗
      I have not been doing increasingly complex things since Opus 4.6 when models got really good.

      My work at my job has stayed the same. But the model quality has varied.

      They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.

      1. johnfn · · focus · HN ↗
        It&#x27;s not about doing more complex things - complexity is more dictated by how large your codebase is, etc.

        &gt; It’s not a crazy conspiracy that the same model can be stupider

        Sorry, I really do think it&#x27;s a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A &quot;nerfing&quot; hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.