‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. nsarrazin · · focus · HN ↗
      Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.

      It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

      1. waterproof · · focus · HN ↗
        I find that I learn to &quot;trust&quot; a model to get certain things right, as I would trust a colleague. So, as `expectation` increases, my prompting and context management gets sloppier.

        `percieved_performance = actual_perf&#x2F;expectation`

        `expectation` is an increasing function over time.

        `actual_perf` is a stochastic function of the model&#x27;s true ability, context, etc. -&gt; a recipe for some bad sessions.

        As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive &quot;runs&quot; because our brains love to find patterns.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.