‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. Grimblewald · · focus · HN ↗
      I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I&#x27;ve been using LLMs heavily even before ada&#x2F;babbage&#x2F;davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
      1. eulgro · · focus · HN ↗
        Your comment makes no sense. How and why would a local model be nerfed anyway...?
        1. r_lee · · focus · HN ↗
          he&#x27;s saying that he notices a difference between local (not nerfable) and hosted ones, so that it&#x27;s not as likely to be just placebo
          1. jacquesm · · focus · HN ↗
            That and &#x27;loss&#x27; may have been intended to be &#x27;launch&#x27;.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.