‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. Razengan · · focus · HN ↗
      Theory (Conjecture? Hypothesis?): What we notice as &quot;model nerfing&quot; is the company diverting compute to training&#x2F;running new unreleased models..

      Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of &quot;Astra&quot; more than a month before it was officially announced

      1. Centigonal · · focus · HN ↗
        wouldn&#x27;t less compute result in slower inference, rather than worse performance?
        1. btown · · focus · HN ↗
          The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.
          1. nightpool · · focus · HN ↗
            why is that more likely?
            1. zxilly · · focus · HN ↗
              Because they already did so. The model in Codex will get lower `juice` than API version.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.