‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. Rapzid · · focus · HN ↗
      The vibe bro science is this always happens on every release, every Tuesday, and twice on Sunday.

      Of course it&#x27;s almost entirely unsubstantiated BS.

      1. fbrncci · · focus · HN ↗
        Well now it’s being substantiated!
        1. Rapzid · · focus · HN ↗
          Or rather it&#x27;s being.. Unsubstantiated. The Nerf conspiracy isn&#x27;t that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI&#x2F;Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
          1. somenameforme · · focus · HN ↗
            Yeah it&#x27;s just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
            1. Rapzid · · focus · HN ↗
              Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it&#x27;s released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.

              So this is a case of extraordinary claims requiring extraordinary evidence.

              And even though it&#x27;s super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.

              1. somenameforme · · focus · HN ↗
                If you&#x27;re going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you&#x27;re not debating in good faith. Let alone when we&#x27;re speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we&#x27;re speaking of companies who are all racing to IPO where short term results matter more than just about anything.

                Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.

                1. Rapzid · · focus · HN ↗
                  &gt; an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account

                  I&#x27;m sorry, but yeah. The official account is a harness regression and some platform bugs.

                  Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then &quot;nerfing&quot; them to save money and shed load. Where is the evidence?!

                  I&#x27;m not saying it&#x27;s illegal, per say, so don&#x27;t come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don&#x27;t even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to &quot;nerfing&quot; shenanigans.

                  So where is the evidence?!

                  1. airstrike · · focus · HN ↗
                    You&#x27;re giving way too much credit to &quot;enterprises&quot; both noticing and publicly airing out their dissatisfaction

                    Not too mention these companies could easily offer one product to enteprises and another to everyone else

                    Model nerfing is real

                    1. Rapzid · · focus · HN ↗
                      Uh huh. OMG you&#x27;re so right, it&#x27;s soooo real ;) ;) ;)

                      I was so certain it wasn&#x27;t, based on the complete lack of evidence.

                      But then you said it&#x27;s real. NVM, I don&#x27;t need evidence! Somebody said it&#x27;s real!

                      This place has fallen off.

                      1. grim_io · · focus · HN ↗
                        Reddit cross contamination. Over there it&#x27;s a rite of passage to accept it as a fact.
                      2. airstrike · · focus · HN ↗
                        This place has fallen off so long ago that people like you consider yourself old timers but don&#x27;t even follow guidelines

                        Bad evidence is worse than no evidence.

                        And you failed to address the specific criticism I made to your point, instead going for an ad hominem &#x2F; poisoning the well.

                        In sum, you&#x27;re out of line, woefully misled, and lacking in logical thinking. Three strikes, you&#x27;re out.

              2. mrandish · · focus · HN ↗
                &gt; They sell HEAPS through their API access

                The claim is that they nerf subscription accounts not API.

                1. applicative · · focus · HN ↗
                  My impression was that, at least with Anthropic, the point of my subsidized subscription is that I&#x27;ll convince my employer to get an API account. If anything, the motive would instead be to ensh*ttify on the corporations already committed. Those of us inducted as boosters would get the fluffed up product.
              3. troupo · · focus · HN ↗
                &gt; And even though it&#x27;s super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.

                Until shit like this: <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;april-23-postmortem" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;april-23-postmortem

                Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks &quot;fixing it&quot;.

                Or shit like this: <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;a-postmortem-of-three-recent-issues" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;a-postmortem-of-three-...

                Anthropic is in a perpetual state of &quot;oops, these &#x27;bugs&#x27; degraded our model quality&quot; and only admit the issues when it&#x27;s immediately obvious and visibly affects a large number of customers.

                Otherwise all open benchmarks can be (and are) gamed. And it&#x27;s quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.