‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. HawtAds · · focus · HN ↗
      It&#x27;s very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it&#x27;s a just matter of a single GPU runtime layer bug&#x2F;update to break inference outputs. The model weights don&#x27;t necessarily change&#x2F;get quantized.
      1. bitexploder · · focus · HN ↗
        I suspect they play with their quants and perform weight sensitive tensor&#x2F;parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.
        1. Rapzid · · focus · HN ↗
          I refuse to believe they &quot;play with their quants&quot; once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren&#x27;t just used through claude&#x2F;codex, they are used through API access and it&#x27;s quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.

          Note: I know what quantization is so don&#x27;t hold back.

          1. r_lee · · focus · HN ↗
            I would guess that if they do use such methods, it&#x27;d be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there&#x27;s excess capacity

            like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn&#x27;t be surprised if they&#x27;d rather try to make inference faster that way than try to just limit people, at least for those on subscriptions

          2. dannyw · · focus · HN ↗
            Inference isn’t flat 24x7, peak hours have more usage, but you buy&#x2F;rent servers; not servers only for peak hours.

            At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.

            API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.

            1. Rapzid · · focus · HN ↗
              During peak hours requests queue and inference slows. During off-peak they can move systems over to training.

              Where is the evidence they are &quot;nerfing&quot; the models due to request volume?

              Edit: I don&#x27;t know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.

              1. sampullman · · focus · HN ↗
                It&#x27;s all just conjecture, your hypothesis about moving systems equally so.

                But you seem adamant that there&#x27;s no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.

                It&#x27;s tricky to prove either way, but the chance is not zero.

                1. Rapzid · · focus · HN ↗
                  Global inference routing isn&#x27;t conjecture.

                  Nerfing conspiracy doesn&#x27;t need to be proven false. Where is the evidence it&#x27;s true?

                2. Dylan16807 · · focus · HN ↗
                  &gt; you seem adamant that there&#x27;s no chance

                  They just want some evidence. It should be pretty easy to measure, shouldn&#x27;t it?

                  1. sampullman · · focus · HN ↗
                    It should be possible, but maybe not trivial with A&#x2F;B testing&#x2F;etc.
            2. [deleted] · · focus · HN ↗

              [deleted]

            3. ashdksnndck · · focus · HN ↗
              Couldn’t the labs time-shift training and other batch workloads to make up for regular changes in inference demand?
          3. bitexploder · · focus · HN ↗
            Trimming parameter size that can be reduced while surviving regression evals. They have so much data they know exactly where to shave the models. Most people will never see it in their work loads. It won’t affect core benches because that is part of the regression evaluation.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.