‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. HawtAds · · focus · HN ↗
      It&#x27;s very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it&#x27;s a just matter of a single GPU runtime layer bug&#x2F;update to break inference outputs. The model weights don&#x27;t necessarily change&#x2F;get quantized.
      1. bitexploder · · focus · HN ↗
        I suspect they play with their quants and perform weight sensitive tensor&#x2F;parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.
        1. Rapzid · · focus · HN ↗
          I refuse to believe they &quot;play with their quants&quot; once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren&#x27;t just used through claude&#x2F;codex, they are used through API access and it&#x27;s quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.

          Note: I know what quantization is so don&#x27;t hold back.

          1. bitexploder · · focus · HN ↗
            Trimming parameter size that can be reduced while surviving regression evals. They have so much data they know exactly where to shave the models. Most people will never see it in their work loads. It won’t affect core benches because that is part of the regression evaluation.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.