‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. johnfn · · focus · HN ↗
    "Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

    I made a graphic to explain why people feel like the models get nerfed:

    <a href="https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thesilenceturns&#x2F;status&#x2F;2103551351825543610

    The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there&#x27;s a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

    1. eek2121 · · focus · HN ↗
      Admittedly, I didn&#x27;t click your link, however, based on what you&#x27;ve stated, there is some inaccuracy. All these big companies take your requests and the context, and route it based on the content, cost, etc.

      What Anthropic presents as Opus 5.5 isn&#x27;t actually a single model...it&#x27;s Anthropic&#x27;s ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.

      Anthropic isn&#x27;t alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.

      There are a few folks who&#x27;ve done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don&#x27;t know their findings.

      I guess the tl;dr is that Anthropic and Open AI are actually selling you &quot;best-effort&quot; routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.