‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. user3939382 · · focus · HN ↗
      Anthropic A&#x2F;Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly&#x2F;5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
      1. jacquesm · · focus · HN ↗
        How did pissing off your customers ever become a business model?

        I can&#x27;t imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don&#x27;t control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don&#x27;t care because eventually it worked. It&#x27;s almost an ideal vehicle to scam people.

        Imagine the power company being able to decide how much you consume and at which price point.

        1. csomar · · focus · HN ↗
          I think it&#x27;s sinister, but not for the reasons you&#x27;re thinking. I think they&#x27;re just wildly unprofitable on subscriptions. The idea that most customers won&#x27;t use their full quota is plain wrong: most people are maxing out their subs, or even reselling whatever quota they have left.

          When you&#x27;re running something at a loss, you can mistreat your customers and they&#x27;ll still stick around (I&#x27;m an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.

          My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they&#x27;ll either drop the subscription model or it&#x27;ll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.

          It&#x27;s not clear what API tokens actually cost them, but I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn&#x27;t possible, even if they&#x27;re delivering real business value (coding, research, etc.). In other words, at API prices I&#x27;d just stop using AI, and I suspect most other developers would too.

          1. airspresso · · focus · HN ↗
            &gt; I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars).

            That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won&#x27;t be fast and won&#x27;t be frontier intelligence, but definitely useful.

            1. zozbot234 · · focus · HN ↗
              Any model running on &quot;single digit thousands of dollars&quot; hardware will either be below SOTA (even for local models) or not even close to fast enough for real-time agentic work. Even the latest so-called &quot;flash&quot; models are large enough that doing real work usably with those on a lower-cost platform is at least dicey. You can fire off non-interactive work and do especially simple Q&amp;A&#x2F;chat (which is vastly more token-efficient than anything agentic - though even then latency will be high for anything genuinely SOTA) but that&#x27;s about it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.