‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. user3939382 · · focus · HN ↗
      Anthropic A&#x2F;Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly&#x2F;5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
      1. braingravy · · focus · HN ↗
        Pretty amazing to see enshitification happen live with a product still in development… Truly web 4.0
        1. LimitExperience · · focus · HN ↗

          [dead]

      2. apitman · · focus · HN ↗
        Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
        1. ffsm8 · · focus · HN ↗
          Claude code supposedly has otel you can set via env. I haven&#x27;t set it up, so I&#x27;m just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost

          It&#x27;s meant for their test env I think, so is not documented to my knowledge

          1. adastra22 · · focus · HN ↗
            Has otel? What is that?
            1. reubenmorais · · focus · HN ↗
              OpenTelemetry
          2. TeMPOraL · · focus · HN ↗
            It&#x27;s for corporate users who want to track how the product is used internally, and it was documented at least some time ago, quite extensively even.
            1. ffsm8 · · focus · HN ↗
              youre right!

              <a href="https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;monitoring-usage#usage-monitoring" rel="nofollow">https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;monitoring-usage#usage-monit...

              thanks for correcting me on that regard

        2. Aeolun · · focus · HN ↗
          Tokens used &#x2F; percentage change is a pretty obvious metric. They give you both, but they don’t do the math for you.
          1. apitman · · focus · HN ↗
            It&#x27;s obvious unless you have multiple requests from different models in flight at the same time, and the sum total usage comes out to less than a single percentage.
      3. jacquesm · · focus · HN ↗
        How did pissing off your customers ever become a business model?

        I can&#x27;t imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don&#x27;t control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don&#x27;t care because eventually it worked. It&#x27;s almost an ideal vehicle to scam people.

        Imagine the power company being able to decide how much you consume and at which price point.

        1. herval · · focus · HN ↗
          &gt; How did pissing off your customers ever become a business model?

          Airlines, banks, health insurance…

          1. tccole · · focus · HN ↗
            So very low margin businesses with hogh amounts of regulations.
            1. petesergeant · · focus · HN ↗
              Banks and health insurance are much more consumer friendly outside of the US, usually because of regulation. Turns out you can just tell banks “make transfers cheap and essentially instant” and they’ll do it, rather the bullshit they have in the US.
              1. herval · · focus · HN ↗
                you&#x27;d be surprised. I&#x27;ve yet to live in a place where either is consumer-friendly...
            2. herval · · focus · HN ↗
              makes you wonder why openai&#x2F;chatgpt&#x2F;xai are in bed with government so much...
        2. pixelready · · focus · HN ↗
          Step 1: Oligopoly Step 2: Regulatory Capture Step 3: Profit
          1. miohtama · · focus · HN ↗
            Only if we did not have these cheap illegal Chinese models
            1. msdz · · focus · HN ↗
              That is why

              &gt; Step 2: Regulatory Capture

              is being worked towards.

            2. TeMPOraL · · focus · HN ↗
              &gt; illegal

              Are they though? Or is it just what some companies would want them to be?

        3. none_to_remain · · focus · HN ↗
          I find it amazingly rich that they bill you for &quot;&quot;&quot;thinking&quot;&quot;&quot; tokens and now you don&#x27;t even get to see them, they&#x27;re gonna train the thing to sing &quot;99 Bottles of Beer on the Wall&quot; to itself before it starts work.
          1. Turskarama · · focus · HN ↗
            They don&#x27;t want to waste tokens on purpose, what they&#x27;re actually hiding is when the model wastes tokens on obviously stupid &quot;thoughts&quot;.
            1. adastra22 · · focus · HN ↗
              No they are hiding the chain of thought to make distillation harder.
              1. slim · · focus · HN ↗
                It&#x27;s fascinating that you all think accounting is real and it did not come to your mind that they could make up numbers when billing
                1. TeMPOraL · · focus · HN ↗
                  They could, but as you see here, people are very eager to create dashboards and trackers that do external accounting by proxy, so they can&#x27;t just &quot;make up numbers&quot; without the customers noticing and making a fuss.
        4. topspin · · focus · HN ↗
          It&#x27;s worked for online PvP gaming for a long time. Nerf stuff the min-maxers &quot;earned&quot; through game mechanics and sell over-powered &quot;premium&quot; things to everyone else to pwn them. Then nerf the old premium stuff and make new premium stuff. Forever.

          I don&#x27;t know if that&#x27;s the actual origin of the term nerf, but it was the first time I&#x27;d heard it.

          1. done_lurking · · focus · HN ↗
            I think the origin of the word &quot;nerf&quot; as a verb came from the Nerf brand of toy guns. The idea being that &quot;Nerfing&quot; something is to turn it into a harmless version of itself.
        5. CodesInChaos · · focus · HN ↗
          Another way Antropic misleads its customers is the description of the max plans. They are advertised as having 5x&#x2F;20x the 5h quota as Pro. But the description says nothing about how the weekly quota scales, leaving customers to infer it scales the same way. But from what I&#x27;ve heard, the weekly quota is only 3.5x&#x2F;7x that of Pro.
        6. csomar · · focus · HN ↗
          I think it&#x27;s sinister, but not for the reasons you&#x27;re thinking. I think they&#x27;re just wildly unprofitable on subscriptions. The idea that most customers won&#x27;t use their full quota is plain wrong: most people are maxing out their subs, or even reselling whatever quota they have left.

          When you&#x27;re running something at a loss, you can mistreat your customers and they&#x27;ll still stick around (I&#x27;m an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.

          My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they&#x27;ll either drop the subscription model or it&#x27;ll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.

          It&#x27;s not clear what API tokens actually cost them, but I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn&#x27;t possible, even if they&#x27;re delivering real business value (coding, research, etc.). In other words, at API prices I&#x27;d just stop using AI, and I suspect most other developers would too.

          1. TeMPOraL · · focus · HN ↗
            &gt; It&#x27;s not clear what API tokens actually cost them, but I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn&#x27;t possible

            Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.

            Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you&#x27;re producing for&#x2F;serving millions of people. This is not unusual.

            &gt; In other words, at API prices I&#x27;d just stop using AI, and I suspect most other developers would too.

            Many say that, but I sincerely doubt they&#x27;d actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring &#x2F; bullshit parts of daily work to give up on merely 3-5x price increase.

            1. csomar · · focus · HN ↗
              &gt; Datacenters have massive economies of scale.

              Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don&#x27;t know, we can only speculate.

              &gt; Many say that, but I sincerely doubt they&#x27;d actually follow through.

              I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I&#x27;ll be using it less.

              Look, I am not saying that there isn&#x27;t a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don&#x27;t have a business model here.

              1. FeepingCreature · · focus · HN ↗
                &gt; Sure. Issue is, no one is providing on how much it actually costs to burn these tokens.

                You can assume Openrouter open-model providers serve at or above margin, because there&#x27;s no branding so there&#x27;s no reason to do it unless you can be profitable. If the Anthropic models are anywhere in that ballpark, they&#x27;re very comfortably profitable on API.

          2. airspresso · · focus · HN ↗
            &gt; I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars).

            That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won&#x27;t be fast and won&#x27;t be frontier intelligence, but definitely useful.

            1. zozbot234 · · focus · HN ↗
              Any model running on &quot;single digit thousands of dollars&quot; hardware will either be below SOTA (even for local models) or not even close to fast enough for real-time agentic work. Even the latest so-called &quot;flash&quot; models are large enough that doing real work usably with those on a lower-cost platform is at least dicey. You can fire off non-interactive work and do especially simple Q&amp;A&#x2F;chat (which is vastly more token-efficient than anything agentic - though even then latency will be high for anything genuinely SOTA) but that&#x27;s about it.
        7. cavoirom · · focus · HN ↗
          Their fate is coming. Until the open-source models will be usable in machine with 256GB memory, they are done. Their behavior is unacceptable (Anthropic) recently but it won&#x27;t last long.
        8. icepush · · focus · HN ↗
          You can put stuff like &quot;make sure your reply is between 800 and 900 tokens&quot; at the end of your prompt and the vast majority of the time it will do so.
      4. CodesInChaos · · focus · HN ↗
        Could be load dependent, not an A&#x2F;B test.

        Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?

        I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.

        1. user3939382 · · focus · HN ↗
          I say A&#x2F;B because when it toggles it does so for days.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.