‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. Grimblewald · · focus · HN ↗
      I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I&#x27;ve been using LLMs heavily even before ada&#x2F;babbage&#x2F;davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
      1. eulgro · · focus · HN ↗
        Your comment makes no sense. How and why would a local model be nerfed anyway...?
        1. r_lee · · focus · HN ↗
          he&#x27;s saying that he notices a difference between local (not nerfable) and hosted ones, so that it&#x27;s not as likely to be just placebo
          1. jacquesm · · focus · HN ↗
            That and &#x27;loss&#x27; may have been intended to be &#x27;launch&#x27;.
          2. martin- · · focus · HN ↗
            But if it is placebo, obviously he wouldn&#x27;t notice any placebo change for local models, since he KNOWS he is using an immutable local model. That comparison only works if he doesn&#x27;t know what model he is using.
            1. Grimblewald · · focus · HN ↗
              I suppose you have a point, there is room for bias in perception and it could explain my sensed degradation of service, however, it doesnt explain the failure of tests, which isn&#x27;t tied to my internal perception.
    2. Rapzid · · focus · HN ↗
      The vibe bro science is this always happens on every release, every Tuesday, and twice on Sunday.

      Of course it&#x27;s almost entirely unsubstantiated BS.

      1. fbrncci · · focus · HN ↗
        Well now it’s being substantiated!
        1. Rapzid · · focus · HN ↗
          Or rather it&#x27;s being.. Unsubstantiated. The Nerf conspiracy isn&#x27;t that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI&#x2F;Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
          1. somenameforme · · focus · HN ↗
            Yeah it&#x27;s just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
            1. Rapzid · · focus · HN ↗
              Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it&#x27;s released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.

              So this is a case of extraordinary claims requiring extraordinary evidence.

              And even though it&#x27;s super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.

              1. somenameforme · · focus · HN ↗
                If you&#x27;re going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you&#x27;re not debating in good faith. Let alone when we&#x27;re speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we&#x27;re speaking of companies who are all racing to IPO where short term results matter more than just about anything.

                Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.

                1. Rapzid · · focus · HN ↗
                  &gt; an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account

                  I&#x27;m sorry, but yeah. The official account is a harness regression and some platform bugs.

                  Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then &quot;nerfing&quot; them to save money and shed load. Where is the evidence?!

                  I&#x27;m not saying it&#x27;s illegal, per say, so don&#x27;t come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don&#x27;t even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to &quot;nerfing&quot; shenanigans.

                  So where is the evidence?!

                  1. airstrike · · focus · HN ↗
                    You&#x27;re giving way too much credit to &quot;enterprises&quot; both noticing and publicly airing out their dissatisfaction

                    Not too mention these companies could easily offer one product to enteprises and another to everyone else

                    Model nerfing is real

                    1. Rapzid · · focus · HN ↗
                      Uh huh. OMG you&#x27;re so right, it&#x27;s soooo real ;) ;) ;)

                      I was so certain it wasn&#x27;t, based on the complete lack of evidence.

                      But then you said it&#x27;s real. NVM, I don&#x27;t need evidence! Somebody said it&#x27;s real!

                      This place has fallen off.

                      1. grim_io · · focus · HN ↗
                        Reddit cross contamination. Over there it&#x27;s a rite of passage to accept it as a fact.
                      2. airstrike · · focus · HN ↗
                        This place has fallen off so long ago that people like you consider yourself old timers but don&#x27;t even follow guidelines

                        Bad evidence is worse than no evidence.

                        And you failed to address the specific criticism I made to your point, instead going for an ad hominem &#x2F; poisoning the well.

                        In sum, you&#x27;re out of line, woefully misled, and lacking in logical thinking. Three strikes, you&#x27;re out.

              2. mrandish · · focus · HN ↗
                &gt; They sell HEAPS through their API access

                The claim is that they nerf subscription accounts not API.

                1. applicative · · focus · HN ↗
                  My impression was that, at least with Anthropic, the point of my subsidized subscription is that I&#x27;ll convince my employer to get an API account. If anything, the motive would instead be to ensh*ttify on the corporations already committed. Those of us inducted as boosters would get the fluffed up product.
              3. troupo · · focus · HN ↗
                &gt; And even though it&#x27;s super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.

                Until shit like this: <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;april-23-postmortem" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;april-23-postmortem

                Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks &quot;fixing it&quot;.

                Or shit like this: <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;a-postmortem-of-three-recent-issues" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;engineering&#x2F;a-postmortem-of-three-...

                Anthropic is in a perpetual state of &quot;oops, these &#x27;bugs&#x27; degraded our model quality&quot; and only admit the issues when it&#x27;s immediately obvious and visibly affects a large number of customers.

                Otherwise all open benchmarks can be (and are) gamed. And it&#x27;s quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.

            2. lxgr · · focus · HN ↗
              It&#x27;s at least highly implausible. Why would they engage in pretty uncontroversially illegal deception&#x2F;fraud if they have so many other legal ways of gaming benchmarks, selling more tokens etc. available to them?

              It&#x27;s like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.

              1. somenameforme · · focus · HN ↗
                They wouldn&#x27;t be breaking any laws at all. It&#x27;d be simply tuning their output as they see fit. And that&#x27;s if you could actually prove it was even intentional, when in reality they have endless plausible deniability of &#x27;oh it was just a technical glitch we&#x27;ve since corrected.&#x27;

                Banks have far greater transparency and legal requirements. LLM companies are just delivering a black box that they have complete control over. And given the regular &#x27;How&#x27;s Claude doing this session?&#x27; stuff, it&#x27;s almost certain that they&#x27;re A-B testing various tweaks on a per session basis.

            3. applicative · · focus · HN ↗
              what&#x27;s unethical about downloading the internet for machine training?
          2. stackghost · · focus · HN ↗
            To me that sounds exactly on-brand for Big Tech in general and Sam Altman in particular.
            1. jackmott42 · · focus · HN ↗

              [dead]

              1. stackghost · · focus · HN ↗
                I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this &quot;conspiracy&quot; theory is perfectly plausible to lots of people

                &gt; in short, yall dumb, shut up.

                no u

                1. Rapzid · · focus · HN ↗
                  It&#x27;s speaks to how low the bar has dropped.

                  Everyone wants to be a software engineer, until it&#x27;s time to do software engineering shit.

                  You know, like scientific method shit we learned in 5th&#x2F;6th grade.

                  It&#x27;s the great bro science incursion.

                  1. stackghost · · focus · HN ↗
                    &gt;Everyone wants to be a software engineer, until it&#x27;s time to do software engineering shit.

                    The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.

        2. scrollop · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;

          this one has been around for over a year

    3. Razengan · · focus · HN ↗
      Theory (Conjecture? Hypothesis?): What we notice as &quot;model nerfing&quot; is the company diverting compute to training&#x2F;running new unreleased models..

      Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of &quot;Astra&quot; more than a month before it was officially announced

      1. Centigonal · · focus · HN ↗
        wouldn&#x27;t less compute result in slower inference, rather than worse performance?
        1. latentsea · · focus · HN ↗
          They could potentially quantize the model and run it at lower quality taking less VRAM.
        2. btown · · focus · HN ↗
          The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.
          1. nightpool · · focus · HN ↗
            why is that more likely?
            1. zxilly · · focus · HN ↗
              Because they already did so. The model in Codex will get lower `juice` than API version.
          2. JohnBooty · · focus · HN ↗
            I assume there&#x27;s classification going on where a really basic &quot;Hi how are you?&quot; style request sent to a high-effort instance can be routed to a lower-level instance. This... is pretty much fine with me, assuming they do a good job of it.

            I would also assume they use nebulous labels like &quot;Medium Effort&quot; or &quot;High Effort&quot; map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?

            I mean, there&#x27;s a reason why they call it &quot;High Effort&quot; and not &quot;Exactly 5 Minutes of GPU Time on Exactly 10 GPUs.&quot; They want to be able to move those sliders and tweak those knobs.

            1. btown · · focus · HN ↗
              The problem is that if benchmarks are run at a tight classifier that says &quot;a request for high effort means check-under-every-stone regardless of simplicity&quot; but a user request is run on a different classifier where &quot;high means maybe high, maybe medium, maybe even low, even for meaningful tasks&quot; then you&#x27;re not getting the model that you saw in the benchmarks.

              And, while you might be billed fewer tokens as a result (because the lower thinking would result in less investigatory work), you might not know this is happening, and know to dial up effort accordingly - you&#x27;d simply get a worse work product. And certainly, Anthropic&#x27;s incentive for anyone on a subscription is to push this as aggressively as they can, so people use less of that subscription.

              Sadly, I&#x27;d also expect that the OP&#x27;s benchmark will be detected as a test of model capabilities, and thus be given a high classification so that this strategy remains undetected.

        3. poizan42 · · focus · HN ↗
          My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.
      2. jackmott42 · · focus · HN ↗
        There is no nerfing, look at the data before coming up with a theory as to why the nerfing that isn&#x27;t even happening is happening.

        fuck

        1. [deleted] · · focus · HN ↗

          [deleted]

    4. comboy · · focus · HN ↗
      But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.
    5. andriy_koval · · focus · HN ↗
      Usage bench is also very useful! Thank you for doing this!
      1. hamandcheese · · focus · HN ↗
        ...is it? I&#x27;m looking and it seems like it doesn&#x27;t have any data. It might be useful if they keep it up.
        1. andriy_koval · · focus · HN ↗
          Yeah, I guess they started it today..
    6. user3939382 · · focus · HN ↗
      Anthropic A&#x2F;Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly&#x2F;5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
      1. braingravy · · focus · HN ↗
        Pretty amazing to see enshitification happen live with a product still in development… Truly web 4.0
        1. LimitExperience · · focus · HN ↗

          [dead]

      2. apitman · · focus · HN ↗
        Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
        1. ffsm8 · · focus · HN ↗
          Claude code supposedly has otel you can set via env. I haven&#x27;t set it up, so I&#x27;m just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost

          It&#x27;s meant for their test env I think, so is not documented to my knowledge

          1. adastra22 · · focus · HN ↗
            Has otel? What is that?
            1. reubenmorais · · focus · HN ↗
              OpenTelemetry
          2. TeMPOraL · · focus · HN ↗
            It&#x27;s for corporate users who want to track how the product is used internally, and it was documented at least some time ago, quite extensively even.
            1. ffsm8 · · focus · HN ↗
              youre right!

              <a href="https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;monitoring-usage#usage-monitoring" rel="nofollow">https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;monitoring-usage#usage-monit...

              thanks for correcting me on that regard

        2. Aeolun · · focus · HN ↗
          Tokens used &#x2F; percentage change is a pretty obvious metric. They give you both, but they don’t do the math for you.
          1. apitman · · focus · HN ↗
            It&#x27;s obvious unless you have multiple requests from different models in flight at the same time, and the sum total usage comes out to less than a single percentage.
      3. jacquesm · · focus · HN ↗
        How did pissing off your customers ever become a business model?

        I can&#x27;t imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don&#x27;t control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don&#x27;t care because eventually it worked. It&#x27;s almost an ideal vehicle to scam people.

        Imagine the power company being able to decide how much you consume and at which price point.

        1. herval · · focus · HN ↗
          &gt; How did pissing off your customers ever become a business model?

          Airlines, banks, health insurance…

          1. tccole · · focus · HN ↗
            So very low margin businesses with hogh amounts of regulations.
            1. petesergeant · · focus · HN ↗
              Banks and health insurance are much more consumer friendly outside of the US, usually because of regulation. Turns out you can just tell banks “make transfers cheap and essentially instant” and they’ll do it, rather the bullshit they have in the US.
              1. herval · · focus · HN ↗
                you&#x27;d be surprised. I&#x27;ve yet to live in a place where either is consumer-friendly...
            2. herval · · focus · HN ↗
              makes you wonder why openai&#x2F;chatgpt&#x2F;xai are in bed with government so much...
        2. pixelready · · focus · HN ↗
          Step 1: Oligopoly Step 2: Regulatory Capture Step 3: Profit
          1. miohtama · · focus · HN ↗
            Only if we did not have these cheap illegal Chinese models
            1. msdz · · focus · HN ↗
              That is why

              &gt; Step 2: Regulatory Capture

              is being worked towards.

            2. TeMPOraL · · focus · HN ↗
              &gt; illegal

              Are they though? Or is it just what some companies would want them to be?

        3. none_to_remain · · focus · HN ↗
          I find it amazingly rich that they bill you for &quot;&quot;&quot;thinking&quot;&quot;&quot; tokens and now you don&#x27;t even get to see them, they&#x27;re gonna train the thing to sing &quot;99 Bottles of Beer on the Wall&quot; to itself before it starts work.
          1. Turskarama · · focus · HN ↗
            They don&#x27;t want to waste tokens on purpose, what they&#x27;re actually hiding is when the model wastes tokens on obviously stupid &quot;thoughts&quot;.
            1. adastra22 · · focus · HN ↗
              No they are hiding the chain of thought to make distillation harder.
              1. slim · · focus · HN ↗
                It&#x27;s fascinating that you all think accounting is real and it did not come to your mind that they could make up numbers when billing
                1. TeMPOraL · · focus · HN ↗
                  They could, but as you see here, people are very eager to create dashboards and trackers that do external accounting by proxy, so they can&#x27;t just &quot;make up numbers&quot; without the customers noticing and making a fuss.
        4. topspin · · focus · HN ↗
          It&#x27;s worked for online PvP gaming for a long time. Nerf stuff the min-maxers &quot;earned&quot; through game mechanics and sell over-powered &quot;premium&quot; things to everyone else to pwn them. Then nerf the old premium stuff and make new premium stuff. Forever.

          I don&#x27;t know if that&#x27;s the actual origin of the term nerf, but it was the first time I&#x27;d heard it.

          1. done_lurking · · focus · HN ↗
            I think the origin of the word &quot;nerf&quot; as a verb came from the Nerf brand of toy guns. The idea being that &quot;Nerfing&quot; something is to turn it into a harmless version of itself.
        5. CodesInChaos · · focus · HN ↗
          Another way Antropic misleads its customers is the description of the max plans. They are advertised as having 5x&#x2F;20x the 5h quota as Pro. But the description says nothing about how the weekly quota scales, leaving customers to infer it scales the same way. But from what I&#x27;ve heard, the weekly quota is only 3.5x&#x2F;7x that of Pro.
        6. csomar · · focus · HN ↗
          I think it&#x27;s sinister, but not for the reasons you&#x27;re thinking. I think they&#x27;re just wildly unprofitable on subscriptions. The idea that most customers won&#x27;t use their full quota is plain wrong: most people are maxing out their subs, or even reselling whatever quota they have left.

          When you&#x27;re running something at a loss, you can mistreat your customers and they&#x27;ll still stick around (I&#x27;m an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.

          My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they&#x27;ll either drop the subscription model or it&#x27;ll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.

          It&#x27;s not clear what API tokens actually cost them, but I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn&#x27;t possible, even if they&#x27;re delivering real business value (coding, research, etc.). In other words, at API prices I&#x27;d just stop using AI, and I suspect most other developers would too.

          1. TeMPOraL · · focus · HN ↗
            &gt; It&#x27;s not clear what API tokens actually cost them, but I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn&#x27;t possible

            Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.

            Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you&#x27;re producing for&#x2F;serving millions of people. This is not unusual.

            &gt; In other words, at API prices I&#x27;d just stop using AI, and I suspect most other developers would too.

            Many say that, but I sincerely doubt they&#x27;d actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring &#x2F; bullshit parts of daily work to give up on merely 3-5x price increase.

            1. csomar · · focus · HN ↗
              &gt; Datacenters have massive economies of scale.

              Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don&#x27;t know, we can only speculate.

              &gt; Many say that, but I sincerely doubt they&#x27;d actually follow through.

              I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I&#x27;ll be using it less.

              Look, I am not saying that there isn&#x27;t a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don&#x27;t have a business model here.

              1. FeepingCreature · · focus · HN ↗
                &gt; Sure. Issue is, no one is providing on how much it actually costs to burn these tokens.

                You can assume Openrouter open-model providers serve at or above margin, because there&#x27;s no branding so there&#x27;s no reason to do it unless you can be profitable. If the Anthropic models are anywhere in that ballpark, they&#x27;re very comfortably profitable on API.

          2. airspresso · · focus · HN ↗
            &gt; I looked into running a local model, and it&#x27;s way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars).

            That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won&#x27;t be fast and won&#x27;t be frontier intelligence, but definitely useful.

            1. zozbot234 · · focus · HN ↗
              Any model running on &quot;single digit thousands of dollars&quot; hardware will either be below SOTA (even for local models) or not even close to fast enough for real-time agentic work. Even the latest so-called &quot;flash&quot; models are large enough that doing real work usably with those on a lower-cost platform is at least dicey. You can fire off non-interactive work and do especially simple Q&amp;A&#x2F;chat (which is vastly more token-efficient than anything agentic - though even then latency will be high for anything genuinely SOTA) but that&#x27;s about it.
        7. cavoirom · · focus · HN ↗
          Their fate is coming. Until the open-source models will be usable in machine with 256GB memory, they are done. Their behavior is unacceptable (Anthropic) recently but it won&#x27;t last long.
        8. icepush · · focus · HN ↗
          You can put stuff like &quot;make sure your reply is between 800 and 900 tokens&quot; at the end of your prompt and the vast majority of the time it will do so.
      4. CodesInChaos · · focus · HN ↗
        Could be load dependent, not an A&#x2F;B test.

        Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?

        I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.

        1. user3939382 · · focus · HN ↗
          I say A&#x2F;B because when it toggles it does so for days.
    7. avazhi · · focus · HN ↗
      Nerfbench isn&#x27;t helpful if it&#x27;s 3 days old.
    8. bdlowery · · focus · HN ↗
      &gt; This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about.

      this bench was just released, it couldn&#x27;t have detected opus 4.6 degradation.

    9. nullbio · · focus · HN ↗
      Do they use private benchmarks? Because if not, it could be selectively nerfed.

      I also wonder if cache could be used to throw these off as well, where it&#x27;s serving un-nerfed cache results for context windows that are identical to ones they&#x27;ve previously had for benchmark requests.

      Seems like the only way to do it well would be to have some randomness involved that couldn&#x27;t be cheated on - but you&#x27;d want to do it in a way that doesn&#x27;t throw out the benchmarks too much, so your results can be compared still.

      1. Gabrys1 · · focus · HN ↗
        Thankfully, we can now use AI to design a test that tests AI. And the AI company can use AI to detect the test and cheat. And we can then use AI to implement anti-cheat.

        All that energy wasted... could just drive a big V8 instead and make less money for the big tech

    10. scrollop · · focus · HN ↗
      There&#x27;s also this one which has been around for a while

      <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;claude-code&#x2F;

      1. sscaryterry · · focus · HN ↗
        The tracker for Codex resonates with me. Its thick as pig shit the last few days: <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;
        1. phoghed · · focus · HN ↗
          Seems like a harness change (maybe bug) rather than a model change, the input tokens dropped quite a bit right when the degradation happened.
        2. lxgr · · focus · HN ↗
          &gt; We use the latest available Codex release with GPT-6 Sol.

          This alone makes the benchmark unsound.

        3. shawabawa3 · · focus · HN ↗
          &gt; &quot;We are collecting a new GPT-6 Sol&#x2F;high baseline from runs beginning September 24, 2026. Degradation detection is paused.&quot;
    11. rplnt · · focus · HN ↗
      &gt; I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

      I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.

      It&#x27;s been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence &quot;believe&quot;.

      1. transcriptase · · focus · HN ↗
        I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.

        People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.

        1. setopt · · focus · HN ↗
          Regarding lazy mode, I recall ChatGPT sometimes almost refusing to do a web search despite me asking explicitly for it, instead replying with speculation about what the search results likely would tell us. If I pretend to be angry that it didn’t search the web it would however do it. Haven’t noticed this in a while either.
          1. natpalmer1776 · · focus · HN ↗
            Oh god I just realized these are the same types of stories passed down to me by sysadmins of yore about when microsoft did XYZ. Am I… old now?
          2. lukan · · focus · HN ↗
            &quot;If I pretend to be angry that it didn’t search the web it would however do it.&quot;

            You have to pretend to be angry in such situations?

        2. katzenq · · focus · HN ↗
          It should have stayed that way. Be a good search engine and encourage the human to do the work themselves.
          1. AbsurdCensor · · focus · HN ↗
            Wouldn&#x27;t that be like a calculator saying &#x27;pull out a math book&#x27; when you try running calculations?
            1. katzenq · · focus · HN ↗
              It&#x27;s like doing the calculations yourself (with whatever tools you want) and knowing what you&#x27;re doing, steering the process yourself, rather than begging someone to give you the result so you can then show it off like you did the work. I can understand and explain what my calculator is doing and I&#x27;m not offloading decisions to it. Using a language model to shit out projects that you can&#x27;t fully understand the structure of yourself is insane. Ceding control of any design decisions to a language model is insane. They&#x27;re fine as second-order autocomplete and semantic search engines. Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can&#x27;t build humane systems.
              1. sejje · · focus · HN ↗
                &gt; Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can&#x27;t build humane systems.

                I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.

                I&#x27;m mostly building my own UIs lately, but I find no problem with the development loop. I&#x27;m building much better UIs, because it&#x27;s way easier to test things, and scrap things that I thought would work, but don&#x27;t. It all happens in a matter of minutes.

                Humans should use more llms.

      2. smurf9852 · · focus · HN ↗
        Friend of mine works for a corp that is one of the top spenders on Claude models. He complained about these nerfs during peak demand. Their Anthropic contact changed something and it did not happen since.
      3. silversmith · · focus · HN ↗
        Anecdote - I work in a time zone offset from continental US. The performance of anthropic models would noticeably drop, around the time US work day started. It was so bad around 4.x time that multiple colleagues re-arranged their schedule to have least overlap with US work day. Admittedly it&#x27;s been better recently.
      4. Barbing · · focus · HN ↗
        &gt; It&#x27;s been a few months since I last recall this though

        Anthropic cut a deal with SpaceXAI in May - $1.25b&#x2F;mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.

        <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;higher-limits-spacex" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;higher-limits-spacex

        1. rplnt · · focus · HN ↗
          &gt; dishonest

          This is what pissed me off the most. Make it slower, rate limit it, move the credits to other time slots, idk.. but returning BAD results? That&#x27;s the worst approach you could take.

      5. jclardy · · focus · HN ↗
        I do remember times in the past where when I was up super early (4am EST) I would get super high quality results, then in mid afternoon EST it seemed to be degraded.
    12. zsoltkacsandi · · focus · HN ↗
      Or nerfed version of models rolled out gradually.
    13. rednb · · focus · HN ↗
      I&#x27;d take this kind of benchmark with a grain of salt. At this point, I have a set of comprehensive guidelines covering both backend and frontend work, and for the frontend we go as far as explaining what we a good design is in our visual system, and even how to conduct a visual review when screenshots are handed to the model.

      Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i&#x27;ve transitioned to using DS 4.1 as a daily driver and am very satisfied.

      My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.

      1. dotancohen · · focus · HN ↗
        You use DS through Open Router? Which harness?

        I&#x27;d love to hear more, I&#x27;m considering jumping ship. I&#x27;m running a Debian desktop if that&#x27;s a concern.

        1. rednb · · focus · HN ↗
          I use the direct API from deepseek using Opencode, no problem to report, works like a charm.

          Except maybe that the model often believes that he is running out of context, and needs to rush so i occasionally need to jump in to tell it that it still has plenty of room left.

          But this does not degrade the quality of my overall experience in a meaningful way.

          1. dotancohen · · focus · HN ↗
            Thank you. I might just check that out.
      2. [deleted] · · focus · HN ↗

        [deleted]

    14. KronisLV · · focus · HN ↗
      We could also use something that tracks concrete token amounts each tier gives you, in case they ever mess with it - and also maybe even the tokens needed to accomplish a particular benchmark, to see how much you can actually get done.
    15. zerop · · focus · HN ↗
      What could be the reason to nerf?
      1. StableAlkyne · · focus · HN ↗
        It&#x27;s cheaper to run a quantization of a model, but its quality is reduced.

        For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.

        The game is to balance between the savings of quantization and making the model dumb enough the people notice

        1. sscaryterry · · focus · HN ↗
          Someone always knows about where the bodies are buried. These shenanigans always end up surfacing eventually.
          1. ForHackernews · · focus · HN ↗
            If they surface after the IPO, everyone who matters will have already gotten paid.

            <a href="https:&#x2F;&#x2F;www.fool.com&#x2F;investing&#x2F;2026&#x2F;07&#x2F;25&#x2F;spacexs-performance-looks-almost-identical-to-past&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.fool.com&#x2F;investing&#x2F;2026&#x2F;07&#x2F;25&#x2F;spacexs-performanc...

          2. sigbottle · · focus · HN ↗
            But as we also see in this thread, all evidence is dismissed and there&#x27;s no good faith discussion. All data from the other side is lies and contamination. In that environment, it&#x27;s power who decides who wins.
    16. quikoa · · focus · HN ↗
      Wouldn&#x27;t it be trivial to detect benchmarking if the same requests are running on a fixed interval?
    17. nsarrazin · · focus · HN ↗
      Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.

      It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

      1. waterproof · · focus · HN ↗
        I find that I learn to &quot;trust&quot; a model to get certain things right, as I would trust a colleague. So, as `expectation` increases, my prompting and context management gets sloppier.

        `percieved_performance = actual_perf&#x2F;expectation`

        `expectation` is an increasing function over time.

        `actual_perf` is a stochastic function of the model&#x27;s true ability, context, etc. -&gt; a recipe for some bad sessions.

        As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive &quot;runs&quot; because our brains love to find patterns.

      2. causal · · focus · HN ↗
        Yeah I bet most of us remember GPT-4 a lot more fondly than we would if we were to return to it today.
        1. svachalek · · focus · HN ↗
          Absolutely. Objectively speaking it was far less consistent and capable than even small local models today.
      3. goodmythical · · focus · HN ↗
        This reminds me of the fact that true random does not feel random to users due to the clumpiness that the average person does not anticipate existing in true random.

        e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are &quot;not random&quot;

        1. shagie · · focus · HN ↗
          Some more on this...

          How to Shuffle Songs? - <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20220215030739&#x2F;https:&#x2F;&#x2F;engineering.atspotify.com&#x2F;2014&#x2F;02&#x2F;how-to-shuffle-songs&#x2F;" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20220215030739&#x2F;https:&#x2F;&#x2F;engineeri... ( <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38330877">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38330877 78 points, 65 comments)

          Took a little bit of digging to find it - I remembered the graphic at the top and found a blog post that copied it and linked to the blog post, but the blog post isn&#x27;t there anymore... so web archive.

          The current version of the blog post is from 2025 - <a href="https:&#x2F;&#x2F;engineering.atspotify.com&#x2F;2025&#x2F;11&#x2F;shuffle-making-random-feel-more-human" rel="nofollow">https:&#x2F;&#x2F;engineering.atspotify.com&#x2F;2025&#x2F;11&#x2F;shuffle-making-ran... (which didn&#x27;t get any traction on HN)

        2. svachalek · · focus · HN ↗
          I&#x27;ve been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc. Just could not find it. Streak continued to another roll, another roll... still couldn&#x27;t find it. Then the streak just broke. Apparently just random being random.
          1. goodmythical · · focus · HN ↗
            I find it useful to consider something like shaken rice. If you take a 10x10 grid and shake 100 grains of rice on it, you&#x27;ll find that some cells contain no rice while others contain as many as 5 or 6. Run the experiment enough and you&#x27;ll converge on each cell getting one grain per run, but any individual sample will likely be clumpy and the state in which each only has one will occur infrequently.

            Also, consider that in flipping 10 coins, you&#x27;ll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you&#x27;ll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...

            Widening the range from &quot;rolling exactly 15&quot; to &quot;rolls 15 or 16&quot; or &quot;rolls between 14-17&quot; makes the strings even more likely as you&#x27;re doubling the success rate from &quot;only 9 15s&quot; to the &quot;any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s&quot; space.

            To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:

            For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:

            General Variables: N = total number of rolls&#x2F;trials k = target streak length p = probability of getting the target outcome (e.g., 1&#x2F;20 for a specific roll on d20 or 1&#x2F;10 for two specific results) q = probability of getting any other outcome (1 - p)

            Expected runs of AT LEAST length k: E(runs &gt;= k) = p^k * (1 + (N - k) * q)

            Expected runs of EXACT length k: E(exact k) = p^k * q * (2 + (N - k - 1) * q)

            Personally, I find that &#x27;sticky&#x27; dice always provide a nice narrative device, at least in narrative games. A character who&#x27;s player can&#x27;t seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.

          2. x______________ · · focus · HN ↗
            &gt; I&#x27;ve been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc

            Been through that this week as well with 100% success on 40% odds over multiple iterations on my game.. I tend to not dig into random but just rather ensure it works &#x27;as closely to intended&#x27; as possible..

      4. booty · · focus · HN ↗
        For years I used to buy ASICS running shoes. Every year they released a new model of each shoe: &quot;Nimbus 23&quot;, then &quot;Nimbus 24&quot; the next year, etc. And every year people would complain in the user reviews about how each shoe was worse than the last.

        I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They&#x27;ve been getting continuously worse for 24 consecutive years!

        Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn&#x27;t have much to say about them.

        (The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it&#x27;s cushy -- you&#x27;re comparing a new sneaker to your old sneaker where the foam had lost its bounce...)

        1. unshavedyak · · focus · HN ↗
          That&#x27;s kinda me with respect to Claude. Generally i&#x27;ve had no issues and just kept pluggin&#x27; along.

          The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i&#x27;d be on OpenAI by now.

        2. yencabulator · · focus · HN ↗
          It&#x27;s pretty typical that physical-goods manufacturing &quot;optimizes the process&quot; to cut costs during years 1 &amp; 2 of manufacturing.

          Ikea is notorious for this: The early Billy bookcase had heavier veneer and sturdier construction early on, and was actually a really good purchase for the money. The later years replaced veneer with paper foil, used thinner shelves, frames, and backing panels, and was just significantly weaker.

          1. ComputerGuru · · focus · HN ↗
            Amazon Basics is incredible in this regard, they’ve optimized SKU identification down to a pipeline. They’ll essentially randomly pick items off their internal list of highest netting sales and test them to see how dependent they are on brand name recognition and price-quality signalling. To do this as efficiently as possible, they simply purchase a few hundred units of a high quality product in the space, stick it in an Amazon Basics box and list it on their site under their Amazon Basixs brand at a price they feel they can achieve via white labeling, and wait to see how it sells. The use of high quality items (with quality above what can actually be had at the listed price point for the duration of the experiment) means they are really only testing the user base’s willingness to forgo a brand name for the category in exchange for a discount. If it sells well, they then work on sourcing it in bulk as a white labeled item “for real”, while if it sells poorly they simply delist and move on.

            I (used to) buy pre-spliced&#x2F;terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3&#x2F;OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.

            To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.

            1. dragontamer · · focus · HN ↗
              I have a similar story.

              Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand &#x27;Eneloop&#x27;. Extremely good specs all around

              Recently though, they are still called Amazon Basics but no longer test like Eneloop. They&#x27;ve changed manufacturers for the worse and are hoping no one notices...

              1. hn_acc1 · · focus · HN ↗
                I did a deep dive into NiMH batteries a few years ago and concluded that most people felt similarly: you can often get &quot;good&quot; (similar to Eneloop) specs for a short run from almost any manufacturer at the outset (see Ikea Ladda batteries - suspected of relabeled Eneloop for a while, but now not as good), but consistent quality is pretty much only Eneloop or other name brand, with Eneloop generally being the best.

                In the interests of saving my sanity and time (it&#x27;s not free!) having to chase down which batch of which brand is &quot;good&quot; at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn&#x27;t worth it to me (I get paid reasonably well).

                1. dragontamer · · focus · HN ↗
                  Yeah, name brands like Energizers are consistent but slightly worse than Eneloop at roughly the same price.

                  I actually made a battery tester as a hobby electronics project. Fully dumps the energy while measuring mAh and seeing some measurements of internal resistance.

                  Something I did notice was that crappy battery chargers can permanently damage even Eneloops. So don&#x27;t cheap out on the chargers either.

                  The high speed chargers (2 hours or less) are on the edge of what is safe and could permanently damage a cell. The 4 hours or slower chargers are just way more reliable. There&#x27;s probably someone out there testing different chargers (trying to find the safe fast chargers) but given how cheap NiHMs are (even the &quot;expensive&quot; Eneloops), it&#x27;s just easier to use 4 hour or 8 hour chargers and have masses and masses of extra NiHMs laying around.

                  1. ComputerGuru · · focus · HN ↗
                    I agree about the chargers!

                    Im sure you know this, but do note that fully dumping the charge can a) damage the battery and dramatically shorten its lifespan, b) give you different results depending on your discharge rate and the batteries you are testing, as they all have different current-dependent discharge curves.

                    1. dragontamer · · focus · HN ↗
                      Oh yeah. I stop at 0.95V. which I consider fully discharged (but not so low that I&#x27;m damaging the cells). I&#x27;m about 10% off the manufacturers claimed specs but I&#x27;d rather not push the limits of testing. That&#x27;s low enough to differentiate between good and bad cells

                      The discharge rate is currently regulated by only a 2.2 Ohm resistor. The next version of this circuit will be a constant current drain circuit to make all tests draw the same mA across the whole test.

                      Right now more current is drawn at 1.35V start and far less current is drawn at 0.95V end of test.

                      --------

                      The part that they don&#x27;t tell you is that 0.95V isn&#x27;t one point. When you disconnect the cell, it charges back up (kinda like a capacitor). This process can take multiple minutes (!!!!). I&#x27;ve defined the end as being lower than 0.95V for more than 30 seconds, even if disconnected.

              2. ComputerGuru · · focus · HN ↗
                No, that’s just white labeling. Like you said though, they can switch manufacturers and pull the rug from underneath you at any time.

                My example is like if you opened the Amazon Basics box and found Eneloop branded white&#x2F;black&#x2F;blue batteries directly.

        3. MikeTheGreat · · focus · HN ↗
          I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).

          Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.

          There&#x27;s also &quot;the Schlitz Mistake&quot;, which I&#x27;ve heard summarized as &quot;most customers won&#x27;t notice if you take your product&#x27;s quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)&quot;

          At this point I kinda assume that any company releasing year updates to a physical product that _doesn&#x27;t_ take the opportunity to trim costs &#x2F; reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...

          1. aesthesia · · focus · HN ↗
            I&#x27;m curious how common such lawsuits are. I don&#x27;t think I&#x27;ve heard of any specific instances where shareholders sued because a company didn&#x27;t make the product worse.
            1. MikeTheGreat · · focus · HN ↗
              On the one hand my assumption is mostly hyperbole

              and

              On the other hand that idea that &quot;companies exist to make money for shareholders, to ONLY make money for shareholders, and doing any other than maximizing shareholder returns is bad&quot; is pretty commonly accepted (and, I believe, enshrined in US law)

              1. skinfaxi · · focus · HN ↗
                Your belief is misplaced and incorrect.
          2. michaelmrose · · focus · HN ↗

            [dead]

          3. ehe78qhe · · focus · HN ↗
            &gt; Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.

            It&#x27;s been much longer than that: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Toblerone#2016_size_changes" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Toblerone#2016_size_changes

            1. xnorswap · · focus · HN ↗
              They handled that comically badly, the proportions meant the gaps went from being perceived as small to being a feature:

              <a href="https:&#x2F;&#x2F;www.bbc.co.uk&#x2F;news&#x2F;uk-44910195" rel="nofollow">https:&#x2F;&#x2F;www.bbc.co.uk&#x2F;news&#x2F;uk-44910195

              Visually it looks like they took out half the peaks.

              Your mind tells you that for every gap there used to be a peak there, regardless of the truth.

          4. m10i · · focus · HN ↗
            &gt; &quot;most customers won&#x27;t notice if you take your product&#x27;s quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)&quot;

            Summarizes the video game industry pretty well

          5. pnt12 · · focus · HN ↗
            Nitpick: I think that&#x27;s called skimpflation (worse quality), a friend of shrinkflation (less quantity) or inflation (higher prices).
            1. MikeTheGreat · · focus · HN ↗
              Today I learned a new word! Thanks!
          6. booty · · focus · HN ↗
            &gt; I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).

            That&#x27;s certainly common!

            I think ASICS&#x27; running shoes were a fun example of where this was probably not the case.

            - I certainly didn&#x27;t notice a difference in that time, though I&#x27;m admittedly not much of an actual runner

            - Serious runners might notice small technical differences, but the negative user reviews didn&#x27;t seem to indicate those were the people making the complaints

            - The overall user reviews remained positive

            - The competition in the shoe market is incredibly fierce; I&#x27;m not sure a brand could tank their quality and survive for long

            - Let&#x27;s not forget the other big variable: the wearers&#x27; bodies, particularly their feet. Now, those definitely do change over time -- most often for the worse, sadly!

            - I doubt anybody was blind A&#x2F;Bing a pair of Nimbus 20 against a pair of Nimbus 21 or 22 or 23 or 24. At best, a longtime Nimbus buyer is probably comparing a brand new pair of e.g. Nimbus 24 against their degraded but broken-in Nimbus 23 and their memory of how the Nimbus 23 felt when new. (And their body is a year or two older at that point..)

            - Because it&#x27;s such a long-running line of shoes, there could certainly be year-to-year variations... but it&#x27;s hard to imagine there was an actual 5 or 10 or 25 year downward slope. I mean, otherwise at that point the shoes would just be instantly injuring you or falling apart in a week

            - Also because it&#x27;s such a long-running product line and (aside from bleeding-edge professional marathon&#x2F;track shoes) sneaker manufacturing in general kind of seems like a solved problem... it seems like all of the possible cost optimizations have already been optimized. I don&#x27;t really think there&#x27;s much of a manufacturing or bill-of-materials cost difference between $20 running shoes and $200 running shoes anyway -- I&#x27;d be pretty surprised if &quot;quality shrinkflation&quot; was really much of a viable way for ASICS to save a few pennies.

        4. ck2 · · focus · HN ↗
          actually they were getting worse each year

          most running shoe series get heavier year to year as manufacturers turn to cheaper materials and add cushioning to try to attract more adopters

          it&#x27;s almost universal, very few manufacturers seem to be able to resist tampering

          (heavier shoes are slower, every three ounces is equal to another vo2max point lost)

          1. reducesuffering · · focus · HN ↗
            You&#x27;re literally doing what GP is describing. We have objective data on running shoes on: <a href="https:&#x2F;&#x2F;runrepeat.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;runrepeat.com&#x2F;

            The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.

            1. ck2 · · focus · HN ↗
              nope

              you mean NEW MODELS are being introduced with lighter faster foams

              not the same model year to year

              modern example: Saucony Endorphin Speed

              v1 in 2020 was award winning

              v2 in 2021 was almost the same, more praise

              v3 bleh

              v4 v5 bleh bleh

              they cannot resist tampering

        5. dgacmu · · focus · HN ↗
          Except for that one year when they completely swapped the meaning of the Cumulus line, which I think was 2008 with the cumulus 9 to 10 transition. It went from a neutral shoe that was good for people with high arches to more of a stiff stability shoe. The complainers aren&#x27;t _always_ crazy. :-) (That doesn&#x27;t mean the cumulus 10 was worse, of course, it just was a more substantial change that affected the type of runner the shoe was designed for.)

          Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I&#x27;m old and heavier and run in the Nimbus and am happy again. But you&#x27;re right, of course, that most of the model changes are just fine and people like to complain.

      5. oooyay · · focus · HN ↗
        &gt; Perceived performance is actual performance over expectations and the latter just keeps increasing over time.

        This is true of all reliability and performance paradigms, incidentally

    18. shados · · focus · HN ↗
      The whole nerfing narrative puts in the spotlight now crazy supertitions come into being. The group think every day that everything is falling apart is crazy.
    19. loopydosuette · · focus · HN ↗
      you can&#x27;t explain honeymoon effects. like wow wow wow and then suddenly: same task, lesser performance is a misperception?

      my fair lady gained a bit weight and the bjs lack variety? [ I&#x27;m certain that that&#x27;s a quote from some time ago by some commenter in a thread with a similar or even the same context ...but my Amnesia (T▽T ):・゚:・゚] fuck off.

      you open two files, before and after you notice a nerf, and from worse comments to logical oversights, it&#x27;s all damn obvious.

      don&#x27;t normalize this make believe bullshit and misleading people who you think barely understand what they see anyway ...

      you wouldn&#x27;t even know if models had somehow timed nerfs hardcoded into them, however much control over the stack you have.

      ridiculous

    20. Aperocky · · focus · HN ↗
      So that&#x27;s why even Qwen-3.8-27B caught up with Opus4.6, it was nerfed to the ground.
    21. stbenjam · · focus · HN ↗
      These guys are on twitter angry about the rate limit decrease and allegedly cancelled all their OpenAI accounts. Wonder how they&#x27;ll maintain this.

      I am quite convinced that the whole nerfing phenomenon is 90% AI psychosis. I have the word muted on X.

      1. cyberes · · focus · HN ↗
        That&#x27;s the wrong use of the word
        1. Barbing · · focus · HN ↗
          Maybe just the wrong ballpark (like 1-10%).
        2. stbenjam · · focus · HN ↗
          what word?
    22. throwitaway222 · · focus · HN ↗
      Often times people think of &quot;nerfs&quot; as my first prompt (which was greenfield - no or little code existed) used 5% of my plan usage. And then 2 weeks later (as the agent is busy reading hundreds of .rs and .ts files it previously generated) the user complains the usage is going down 30% for a single prompt instead of 5%. Attributing this to a &quot;NERF&quot; makes little sense because it&#x27;s the same model.
    23. jtrn · · focus · HN ↗
      If this is actually even close to reliable tracking, it&#x27;s one of the most awesome benchmarks I&#x27;ve seen. I gave up on feeling the zeitgeist for what people were saying.
    24. jdthedisciple · · focus · HN ↗
      Well looks like thus far Astra seems to be getting anything but nerfed, given that its score is actually rising
    25. futune · · focus · HN ↗
      Onyxia deep breathes more this patch.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.