‹ BackHN Continuity

Thread

Livenerf: Has Opus 5.5 been nerfed yet?

922 points · 392 comments · bryan0

  1. jug · · focus · HN ↗
    We also have Nerf Bench:

    <a href="https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench" rel="nofollow">https:&#x2F;&#x2F;www.bridgebench.ai&#x2F;nerf-bench

    They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They&#x27;re currently tracking Opus 5.5 and GPT-6 Astra.

    This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it&#x27;s often about honeymoon effects.

    1. nsarrazin · · focus · HN ↗
      Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.

      It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

      1. waterproof · · focus · HN ↗
        I find that I learn to &quot;trust&quot; a model to get certain things right, as I would trust a colleague. So, as `expectation` increases, my prompting and context management gets sloppier.

        `percieved_performance = actual_perf&#x2F;expectation`

        `expectation` is an increasing function over time.

        `actual_perf` is a stochastic function of the model&#x27;s true ability, context, etc. -&gt; a recipe for some bad sessions.

        As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive &quot;runs&quot; because our brains love to find patterns.

      2. causal · · focus · HN ↗
        Yeah I bet most of us remember GPT-4 a lot more fondly than we would if we were to return to it today.
        1. svachalek · · focus · HN ↗
          Absolutely. Objectively speaking it was far less consistent and capable than even small local models today.
      3. goodmythical · · focus · HN ↗
        This reminds me of the fact that true random does not feel random to users due to the clumpiness that the average person does not anticipate existing in true random.

        e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are &quot;not random&quot;

        1. shagie · · focus · HN ↗
          Some more on this...

          How to Shuffle Songs? - <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20220215030739&#x2F;https:&#x2F;&#x2F;engineering.atspotify.com&#x2F;2014&#x2F;02&#x2F;how-to-shuffle-songs&#x2F;" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20220215030739&#x2F;https:&#x2F;&#x2F;engineeri... ( <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38330877">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=38330877 78 points, 65 comments)

          Took a little bit of digging to find it - I remembered the graphic at the top and found a blog post that copied it and linked to the blog post, but the blog post isn&#x27;t there anymore... so web archive.

          The current version of the blog post is from 2025 - <a href="https:&#x2F;&#x2F;engineering.atspotify.com&#x2F;2025&#x2F;11&#x2F;shuffle-making-random-feel-more-human" rel="nofollow">https:&#x2F;&#x2F;engineering.atspotify.com&#x2F;2025&#x2F;11&#x2F;shuffle-making-ran... (which didn&#x27;t get any traction on HN)

        2. svachalek · · focus · HN ↗
          I&#x27;ve been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc. Just could not find it. Streak continued to another roll, another roll... still couldn&#x27;t find it. Then the streak just broke. Apparently just random being random.
          1. goodmythical · · focus · HN ↗
            I find it useful to consider something like shaken rice. If you take a 10x10 grid and shake 100 grains of rice on it, you&#x27;ll find that some cells contain no rice while others contain as many as 5 or 6. Run the experiment enough and you&#x27;ll converge on each cell getting one grain per run, but any individual sample will likely be clumpy and the state in which each only has one will occur infrequently.

            Also, consider that in flipping 10 coins, you&#x27;ll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you&#x27;ll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...

            Widening the range from &quot;rolling exactly 15&quot; to &quot;rolls 15 or 16&quot; or &quot;rolls between 14-17&quot; makes the strings even more likely as you&#x27;re doubling the success rate from &quot;only 9 15s&quot; to the &quot;any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s&quot; space.

            To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:

            For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:

            General Variables: N = total number of rolls&#x2F;trials k = target streak length p = probability of getting the target outcome (e.g., 1&#x2F;20 for a specific roll on d20 or 1&#x2F;10 for two specific results) q = probability of getting any other outcome (1 - p)

            Expected runs of AT LEAST length k: E(runs &gt;= k) = p^k * (1 + (N - k) * q)

            Expected runs of EXACT length k: E(exact k) = p^k * q * (2 + (N - k - 1) * q)

            Personally, I find that &#x27;sticky&#x27; dice always provide a nice narrative device, at least in narrative games. A character who&#x27;s player can&#x27;t seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.

          2. x______________ · · focus · HN ↗
            &gt; I&#x27;ve been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc

            Been through that this week as well with 100% success on 40% odds over multiple iterations on my game.. I tend to not dig into random but just rather ensure it works &#x27;as closely to intended&#x27; as possible..

      4. booty · · focus · HN ↗
        For years I used to buy ASICS running shoes. Every year they released a new model of each shoe: &quot;Nimbus 23&quot;, then &quot;Nimbus 24&quot; the next year, etc. And every year people would complain in the user reviews about how each shoe was worse than the last.

        I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They&#x27;ve been getting continuously worse for 24 consecutive years!

        Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn&#x27;t have much to say about them.

        (The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it&#x27;s cushy -- you&#x27;re comparing a new sneaker to your old sneaker where the foam had lost its bounce...)

        1. unshavedyak · · focus · HN ↗
          That&#x27;s kinda me with respect to Claude. Generally i&#x27;ve had no issues and just kept pluggin&#x27; along.

          The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i&#x27;d be on OpenAI by now.

        2. yencabulator · · focus · HN ↗
          It&#x27;s pretty typical that physical-goods manufacturing &quot;optimizes the process&quot; to cut costs during years 1 &amp; 2 of manufacturing.

          Ikea is notorious for this: The early Billy bookcase had heavier veneer and sturdier construction early on, and was actually a really good purchase for the money. The later years replaced veneer with paper foil, used thinner shelves, frames, and backing panels, and was just significantly weaker.

          1. ComputerGuru · · focus · HN ↗
            Amazon Basics is incredible in this regard, they’ve optimized SKU identification down to a pipeline. They’ll essentially randomly pick items off their internal list of highest netting sales and test them to see how dependent they are on brand name recognition and price-quality signalling. To do this as efficiently as possible, they simply purchase a few hundred units of a high quality product in the space, stick it in an Amazon Basics box and list it on their site under their Amazon Basixs brand at a price they feel they can achieve via white labeling, and wait to see how it sells. The use of high quality items (with quality above what can actually be had at the listed price point for the duration of the experiment) means they are really only testing the user base’s willingness to forgo a brand name for the category in exchange for a discount. If it sells well, they then work on sourcing it in bulk as a white labeled item “for real”, while if it sells poorly they simply delist and move on.

            I (used to) buy pre-spliced&#x2F;terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3&#x2F;OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.

            To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.

            1. dragontamer · · focus · HN ↗
              I have a similar story.

              Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand &#x27;Eneloop&#x27;. Extremely good specs all around

              Recently though, they are still called Amazon Basics but no longer test like Eneloop. They&#x27;ve changed manufacturers for the worse and are hoping no one notices...

              1. hn_acc1 · · focus · HN ↗
                I did a deep dive into NiMH batteries a few years ago and concluded that most people felt similarly: you can often get &quot;good&quot; (similar to Eneloop) specs for a short run from almost any manufacturer at the outset (see Ikea Ladda batteries - suspected of relabeled Eneloop for a while, but now not as good), but consistent quality is pretty much only Eneloop or other name brand, with Eneloop generally being the best.

                In the interests of saving my sanity and time (it&#x27;s not free!) having to chase down which batch of which brand is &quot;good&quot; at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn&#x27;t worth it to me (I get paid reasonably well).

                1. dragontamer · · focus · HN ↗
                  Yeah, name brands like Energizers are consistent but slightly worse than Eneloop at roughly the same price.

                  I actually made a battery tester as a hobby electronics project. Fully dumps the energy while measuring mAh and seeing some measurements of internal resistance.

                  Something I did notice was that crappy battery chargers can permanently damage even Eneloops. So don&#x27;t cheap out on the chargers either.

                  The high speed chargers (2 hours or less) are on the edge of what is safe and could permanently damage a cell. The 4 hours or slower chargers are just way more reliable. There&#x27;s probably someone out there testing different chargers (trying to find the safe fast chargers) but given how cheap NiHMs are (even the &quot;expensive&quot; Eneloops), it&#x27;s just easier to use 4 hour or 8 hour chargers and have masses and masses of extra NiHMs laying around.

                  1. ComputerGuru · · focus · HN ↗
                    I agree about the chargers!

                    Im sure you know this, but do note that fully dumping the charge can a) damage the battery and dramatically shorten its lifespan, b) give you different results depending on your discharge rate and the batteries you are testing, as they all have different current-dependent discharge curves.

                    1. dragontamer · · focus · HN ↗
                      Oh yeah. I stop at 0.95V. which I consider fully discharged (but not so low that I&#x27;m damaging the cells). I&#x27;m about 10% off the manufacturers claimed specs but I&#x27;d rather not push the limits of testing. That&#x27;s low enough to differentiate between good and bad cells

                      The discharge rate is currently regulated by only a 2.2 Ohm resistor. The next version of this circuit will be a constant current drain circuit to make all tests draw the same mA across the whole test.

                      Right now more current is drawn at 1.35V start and far less current is drawn at 0.95V end of test.

                      --------

                      The part that they don&#x27;t tell you is that 0.95V isn&#x27;t one point. When you disconnect the cell, it charges back up (kinda like a capacitor). This process can take multiple minutes (!!!!). I&#x27;ve defined the end as being lower than 0.95V for more than 30 seconds, even if disconnected.

              2. ComputerGuru · · focus · HN ↗
                No, that’s just white labeling. Like you said though, they can switch manufacturers and pull the rug from underneath you at any time.

                My example is like if you opened the Amazon Basics box and found Eneloop branded white&#x2F;black&#x2F;blue batteries directly.

        3. MikeTheGreat · · focus · HN ↗
          I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).

          Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.

          There&#x27;s also &quot;the Schlitz Mistake&quot;, which I&#x27;ve heard summarized as &quot;most customers won&#x27;t notice if you take your product&#x27;s quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)&quot;

          At this point I kinda assume that any company releasing year updates to a physical product that _doesn&#x27;t_ take the opportunity to trim costs &#x2F; reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...

          1. aesthesia · · focus · HN ↗
            I&#x27;m curious how common such lawsuits are. I don&#x27;t think I&#x27;ve heard of any specific instances where shareholders sued because a company didn&#x27;t make the product worse.
            1. MikeTheGreat · · focus · HN ↗
              On the one hand my assumption is mostly hyperbole

              and

              On the other hand that idea that &quot;companies exist to make money for shareholders, to ONLY make money for shareholders, and doing any other than maximizing shareholder returns is bad&quot; is pretty commonly accepted (and, I believe, enshrined in US law)

              1. skinfaxi · · focus · HN ↗
                Your belief is misplaced and incorrect.
          2. michaelmrose · · focus · HN ↗

            [dead]

          3. ehe78qhe · · focus · HN ↗
            &gt; Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.

            It&#x27;s been much longer than that: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Toblerone#2016_size_changes" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Toblerone#2016_size_changes

            1. xnorswap · · focus · HN ↗
              They handled that comically badly, the proportions meant the gaps went from being perceived as small to being a feature:

              <a href="https:&#x2F;&#x2F;www.bbc.co.uk&#x2F;news&#x2F;uk-44910195" rel="nofollow">https:&#x2F;&#x2F;www.bbc.co.uk&#x2F;news&#x2F;uk-44910195

              Visually it looks like they took out half the peaks.

              Your mind tells you that for every gap there used to be a peak there, regardless of the truth.

          4. m10i · · focus · HN ↗
            &gt; &quot;most customers won&#x27;t notice if you take your product&#x27;s quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)&quot;

            Summarizes the video game industry pretty well

          5. pnt12 · · focus · HN ↗
            Nitpick: I think that&#x27;s called skimpflation (worse quality), a friend of shrinkflation (less quantity) or inflation (higher prices).
            1. MikeTheGreat · · focus · HN ↗
              Today I learned a new word! Thanks!
          6. booty · · focus · HN ↗
            &gt; I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).

            That&#x27;s certainly common!

            I think ASICS&#x27; running shoes were a fun example of where this was probably not the case.

            - I certainly didn&#x27;t notice a difference in that time, though I&#x27;m admittedly not much of an actual runner

            - Serious runners might notice small technical differences, but the negative user reviews didn&#x27;t seem to indicate those were the people making the complaints

            - The overall user reviews remained positive

            - The competition in the shoe market is incredibly fierce; I&#x27;m not sure a brand could tank their quality and survive for long

            - Let&#x27;s not forget the other big variable: the wearers&#x27; bodies, particularly their feet. Now, those definitely do change over time -- most often for the worse, sadly!

            - I doubt anybody was blind A&#x2F;Bing a pair of Nimbus 20 against a pair of Nimbus 21 or 22 or 23 or 24. At best, a longtime Nimbus buyer is probably comparing a brand new pair of e.g. Nimbus 24 against their degraded but broken-in Nimbus 23 and their memory of how the Nimbus 23 felt when new. (And their body is a year or two older at that point..)

            - Because it&#x27;s such a long-running line of shoes, there could certainly be year-to-year variations... but it&#x27;s hard to imagine there was an actual 5 or 10 or 25 year downward slope. I mean, otherwise at that point the shoes would just be instantly injuring you or falling apart in a week

            - Also because it&#x27;s such a long-running product line and (aside from bleeding-edge professional marathon&#x2F;track shoes) sneaker manufacturing in general kind of seems like a solved problem... it seems like all of the possible cost optimizations have already been optimized. I don&#x27;t really think there&#x27;s much of a manufacturing or bill-of-materials cost difference between $20 running shoes and $200 running shoes anyway -- I&#x27;d be pretty surprised if &quot;quality shrinkflation&quot; was really much of a viable way for ASICS to save a few pennies.

        4. ck2 · · focus · HN ↗
          actually they were getting worse each year

          most running shoe series get heavier year to year as manufacturers turn to cheaper materials and add cushioning to try to attract more adopters

          it&#x27;s almost universal, very few manufacturers seem to be able to resist tampering

          (heavier shoes are slower, every three ounces is equal to another vo2max point lost)

          1. reducesuffering · · focus · HN ↗
            You&#x27;re literally doing what GP is describing. We have objective data on running shoes on: <a href="https:&#x2F;&#x2F;runrepeat.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;runrepeat.com&#x2F;

            The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.

            1. ck2 · · focus · HN ↗
              nope

              you mean NEW MODELS are being introduced with lighter faster foams

              not the same model year to year

              modern example: Saucony Endorphin Speed

              v1 in 2020 was award winning

              v2 in 2021 was almost the same, more praise

              v3 bleh

              v4 v5 bleh bleh

              they cannot resist tampering

        5. dgacmu · · focus · HN ↗
          Except for that one year when they completely swapped the meaning of the Cumulus line, which I think was 2008 with the cumulus 9 to 10 transition. It went from a neutral shoe that was good for people with high arches to more of a stiff stability shoe. The complainers aren&#x27;t _always_ crazy. :-) (That doesn&#x27;t mean the cumulus 10 was worse, of course, it just was a more substantial change that affected the type of runner the shoe was designed for.)

          Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I&#x27;m old and heavier and run in the Nimbus and am happy again. But you&#x27;re right, of course, that most of the model changes are just fine and people like to complain.

      5. oooyay · · focus · HN ↗
        &gt; Perceived performance is actual performance over expectations and the latter just keeps increasing over time.

        This is true of all reliability and performance paradigms, incidentally

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.