They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
For years I used to buy ASICS running shoes. Every year they released a new model of each shoe: "Nimbus 23", then "Nimbus 24" the next year, etc. And every year people would complain in the user reviews about how each shoe was worse than the last.
I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They've been getting continuously worse for 24 consecutive years!
Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn't have much to say about them.
(The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it's cushy -- you're comparing a new sneaker to your old sneaker where the foam had lost its bounce...)
It's pretty typical that physical-goods manufacturing "optimizes the process" to cut costs during years 1 & 2 of manufacturing.
Ikea is notorious for this: The early Billy bookcase had heavier veneer and sturdier construction early on, and was actually a really good purchase for the money. The later years replaced veneer with paper foil, used thinner shelves, frames, and backing panels, and was just significantly weaker.
Amazon Basics is incredible in this regard, they’ve optimized SKU identification down to a pipeline. They’ll essentially randomly pick items off their internal list of highest netting sales and test them to see how dependent they are on brand name recognition and price-quality signalling. To do this as efficiently as possible, they simply purchase a few hundred units of a high quality product in the space, stick it in an Amazon Basics box and list it on their site under their Amazon Basixs brand at a price they feel they can achieve via white labeling, and wait to see how it sells. The use of high quality items (with quality above what can actually be had at the listed price point for the duration of the experiment) means they are really only testing the user base’s willingness to forgo a brand name for the category in exchange for a discount. If it sells well, they then work on sourcing it in bulk as a white labeled item “for real”, while if it sells poorly they simply delist and move on.
I (used to) buy pre-spliced/terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3/OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.
To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.
Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand 'Eneloop'. Extremely good specs all around
Recently though, they are still called Amazon Basics but no longer test like Eneloop. They've changed manufacturers for the worse and are hoping no one notices...
I did a deep dive into NiMH batteries a few years ago and concluded that most people felt similarly: you can often get "good" (similar to Eneloop) specs for a short run from almost any manufacturer at the outset (see Ikea Ladda batteries - suspected of relabeled Eneloop for a while, but now not as good), but consistent quality is pretty much only Eneloop or other name brand, with Eneloop generally being the best.
In the interests of saving my sanity and time (it's not free!) having to chase down which batch of which brand is "good" at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn't worth it to me (I get paid reasonably well).
Yeah, name brands like Energizers are consistent but slightly worse than Eneloop at roughly the same price.
I actually made a battery tester as a hobby electronics project. Fully dumps the energy while measuring mAh and seeing some measurements of internal resistance.
Something I did notice was that crappy battery chargers can permanently damage even Eneloops. So don't cheap out on the chargers either.
The high speed chargers (2 hours or less) are on the edge of what is safe and could permanently damage a cell. The 4 hours or slower chargers are just way more reliable. There's probably someone out there testing different chargers (trying to find the safe fast chargers) but given how cheap NiHMs are (even the "expensive" Eneloops), it's just easier to use 4 hour or 8 hour chargers and have masses and masses of extra NiHMs laying around.
Im sure you know this, but do note that fully dumping the charge can a) damage the battery and dramatically shorten its lifespan, b) give you different results depending on your discharge rate and the batteries you are testing, as they all have different current-dependent discharge curves.
Oh yeah. I stop at 0.95V. which I consider fully discharged (but not so low that I'm damaging the cells). I'm about 10% off the manufacturers claimed specs but I'd rather not push the limits of testing. That's low enough to differentiate between good and bad cells
The discharge rate is currently regulated by only a 2.2 Ohm resistor. The next version of this circuit will be a constant current drain circuit to make all tests draw the same mA across the whole test.
Right now more current is drawn at 1.35V start and far less current is drawn at 0.95V end of test.
--------
The part that they don't tell you is that 0.95V isn't one point. When you disconnect the cell, it charges back up (kinda like a capacitor). This process can take multiple minutes (!!!!). I've defined the end as being lower than 0.95V for more than 30 seconds, even if disconnected.
I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).
Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
There's also "the Schlitz Mistake", which I've heard summarized as "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
At this point I kinda assume that any company releasing year updates to a physical product that _doesn't_ take the opportunity to trim costs / reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...
I'm curious how common such lawsuits are. I don't think I've heard of any specific instances where shareholders sued because a company didn't make the product worse.
On the other hand that idea that "companies exist to make money for shareholders, to ONLY make money for shareholders, and doing any other than maximizing shareholder returns is bad" is pretty commonly accepted (and, I believe, enshrined in US law)
> Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
It's been much longer than that: <a href="https://en.wikipedia.org/wiki/Toblerone#2016_size_changes" rel="nofollow">https://en.wikipedia.org/wiki/Toblerone#2016_size_changes
> "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
> I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).
That's certainly common!
I think ASICS' running shoes were a fun example of where this was probably not the case.
- I certainly didn't notice a difference in that time, though I'm admittedly not much of an actual runner
- Serious runners might notice small technical differences, but the negative user reviews didn't seem to indicate those were the people making the complaints
- The overall user reviews remained positive
- The competition in the shoe market is incredibly fierce; I'm not sure a brand could tank their quality and survive for long
- Let's not forget the other big variable: the wearers' bodies, particularly their feet. Now, those definitely do change over time -- most often for the worse, sadly!
- I doubt anybody was blind A/Bing a pair of Nimbus 20 against a pair of Nimbus 21 or 22 or 23 or 24. At best, a longtime Nimbus buyer is probably comparing a brand new pair of e.g. Nimbus 24 against their degraded but broken-in Nimbus 23 and their memory of how the Nimbus 23 felt when new. (And their body is a year or two older at that point..)
- Because it's such a long-running line of shoes, there could certainly be year-to-year variations... but it's hard to imagine there was an actual 5 or 10 or 25 year downward slope. I mean, otherwise at that point the shoes would just be instantly injuring you or falling apart in a week
- Also because it's such a long-running product line and (aside from bleeding-edge professional marathon/track shoes) sneaker manufacturing in general kind of seems like a solved problem... it seems like all of the possible cost optimizations have already been optimized. I don't really think there's much of a manufacturing or bill-of-materials cost difference between $20 running shoes and $200 running shoes anyway -- I'd be pretty surprised if "quality shrinkflation" was really much of a viable way for ASICS to save a few pennies.
You're literally doing what GP is describing. We have objective data on running shoes on: <a href="https://runrepeat.com/" rel="nofollow">https://runrepeat.com/
The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.
Except for that one year when they completely swapped the meaning of the Cumulus line, which I think was 2008 with the cumulus 9 to 10 transition. It went from a neutral shoe that was good for people with high arches to more of a stiff stability shoe. The complainers aren't _always_ crazy. :-) (That doesn't mean the cumulus 10 was worse, of course, it just was a more substantial change that affected the type of runner the shoe was designed for.)
Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I'm old and heavier and run in the Nimbus and am happy again. But you're right, of course, that most of the model changes are just fine and people like to complain.
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
nsarrazin · · focus · HN ↗
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
booty · · focus · HN ↗
I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They've been getting continuously worse for 24 consecutive years!
Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn't have much to say about them.
(The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it's cushy -- you're comparing a new sneaker to your old sneaker where the foam had lost its bounce...)
unshavedyak · · focus · HN ↗
The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i'd be on OpenAI by now.
yencabulator · · focus · HN ↗
Ikea is notorious for this: The early Billy bookcase had heavier veneer and sturdier construction early on, and was actually a really good purchase for the money. The later years replaced veneer with paper foil, used thinner shelves, frames, and backing panels, and was just significantly weaker.
ComputerGuru · · focus · HN ↗
I (used to) buy pre-spliced/terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3/OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.
To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.
dragontamer · · focus · HN ↗
Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand 'Eneloop'. Extremely good specs all around
Recently though, they are still called Amazon Basics but no longer test like Eneloop. They've changed manufacturers for the worse and are hoping no one notices...
hn_acc1 · · focus · HN ↗
In the interests of saving my sanity and time (it's not free!) having to chase down which batch of which brand is "good" at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn't worth it to me (I get paid reasonably well).
dragontamer · · focus · HN ↗
I actually made a battery tester as a hobby electronics project. Fully dumps the energy while measuring mAh and seeing some measurements of internal resistance.
Something I did notice was that crappy battery chargers can permanently damage even Eneloops. So don't cheap out on the chargers either.
The high speed chargers (2 hours or less) are on the edge of what is safe and could permanently damage a cell. The 4 hours or slower chargers are just way more reliable. There's probably someone out there testing different chargers (trying to find the safe fast chargers) but given how cheap NiHMs are (even the "expensive" Eneloops), it's just easier to use 4 hour or 8 hour chargers and have masses and masses of extra NiHMs laying around.
ComputerGuru · · focus · HN ↗
Im sure you know this, but do note that fully dumping the charge can a) damage the battery and dramatically shorten its lifespan, b) give you different results depending on your discharge rate and the batteries you are testing, as they all have different current-dependent discharge curves.
dragontamer · · focus · HN ↗
The discharge rate is currently regulated by only a 2.2 Ohm resistor. The next version of this circuit will be a constant current drain circuit to make all tests draw the same mA across the whole test.
Right now more current is drawn at 1.35V start and far less current is drawn at 0.95V end of test.
--------
The part that they don't tell you is that 0.95V isn't one point. When you disconnect the cell, it charges back up (kinda like a capacitor). This process can take multiple minutes (!!!!). I've defined the end as being lower than 0.95V for more than 30 seconds, even if disconnected.
ComputerGuru · · focus · HN ↗
My example is like if you opened the Amazon Basics box and found Eneloop branded white/black/blue batteries directly.
MikeTheGreat · · focus · HN ↗
Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
There's also "the Schlitz Mistake", which I've heard summarized as "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
At this point I kinda assume that any company releasing year updates to a physical product that _doesn't_ take the opportunity to trim costs / reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...
aesthesia · · focus · HN ↗
MikeTheGreat · · focus · HN ↗
and
On the other hand that idea that "companies exist to make money for shareholders, to ONLY make money for shareholders, and doing any other than maximizing shareholder returns is bad" is pretty commonly accepted (and, I believe, enshrined in US law)
skinfaxi · · focus · HN ↗
michaelmrose · · focus · HN ↗
[dead]
ehe78qhe · · focus · HN ↗
It's been much longer than that: <a href="https://en.wikipedia.org/wiki/Toblerone#2016_size_changes" rel="nofollow">https://en.wikipedia.org/wiki/Toblerone#2016_size_changes
xnorswap · · focus · HN ↗
<a href="https://www.bbc.co.uk/news/uk-44910195" rel="nofollow">https://www.bbc.co.uk/news/uk-44910195
Visually it looks like they took out half the peaks.
Your mind tells you that for every gap there used to be a peak there, regardless of the truth.
m10i · · focus · HN ↗
Summarizes the video game industry pretty well
pnt12 · · focus · HN ↗
MikeTheGreat · · focus · HN ↗
booty · · focus · HN ↗
That's certainly common!
I think ASICS' running shoes were a fun example of where this was probably not the case.
- I certainly didn't notice a difference in that time, though I'm admittedly not much of an actual runner
- Serious runners might notice small technical differences, but the negative user reviews didn't seem to indicate those were the people making the complaints
- The overall user reviews remained positive
- The competition in the shoe market is incredibly fierce; I'm not sure a brand could tank their quality and survive for long
- Let's not forget the other big variable: the wearers' bodies, particularly their feet. Now, those definitely do change over time -- most often for the worse, sadly!
- I doubt anybody was blind A/Bing a pair of Nimbus 20 against a pair of Nimbus 21 or 22 or 23 or 24. At best, a longtime Nimbus buyer is probably comparing a brand new pair of e.g. Nimbus 24 against their degraded but broken-in Nimbus 23 and their memory of how the Nimbus 23 felt when new. (And their body is a year or two older at that point..)
- Because it's such a long-running line of shoes, there could certainly be year-to-year variations... but it's hard to imagine there was an actual 5 or 10 or 25 year downward slope. I mean, otherwise at that point the shoes would just be instantly injuring you or falling apart in a week
- Also because it's such a long-running product line and (aside from bleeding-edge professional marathon/track shoes) sneaker manufacturing in general kind of seems like a solved problem... it seems like all of the possible cost optimizations have already been optimized. I don't really think there's much of a manufacturing or bill-of-materials cost difference between $20 running shoes and $200 running shoes anyway -- I'd be pretty surprised if "quality shrinkflation" was really much of a viable way for ASICS to save a few pennies.
ck2 · · focus · HN ↗
most running shoe series get heavier year to year as manufacturers turn to cheaper materials and add cushioning to try to attract more adopters
it's almost universal, very few manufacturers seem to be able to resist tampering
(heavier shoes are slower, every three ounces is equal to another vo2max point lost)
reducesuffering · · focus · HN ↗
The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.
ck2 · · focus · HN ↗
you mean NEW MODELS are being introduced with lighter faster foams
not the same model year to year
modern example: Saucony Endorphin Speed
v1 in 2020 was award winning
v2 in 2021 was almost the same, more praise
v3 bleh
v4 v5 bleh bleh
they cannot resist tampering
dgacmu · · focus · HN ↗
Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I'm old and heavier and run in the Nimbus and am happy again. But you're right, of course, that most of the model changes are just fine and people like to complain.