Livenerf: Has Opus 5.5 been nerfed yet?
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Livenerf: Has Opus 5.5 been nerfed yet?
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
gigatexal · · focus · HN ↗
aabhay · · focus · HN ↗
refulgentis · · focus · HN ↗
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
bbg2401 · · focus · HN ↗
It’s frustrating to observe communities made up of smart, professional individuals as they behave like spoiled children on the day after Christmas when new toy novelty has begun to wane.
I understand it’s relatively harmless but for goodness sake, take a step back and appreciate what you have instead of immediately wanting the thrill of a newer model. Slow down and do deliberate work to get the most out of these amazing tools. Don’t just live off the thrill of finding something marginally better than what you have.
mattgreenrocks · · focus · HN ↗
vikramkr · · focus · HN ↗
_pushes model further_
"Ugh why is this model failing now even though I'm asking it to do harder things and also got sloppier with my prompts because I got used to it being able to figure stuff out"
aloukissas · · focus · HN ↗
ninjahawk1 · · focus · HN ↗
Also, it isn’t comparing one 10-day period once and calling it done. The window rolls forward daily, and a change has to clear the pre-registered 99% threshold in two consecutive windows before it’s flagged.
Ten days isn’t sacred, though. Once there’s enough longitudinal data, one of the things I want to evaluate is whether that window length is actually well calibrated or should be changed in a future version.
judge2020 · · focus · HN ↗
solenoid0937 · · focus · HN ↗
solfox · · focus · HN ↗
nba456_ · · focus · HN ↗
mcmcmc · · focus · HN ↗
omani · · focus · HN ↗
nba456_ · · focus · HN ↗
AnimalMuppet · · focus · HN ↗
solenoid0937 · · focus · HN ↗
voiceeh · · focus · HN ↗
nba456_ · · focus · HN ↗
doginasuit · · focus · HN ↗
redanddead · · focus · HN ↗
Kiro · · focus · HN ↗
nimchimpsky · · focus · HN ↗
[dead]
Craighead · · focus · HN ↗
bradfa · · focus · HN ↗
jyoung8607 · · focus · HN ↗
If so, please share. This should be measurable, and I'm glad this project is measuring it.
Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.
dannyw · · focus · HN ↗
dude250711 · · focus · HN ↗
empath75 · · focus · HN ↗
wccrawford · · focus · HN ↗
raincole · · focus · HN ↗
jascha_eng · · focus · HN ↗
It would be economical suicide from anthropic and OpenAI to actually need models intentionally.
But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.
sumedh · · focus · HN ↗
Links have been provided by others in this post.
Kiro · · focus · HN ↗
sumedh · · focus · HN ↗
Ant denied them at first though.
solfox · · focus · HN ↗
After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.
I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
onemoresoop · · focus · HN ↗
lxgr · · focus · HN ↗
whs · · focus · HN ↗
madeofpalk · · focus · HN ↗
ENGNR · · focus · HN ↗
Computer0 · · focus · HN ↗
dannyw · · focus · HN ↗
gr_norm · · focus · HN ↗
The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.
nico · · focus · HN ↗
The quality of the output/work seems the same, but the speed at which is gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
Computer0 · · focus · HN ↗
pkaye · · focus · HN ↗
nsonha · · focus · HN ↗
jacquesm · · focus · HN ↗
layla5alive · · focus · HN ↗
dannyw · · focus · HN ↗
I’m sure it’s lowkey intentional, probably encourages users to spin up new chats; hence less context.
mlh496 · · focus · HN ↗
fendy3002 · · focus · HN ↗
ricardobeat · · focus · HN ↗
dagss · · focus · HN ↗
Have a look into Docker sbx for instance.
LeoPanthera · · focus · HN ↗
ninjahawk1 · · focus · HN ↗
[dead]
colordrops · · focus · HN ↗
ninjahawk1 · · focus · HN ↗
[dead]
srnvs · · focus · HN ↗
[dead]
Razengan · · focus · HN ↗
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
jug · · focus · HN ↗
<a href="https://www.bridgebench.ai/nerf-bench" rel="nofollow">https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Grimblewald · · focus · HN ↗
eulgro · · focus · HN ↗
r_lee · · focus · HN ↗
jacquesm · · focus · HN ↗
martin- · · focus · HN ↗
Grimblewald · · focus · HN ↗
Rapzid · · focus · HN ↗
Of course it's almost entirely unsubstantiated BS.
fbrncci · · focus · HN ↗
Rapzid · · focus · HN ↗
somenameforme · · focus · HN ↗
Rapzid · · focus · HN ↗
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
somenameforme · · focus · HN ↗
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
Rapzid · · focus · HN ↗
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
airstrike · · focus · HN ↗
Not too mention these companies could easily offer one product to enteprises and another to everyone else
Model nerfing is real
Rapzid · · focus · HN ↗
I was so certain it wasn't, based on the complete lack of evidence.
But then you said it's real. NVM, I don't need evidence! Somebody said it's real!
This place has fallen off.
grim_io · · focus · HN ↗
airstrike · · focus · HN ↗
mrandish · · focus · HN ↗
The claim is that they nerf subscription accounts not API.
applicative · · focus · HN ↗
troupo · · focus · HN ↗
Until shit like this: <a href="https://www.anthropic.com/engineering/april-23-postmortem" rel="nofollow">https://www.anthropic.com/engineering/april-23-postmortem
Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks "fixing it".
Or shit like this: <a href="https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues" rel="nofollow">https://www.anthropic.com/engineering/a-postmortem-of-three-...
Anthropic is in a perpetual state of "oops, these 'bugs' degraded our model quality" and only admit the issues when it's immediately obvious and visibly affects a large number of customers.
Otherwise all open benchmarks can be (and are) gamed. And it's quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.
lxgr · · focus · HN ↗
It's like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.
somenameforme · · focus · HN ↗
Banks have far greater transparency and legal requirements. LLM companies are just delivering a black box that they have complete control over. And given the regular 'How's Claude doing this session?' stuff, it's almost certain that they're A-B testing various tweaks on a per session basis.
applicative · · focus · HN ↗
stackghost · · focus · HN ↗
jackmott42 · · focus · HN ↗
stackghost · · focus · HN ↗
> in short, yall dumb, shut up.
no u
Rapzid · · focus · HN ↗
Everyone wants to be a software engineer, until it's time to do software engineering shit.
You know, like scientific method shit we learned in 5th/6th grade.
It's the great bro science incursion.
stackghost · · focus · HN ↗
The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.
scrollop · · focus · HN ↗
this one has been around for over a year
Razengan · · focus · HN ↗
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
Centigonal · · focus · HN ↗
latentsea · · focus · HN ↗
btown · · focus · HN ↗
nightpool · · focus · HN ↗
zxilly · · focus · HN ↗
JohnBooty · · focus · HN ↗
I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?
I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs."
btown · · focus · HN ↗
And, while you might be billed fewer tokens as a result (because the lower thinking would result in less investigatory work), you might not know this is happening, and know to dial up effort accordingly - you'd simply get a worse work product. And certainly, Anthropic's incentive for anyone on a subscription is to push this as aggressively as they can, so people use less of that subscription.
Sadly, I'd also expect that the OP's benchmark will be detected as a test of model capabilities, and thus be given a high classification so that this strategy remains undetected.
poizan42 · · focus · HN ↗
jackmott42 · · focus · HN ↗
fuck
Razengan · · focus · HN ↗
comboy · · focus · HN ↗
andriy_koval · · focus · HN ↗
hamandcheese · · focus · HN ↗
andriy_koval · · focus · HN ↗
user3939382 · · focus · HN ↗
braingravy · · focus · HN ↗
LimitExperience · · focus · HN ↗
[dead]
apitman · · focus · HN ↗
ffsm8 · · focus · HN ↗
adastra22 · · focus · HN ↗
reubenmorais · · focus · HN ↗
TeMPOraL · · focus · HN ↗
ffsm8 · · focus · HN ↗
<a href="https://code.claude.com/docs/en/monitoring-usage" rel="nofollow">https://code.claude.com/docs/en/monitoring-usage
thanks for correcting me on that regard
Aeolun · · focus · HN ↗
apitman · · focus · HN ↗
jacquesm · · focus · HN ↗
I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.
Imagine the power company being able to decide how much you consume and at which price point.
herval · · focus · HN ↗
Airlines, banks, health insurance…
tccole · · focus · HN ↗
petesergeant · · focus · HN ↗
herval · · focus · HN ↗
herval · · focus · HN ↗
pixelready · · focus · HN ↗
miohtama · · focus · HN ↗
msdz · · focus · HN ↗
> Step 2: Regulatory Capture
is being worked towards.
TeMPOraL · · focus · HN ↗
Are they though? Or is it just what some companies would want them to be?
none_to_remain · · focus · HN ↗
Turskarama · · focus · HN ↗
adastra22 · · focus · HN ↗
slim · · focus · HN ↗
TeMPOraL · · focus · HN ↗
topspin · · focus · HN ↗
I don't know if that's the actual origin of the term nerf, but it was the first time I'd heard it.
done_lurking · · focus · HN ↗
CodesInChaos · · focus · HN ↗
csomar · · focus · HN ↗
When you're running something at a loss, you can mistreat your customers and they'll still stick around (I'm an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.
My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they'll either drop the subscription model or it'll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.
It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible, even if they're delivering real business value (coding, research, etc.). In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
TeMPOraL · · focus · HN ↗
Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.
Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you're producing for/serving millions of people. This is not unusual.
> In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Many say that, but I sincerely doubt they'd actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring / bullshit parts of daily work to give up on merely 3-5x price increase.
csomar · · focus · HN ↗
Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don't know, we can only speculate.
> Many say that, but I sincerely doubt they'd actually follow through.
I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I'll be using it less.
Look, I am not saying that there isn't a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don't have a business model here.
FeepingCreature · · focus · HN ↗
You can assume Openrouter open-model providers serve at or above margin, because there's no branding so there's no reason to do it unless you can be profitable. If the Anthropic models are anywhere in that ballpark, they're very comfortably profitable on API.
airspresso · · focus · HN ↗
That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won't be fast and won't be frontier intelligence, but definitely useful.
zozbot234 · · focus · HN ↗
cavoirom · · focus · HN ↗
icepush · · focus · HN ↗
CodesInChaos · · focus · HN ↗
Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?
I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.
user3939382 · · focus · HN ↗
avazhi · · focus · HN ↗
bdlowery · · focus · HN ↗
this bench was just released, it couldn't have detected opus 4.6 degradation.
nullbio · · focus · HN ↗
I also wonder if cache could be used to throw these off as well, where it's serving un-nerfed cache results for context windows that are identical to ones they've previously had for benchmark requests.
Seems like the only way to do it well would be to have some randomness involved that couldn't be cheated on - but you'd want to do it in a way that doesn't throw out the benchmarks too much, so your results can be compared still.
Gabrys1 · · focus · HN ↗
All that energy wasted... could just drive a big V8 instead and make less money for the big tech
scrollop · · focus · HN ↗
<a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/
sscaryterry · · focus · HN ↗
phoghed · · focus · HN ↗
lxgr · · focus · HN ↗
This alone makes the benchmark unsound.
shawabawa3 · · focus · HN ↗
rplnt · · focus · HN ↗
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
transcriptase · · focus · HN ↗
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
setopt · · focus · HN ↗
natpalmer1776 · · focus · HN ↗
lukan · · focus · HN ↗
You have to pretend to be angry in such situations?
katzenq · · focus · HN ↗
AbsurdCensor · · focus · HN ↗
katzenq · · focus · HN ↗
sejje · · focus · HN ↗
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
Humans should use more llms.
smurf9852 · · focus · HN ↗
silversmith · · focus · HN ↗
Barbing · · focus · HN ↗
Anthropic cut a deal with SpaceXAI in May - $1.25b/mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.
<a href="https://www.anthropic.com/news/higher-limits-spacex" rel="nofollow">https://www.anthropic.com/news/higher-limits-spacex
rplnt · · focus · HN ↗
This is what pissed me off the most. Make it slower, rate limit it, move the credits to other time slots, idk.. but returning BAD results? That's the worst approach you could take.
jclardy · · focus · HN ↗
zsoltkacsandi · · focus · HN ↗
rednb · · focus · HN ↗
Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i've transitioned to using DS 4.1 as a daily driver and am very satisfied.
My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.
dotancohen · · focus · HN ↗
rednb · · focus · HN ↗
Except maybe that the model often believes that he is running out of context, and needs to rush so i occasionally need to jump in to tell it that it still has plenty of room left.
But this does not degrade the quality of my overall experience in a meaningful way.
dotancohen · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
KronisLV · · focus · HN ↗
zerop · · focus · HN ↗
StableAlkyne · · focus · HN ↗
For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.
The game is to balance between the savings of quantization and making the model dumb enough the people notice
sscaryterry · · focus · HN ↗
ForHackernews · · focus · HN ↗
<a href="https://www.fool.com/investing/2026/07/25/spacexs-performance-looks-almost-identical-to-past/" rel="nofollow">https://www.fool.com/investing/2026/07/25/spacexs-performanc...
sigbottle · · focus · HN ↗
quikoa · · focus · HN ↗
nsarrazin · · focus · HN ↗
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
waterproof · · focus · HN ↗
`percieved_performance = actual_perf/expectation`
`expectation` is an increasing function over time.
`actual_perf` is a stochastic function of the model's true ability, context, etc. -> a recipe for some bad sessions.
As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive "runs" because our brains love to find patterns.
causal · · focus · HN ↗
svachalek · · focus · HN ↗
goodmythical · · focus · HN ↗
e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"
shagie · · focus · HN ↗
svachalek · · focus · HN ↗
goodmythical · · focus · HN ↗
Also, consider that in flipping 10 coins, you'll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you'll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...
Widening the range from "rolling exactly 15" to "rolls 15 or 16" or "rolls between 14-17" makes the strings even more likely as you're doubling the success rate from "only 9 15s" to the "any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s" space.
To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:
For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:
General Variables: N = total number of rolls/trials k = target streak length p = probability of getting the target outcome (e.g., 1/20 for a specific roll on d20 or 1/10 for two specific results) q = probability of getting any other outcome (1 - p)
Expected runs of AT LEAST length k: E(runs >= k) = p^k (1 + (N - k) * q)
Expected runs of EXACT length k: E(exact k) = p^k * q * (2 + (N - k - 1) * q)
Personally, I find that 'sticky' dice always provide a nice narrative device, at least in narrative games. A character who's player can't seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.
x______________ · · focus · HN ↗
Been through that this week as well with 100% success on 40% odds over multiple iterations on my game.. I tend to not dig into random but just rather ensure it works 'as closely to intended' as possible..
booty · · focus · HN ↗
I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They've been getting continuously worse for 24 consecutive years!
Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn't have much to say about them.
(The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it's cushy -- you're comparing a new sneaker to your old sneaker where the foam had lost its bounce...)
unshavedyak · · focus · HN ↗
The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i'd be on OpenAI by now.
yencabulator · · focus · HN ↗
ComputerGuru · · focus · HN ↗
I (used to) buy pre-spliced/terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3/OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.
To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.
dragontamer · · focus · HN ↗
Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand 'Eneloop'. Extremely good specs all around
Recently though, they are still called Amazon Basics but no longer test like Eneloop. They've changed manufacturers for the worse and are hoping no one notices...
hn_acc1 · · focus · HN ↗
In the interests of saving my sanity and time (it's not free!) having to chase down which batch of which brand is "good" at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn't worth it to me (I get paid reasonably well).
dragontamer · · focus · HN ↗
ComputerGuru · · focus · HN ↗
Im sure you know this, but do note that fully dumping the charge can a) damage the battery and dramatically shorten its lifespan, b) give you different results depending on your discharge rate and the batteries you are testing, as they all have different current-dependent discharge curves.
dragontamer · · focus · HN ↗
ComputerGuru · · focus · HN ↗
My example is like if you opened the Amazon Basics box and found Eneloop branded white/black/blue batteries directly.
MikeTheGreat · · focus · HN ↗
Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
There's also "the Schlitz Mistake", which I've heard summarized as "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
At this point I kinda assume that any company releasing year updates to a physical product that _doesn't_ take the opportunity to trim costs / reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...
aesthesia · · focus · HN ↗
MikeTheGreat · · focus · HN ↗
and
On the other hand that idea that "companies exist to make money for shareholders, to ONLY make money for shareholders, and doing any other than maximizing shareholder returns is bad" is pretty commonly accepted (and, I believe, enshrined in US law)
skinfaxi · · focus · HN ↗
michaelmrose · · focus · HN ↗
[dead]
ehe78qhe · · focus · HN ↗
It's been much longer than that: <a href="https://en.wikipedia.org/wiki/Toblerone#2016_size_changes" rel="nofollow">https://en.wikipedia.org/wiki/Toblerone#2016_size_changes
xnorswap · · focus · HN ↗
<a href="https://www.bbc.co.uk/news/uk-44910195" rel="nofollow">https://www.bbc.co.uk/news/uk-44910195
Visually it looks like they took out half the peaks.
m10i · · focus · HN ↗
Summarizes the video game industry pretty well
pnt12 · · focus · HN ↗
MikeTheGreat · · focus · HN ↗
booty · · focus · HN ↗
That's certainly common!
I think ASICS' running shoes were a fun example of where this was probably not the case.
- I certainly didn't notice a difference in that time, though I'm admittedly not much of an actual runner
- Serious runners might notice small technical differences, but the negative user reviews didn't seem to indicate those were the people making the complaints
- The overall user reviews remained positive
- The competition in the shoe market is incredibly fierce; I'm not sure a brand could tank their quality and survive for long
- Let's not forget the other big variable: the wearers' bodies, particularly their feet. Now, those definitely do change over time -- most often for the worse, sadly!
- I doubt anybody was blind A/Bing a pair of Nimbus 20 against a pair of Nimbus 21 or 22 or 23 or 24. At best, a longtime Nimbus buyer is probably comparing a brand new pair of e.g. Nimbus 24 against their degraded but broken-in Nimbus 23 and their memory of how the Nimbus 23 felt when new. (And their body is a year or two older at that point..)
- Because it's such a long-running line of shoes, there could certainly be year-to-year variations... but it's hard to imagine there was an actual 5 or 10 or 25 year downward slope. I mean, otherwise at that point the shoes would just be instantly injuring you or falling apart in a week
- Also because it's such a long-running product line and (aside from bleeding-edge professional marathon/track shoes) sneaker manufacturing in general kind of seems like a solved problem... it seems like all of the possible cost optimizations have already been optimized. I don't really think there's much of a manufacturing or bill-of-materials cost difference between $20 running shoes and $200 running shoes anyway -- I'd be pretty surprised if "quality shrinkflation" was really much of a viable way for ASICS to save a few pennies.
ck2 · · focus · HN ↗
most running shoe series get heavier year to year as manufacturers turn to cheaper materials and add cushioning to try to attract more adopters
it's almost universal, very few manufacturers seem to be able to resist tampering
(heavier shoes are slower, every three ounces is equal to another vo2max point lost)
reducesuffering · · focus · HN ↗
The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.
ck2 · · focus · HN ↗
you mean NEW MODELS are being introduced with lighter faster foams
not the same model year to year
modern example: Saucony Endorphin Speed
v1 in 2020 was award winning
v2 in 2021 was almost the same, more praise
v3 bleh
v4 v5 bleh bleh
they cannot resist tampering
dgacmu · · focus · HN ↗
Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I'm old and heavier and run in the Nimbus and am happy again. But you're right, of course, that most of the model changes are just fine and people like to complain.
oooyay · · focus · HN ↗
This is true of all reliability and performance paradigms, incidentally
shados · · focus · HN ↗
loopydosuette · · focus · HN ↗
you open two files, before and after you notice a nerf, and from worse comments to logical oversights, it's all damn obvious.
don't normalize this make believe bullshit and misleading people who you think barely understand what they see anyway ...
you wouldn't even know if models had somehow timed nerfs hardcoded into them, however much control over the stack you have.
ridiculous
Aperocky · · focus · HN ↗
stbenjam · · focus · HN ↗
I am quite convinced that the whole nerfing phenomenon is 90% AI psychosis. I have the word muted on X.
cyberes · · focus · HN ↗
Barbing · · focus · HN ↗
stbenjam · · focus · HN ↗
throwitaway222 · · focus · HN ↗
jtrn · · focus · HN ↗
jdthedisciple · · focus · HN ↗
futune · · focus · HN ↗
octoberfranklin · · focus · HN ↗
Open models are the endgame.
paradox460 · · focus · HN ↗
octoberfranklin · · focus · HN ↗
The classifier is a model; it examines the actual prompt.
layla5alive · · focus · HN ↗
xpct · · focus · HN ↗
redanddead · · focus · HN ↗
johnfn · · focus · HN ↗
I made a graphic to explain why people feel like the models get nerfed:
<a href="https://x.com/thesilenceturns/status/2103551351825543610" rel="nofollow">https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
486sx33 · · focus · HN ↗
[dead]
hbn · · focus · HN ↗
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
Computer0 · · focus · HN ↗
johnfn · · focus · HN ↗
> It’s not a crazy conspiracy that the same model can be stupider
Sorry, I really do think it's a conspiracy. If this were the case, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.
frde_me · · focus · HN ↗
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
usef- · · focus · HN ↗
Grimblewald · · focus · HN ↗
My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.
How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?
gobdovan · · focus · HN ↗
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
johnfn · · focus · HN ↗
prodigycorp · · focus · HN ↗
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
simonw · · focus · HN ↗
Where?
prodigycorp · · focus · HN ↗
consumer451 · · focus · HN ↗
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
[0] <a href="https://status.claude.com/incidents/72f99lh1cj2c" rel="nofollow">https://status.claude.com/incidents/72f99lh1cj2c
[1] <a href="https://anthropic.com/engineering/a-postmortem-of-three-recent-issues" rel="nofollow">https://anthropic.com/engineering/a-postmortem-of-three-rece...
[2] <a href="https://texxr.com/handle/claudedevs" rel="nofollow">https://texxr.com/handle/claudedevs
source: <a href="https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2" rel="nofollow">https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2
what · · focus · HN ↗
There are no bugs, just happy little accidents.
consumer451 · · focus · HN ↗
Spooky23 · · focus · HN ↗
p-e-w · · focus · HN ↗
erinnh · · focus · HN ↗
p-e-w · · focus · HN ↗
Christ this forum has become intellectually dishonest.
jibalt · · focus · HN ↗
xnorswap · · focus · HN ↗
Firefox never really recovered here after that, anything Firefox did, for years, was met with, "but remember pocket".
jibalt · · focus · HN ↗
> There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs
erinnh · · focus · HN ↗
So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.
prodigycorp · · focus · HN ↗
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”.
winwang · · focus · HN ↗
computerex · · focus · HN ↗
Kiro · · focus · HN ↗
QwenGlazer9000 · · focus · HN ↗
Unintentional tbf.
weird-eye-issue · · focus · HN ↗
ffsm8 · · focus · HN ↗
fwiw, i think almost all regressions are down to a/b testing in the harness by anthropic, but it is objectively indistinguishable beyond "the coding agent ceases to be usable" and i'm back to traditional coding for a few hours until its back to normal again
weird-eye-issue · · focus · HN ↗
It absolutely matters because something like Claude Code has no guarantee that there won't be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...
thephyber · · focus · HN ↗
The /r/antigravity SubReddit is full of users who very much notice bugs with the tool/agent. We should be thankful that Claude Code is pretty stable by comparison.
weedfroglozenge · · focus · HN ↗
Rapzid · · focus · HN ↗
swader999 · · focus · HN ↗
johnfn · · focus · HN ↗
I am more skeptical about the compute provider claim - do you have any evidence of that?
r_lee · · focus · HN ↗
and there's sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence
rhdunn · · focus · HN ↗
vikramkr · · focus · HN ↗
HawtAds · · focus · HN ↗
bitexploder · · focus · HN ↗
Rapzid · · focus · HN ↗
Note: I know what quantization is so don't hold back.
r_lee · · focus · HN ↗
like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions
dannyw · · focus · HN ↗
At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
Rapzid · · focus · HN ↗
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
sampullman · · focus · HN ↗
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
Rapzid · · focus · HN ↗
Nerfing conspiracy doesn't need to be proven false. Where is the evidence it's true?
Dylan16807 · · focus · HN ↗
They just want some evidence. It should be pretty easy to measure, shouldn't it?
sampullman · · focus · HN ↗
wavemode · · focus · HN ↗
ashdksnndck · · focus · HN ↗
bitexploder · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
[dead]
mwigdahl · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
You're a genius, did you think of that yourself? Or did you local token dealer teach you that?
kdkdjcjejxowjdj · · focus · HN ↗
[dead]
physicallyIllfr · · focus · HN ↗
cheevly · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
perching_aix · · focus · HN ↗
cindyllm · · focus · HN ↗
[dead]
eek2121 · · focus · HN ↗
What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.
Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.
There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.
I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.
winwang · · focus · HN ↗
khalic · · focus · HN ↗
You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?
lxgr · · focus · HN ↗
khalic · · focus · HN ↗
lxgr · · focus · HN ↗
khalic · · focus · HN ↗
lxgr · · focus · HN ↗
winwang · · focus · HN ↗
fendy3002 · · focus · HN ↗
topspin · · focus · HN ↗
That sentence... This conversation is indistinguishable from a billion conversations had around multi-player online gaming.
areoform · · focus · HN ↗
Image 31 is the list of 2024's Library of Congress DMCA circumvention exemptions. Apparently, US law is too spicy to talk about, <a href="https://en.wikipedia.org/wiki/Digital_Millennium_Copyright_Act#2024_rulemaking" rel="nofollow">https://en.wikipedia.org/wiki/Digital_Millennium_Copyright_A...
Based on my testing, the model cannot be used for anything related to chemistry, biology, and vSLAM. I recommend asking Opus 5.5 about 200 to 300 level undergrad biology.
That is a nerf.
itemize123 · · focus · HN ↗
nullbio · · focus · HN ↗
scrollop · · focus · HN ↗
<a href="https://marginlab.ai/trackers/claude-code/" rel="nofollow">https://marginlab.ai/trackers/claude-code/
This site has been documenting it for a while
cbg0 · · focus · HN ↗
wongarsu · · focus · HN ↗
<a href="https://marginlab.ai/trackers/claude-code-historical-performance/" rel="nofollow">https://marginlab.ai/trackers/claude-code-historical-perform...
lxgr · · focus · HN ↗
> We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).
Changing the harness can have a big impact on performance even when leaving the model completely unchanged.
wongarsu · · focus · HN ↗
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
lxgr · · focus · HN ↗
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.
ricardobeat · · focus · HN ↗
sspiff · · focus · HN ↗
The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.
I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.
stratos123 · · focus · HN ↗
s08148692 · · focus · HN ↗
bigmadshoe · · focus · HN ↗
rybosworld · · focus · HN ↗
A/B testing alone would result in a performance nerf for one group.
8-bit quantized models will barely show degradation on benchmarks. The performance is reliably at 99% of the non-quantized model. 4-bit quantization retains somewhere around 95-98% performance on benchmarks. But if you've ever used a 4-bit model, it feels lobotomized.
And just consider what a compny serving these models would do if they were at capacity. Would they stop serving the model altogether? Of course they wouldn't...
Denying that models experience purposeful degradation is gaslighting.
applicative · · focus · HN ↗
xlayn · · focus · HN ↗
Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..
I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1
I had the exact same feeling every time they have a new big release
solarkraft · · focus · HN ↗
Madmallard · · focus · HN ↗
smbullet · · focus · HN ↗
itopaloglu83 · · focus · HN ↗
Even in fresh sessions with minimal context like a file with a couple hundreds of lines of code, it would frequently ignore clear instructions, avoid work, and even sometimes “think” things like “looking for ways to code without approval”.
The model is just tuned for long running and doesn’t like to work with a human in the loop, or work back and forth.
bethekidyouwant · · focus · HN ↗
system2 · · focus · HN ↗
zeroonetwothree · · focus · HN ↗
ninjahawk1 · · focus · HN ↗
[dead]
vikramkr · · focus · HN ↗
apt-apt-apt-apt · · focus · HN ↗
j45 · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
Truly impressively sloppy.
Cider9986 · · focus · HN ↗
winwang · · focus · HN ↗
FranzFerdiNaN · · focus · HN ↗
Its quite worrisome honestly and i feel like that 'im in danger' Simpsons meme more and more. Right now i still have a lot of domain expertise which helps in knowing which questions to ask and which problems to solve, but well, i wonder how long that is going to save my job.
lqstuart · · focus · HN ↗
avazhi · · focus · HN ↗
First two days of this thing was like working with Einstein, then about 24-36 hours ago I started getting frustrated at bullshit that hadn't been a problem before. It was so egregious that I checked to make sure I was still on Opus 5.5 Max.
gaigalas · · focus · HN ↗
My gut tells me this involves an undisclosed, never-released grandparent model (higher-class than Fable/Astra level, roughly unsellable due to unfeasible cost). That grandparent model is distilled into lower models, of which Opus 5.5 might be an instance of.
That also guarantees protection against distilling a core business. You never make your prime weights available to the public, you only make distillings themselves available.
The downside of this strategy is that you spend a lot of compute on something that you never release, but it might be just the right play (for now) for closed weight companies.
It's a gut feeling, I have zero hard evidence to back it up.
folayii · · focus · HN ↗
[dead]
sheepscreek · · focus · HN ↗
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
zahlman · · focus · HN ↗
I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.
ben_w · · focus · HN ↗
What "stack" do you have in mind here?
An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. <a href="https://developers.openai.com/api/docs/models/gpt-4o" rel="nofollow">https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and <a href="https://platform.claude.com/docs/en/about-claude/model-deprecations" rel="nofollow">https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.
TeMPOraL · · focus · HN ↗
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
ben_w · · focus · HN ↗
The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.
> Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
Good points.
* the two exceptions being automated safety downgrade for dangerous topics, and "Auto-switch to Thinking" as a used-specified option in ChatGPT
dannyw · · focus · HN ↗
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
jeffybefffy519 · · focus · HN ↗
sheepscreek · · focus · HN ↗
And because token generation is a sequential process, all the code to orchestrate tensor parallelism (spreading a single request across multiple GPUs) is non-trivial. Each new token depends on the previous ones.
Combine that with KV caching optimizations, loading/unloading from cheaper cache storage, session management, load balancing, parallel sessions, and people updating them everyday plus testing optimizations to kernels, it’s a lot of moving parts.
We haven’t even discussed coordinating stuff across different data centres, failover mechanism, and what not.
jibalt · · focus · HN ↗
I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".
> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
So, not the sort of thing referred to.
> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.
blurbleblurble · · focus · HN ↗
kingcauchy · · focus · HN ↗
areoform · · focus · HN ↗
It's just that the nerfs are branded as "Safeguards." For example, Opus 5.5 can't reply to this message,
Image of the terminal response, <a href="https://i.postimg.cc/RVZx552k/image.png" rel="nofollow">https://i.postimg.cc/RVZx552k/image.pngImage 31 is the list of 2024's Library of Congress DMCA circumvention exemptions. Apparently, US law is too spicy to talk about, <a href="https://en.wikipedia.org/wiki/Digital_Millennium_Copyright_Act#2024_rulemaking" rel="nofollow">https://en.wikipedia.org/wiki/Digital_Millennium_Copyright_A...
Based on my testing, the model cannot be used for anything related to chemistry, biology, and vSLAM. I recommend asking Opus 5.5 about 200 to 300 level undergrad biology.
That is a nerf in every sense of the word.
intellyinstinct · · focus · HN ↗
[dead]
hyperionultra · · focus · HN ↗
At scale where anthro and cgpt operates - every single token matters.
[deleted] · · focus · HN ↗
[deleted]
perching_aix · · focus · HN ↗
nomilk · · focus · HN ↗
codr1 · · focus · HN ↗
killingtime74 · · focus · HN ↗
nbardy · · focus · HN ↗
Step 1. Model can’t do something challenging Step 2. You try a bunch and fail Step 3. Anthropic trains on your usage data. Your current code base and current problem are now in domain Step 4. Model comes out and you’re shocked when it can tackle the thing you were stuck on Step 4. Codebase drifts significantly and you try new problems you thought were a similar level. Your code is less familiar and the problem doesn’t have a bunch of failure cases in the train set. Feels of it being worse on similar problems
oh_my_goodness · · focus · HN ↗
rednb · · focus · HN ↗
This is definitely not about dealing with the frontier of AI. I wasn't been part of the nerfing chord, but Astra changed my mind. Quality got me to upgrade from Pro x5 to Pro x20 on launch day. A couple of days later was dumb af, horrendous code quality etc...
Something fishy, or at least unethical is going on. Not sure it impacts API users though.
sample369 · · focus · HN ↗
[dead]
Traubenfuchs · · focus · HN ↗
rw2 · · focus · HN ↗
msejas · · focus · HN ↗
Now I dismiss it every time and the quality is more consistent.
Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
itopaloglu83 · · focus · HN ↗
And strangely, expressing frustration multiple times in a row would reliably trigger a feedback popup as well.
okwhateverdude · · focus · HN ↗
That said, given the propensity for mature code bases to have "fuck" in commit messages/comments and those are typically of higher quality, I curse up a storm when the clankers make mistakes, if only to try put more quality-code valence into context. <a href="https://news.ycombinator.com/item?id=36584464">https://news.ycombinator.com/item?id=36584464
ClikeX · · focus · HN ↗
wunderlotus · · focus · HN ↗
jcutrell · · focus · HN ↗
Galilyou · · focus · HN ↗
xnorswap · · focus · HN ↗
Random performance is random, your brain will jump through hurdles to fit patterns where there aren't any.
unglaublich · · focus · HN ↗
In fact, since they have some rule based system (if bioenegineering or security, route to degraded model) it would be almost trivial to add 'user has filed feedback' to it.
Not saying this is what happens, but just that it's not as insane as it sounds.
ftchd · · focus · HN ↗
smashed · · focus · HN ↗
And it's also the kind of solution a misaligned agent would implement:
Make sure users are happy about their experience but also optimize resources.
The obvious solution is to route users based on feedback.
bb123 · · focus · HN ↗
jayd16 · · focus · HN ↗
Aurornis · · focus · HN ↗
SilverSlash · · focus · HN ↗
Aurornis · · focus · HN ↗
redanddead · · focus · HN ↗
grim_io · · focus · HN ↗
Using feedback on your sessions makes that session, and presumably all attached data, fair game for training.
margalabargala · · focus · HN ↗
grim_io · · focus · HN ↗
margalabargala · · focus · HN ↗
They explicitly ask, after the rating, whether they can look at your chat. You can just say "no" to that, if you believe that they are following the rules they say they are.
Your original statement of "using feedback gives them permission to train" is just plain false.
grim_io · · focus · HN ↗
iamjackg · · focus · HN ↗
Aachen · · focus · HN ↗
dofm · · focus · HN ↗
Great future everyone has chosen for us.
n3storm · · focus · HN ↗
khalic · · focus · HN ↗
lxgr · · focus · HN ↗
dev_l1x_be · · focus · HN ↗
bandrami · · focus · HN ↗
kroaton · · focus · HN ↗
k__ · · focus · HN ↗
aragornii · · focus · HN ↗
I wouldn't say Opus has been nerfed. They are just two different models but, at the end of the day, GPT Sol is the model I trust the most, at least for now.
hbroom · · focus · HN ↗
gitghxst · · focus · HN ↗
srnvs · · focus · HN ↗
[dead]
semiquaver · · focus · HN ↗
My company uses Claude models exclusively via Azure and AWS bedrock, which have their own licensed copies of the weights.
If all these people are so convinced Anthropic is nerfing models, have they tried other inference providers? Do they think the nerfing is coordinated across independent providers? Why wouldn’t any of these nerf-benches use these comparison points or even talk about them?
I think the most likely conclusion by far is that this is a psychological phenomenon.
dooglius · · focus · HN ↗
lxgr · · focus · HN ↗
giancarlostoro · · focus · HN ↗
ahknight · · focus · HN ↗
saejox · · focus · HN ↗
As my weekly/hourly quota is nearing its end, does Anthropic start to serve me a worse model?
Am i personally being throttled? testing against API is meaningless.
Tadpole9181 · · focus · HN ↗
navjeetgill307 · · focus · HN ↗
jtrn · · focus · HN ↗
ChristmasTomer · · focus · HN ↗
A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.
chrisss395 · · focus · HN ↗
kayhantolga · · focus · HN ↗
AtNightWeCode · · focus · HN ↗
tomaskafka · · focus · HN ↗
This week I have done considerably lighter work, and 2 days in, 54 % of Max usage is gone.
I hate this bait and switch cycle, and seeing how OpenAI cut its limits (5x raise from 0.1x api price to 0.5x api price for their max plan, so new $500 plan gets less usage than previous $200 plan) I think this will get worse still.
hgo · · focus · HN ↗
tomaskafka · · focus · HN ↗
notoriousjpg · · focus · HN ↗
augunrik · · focus · HN ↗