Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
I wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
I max out my Claude Max plan every week, and I can measure the output, and for me it's stayed fairly constant, subject to the various "bonuses" whenever Anthropic is feeling the competitive pressure.
Well to be fair the output you get has improved greatly. You can still get billions and billions of Luna/Sonnet tokens within your subscription comfortably (doesn’t feel fair to compare Luna to Haiku). Sol/Astra/Fable… yeah, they’ll chew through your credit.
Last time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC.
Where'd you get 30 years from? Show your work.
But you're reaching for Kimi, no? A model from an entirely otherwise-redundant company. You're already using Open AI models, so Kimi seems even more of a reach.
I'm asking, not arguing, because I'd like to understand. Is Astra so much more expensive for those tasks, and are they frequent?
If he has 2x 20x OpenAI, that means he's running heavy jobs that burn through usage. So Kimi must be there to reduce OpenAI usage. With Astra + Fable, I burn through my 5x OpenAI and 20x Claude real fast. I do have a backup GLM sub but I've never had to use it. So the prediction that AI would become more expensive seems to be panning out. Partly outweighed by better models of course.
Yep, no problem at all. The only drawback is sessions cannot be shared directly between accounts, so if you're in the middle of something you'll have to do some extra work. To that end I have a handoff skill to persist state to a local markdown and a resume skill to load that state into a new session.
My calculation (that is 3x 20x subscriptions) it will pay off after ~18.5 months, if I were to stop my subscriptions today. I will likely keep at least one though, so it's more like 27.5 months for a payoff. I am not doing it to save money though. I am doing it for security of supply. Constant quality - I know my model isn't getting lobotomized etc. - and open weight models mostly just lack very little behind. I'm sure I will have affordable, fast and efficient sol 5.6 capabilities with open weight models on my sparks within 12-18 months easily.
flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.
Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago.
I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)
Full GLM-5.3 needs a beast of a system, but you can run GLM-5.3 Flash on the 2x Spark setup the GP comment mentioned. If benchmarks are anything to go by, Flash is like having a local Terra-tier coding model: <a href="https://artificialanalysis.ai/models/comparisons?compare=glm-5-3-flash%2Cglm-5-3%2Cclaude-sonnet-5" rel="nofollow">https://artificialanalysis.ai/models/comparisons?compare=glm...
I’m also pretty happy with GLM 5.3 Flash (for coding, navigation and german language it sucks at). Incredible that you can run it on a fairly practical (seeming) home setup.
But here’s the standard question: At what speeds/other limiting factors?
Same here. Moved from Opus to GLM 5.2 to 5.3 and I've been pretty happy with the result. Mainly, it doesn't hallucinate and convince itself of mistake so it's good at retrieving information or asking the user for it. Opus and Fable always state something, then try to "prove" it but end up convincing themselves of the wrong thing. Having subagents for retrieval and validation helped but were not enough.
It will also fully ignore you if it has the slightest belief (not even a hint) that it knows what you want better than you and just start doing things.
This is also why I think it's baffling that they switched to auto mode by default. It's becoming harder to use Claude at least to help with improving at coding.
If I ask something like:
"I'm building a simple X as a learning exercise, I'm writing the code so please only answer the question I'm asking and don't try to solve the problem directly. How does ..."
There's a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a "design decision" that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.
IMO this comes down to your harness. Any frontier model from a huge shop has an inherent benefit in the system you're using it in. Search, memory, skills, integrations you don't realize even exist make them much more powerful. It is some effort but I recommend trying Hermes Agent and setting it up fully, that's the closest you'll get to a more complete experience.
Some of us are stuck on subscriptions and our executives will never give us API access.
But also, everyone says "it's the harness" and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.
Qwen 3.8 Flash-Next is not dumb. If you've used it and that was your experience, your workload is either ultra-ultra-sophisticated or you're dealing with a broken quant/buggy chat template/other issue. That model is a smart, reliable workhorse.
it was an extremely simply workload with different off the shelf harnesses, they just all sucked when you compare it to a paid hosted model. It was fine for classifying stuff or summarizing though, but missed technical details.
> Isn't Anthropic the biggest competitor Sam Altman has?
Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".
The reason they're teaming up is the real competition is, as in many other domains, China.
I assume you’ve calculated expected cost vs subscription.
How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.
reliability and self reliance is worth a lot to most. Heck, you could be the best in the world at what you do, but if you're unreliable you wont find stable employment. So, not having some amoral shady company errode model quality out from under you constantly is also worth a lot more than simple cost balancing calculations can capture.
I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.
I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.
Picked a local model I was happy with, then figured out what is required to get it running at acceptable tok/s, where I settled on a amd "ai" variant NUC with 96gb vram avaliable to gpu. This arrived recently, and qwen3.8 27b runs fast enough for me on that, with full context, and plenty of paralell streams. That said i'm in a fortunate position, so also got an rtx6000 to continue rlhf based finetunes in data I amass over the years from myself and friendly highly knowledgable/skilled friends, since that is in essence what makes fronteir labs models better, so i figure, why shouldnt we benefit from our expertise and input direcerly, instead of having it sold back to me by some amoral company? If that eventuates in a model that genuinly beats current, obviously we'd give back to the community by releasing that.
Every SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.
I have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.
I’m not sure why this point keeps coming up — if your service/product is so popular that it’s capacity-constrained, then the answer is to raise prices, not degrade service, because the demand should be inelastic.
Raising prices also has second order effects, like consumer and business expectations around how widespread the tech can be. Valuations depend on it being reasonably affordable to roll out on a much more massive scale than today. If people get the impression that it seems too limited to very rich people (200 is affordable for a North American / Western European professional), the impression about the trajectory will change.
The Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code.
Fable seems to be following the same enshittification arc of other Anthropic models.
Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.
I’m so behind on this topic but I find it interesting how quickly things change. I feel like just yesterday I way hearing how anthropic is far and away better than OAI, and now this.
I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments
This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason
There must be some benefit if all the providers are doing it independently.
GPT5.6-Sol on Max thinking just became regarded as of a few days ago.
The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).
Again, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant?
Kinda. Off the top of my head, DeepSeek and their thinking model was pretty new and interesting. Multi input models are also newish (combined input of text, image, video, audio, etc). Then there's Jev, a recently release that has a lot of people talking. It isn't really an LLM, but also is one.
Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.
Overall models have become cheaper to run and smarter per token.
They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record.
Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves.
See Shepard tone. Similarly model releases could be engineered to appear that they’re always getting better by slowly degrading and upgrading at the right time. That plus hitting some benchmarks and making a lot of noise around that.
We are being A/B tested on and there is nothing you can do about it.
Oh, yes there is. DeepSeek 4.1 Flash on max thinking can simply be dropped into Claude Code. Close your eyes as the chain-of-thought traffic scrolls by and you can easily fool yourself into thinking you're still running Opus, in terms of both cognition and throughput.
To be fair, matching Opus's throughput costs about as much as a new car, but cars suck nowadays and you didn't want a new one anyway, right...? Failing that, rent a cloud server, one that you control.
They obviously test various quants and other serving cost saving strategies. Models like Fable are probably trillions parameters with hundreds of billions active MoE. They probably try to squeeze and quant each piece until people notice.
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
Not unless your competitors do the same, or else you will only be perceived as falling behind others.
Yep, that's what they've been doing for a long while now. Also the amount of tokens you get per sub varies drastically from month to month. Needs to be regulated.
> to create a perceived improvement when in reality there isn’t really one?
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
* Release new model that scores an arbitrary 100 on a benchmark
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
>gaslight them into thinking it never changed or that it's just a harness problem
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.
Because frontier models are completely opaque. Doing a controlled test of "the same model" months apart is simply impossible if you don't work for that provider (and even then, may not be feasible). We know from external observation that model performance changes minute to minute, day to day and week to week for a variety of reasons: load balancing, inference hardware, and shared RAM pool to dozens of internal software settings each of which impact cost, latency, time-to-first-token, quality, veracity, tool use, etc.
Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.
This is a good point I hadn’t considered, thank you.
Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?
Sorry if it’s a dumb question, I don’t really know much about the topic.
Also, in a world where there are several models competing with each other for public perception of which is best, that seems like an extremely bad move.
you mean like a Shepards Tone (<a href="https://en.wikipedia.org/wiki/Shepard_tone" rel="nofollow">https://en.wikipedia.org/wiki/Shepard_tone); i wouldn't doubt they slowly tweak quants to try to eke out.
there's also probably load balancers that downgrade models during high use.
This would only provides a benefit if we're approaching some sort of theoretical limit of how good LLMs can be with the current approaches and data.
Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?
So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.
There doesn't need to be an explicit conspiracy. It only needs all the US frontier labs to be facing the same economic pressure (logarithmic improvement / $). It's a pretty obvious strategy - it's not like the large labs can magic up huge volumes of extra compute as demand comes online; there are almost certainly tweaking model performance to occupy the compute available and margin/cash burn targets.
That was one of the main points of the movie 'A Beautiful Mind' - that actors can coordinate without any explicit communication.
No, what they are doing is trying to optimize inference to increase margins which leads to degradations. Model deployment is not like websites, you can continuously tune performance based on usage, new memory optimizations, etc.
Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
I get the perspective from which you're making that statement, but the industry keeps moving its own goalposts.
Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.
I would define it as a quick-moving market, but not nearly moving enough for the fantastic claims they make to justify ever-increasing funding.
AI, if not AGI, has certainly become uniquely dangerous in the past 6 months though. We have a lot of evidence to that effect.
The best case scenario is a situation like Y2K: a ton of people coordinate and work hard to produce no perceptible change, because unlike catastrophe, averting catastrophe feels boring.
>Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.
Absolutely none of this points to “stagnant.” Stagnant is a terrible description of the AI industry.
And yet they have only improved marginally in my use cases since around Opus 4.5.
The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).
For my use cases, we are definitely on the flatter part of the curve at the moment.
Same experience here, anything frontier human knowledge wise, same if not a regression. For human understanding and emotional intelligence, for many tasks regressiin is so bad that many near anchient llama era models now beat frontier anthropic/oai models. Notable exceptions to capability rot seem to be qwen models, and previously deepseek but the latest gen of models has started showing the same rot. General writing quality is down significantly accross the board, often it is outright ass. For example, I didnt mind reading 4.5's outout, but opus 5 makes me goddamn near violent, its fucking insufferable.
>For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
There was a coding horror story I read some years ago where a developer bragged that he improved performance by artificially increasing iterations on some critical path in an app and then lowering the iterations occasionally while bragging to management about squeezing out more performance.
Kind of reminds me of that, but with more smoke and mirrors
AI just solved a millennium problem two weeks ago. "The pace is insane. And there is no reason to be this fast." to quote Terence Tao word by word.
In addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more resources adaptively. 'Driver optimizations' that detected benchmarks in real-time were deployed in the first 'GPU Wars'.
> I have no idea is the actual frontier is stagnating.
Like a lot of complex, rapidly evolving tech, the truth is it's probably rapidly accelerating on some measures for a few and stagnating on many others for most - hence the divergence in user reports. It's depends on how you use it, for what problems, how rigorously you assess the output and whether you happen to be on a server bank, RAM pool or shard at this moment which hasn't yet been sufficiently 'cost optimized' by the margin algorithms. They don't call them load balancers anymore. They're Margin Balancers.
I've subjectively detected this in previous codex releases where the 2 days before release of a new model the agent went from great to me pulling my hair out yelling at it. I think it's just a win-win for them. They need to ramp up basic capacity and usage on the new model, what better way to free up capacity than to reduce the effort with the current gen. There's not deep detectors that they've regressed (especially before this guy showed solid data).
talon8635 · · focus · HN ↗
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
AmazingTurtle · · focus · HN ↗
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
ramesh31 · · focus · HN ↗
Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
torben-friis · · focus · HN ↗
rybosworld · · focus · HN ↗
vidarh · · focus · HN ↗
CookieCrisp · · focus · HN ↗
rybosworld · · focus · HN ↗
d1sxeyes · · focus · HN ↗
zeroonetwothree · · focus · HN ↗
fragmede · · focus · HN ↗
Where'd you get 30 years from? Show your work.
wilj · · focus · HN ↗
mike_d · · focus · HN ↗
The problem is they can't fit any frontier level open models.
sandblast · · focus · HN ↗
knollimar · · focus · HN ↗
dotancohen · · focus · HN ↗
I'm asking, not arguing, because I'd like to understand. Is Astra so much more expensive for those tasks, and are they frequent?
edg5000 · · focus · HN ↗
jwpapi · · focus · HN ↗
brandall10 · · focus · HN ↗
jwpapi · · focus · HN ↗
brandall10 · · focus · HN ↗
tiagod · · focus · HN ↗
AmazingTurtle · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
AtHeartEngineer · · focus · HN ↗
boardwaalk · · focus · HN ↗
cyanydeez · · focus · HN ↗
Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?
I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.
I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.
So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?
w0m · · focus · HN ↗
cyanydeez · · focus · HN ↗
poslathian · · focus · HN ↗
Insanity · · focus · HN ↗
bigyabai · · focus · HN ↗
solarkraft · · focus · HN ↗
But here’s the standard question: At what speeds/other limiting factors?
Toslink · · focus · HN ↗
[dead]
Sayrus · · focus · HN ↗
jan_m_savage · · focus · HN ↗
_blk · · focus · HN ↗
Co-Authored By: Haiku 4.5
greenavocado · · focus · HN ↗
SpaceNugget · · focus · HN ↗
If I ask something like: "I'm building a simple X as a learning exercise, I'm writing the code so please only answer the question I'm asking and don't try to solve the problem directly. How does ..." There's a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a "design decision" that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.
_s_a_m_ · · focus · HN ↗
wronglebowski · · focus · HN ↗
pdimitar · · focus · HN ↗
But also, everyone says "it's the harness" and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.
srcreigh · · focus · HN ↗
anon373839 · · focus · HN ↗
hhh · · focus · HN ↗
julianlam · · focus · HN ↗
Yeah, expecting the world when all you have is a 8GB graphics card? You're going to be disappointed.
16GB is table stakes (IQ3_XSS). 32 GB is better.
esseph · · focus · HN ↗
redanddead · · focus · HN ↗
Yet… even Altman called out Anthropic for serving dumbed down models.
Shits weird man
tasuki · · focus · HN ↗
Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?
redanddead · · focus · HN ↗
TacticalCoder · · focus · HN ↗
Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".
The reason they're teaming up is the real competition is, as in many other domains, China.
andsoitis · · focus · HN ↗
How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.
Grimblewald · · focus · HN ↗
I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.
I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.
andsoitis · · focus · HN ↗
Grimblewald · · focus · HN ↗
ncr100 · · focus · HN ↗
tsunamifury · · focus · HN ↗
briffle · · focus · HN ↗
Denzel · · focus · HN ↗
pixl97 · · focus · HN ↗
A very small raise in prices may cause a very large loss in customers that you risk never getting back.
For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.
Denzel · · focus · HN ↗
bonoboTP · · focus · HN ↗
fnordpiglet · · focus · HN ↗
Fable seems to be following the same enshittification arc of other Anthropic models.
Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.
fragmede · · focus · HN ↗
On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.
fnordpiglet · · focus · HN ↗
talon8635 · · focus · HN ↗
I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments
progval · · focus · HN ↗
bmicraft · · focus · HN ↗
Maxatar · · focus · HN ↗
<a href="https://www.tomshardware.com/features/crucial-p2-ssd-qlc-flash-swap-downgrade" rel="nofollow">https://www.tomshardware.com/features/crucial-p2-ssd-qlc-fla...
<a href="https://www.tomshardware.com/news/wd-blue-sn550-ssd-performance-cut-in-half-slc-runs-out" rel="nofollow">https://www.tomshardware.com/news/wd-blue-sn550-ssd-performa...
<a href="https://www.tomshardware.com/news/adata-switches-nand-on-sx8200-pro-ssd-performance-impacted" rel="nofollow">https://www.tomshardware.com/news/adata-switches-nand-on-sx8...
nxc18 · · focus · HN ↗
GPT5.6-Sol on Max thinking just became regarded as of a few days ago.
The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).
The cycle repeats.
[deleted] · · focus · HN ↗
[deleted]
[deleted] · · focus · HN ↗
[deleted]
talon8635 · · focus · HN ↗
Thanks for your insight
dalenw · · focus · HN ↗
Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.
Overall models have become cheaper to run and smarter per token.
mrandish · · focus · HN ↗
They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record.
Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves.
onemoresoop · · focus · HN ↗
claydugo · · focus · HN ↗
We are being A/B tested on and there is nothing you can do about it.
CamperBob2 · · focus · HN ↗
Oh, yes there is. DeepSeek 4.1 Flash on max thinking can simply be dropped into Claude Code. Close your eyes as the chain-of-thought traffic scrolls by and you can easily fool yourself into thinking you're still running Opus, in terms of both cognition and throughput.
To be fair, matching Opus's throughput costs about as much as a new car, but cars suck nowadays and you didn't want a new one anyway, right...? Failing that, rent a cloud server, one that you control.
bitexploder · · focus · HN ↗
kleiba2 · · focus · HN ↗
Not unless your competitors do the same, or else you will only be perceived as falling behind others.
talon8635 · · focus · HN ↗
If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide.
I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me
senordevnyc · · focus · HN ↗
wmf · · focus · HN ↗
This isn't what we see in benchmarks.
rfgplk · · focus · HN ↗
ekjhgkejhgk · · focus · HN ↗
Yes, the AI technology is known primarily for how stagant it is.
talon8635 · · focus · HN ↗
I have no idea, just had a thought and put it out there
whatever1 · · focus · HN ↗
I would do a test to verify my suspicions.
talon8635 · · focus · HN ↗
I’m actually so far removed from this tech that I couldn’t run such a test myself lol
zarmin · · focus · HN ↗
Aurornis · · focus · HN ↗
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
pixl97 · · focus · HN ↗
well_ackshually · · focus · HN ↗
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Bis repetitae.
scrollop · · focus · HN ↗
I imagine some people have their own personal in depth benchmarks they could do this for.
well_ackshually · · focus · HN ↗
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.
mrandish · · focus · HN ↗
Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.
Aurornis · · focus · HN ↗
Where?
I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.
mobelkh · · focus · HN ↗
Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.
talon8635 · · focus · HN ↗
Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?
Sorry if it’s a dumb question, I don’t really know much about the topic.
AnimalMuppet · · focus · HN ↗
ruszki · · focus · HN ↗
trenchgun · · focus · HN ↗
arational · · focus · HN ↗
holoduke · · focus · HN ↗
cyanydeez · · focus · HN ↗
there's also probably load balancers that downgrade models during high use.
Groxx · · focus · HN ↗
jaredklewis · · focus · HN ↗
Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?
So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.
ponkpanda · · focus · HN ↗
That was one of the main points of the movie 'A Beautiful Mind' - that actors can coordinate without any explicit communication.
gizmodo59 · · focus · HN ↗
loteck · · focus · HN ↗
siva7 · · focus · HN ↗
<a href="https://thedailywtf.com/articles/The-Speedup-Loop" rel="nofollow">https://thedailywtf.com/articles/The-Speedup-Loop
sbxfree · · focus · HN ↗
a2dam · · focus · HN ↗
Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
Lalabadie · · focus · HN ↗
Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.
I would define it as a quick-moving market, but not nearly moving enough for the fantastic claims they make to justify ever-increasing funding.
a2dam · · focus · HN ↗
The best case scenario is a situation like Y2K: a ton of people coordinate and work hard to produce no perceptible change, because unlike catastrophe, averting catastrophe feels boring.
nonethewiser · · focus · HN ↗
Absolutely none of this points to “stagnant.” Stagnant is a terrible description of the AI industry.
koyote · · focus · HN ↗
The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).
For my use cases, we are definitely on the flatter part of the curve at the moment.
a2dam · · focus · HN ↗
Grimblewald · · focus · HN ↗
talon8635 · · focus · HN ↗
I’m not saying it is stagnant. I’m saying for a hypothetical industry that was (maybe that fits AI, maybe not, I have zero authority to say myself)…
nonethewiser · · focus · HN ↗
Are you really saying AI is a stagnant industry?
talon8635 · · focus · HN ↗
KingMob · · focus · HN ↗
Unlike you, however...
yareally · · focus · HN ↗
Kind of reminds me of that, but with more smoke and mirrors
le-mark · · focus · HN ↗
raincole · · focus · HN ↗
HN: Well, must be a stagnant industry...
talon8635 · · focus · HN ↗
mrandish · · focus · HN ↗
In addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more resources adaptively. 'Driver optimizations' that detected benchmarks in real-time were deployed in the first 'GPU Wars'.
> I have no idea is the actual frontier is stagnating.
Like a lot of complex, rapidly evolving tech, the truth is it's probably rapidly accelerating on some measures for a few and stagnating on many others for most - hence the divergence in user reports. It's depends on how you use it, for what problems, how rigorously you assess the output and whether you happen to be on a server bank, RAM pool or shard at this moment which hasn't yet been sufficiently 'cost optimized' by the margin algorithms. They don't call them load balancers anymore. They're Margin Balancers.
jfoster · · focus · HN ↗
Could be what happens next, though.
sporkland · · focus · HN ↗