I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.
Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."
The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.
Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
Can you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?
Dollars to donuts, they are speculating, and not privy to inside information on the topic.
However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
An AI lab will never volunteer the information because it opens them up to lawsuits if they are purposely degrading service and not letting users know.
They can limit how hard the model thinks for a given effort. Suddenly xhigh only thinks as hard as high did, and high shifts down to medium effort, and so on.
They can also serve quantized models. And this has the benefit of practically not showing up in benchmarks at all even if the user experience is obviously degraded.
The other major thing the labs do is silently drop the usage limits. This has become very noticeable for codex users who are suddenly burning through their weekly usage in a few hours.
Yea, if you ever run your own models on a GPU there are a whole ton of different dials you can adjust that drastically affect compute use, memory use, and output token quality, and number of tokens held in memory.
If anyone reading has a GPU it's worthwhile just messing with a smaller model for a bit to watch how the settings affect output.
I don't work at Anthropic, but I would assume they could serve smaller quantizations during peak hours - this effectively controls the "resolution" of the model. They could also control the resolution of the KV cache, which would make the model not necessarily dumber, but worse at understanding the incoming requests.
And finally, you could pass off what was "high" effort as "extra", because why not.
And it's not. A conspiracy theory is what it is.
I have no reason to doubt the claims of the employees at OpenAI and Anthropic who have told us personally multiple times, including here on HN, that they do not degrade the models in order to reduce load.
As for the endlessly long analysis in the OP, it appears it's based on analyzing their random usage data rather than any fixed benchmark. I don't think it makes much sense.
Isn’t the idea that they’re limiting the amount of gpu time normal users get to spend on the “thinking” portion of their query?
I think that’s the claim in the post, that even though no one can see the true chain of thought, that even the “thinking” text that does get exposed to the user is shorter given the same prompts over time. Not saying it’s true but I think that’s the claim. I’ve personally never noticed the alleged “nerfing” with my enterprise use at work or my subscription use at home which is only during off hours.
Waterluvian · · focus · HN ↗
Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."
The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.
Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
prodigycorp · · focus · HN ↗
Opus 5.5 is being served under opus 5 right now.
w1296 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
SequoiaHope · · focus · HN ↗
arcanemachiner · · focus · HN ↗
However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
rybosworld · · focus · HN ↗
They can limit how hard the model thinks for a given effort. Suddenly xhigh only thinks as hard as high did, and high shifts down to medium effort, and so on.
They can also serve quantized models. And this has the benefit of practically not showing up in benchmarks at all even if the user experience is obviously degraded.
The other major thing the labs do is silently drop the usage limits. This has become very noticeable for codex users who are suddenly burning through their weekly usage in a few hours.
pixl97 · · focus · HN ↗
If anyone reading has a GPU it's worthwhile just messing with a smaller model for a bit to watch how the settings affect output.
sznio · · focus · HN ↗
gslepak · · focus · HN ↗
On what basis are you claiming this?
prodigycorp · · focus · HN ↗
pllbnk · · focus · HN ↗
user43928 · · focus · HN ↗
I have no reason to doubt the claims of the employees at OpenAI and Anthropic who have told us personally multiple times, including here on HN, that they do not degrade the models in order to reduce load.
As for the endlessly long analysis in the OP, it appears it's based on analyzing their random usage data rather than any fixed benchmark. I don't think it makes much sense.
pertymcpert · · focus · HN ↗
carljungslabtek · · focus · HN ↗
I think that’s the claim in the post, that even though no one can see the true chain of thought, that even the “thinking” text that does get exposed to the user is shorter given the same prompts over time. Not saying it’s true but I think that’s the claim. I’ve personally never noticed the alleged “nerfing” with my enterprise use at work or my subscription use at home which is only during off hours.