Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
And yet they have only improved marginally in my use cases since around Opus 4.5.
The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).
For my use cases, we are definitely on the flatter part of the curve at the moment.
Same experience here, anything frontier human knowledge wise, same if not a regression. For human understanding and emotional intelligence, for many tasks regressiin is so bad that many near anchient llama era models now beat frontier anthropic/oai models. Notable exceptions to capability rot seem to be qwen models, and previously deepseek but the latest gen of models has started showing the same rot. General writing quality is down significantly accross the board, often it is outright ass. For example, I didnt mind reading 4.5's outout, but opus 5 makes me goddamn near violent, its fucking insufferable.
talon8635 · · focus · HN ↗
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
a2dam · · focus · HN ↗
Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
koyote · · focus · HN ↗
The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).
For my use cases, we are definitely on the flatter part of the curve at the moment.
a2dam · · focus · HN ↗
Grimblewald · · focus · HN ↗