Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
> to create a perceived improvement when in reality there isn’t really one?
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
* Release new model that scores an arbitrary 100 on a benchmark
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
>gaslight them into thinking it never changed or that it's just a harness problem
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.
talon8635 · · focus · HN ↗
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
Aurornis · · focus · HN ↗
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
well_ackshually · · focus · HN ↗
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Bis repetitae.
scrollop · · focus · HN ↗
I imagine some people have their own personal in depth benchmarks they could do this for.
well_ackshually · · focus · HN ↗
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.