I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).
I wonder what their official explanation for this behavior is.
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
>What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.
alexjplant · · focus · HN ↗
I wonder what their official explanation for this behavior is.
Wowfunhappy · · focus · HN ↗
(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
chrsw · · focus · HN ↗
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
sebzim4500 · · focus · HN ↗
As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.