I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).
I wonder what their official explanation for this behavior is.
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?
And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
Are you telling me that companies might defraud people for millions and billions of dollars and pay fines that are 1000% less than their profits?" My goodness, you must live on a hell planet.
Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.
And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
The question isn't whether the penalty would be less than their profit, it's whether the penalty would be less than whatever they make by secretly downgrading the models (or whatever it is you suspect), which remember also causes consumers to get less value out of the product and more likely to cancel.
The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.
alexjplant · · focus · HN ↗
I wonder what their official explanation for this behavior is.
Wowfunhappy · · focus · HN ↗
(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
chrsw · · focus · HN ↗
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
Wowfunhappy · · focus · HN ↗
...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?
And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
pixl97 · · focus · HN ↗
Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.
And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
Wowfunhappy · · focus · HN ↗
The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.
mobelkh · · focus · HN ↗
nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is