Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
SOTA models cost hundreds of millions to train. Did creating the contents of the text corpus they were trained on really cost an equivalent of 20x as much (~10 billions)? I honestly don’t know, but I could imagine it having been significantly less.
This isn’t meant as a moral argument, just musing about the relative cost comparison.
cmiles8 · · focus · HN ↗
Feels like Anthropic crying do as I say not as I do.
jedberg · · focus · HN ↗
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
JackFr · · focus · HN ↗
The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.
layer8 · · focus · HN ↗
This isn’t meant as a moral argument, just musing about the relative cost comparison.
phamilton · · focus · HN ↗
A training set of 15 trillion tokens is 10 trillion words.
A penny a word is cheaper than the cheapest beginner freelance writer.
That makes a training set of 10 trillion words cost $100B.
Lots of assumptions there for sure, but we're certainly in the ballpark you are describing.
ToValueFunfetti · · focus · HN ↗