Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
To a degree. The human produced knowledge is the product of all humanity (no human is an island).
A comparable idea could be that an encyclopedia or maths book is only distilling the things that other people did, and how dare they sell them. But the "only" is doing quite a bit of work. LLMs do not just spawn into existence. There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.
cmiles8 · · focus · HN ↗
Feels like Anthropic crying do as I say not as I do.
jedberg · · focus · HN ↗
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
JackFr · · focus · HN ↗
The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.
jstummbillig · · focus · HN ↗
A comparable idea could be that an encyclopedia or maths book is only distilling the things that other people did, and how dare they sell them. But the "only" is doing quite a bit of work. LLMs do not just spawn into existence. There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.