Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
Distillation doesn't "grab 95% of lab's work", that's ridiculous. At best it's tiny icing on top of the cake that's already there. It's not even necessarily done on a better model (e.g. GLM 4.7 distilled Gemini 2.5, a weaker model), I'm pretty sure A\ and OAI could do (or even do) the same with greater efficiency since they have access to logits, weights, and internal state of open models.
>An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
How is this philosophical? They should release the unsupervised pretrains, at the very least.
cmiles8 · · focus · HN ↗
Feels like Anthropic crying do as I say not as I do.
jedberg · · focus · HN ↗
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
orbital-decay · · focus · HN ↗
>An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
How is this philosophical? They should release the unsupervised pretrains, at the very least.