Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
I'm overall pro-Anthropic and pro-banning open-weights AI, but I agree with the parent commenter; distilling Claude models is not that different from pretraining on web data. It's all basically the same sort of thing.
I think a good litmus test here would be if Anthropic were to not care about distilling their models when the distillers keep the resulting models closed-source and sell tokens via an API. If they cared only about security concerns and not about people profiting off of their work, then they should be publicly fine with this and only protest against it going into open-weights models.
cmiles8 · · focus · HN ↗
Feels like Anthropic crying do as I say not as I do.
jedberg · · focus · HN ↗
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
meowface · · focus · HN ↗
I think a good litmus test here would be if Anthropic were to not care about distilling their models when the distillers keep the resulting models closed-source and sell tokens via an API. If they cared only about security concerns and not about people profiting off of their work, then they should be publicly fine with this and only protest against it going into open-weights models.