Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
>Aren’t Anthropic’s models not just distilling down other people’s work?
Can you elaborate on that? I mean my direct answer would be no, of course not. But why do you think frontier models are distilled? I think maybe there is an equivocation over the word “distillation.”
Frontier labs train on their own pretraining data, human feedback, synthetic data, and research. A distilled model is specifically optimized to reproduce another model's behavior.
Meanwhile R1-Distill-Qwen-32B was distilled from DeepSeek-R1.
If you want to say a frontier model is "distilled" from the world's data and R1-Distill-Qwen-32B is distilled from DeepSeek-R1 then you are equivocating two very different things.
cmiles8 · · focus · HN ↗
Feels like Anthropic crying do as I say not as I do.
nonethewiser · · focus · HN ↗
Can you elaborate on that? I mean my direct answer would be no, of course not. But why do you think frontier models are distilled? I think maybe there is an equivocation over the word “distillation.”
Frontier labs train on their own pretraining data, human feedback, synthetic data, and research. A distilled model is specifically optimized to reproduce another model's behavior.
Meanwhile R1-Distill-Qwen-32B was distilled from DeepSeek-R1.
If you want to say a frontier model is "distilled" from the world's data and R1-Distill-Qwen-32B is distilled from DeepSeek-R1 then you are equivocating two very different things.
nuancebydefault · · focus · HN ↗