Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?
Feels like Anthropic crying do as I say not as I do.
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
Yeah, the whole thing seems like a human centipede of rug pulling. Probably the same as it's always been. Curating AI knowledge should be something that we put our best researchers towards, but realistically I think we wind up with 2-3 highly biased nationalistic models that are constantly copying off each other's notes.
Thanks, that’s the nature of any business. Founders see a way to take existing knowledge and expertise, combine it in some novel or interesting way, and produce a new product
Napster was a fantastic and disruptive product, the likes of which arguably has no equal to this day. But eventually the hammer came down from the courts and it was replaced by streaming services like Netflix, which pay to license materials from their creators.
SOTA models cost hundreds of millions to train. Did creating the contents of the text corpus they were trained on really cost an equivalent of 20x as much (~10 billions)? I honestly don’t know, but I could imagine it having been significantly less.
This isn’t meant as a moral argument, just musing about the relative cost comparison.
Pre vaccination smallpox killed hundreds of millions of people just in the twentieth century [0], the knowledge that allowed for the creation of just that vaccine is worth hundreds trillions of dollars in humans lives, let alone all of `the knowledge and experiences those people were involved in.
The knowledge that created the Haber-Bosch process [1] helps to sustain the majority of the world's populous, add another five hundred trillion dollars for that just to start with.
The creation of the printing press and all written information that allowed it to be built provided dissemination of knowledge beyond the ultra wealthy and is worth a non-finite amount of money.
LLM's are cool math, but they are less than a rounding error in comparison to even the tiniest sliver of human knowledge and technological output.
Shush, this is Hackernews. We all want to have our egos stoked that our industry is the most important in human history and that tech will transform and save us all. Go away with your historical analysis /s
do you also count eg published results of very expensive physics experiments? because once the costs of things like these are taken into account, we are way over 10 billions.
If you look at movies alone that would easily surpass 10s of billions. The cost of most books is probably more nebulous, but books, research, and more all have time and money spent to create them. I would guess the corpus of all media from the 20th century on would be minimally in the hundreds of billions of dollars.
I would argue that producing the complete written corpus on which they at least intend to train (even if some is still out of reach) cost literally everything to produce.
And the monetary cost doesn't even register when weighed against the blood, sweat and tears that went into capturing the authentic experiences of real human beings, whose honest expressions are now at least in some cases getting hoovered up, ingested, and then destroyed for all eternity, for fear that this specific work is the rounding error that might give an equally immoral competitor the edge in the bicycle-riding flamingo race that is currently consuming an absurd amount of the world's creativity and attention.
It probably cost vastly more than training the LLM. You need to consider the time people spent and perhaps weigh it in their hourly wage. Accumulate that across all the training data and it will be a mindboggling amount, compared to training the model.
But if you take it deeper, didn't most of those authors rely on the work of others? Most of human knowledge is small advancements of things we already knew. Often by reorganizing what we already knew.
Is that not what the foundation models are? A new reorganization of existing knowledge?
Often people ignore scale. N=1 is OK, therefore, N=1billion is OK. Same flawed argument as: "It's OK for one police officer to watch one street corner for the purpose of observing crime; therefore it's equally OK to have cameras recording every street corner in the city 24/7, for all purposes. Same thing!"
So then the problem is that Anthropic seems hypocritical when they knowingly insert themselves into this chain, and then complain about people down-chain from them.
To remedy the negative impressions (if they even care to do so) they should do 1 of 2 things:
1) stop complaining about it
2) stop distilling other people's work
They insert themselves into this chain for profit and complain about it. I really think that adds a thick layer to the hypocrisy that people, or at least me, feel is especially distasteful.
No LLM products would exist without the avalanche of largely non-consensual use of IP to create them, full stop. Any of these companies doing this and then turning around and complaining when their IP is "breached" are going to met with a chorus of tiny violins.
It's hypocritical and we shouldn't listen to their bullshit but I wouldn't expect anything else from them.
This generation of frontier models is "good enough." At some point, the Chinese labs will get to where the Americans are right now, they'll race to the bottom, and we will actually see what proliferation of AI looks like.
If Anthropic wants to be a trillion dollar company, they need to make revenue like Google or Apple do. Both of them have near monopolies, Anthropic is only getting further and further away as time goes on.
Of course they're gonna complain until they either figure out a better plan or accept a new valuation which is high but not spectacular.
It is an ancient practice, that when a human creates something, other humans will observe it and learn from it. Every group of humans living together has practiced this in some form for tens of thousands of years if not longer. Even animals do it. It's a natural assumption when making any form of art.
It is not a natural assumption that someone will digitize the artwork and use it to adjust a couple thousand matrix coefficients in a complex computer program. To most people that seems like copying with extra steps. The brain may in some ways resemble a computer, but what sets it apart is that we have always lived with brains. Everything a human does has already anticipated the presence of other brains, while etched circuits on ultrapure silicon crystals are something new.
Don't forget the mothers of all those original authors, as well as everyone who labored to build and sustain the societies which produced writers.
To a degree. The human produced knowledge is the product of all humanity (no human is an island).
A comparable idea could be that an encyclopedia or maths book is only distilling the things that other people did, and how dare they sell them. But the "only" is doing quite a bit of work. LLMs do not just spawn into existence. There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.
cmiles8 · · focus · HN ↗
Feels like Anthropic crying do as I say not as I do.
jedberg · · focus · HN ↗
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
JackFr · · focus · HN ↗
The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.
jdonaldson · · focus · HN ↗
rubicon33 · · focus · HN ↗
mitthrowaway2 · · focus · HN ↗
tsunamifury · · focus · HN ↗
Did Anthropic put work in? Yes. Did they derive their value from Humanity being open with knowledge then try to sell it back? Also yes.
Did they even steal the tech? Also yes.
_DeadFred_ · · focus · HN ↗
nonethewiser · · focus · HN ↗
ihsw · · focus · HN ↗
[dead]
layer8 · · focus · HN ↗
This isn’t meant as a moral argument, just musing about the relative cost comparison.
Ar-Curunir · · focus · HN ↗
How is this even a question.
bmacho · · focus · HN ↗
mekael · · focus · HN ↗
The knowledge that created the Haber-Bosch process [1] helps to sustain the majority of the world's populous, add another five hundred trillion dollars for that just to start with.
The creation of the printing press and all written information that allowed it to be built provided dissemination of knowledge beyond the ultra wealthy and is worth a non-finite amount of money.
LLM's are cool math, but they are less than a rounding error in comparison to even the tiniest sliver of human knowledge and technological output.
[0] <a href="https://pubmed.ncbi.nlm.nih.gov/35143880/" rel="nofollow">https://pubmed.ncbi.nlm.nih.gov/35143880/ [1] <a href="https://cen.acs.org/food/agriculture/The-industrialization-Haber-Bosch-process/101/i26" rel="nofollow">https://cen.acs.org/food/agriculture/The-industrialization-H...
DirkH · · focus · HN ↗
Ar-Curunir · · focus · HN ↗
And also, humans have been doing a lot more work than just mathematics...
phamilton · · focus · HN ↗
A training set of 15 trillion tokens is 10 trillion words.
A penny a word is cheaper than the cheapest beginner freelance writer.
That makes a training set of 10 trillion words cost $100B.
Lots of assumptions there for sure, but we're certainly in the ballpark you are describing.
ToValueFunfetti · · focus · HN ↗
louiskottmann · · focus · HN ↗
The totality of the content on internet is worth several orders of magnitude more.
layer8 · · focus · HN ↗
tpm · · focus · HN ↗
allturtles · · focus · HN ↗
jonhohle · · focus · HN ↗
layer8 · · focus · HN ↗
Image/video models are, but those weren’t the topic.
Enginerrrd · · focus · HN ↗
rsingel · · focus · HN ↗
$500B for all kinds including TV and online
$300B for newsrooms including all staff
$140B for newsroom reporters only
So yeah, I think the price of the information ingested is way higher than training costs
Ohentis · · focus · HN ↗
wonnage · · focus · HN ↗
“bro like, what if we could price the sum total of human knowledge? That wouldn’t be that much, right?”
MathiasPius · · focus · HN ↗
And the monetary cost doesn't even register when weighed against the blood, sweat and tears that went into capturing the authentic experiences of real human beings, whose honest expressions are now at least in some cases getting hoovered up, ingested, and then destroyed for all eternity, for fear that this specific work is the rounding error that might give an equally immoral competitor the edge in the bicycle-riding flamingo race that is currently consuming an absurd amount of the world's creativity and attention.
zelphirkalt · · focus · HN ↗
jedberg · · focus · HN ↗
Is that not what the foundation models are? A new reorganization of existing knowledge?
xdavidliu · · focus · HN ↗
iAMkenough · · focus · HN ↗
kbelder · · focus · HN ↗
wonnage · · focus · HN ↗
ryandrake · · focus · HN ↗
gretch · · focus · HN ↗
So then the problem is that Anthropic seems hypocritical when they knowingly insert themselves into this chain, and then complain about people down-chain from them.
To remedy the negative impressions (if they even care to do so) they should do 1 of 2 things: 1) stop complaining about it 2) stop distilling other people's work
ToucanLoucan · · focus · HN ↗
No LLM products would exist without the avalanche of largely non-consensual use of IP to create them, full stop. Any of these companies doing this and then turning around and complaining when their IP is "breached" are going to met with a chorus of tiny violins.
SR2Z · · focus · HN ↗
This generation of frontier models is "good enough." At some point, the Chinese labs will get to where the Americans are right now, they'll race to the bottom, and we will actually see what proliferation of AI looks like.
If Anthropic wants to be a trillion dollar company, they need to make revenue like Google or Apple do. Both of them have near monopolies, Anthropic is only getting further and further away as time goes on.
Of course they're gonna complain until they either figure out a better plan or accept a new valuation which is high but not spectacular.
KolibriFly · · focus · HN ↗
[dead]
scythe · · focus · HN ↗
It is not a natural assumption that someone will digitize the artwork and use it to adjust a couple thousand matrix coefficients in a complex computer program. To most people that seems like copying with extra steps. The brain may in some ways resemble a computer, but what sets it apart is that we have always lived with brains. Everything a human does has already anticipated the presence of other brains, while etched circuits on ultrapure silicon crystals are something new.
agumonkey · · focus · HN ↗
marshray · · focus · HN ↗
eithed · · focus · HN ↗
dr_dshiv · · focus · HN ↗
briansm · · focus · HN ↗
mc32 · · focus · HN ↗
It’s like FTL. Until someone realizes it, it’s just talk.
jstummbillig · · focus · HN ↗
A comparable idea could be that an encyclopedia or maths book is only distilling the things that other people did, and how dare they sell them. But the "only" is doing quite a bit of work. LLMs do not just spawn into existence. There is a body of work that they feed on, and then there is also very attributable work they do around and on top of that. All labs are struggling around the first order question: Is it okay to use prior work like this? The second order issue is still entirely reasonable to separately have and enforce rules about.
OJFord · · focus · HN ↗