This is a lobby organisation using only the pieces and bits they like to push their own agenda.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
Those lyrics aren't in the model. They're either in the database DeepSeek makes tool calls into (unlikely), or they're out on the Internet and DeepSeek simply retrieved them on your behalf.
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.
Skyy93 · · focus · HN ↗
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
trompetenaccoun · · focus · HN ↗
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
qarl · · focus · HN ↗
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
Trusteando · · focus · HN ↗
consensus1 · · focus · HN ↗
echoangle · · focus · HN ↗
It will give you the complete song lyrics which are under copyright.
riskable · · focus · HN ↗
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.