This is a lobby organisation using only the pieces and bits they like to push their own agenda.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
Most people can't recite a book they've read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.
If you would go on a tv and read out loud a book you have not purchased rights to present, according to the most copyright laws in the world, yes you would and you would get a fine. Similarly if you broadcast a tv-show you have not purchased rights would.
Whether this is right or not is a seperate question.
> If you would go on a tv and read out loud a book
That is not the situation we are discussing. No one is arguing that what you describe is infringement.
What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.
Using analogies like "reading" to describe AI training is quite misleading imho. Training effectively embeds the book's contents into the model's weights. Changing the format doesn't change the content; a better analogy is distributing a book's text within software.
Current copyright laws are simply not prepared for this unprecedented use.
I'm not disagreeing with you, but the comparison is breaking down here. Reading a book is the specific intended purpose. It's why the book exists, not just fair use. If you couldn't remember what was going on in the book as you read it, it would be worthless.
That begs the question though, should it be? If you’re looking up the lyrics to a song you’ve bought, does it really matter, at least philosophically,
how one retrieves them?
It shouldn't be. The idea that writing down the lyrics to a song and putting them on the internet for no financial gain at all should be punishable by some big record label is the kind of over zealous IP protection that 99% of HN users would have been against until something changed a few years ago and this place became a haven for copyright extremists.
The law here is no: If you own a copy of the song, you can reproduce the lyrics in whatever way you want.
What you can't do is distribute the lyrics without the author's permission.
When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary/ISP according to the DMCA (in which case they'd be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author's permission?
Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.
My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they're an ISP. However, if they pulled them out of their own database, they're violating copyright.
Is this a direct inference call to the model or does it use a web search tool to pull the answer? In the former case I guess that could be infringement. In the latter case it is no more than a tool that facilitates it and would be not be infringement any more than Xerox is infringing if you use their copier to copy the front page of the NYT.
Those lyrics aren't in the model. They're either in the database DeepSeek makes tool calls into (unlikely), or they're out on the Internet and DeepSeek simply retrieved them on your behalf.
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.
Human memory decays; similarly, LLMs do not have perfect recall. But even if I reread a book every month for 60 years and use the knowledge in there to build a 25 billion dollar empire, the only thing the author is ever going to get from me is the $10.25 the book cost me. And no court in the western world would say he's due more.
Skyy93 · · focus · HN ↗
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
trompetenaccoun · · focus · HN ↗
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
qarl · · focus · HN ↗
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
Trusteando · · focus · HN ↗
qarl · · focus · HN ↗
Exactly analogous to a human reading the material.
asutekku · · focus · HN ↗
qarl · · focus · HN ↗
Yes - but some people can.
Are they criminals for reading books?
asutekku · · focus · HN ↗
qarl · · focus · HN ↗
You're not making sense.
The problem is the reciting. Not the reading. And hence, not the training.
asutekku · · focus · HN ↗
Whether this is right or not is a seperate question.
qarl · · focus · HN ↗
That is not the situation we are discussing. No one is arguing that what you describe is infringement.
What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.
asutekku · · focus · HN ↗
Current copyright laws are simply not prepared for this unprecedented use.
qarl · · focus · HN ↗
Reading a book embeds its contents into your brain. And yet, that is considered fair use.
I agree, the analogies are meaningless in a legal context. In the legal context, the courts disagree with you.
efreak · · focus · HN ↗
consensus1 · · focus · HN ↗
echoangle · · focus · HN ↗
It will give you the complete song lyrics which are under copyright.
qarl · · focus · HN ↗
The infringement occurs at the time of copy, not at the time of training.
ChickeNES · · focus · HN ↗
consensus1 · · focus · HN ↗
riskable · · focus · HN ↗
What you can't do is distribute the lyrics without the author's permission.
When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary/ISP according to the DMCA (in which case they'd be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author's permission?
Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.
My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they're an ISP. However, if they pulled them out of their own database, they're violating copyright.
consensus1 · · focus · HN ↗
riskable · · focus · HN ↗
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.
tomjen3 · · focus · HN ↗
Trusteando · · focus · HN ↗
[dead]