This is a lobby organisation using only the pieces and bits they like to push their own agenda.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
That begs the question though, should it be? If you’re looking up the lyrics to a song you’ve bought, does it really matter, at least philosophically,
how one retrieves them?
It shouldn't be. The idea that writing down the lyrics to a song and putting them on the internet for no financial gain at all should be punishable by some big record label is the kind of over zealous IP protection that 99% of HN users would have been against until something changed a few years ago and this place became a haven for copyright extremists.
The law here is no: If you own a copy of the song, you can reproduce the lyrics in whatever way you want.
What you can't do is distribute the lyrics without the author's permission.
When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary/ISP according to the DMCA (in which case they'd be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author's permission?
Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.
My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they're an ISP. However, if they pulled them out of their own database, they're violating copyright.
Is this a direct inference call to the model or does it use a web search tool to pull the answer? In the former case I guess that could be infringement. In the latter case it is no more than a tool that facilitates it and would be not be infringement any more than Xerox is infringing if you use their copier to copy the front page of the NYT.
Those lyrics aren't in the model. They're either in the database DeepSeek makes tool calls into (unlikely), or they're out on the Internet and DeepSeek simply retrieved them on your behalf.
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.
Skyy93 · · focus · HN ↗
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
trompetenaccoun · · focus · HN ↗
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
qarl · · focus · HN ↗
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
Trusteando · · focus · HN ↗
consensus1 · · focus · HN ↗
echoangle · · focus · HN ↗
It will give you the complete song lyrics which are under copyright.
qarl · · focus · HN ↗
The infringement occurs at the time of copy, not at the time of training.
ChickeNES · · focus · HN ↗
consensus1 · · focus · HN ↗
riskable · · focus · HN ↗
What you can't do is distribute the lyrics without the author's permission.
When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary/ISP according to the DMCA (in which case they'd be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author's permission?
Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.
My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they're an ISP. However, if they pulled them out of their own database, they're violating copyright.
consensus1 · · focus · HN ↗
riskable · · focus · HN ↗
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.