This is a lobby organisation using only the pieces and bits they like to push their own agenda.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
Training in the US has, in fact, been decided (at least to the extent that anything has currently been decided). Bartz vs Anthropic specifically ruled training an AI model on legally owned copyrighted material is sufficiently transformative[1]:
This order grants summary judgment for Anthropic that the training use was a fair use.
And, it grants that the print-to-digital format change was a fair use for a different reason. But it
denies summary judgment for Anthropic that the pirated library copies must be treated as
training copies.
The document you linked was written a month before the Bartz decision was reached. It's also worth noting even the document you linked says this in its conclusion:
Various uses of copyrighted works in AI training are likely to be transformative. The
extent to which they are fair, however, will depend on what works were used, from what
source, for what purpose, and with what controls on the outputs—all of which can affect the
market.
That is not what the text you quoted is saying. It says that the output may be transformative, which is one criteria, but the other criteria depends on how it is used and what the source is.
Skyy93 · · focus · HN ↗
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
trompetenaccoun · · focus · HN ↗
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
qarl · · focus · HN ↗
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
trompetenaccoun · · focus · HN ↗
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
qarl · · focus · HN ↗
100% of the rulings agree with me.
The piracy is not in question. It is unarguably copyright violation.
But that's not what anyone means in this context. Training is what everyone means.
> The law is the law, there can't be different law for corporations with billions in backing.
I didn't say otherwise. That's a straw man.
latexr · · focus · HN ↗
qarl · · focus · HN ↗
mrdependable · · focus · HN ↗
<a href="https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf" rel="nofollow">https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
tpmoney · · focus · HN ↗
mrdependable · · focus · HN ↗
That is not what the text you quoted is saying. It says that the output may be transformative, which is one criteria, but the other criteria depends on how it is used and what the source is.