100%. First stealing content from creators and pirating TBs of books. And now this. Baffles my mind that these labs can just get away with their 'rogue' agents trying to hack into other systems.
Imagine if a human did that. FBI would be knocking on their door.
The argument for copyright violation is very weak. Yes, AI training reads much copyrighted material. So do researchers of all types. That's why there are academic libraries. A copyright violation only exists if what comes out is a close match to what went in. That's happened, but it was considered a bug and was fixed a year or two ago.
No, the copyright violation is so obvious that people arguing against the object have to spontaneously assert bizarre qualifications for copyright violation that have never existed until LLM companies wanted to freely violate copyrights. Just about a decade ago, people were getting sued for "look and feel." "Blurred Lines" lost a copyright suit for reminding people of a song by another artist.
The metaphor pretended to be reality of computers "learning" being the same as human learning is their last resort, and it is obviously silly. Computers are not human or animal, and learning is something that humans and animals do. If computers are human, then copying a file verbatim to a computer is learning. It's not even worth acknowledging - learning is a metaphor for training LLMs. When I say "you" to an LLM, I'm not referring to anyone, I'm dealing with a UI.
I don't doubt that a lot of AI people have internalized this silliness, which is how they can humor fantasies of how a bunch of programs on different computers explicitly evoked to do particular things might be alive because they can talk with it. I can't talk with my dog, and my dog is alive - so why should having a quality that my dog is incapable of be proof of life? You can certainly conceive of AI that it would be very difficult to say for sure isn't alive, and this certainly is not it. LLMs learn like books speak.
There is no need to invent “bizarre qualifications” for copyright violations, those metaphors exist to explain what is happening in in layman’s terms. Training extracts patterns in the data and encodes them as tiny perturbations in a gazillion weights. That is analogous to “learning” because those patterns represent abstract concepts an relations between them that can be used to “reason” about related topics in response to a prompt.
In the above description, at no point does the actual verbatim content exist anywhere in the model, and copyright law being about rights to reproducing copies (verbatim or substantial portions thereof) does not really apply and does not need any exceptions or qualifications. (If you’re thinking of regurgitation, you should look into studies about it to see how vanishingly rare it is.)
Why would it matter if I stored exact copies of copyrighted works? Storage format is irrelevant. All you need to say is in that last sentence: it’s all about whether the output is considered a derivative work or a too close of a copy.
But that's the thing, LLMs don't store exact copies and neither can they reproduce them. The cases of regurgitation have only been shown to work for a very small handful of extremely popular works.
uxcolumbo · · focus · HN ↗
Imagine if a human did that. FBI would be knocking on their door.
Why are there zero consequences for these labs?
Animats · · focus · HN ↗
pessimizer · · focus · HN ↗
The metaphor pretended to be reality of computers "learning" being the same as human learning is their last resort, and it is obviously silly. Computers are not human or animal, and learning is something that humans and animals do. If computers are human, then copying a file verbatim to a computer is learning. It's not even worth acknowledging - learning is a metaphor for training LLMs. When I say "you" to an LLM, I'm not referring to anyone, I'm dealing with a UI.
I don't doubt that a lot of AI people have internalized this silliness, which is how they can humor fantasies of how a bunch of programs on different computers explicitly evoked to do particular things might be alive because they can talk with it. I can't talk with my dog, and my dog is alive - so why should having a quality that my dog is incapable of be proof of life? You can certainly conceive of AI that it would be very difficult to say for sure isn't alive, and this certainly is not it. LLMs learn like books speak.
keeda · · focus · HN ↗
In the above description, at no point does the actual verbatim content exist anywhere in the model, and copyright law being about rights to reproducing copies (verbatim or substantial portions thereof) does not really apply and does not need any exceptions or qualifications. (If you’re thinking of regurgitation, you should look into studies about it to see how vanishingly rare it is.)
catlifeonmars · · focus · HN ↗
keeda · · focus · HN ↗