At 15m18s he quotes LG Research "Common corpus contains only 20% legally allowed-to-be-used data". But he doesn't comment on how Linux itself legally navigates accepting patches from evidently illegally sourced means. He just says, "We'll let the courts deal with that". So what happens if courts do decide that it is illegal to use LLM output that's strikingly similar to copyrighted training data? Does anybody know of any discussions that have happened about this from within the Linux project?
I'm fully aware of the philosophical arguments about transformative use etc. But I'm more interested in the concrete reality of high profile projects navigating this novel legal territory.
This was something I noticed as well, I was hoping that an audience member would pick them up on it. It seems incredibly risky to commit anything produced by these tools when they were trained on such dodgy data
tombh · · focus · HN ↗
I'm fully aware of the philosophical arguments about transformative use etc. But I'm more interested in the concrete reality of high profile projects navigating this novel legal territory.
whateverboat · · focus · HN ↗
This is not a new problem. AI just changes the scale.
20k · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]