Now when the big LLM companies create their products by accumulating a notable portion of copyrighted creative works (including computer programs) from Internet, it does not count as copyright infringement or competition (or "theft"). It is considered as "fair use". So, why training LLMs on other LLMs is not fair use too?
> If you don’t like that, if you don’t like people to use your products, all you [have to do is] know your customers, and disable the service
Considering the current situation, I understand Mr. Huang point, but that's not how the copyright framework assumed to be working from the beginning. It should protect both small actors (authors) and the big companies from unrestricted use of creative works. Now this mechanism seems to be practically dysfunctional.
And a big portion of this lies on shoulders of proponents of permissive OSS, who defend an idea of (almost) unrestricted use of their source code texts for many years, and long before mass LLM scrapping became a thing.
Eliah_Lakhin · · focus · HN ↗
> If you don’t like that, if you don’t like people to use your products, all you [have to do is] know your customers, and disable the service
Considering the current situation, I understand Mr. Huang point, but that's not how the copyright framework assumed to be working from the beginning. It should protect both small actors (authors) and the big companies from unrestricted use of creative works. Now this mechanism seems to be practically dysfunctional.
And a big portion of this lies on shoulders of proponents of permissive OSS, who defend an idea of (almost) unrestricted use of their source code texts for many years, and long before mass LLM scrapping became a thing.