Microsoft exec called AI scraping 'the largest theft of labor in human history'
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Microsoft exec called AI scraping 'the largest theft of labor in human history'
Unofficial Hacker News client; not affiliated with Y Combinator.
jacquesm · · focus · HN ↗
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
pingou · · focus · HN ↗
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
kshri24 · · focus · HN ↗
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
TeMPOraL · · focus · HN ↗
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
kshri24 · · focus · HN ↗
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
TeMPOraL · · focus · HN ↗
That does not follow in any reasonable way.
kshri24 · · focus · HN ↗
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings).