Microsoft exec called AI scraping 'the largest theft of labor in human history'
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Microsoft exec called AI scraping 'the largest theft of labor in human history'
Unofficial Hacker News client; not affiliated with Y Combinator.
jacquesm · · focus · HN ↗
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
pingou · · focus · HN ↗
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
embedding-shape · · focus · HN ↗
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
ben_w · · focus · HN ↗
Any argument that writers and artists lose from these existing, would remain unchanged.
embedding-shape · · focus · HN ↗
CJefferson · · focus · HN ↗
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
hlynurd · · focus · HN ↗
really not the same entities here
CJefferson · · focus · HN ↗
hlynurd · · focus · HN ↗
cgio · · focus · HN ↗
ipython · · focus · HN ↗
<a href="https://www.napster.com/blog/napster-heads-to-microsoft-build-with-omniagent-api" rel="nofollow">https://www.napster.com/blog/napster-heads-to-microsoft-buil...
torginus · · focus · HN ↗
Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.
When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.
Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.
Same applies for Steam & videogames.
Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.
Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.
This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.
And once nobody cares, people even forget what a quality product is like.
TheOtherHobbes · · focus · HN ↗
Without that, everything gets atomised into lonely individualism. You sit there with your headphones on listening to [Interesting band]. Not only do you not really care because you don't feel personally connected to the music - it's one of literally more than a hundred million content items on Spotify - but you're not sharing the experience.
This seems like the loss of a valuable thing which capitalist economics can't put a price on because it has no concept of value-created-by-shared-experience.
Superficially it's the same as 'sell-content-consumption-item-to-the-mass-market' but it's fundamentally not the same kind of thing.
johnsea · · focus · HN ↗
Maybe change your perspective? Treat Spotify like a valuable audio lexicon. You read about an artist, a song, a time and immediately you can hear what is it about. Incredible!
If Spotify is only treated as a lazy background feelgood provider (while reading Marx;)), no wonder you feel that way. But it's your power/choice to appreciate it (or not), regardless of money.
autoexec · · focus · HN ↗
The same can be said for time. You said yourself it was an investment. There's no need for money to be involved since how much time you spend on a thing will establish the value of that information to you. If you must have a ritualistic act to form a connection to music, let it be the act of listening instead of paying.
hkt · · focus · HN ↗
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
ImHereToVote · · focus · HN ↗
For instance a image/video generating model.
stego-tech · · focus · HN ↗
philipallstar · · focus · HN ↗
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
card_zero · · focus · HN ↗
dbspin · · focus · HN ↗
philipallstar · · focus · HN ↗
While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
TeMPOraL · · focus · HN ↗
dbspin · · focus · HN ↗
This is the most unintentionally hilarious misunderstanding of what the RIAA does, and the power relationship between artists and publishers I've read in years. In practice the RIAA exists to maintain the copyright monopoly of a few major labels. Rent seeking from the non-artist owned catalogues of the enormous majority of musicians who never 'recoup' their initial record deal.
> While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
It's far more seductive (since it's the default) to assume class relations don't exist, and wealth distribution is meritocratic. There may not be 'goodies and baddies', but there absolutely are rentiers and workers, billionaires and plebs.
All the major labels are public companies. Which means it's the very same people - the investment class, who claim ownership and extract wealth from say Warner and Open AI (should it make any money - obviously the whole house of cards could come down first).
xienze · · focus · HN ↗
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
QuantumNomad_ · · focus · HN ↗
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
CJefferson · · focus · HN ↗
1) AI companies all get sued out of existence.
2) AI companies can train networks, those networks don't fall under copyright, so anyone can use them for anything they want without paying.
px43 · · focus · HN ↗
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
CJefferson · · focus · HN ↗
victorbjorklund · · focus · HN ↗
mitxela · · focus · HN ↗
loloquwowndueo · · focus · HN ↗
n_plus_1_acc · · focus · HN ↗
steveBK123 · · focus · HN ↗
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
robinsonb5 · · focus · HN ↗
vitorfblima · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
fzeroracer · · focus · HN ↗
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
Luker88 · · focus · HN ↗
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
barnabee · · focus · HN ↗
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
jappgar · · focus · HN ↗
Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
kshri24 · · focus · HN ↗
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
echoangle · · focus · HN ↗
kshri24 · · focus · HN ↗
The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.
echoangle · · focus · HN ↗
Yeah, which is why it's only used to compare cars etc. among each other. Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.
In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
kshri24 · · focus · HN ↗
Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.
> In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).
> Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.
echoangle · · focus · HN ↗
I was talking about humans vs. cars as an analogy to you comparing GPUs and cars.
Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?
> A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
kshri24 · · focus · HN ↗
I was talking about cars vs horses. HP is Horse-power not human-power.
> Nobody is taxing humans the way cars are taxed
Cars are not free to roam the road. I don't know why it is so hard for you to understand that we use metrics like HP/CC etc to equalize with humans so that automobiles can be taxed just like humans. Without metrics like HP/CC etc there is no way to tax cars. Get it?
> Nobody is taxing humans the way cars are taxed
Duh. It is the opposite. We use deterministic metrics for automobiles to tax them the way WE ALREADY HAVE BEEN TAXING HUMANS for thousands of years. Get it? Cars did not come first. Humans came first.
> so why should the computing speed of a GPU be compared to that of a human?
Because, believe it or not, we built computers to replace humans. The very point of computing was because humans are "SLOW" to do mundane computations, repeatedly, with 100% efficiency and not be subjected to biological functions like wanting to eat, sleep or shit. So there is a direct connection between the computing speed and human replacement. There used to be a time where CPU computing speed was touted in terms of how many humans it replaced... IBM's 1951 Electronic Calculator ad about "150 extra engineers" makes the point. We have ALWAYS built computing as a proxy to human replacement. Heck, even AI is touted to replace humans by the very same people who are training the models.
Now it is quite ridiculous to then turn around and ask why should computing speed of a GPU be compared to that of a human. It is literally the building block of model training/evals/inference. The very basis for automation that is replacing humans. Obviously people are going to relate the two together.
> And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
Come on you are clutching at straws here. It is not about "making sense". It is about using a metric to equalize unequal entities. When I am already saying they are unequal and have no direct relation to each other and any relation can only be arrived at indirectly. FLOPS is just an example I gave. I am not saying we should literally go with the FLOPS example itself. But we can use any metric and equalize it with human work. That's all I am getting it. It is the same argument as Horse-power.
> So you want to compare the learning rate of a GPU to that of a human by comparing their respective FLOPS. Why would FLOPS be a valid proxy for learning ability in humans just because that works out in GPUs,
Because that is the only metric we can use to measure how quickly GPUs can process arithmetic (you can label it "training" or "learning" or whatever name you want). There is no other metric that is deterministic and comparable to something humans do (which is also process arithmetic).
> if the way they learn is fundamentally different?
It does not matter if how they learn is fundamentally different. Automobiles use an engine to move around. Humans use legs. We both are still taxed. Automobiles are taxed on metrics like HP/CC etc. Humans are taxed via betterment taxes while purchasing property and yearly property tax. The point I am making is that it is not free to ride an automobile on the roads which are built for pedestrians fundamentally. Hence why "right of way" is for pedestrians first and foremost. Because roads existed before automobiles or any animal-drawn cart ever existed. We figured out a way to tax automobiles by way of metrics that is deterministic and quantifiable. You may ask why should cars be taxed on HP/CC while humans are not taxed on their legs etc. That is totally missing the point being made.
> Is the effect of someone reading a copyrighted book dependent on how fast they are at doing math in their head?
It is fundamentally math. Every physical law in the Universe is expressed and backed by math. So on a fundamental level, yes "reading" is essentially maths only.
b3lvedere · · focus · HN ↗
I went to the library. Didn't pay a cent.
Now what?
kshri24 · · focus · HN ↗
> Didn't pay a cent.
Taxpayers did pay on your behalf.
Nothing is free. Except ofcourse stealing, which is free.
echoangle · · focus · HN ↗
If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Can we not just stick to calling it copyright infringement?
kshri24 · · focus · HN ↗
No.
> If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Not mine. But the company that made the coffee machine. It is stealing IP.
> Can we not just stick to calling it copyright infringement?
It is just a fancy way of saying you stole someone's IP. You can call it infringement if it makes you feel good. But the act is the same end of the day.
b3lvedere · · focus · HN ↗
kshri24 · · focus · HN ↗
Public domain on the other hand is legally only possible if/when copyright has expired. That means the owner has enjoyed proceeds from copyright protection for more than his own lifetime. That is fair. It is still not comparable.
EDIT: since you tacked on more like "wind, gravity, radioactivity" etc, I would still not classify them as "free". They are invaluable to very existence of life.
"Knowledge passed on" is also after someone (in ancestry) has paid for it through blood, sweat and tears. It isn't "free". "Public domain" is legally recognized form of "knowledge passed on".
b3lvedere · · focus · HN ↗
kshri24 · · focus · HN ↗
They can't claim. That's the point. They are an invaluable resource precisely because they cannot be OWNED by anyone. They are not FREE.
Hence why even corporate entities that deal with solar, wind etc talking about "HARNESSING" energy. They don't talk about OWNERSHIP of energy.
b3lvedere · · focus · HN ↗
TeMPOraL · · focus · HN ↗
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
kshri24 · · focus · HN ↗
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
TeMPOraL · · focus · HN ↗
That does not follow in any reasonable way.
kshri24 · · focus · HN ↗
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings).
Dead_Lemon · · focus · HN ↗
2frrrr · · focus · HN ↗
ElProlactin · · focus · HN ↗
dofm · · focus · HN ↗
I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
the_other · · focus · HN ↗
AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
tripzilch · · focus · HN ↗
I mean it's already happened, right?
I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
saynay · · focus · HN ↗
It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
mina_bridge · · focus · HN ↗
[dead]