‹ BackHN Continuity

Thread

Microsoft exec called AI scraping 'the largest theft of labor in human history'

955 points · 841 comments · pluc

  1. jacquesm · · focus · HN ↗
    It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

    It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.

    1. TeMPOraL · · focus · HN ↗
      > It's the robbery of all of our culture to sell it back to us at a mark-up.

      Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.

      Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.

      And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.

      There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".

      1. cowanon77 · · focus · HN ↗
        I only see 2 consistent world views: either intellectual property is real, or it is a false concept and all information should be free.

        If IP is real, then the AI companies have performed flagrant theft.

        If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.

        The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.

        1. GuB-42 · · focus · HN ↗
          There are nuances.

          AI scraping for the goal of making a commercial LLM service is different from, say, a commercial file sharing platform.

          The first difference is that LLM training is highly transformative. Let's say your LLM ingests the Harry Potter novels during its training. What you get at the other end is not the Harry Potter novels, it is a LLM that can talk to you about Harry Potter, it is not the same thing, and going from one to the other requires a significant amount of work, very expensive work in this case.

          Not only that but there is no direct competition. People won't stop buying the Harry Potter novels because a LLM trained on it exists. If you want to read the books, you buy the books, you don't ask a LLM about it. A file sharing service on the other hand competes directly against the official channels, if you want to read the books, you can download it from this service instead of buying it on the official channels.

          So, about how free you should be to get these weights from the AI companies. If you just share a 1:1 copy of the weights, that's the "file sharing" situation, not transformative, you took their work, didn't do any of your own. Usually considered unacceptable by IP laws.

          Distillation is a more interesting case, you are using a LLM to train your own, it is transformative work, but you may also be competing directly against the LLM you are distilling. So, in a sense it is worse than scraping, but still, despite how much the likes of OpenAI and Anthropic are complaining, it seems to be legal.

          So it is somewhat consistent: 1:1 copy and distribution is not allowed, be it source material or LLM weights, and training is, be it source material or another LLM (as in distillation).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.