‹ BackHN Continuity

Thread

Microsoft director: AI scraping 'the largest theft of labor in human history'

190 points · 52 comments · jonbaer

  1. adamddev1 · · focus · HN ↗
    I wrote something original in a language learning grammar online. I coined a term to describe something about how a certain language with negative language functions.

    I tried talking to an LLM about it and sure and enough, it had scraped it and learned to use that terminology and explanation that I made up. I asked the LLM, "Where did that concept/term come from, who made it up?" It did a bunch of searching and tried to site a whole bunch of other sources which did NOT contain the terms of the explanation I had written. It simply would not cite or mention my source. It kept parroting my material semi-correctly and hiding the source. By any other actor that would be an egregious act of sloppiness, dishonesty, and plagiarism.

    Generative AI is speed-running a widespread corruption of truth. I do not believe any of the gains are worth this.

    1. baranul · · focus · HN ↗
      Your comment illustrates that AI companies are playing a kind of shell game. On the one hand there is theft of IP and violations of copyright, in the middle there is the obscuring of where things come from, and on the other hand there is the selling of it via tokens.

      More countries should debatably follow what Japan is doing[1], where they are creating new laws to force AI companies to show their training data and how they collected it.

      [1]: <a href="https:&#x2F;&#x2F;www.japantimes.co.jp&#x2F;news&#x2F;2026&#x2F;08&#x2F;19&#x2F;japan&#x2F;ai-training-data-disclosure" rel="nofollow">https:&#x2F;&#x2F;www.japantimes.co.jp&#x2F;news&#x2F;2026&#x2F;08&#x2F;19&#x2F;japan&#x2F;ai-traini...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.