‹ BackHN Continuity

Thread

Microsoft exec called AI scraping 'the largest theft of labor in human history'

955 points · 841 comments · pluc

  1. jacquesm · · focus · HN ↗
    It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

    It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.

    1. user43928 · · focus · HN ↗
      It is the largest democratization of knowledge that ever happened.

      The 'sell it back to us' argument falls short in my view.

      Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.

      The comment here seems incredibly pessimistic and quite dramatical.

      1. SecretDreams · · focus · HN ↗
        Wiki already democratized it just fine and was legitimately free for people who know how to read.

        It's asinine that you think the sell it back to us argument falls short.

        Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.

        From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.

        1. user43928 · · focus · HN ↗
          Wikis made knowledge available.

          Making it accessible, understandable, and usable is another matter.

          How LLMs sound is not a fundamental limitation of the technology.

          The current model's poor writing style and tone are currently a main focus of research and I would expect improvements there soon.

          You do not have to train on anything you do not deem up to standard. This supposed poisoning of training data remains a common fantasy.

          1. SecretDreams · · focus · HN ↗
            > Making it accessible, understandable, and usable is another matter.

            I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.

            1. frozenseven · · focus · HN ↗
              Cool. Than don't use LLMs and keep waiting for that prophesied "model collapse". See how that works out for you.
              1. SecretDreams · · focus · HN ↗
                I'm more worried about the collapse of humanity a la idocracy. The model collapse might just be secondary. Even just conversing on HN, that's the vibe I feel we're headed towards. Or, perhaps, I'm just running into more and more posters that have financial ties to the success of LLMs.
                1. frozenseven · · focus · HN ↗
                  Model collapse is an induced phenomenon, not something you'll ever see in practice. Of course naive predictions of an impending 'collapse' go back to at least GPT-3 (that's over 6 years ago), yet models continue to get smarter. And speaking of predictions, I assume you've been saying that AI is "fake" for just as long?
                  1. Anamon · · focus · HN ↗
                    How do you see a possibility of avoiding model collapse? It seems to me to be an unavoidable consequence of two facts I consider pretty much irrefutable:

                    1) LLM content can at best be as good as the source material it was trained on. That's the upper bound. "Out of distribution" output of LLMs is mostly unusable.

                    2) LLM-generated content increasingly drowns out original content, online and elsewhere.

                    This spells "monotonically decreasing content quality" to me.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.