‹ BackHN Continuity

Thread

Microsoft exec called AI scraping 'the largest theft of labor in human history'

955 points · 841 comments · pluc

  1. jacquesm · · focus · HN ↗
    It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

    It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.

    1. user43928 · · focus · HN ↗
      It is the largest democratization of knowledge that ever happened.

      The 'sell it back to us' argument falls short in my view.

      Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.

      The comment here seems incredibly pessimistic and quite dramatical.

      1. oblio · · focus · HN ↗
        > Free versions are abundant,

        Awesome, can I make my own competitive LLM, just like I can make my own open source software?

        > and in some time useful models will ship preinstalled on all mobile phones.

        Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.

        1. user43928 · · focus · HN ↗
          No. I also cannot make a competitive computer, yet can use one for my work to remain competitive.

          Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.

          But what are 10 years in the grand scheme of things?

          Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?

          1. oblio · · focus · HN ↗
            > Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?

            Nah, we should have:

            1. invested less, in a more targeted way

            2. ideally like ARPANET, with the benefits given to all humanity

            3. with fair royalties paid to all (where relevant)

            4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM

            ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top.

            Exactly like the internet.

          2. SecretDreams · · focus · HN ↗
            > and never developed AI, because it will take a decade to disseminate the benefits to everyone?

            Yes, probably. The benefits of dissemination already existed. The internet was free. Libraries are free. You deny people basic reading and comprehension growth by giving a distilled, without thought, answer.

            You say "some time in the future this might be available on our phones for free" and "what's 10 years in the grand scheme of things?"

            We'll, I'll argue that in 10 years from now we may see the full damage of what doing this has done to us and I'd rather not wait for "the what's 10 years in the grand scheme of things" to playout and irreversibly damage an entire generation the way we let the unmitigated and unregulated rollout of social media do the same to the most recent generation.

            The risk/reward of this technology is unproven regarding the long-term effects on developing minds. We are beta testing the bullshit dreams of a couple billionaire techbros on an entire cohort of kids and young adults.

            1. user43928 · · focus · HN ↗
              I couldn't disagree more.

              What I expect to see in 10 years is breakthroughs in science and medicine, with major diseases becoming treatable.

              You can disagree, but on what basis do you think your predictions are more likely to be correct?

              Humanity has turned out fine despite the invention of the book, the TV, and then social media. I reckon kids will develop just fine with AI too.

              1. SecretDreams · · focus · HN ↗
                It's absurd to equate books and social media to being akin. And you're being facetious in your comments.

                What if you are wrong? What are the risks vs rewards?

                1. user43928 · · focus · HN ↗
                  The main risk you have presented is a supposed negative effect on children's development. I am not convinced that there will be a negative effect at all.

                  On the reward side we have potentially curing most major disease, automating labor and freeing humanity from having to work for a living, as well as perhaps generally advancing science at an unprecedented pace. And of course improving education.

                  1. SecretDreams · · focus · HN ↗
                    You not being convinced does not mean we do not need to protect for the risk.

                    I suspect there's a lot of things you're not convinced of. Probably some things that are financially misaligned with your own welfare.

        2. aeon_ai · · focus · HN ↗
          > Can I make my own competitive LLM?

          Yes, yes you can.

          See the number of startups that have finetuned or trained an OSS model to build their business on.

          “But you have to have compute!”

          Ok, and you’re writing OSS on a rock with no internet connection?

          The world has all kinds of barriers, but if anyone has a chance of competing on the LLM from it’s not going to be by getting rid of fair use.

        3. oblio · · focus · HN ↗
          Edit: There.is.no.hope.reality.has.left.the.building.

          Oh, well, I guess I'll just wait for the bubble to pop for the over-excited SF fans to cool down.

          I'm having a TON of flashbacks about cryptocurrency discussions.

      2. colesantiago · · focus · HN ↗
        It is woefully pessimistic indeed.

        There will be new jobs coming from this.

        The bountiful abundance of intelligence is truly the best thing that has happened this decade.

      3. SecretDreams · · focus · HN ↗
        Wiki already democratized it just fine and was legitimately free for people who know how to read.

        It's asinine that you think the sell it back to us argument falls short.

        Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.

        From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.

        1. user43928 · · focus · HN ↗
          Wikis made knowledge available.

          Making it accessible, understandable, and usable is another matter.

          How LLMs sound is not a fundamental limitation of the technology.

          The current model's poor writing style and tone are currently a main focus of research and I would expect improvements there soon.

          You do not have to train on anything you do not deem up to standard. This supposed poisoning of training data remains a common fantasy.

          1. SecretDreams · · focus · HN ↗
            > Making it accessible, understandable, and usable is another matter.

            I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.

            1. frozenseven · · focus · HN ↗
              Cool. Than don't use LLMs and keep waiting for that prophesied "model collapse". See how that works out for you.
              1. SecretDreams · · focus · HN ↗
                I'm more worried about the collapse of humanity a la idocracy. The model collapse might just be secondary. Even just conversing on HN, that's the vibe I feel we're headed towards. Or, perhaps, I'm just running into more and more posters that have financial ties to the success of LLMs.
                1. frozenseven · · focus · HN ↗
                  Model collapse is an induced phenomenon, not something you'll ever see in practice. Of course naive predictions of an impending 'collapse' go back to at least GPT-3 (that's over 6 years ago), yet models continue to get smarter. And speaking of predictions, I assume you've been saying that AI is "fake" for just as long?
                  1. SecretDreams · · focus · HN ↗
                    I use AI for coding regularly. It's not fake. Just most aspects outside of the coding are looking to be a net negative on humanity if we think comprehensive and critical thinking are positive aspects of the human condition.
                  2. Anamon · · focus · HN ↗
                    How do you see a possibility of avoiding model collapse? It seems to me to be an unavoidable consequence of two facts I consider pretty much irrefutable:

                    1) LLM content can at best be as good as the source material it was trained on. That's the upper bound. "Out of distribution" output of LLMs is mostly unusable.

                    2) LLM-generated content increasingly drowns out original content, online and elsewhere.

                    This spells "monotonically decreasing content quality" to me.

      4. Zsfe510asG · · focus · HN ↗
        There are no viable free versions with sufficient computing power. The do-it-yourself AI is dangled as a carrot in front of users to camouflage the lock-in and rent seeking by Big AI. Go prove something new like Navier Stokes (replicating N-S itself no longer counts due to scraping and plagiarism) on your Mac Pro!

        Paid influencers who perpetuate the open narrative are a whole new industry.

        Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".

        1. bhelkey · · focus · HN ↗
          There is no lock in. Anyone with sufficient budget can download the weights for GLM 5.3. On OpenRouter, I count 29 different providers for this model [1].

          In terms of purely local LLMs, one can run GLM 5.3 flash on a beefy workstation.

          [1] <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;z-ai&#x2F;glm-5.3" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;z-ai&#x2F;glm-5.3

      5. canelonesdeverd · · focus · HN ↗
        The comment seems in line with real life, while your catchphrase about democratization is more like wishful thinking.
      6. puelocesar · · focus · HN ↗
        Ah, this remembers me when I was also young and naive.

        Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.