‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. lousken · · focus · HN ↗
    AI companies should pay billions to wayback machine for access
    1. KPGv2 · · focus · HN ↗
      I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.
      1. roblh · · focus · HN ↗
        Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?
        1. [deleted] · · focus · HN ↗

          [deleted]

        2. Joel_Mckay · · focus · HN ↗

          [dead]

        3. mitxela · · focus · HN ↗
          It should, but it doesn't.
        4. KPGv2 · · focus · HN ↗
          no. The topic at hand is distribution of copyrighted material, and you're talking about reproduction and possibly preparation of a derivative work.

          At last in the USA, copyright law defines specific things copyright owners have exclusive rights to:

          - reproduction - preparation of derivative works - distribution of copies to the public - public performance - public display

          The most immediate issue with AI companies is whether they've made infringing reproductions.

          The other possibility is the preparation of derivative works: does an AI response count as derivative of something it's consumed?

          Sorry I'm not going to do the analysis for you, though. I'm no longer a bright, chipper IP law scholar.

      2. jMyles · · focus · HN ↗
        It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.
        1. autoexec · · focus · HN ↗
          I'd have a lot less of a problem with AI if everything that went into their training was public domain and made easily available to anyone for any use. It'd feel less like AI companies were just stealing the work of others and charging for it.
          1. jMyles · · focus · HN ↗
            Seems like a reasonable norm:

            * If you train AI on it, you have to afford public access to it.

            * Nobody can exact violence against anybody else in response to that person providing public access to any data anymore (ie, all bytestrings are public domain).

            That's the world I'd like to try in the coming years.

      3. autoexec · · focus · HN ↗
        They wouldn't be paying for the content, just the bandwidth. Like buying a linux OS on a CD ROM was about the cost of media not profiting off of the software.
      4. 0xDEAFBEAD · · focus · HN ↗
        Isn't that already a big part of reddit's business model?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.