‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. pantsforbirds · · focus · HN ↗
      We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

      Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

      1. subarctic · · focus · HN ↗
        What if they charged money? Is it something you'd pay for?
        1. msephton · · focus · HN ↗
          I'd pay for it, but only if they implemented the changes the community of users have been requesting for years.
          1. carlosjobim · · focus · HN ↗
            No matter what they did, you'd have a new excuse for why you won't pay.
            1. msephton · · focus · HN ↗
              Ah, the old ad hominem attack. How refreshing.

              But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but".

              It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

              1. jakderrida · · focus · HN ↗
                Is it really an ad hom if he doesn't know the hom?

                Their reply is 100% based on the content of your post.

                1. Dylan16807 · · focus · HN ↗
                  If you make up a person to insult then yeah it's still ad hominem.
              2. carlosjobim · · focus · HN ↗
                If you're donating, then you are evidently willing to pay without any "buts". So aren't you arguing against your own actions?
                1. msephton · · focus · HN ↗
                  Not at all. I donate to Internet Archive, but we're talking here about paying for unobstructed access to but one part of their service: Wayback Machine. Two different things.
                  1. carlosjobim · · focus · HN ↗
                    100% of the people who write "I would pay, if..." or "I would pay, but..." are people who are never going to pay even a dime. You might be the exception, and sorry for bunching you up with them. You have paid already by donation.

                    I think that we should all pay when asked for things which we find useful, even if they aren't perfect. If nobody else is offering anything, then we have to take what's being offered. When there's a market, more providers will begin offering their versions.

        2. bonestamp2 · · focus · HN ↗
          I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.
          1. DaSHacka · · focus · HN ↗
            I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.
            1. g-b-r · · focus · HN ↗
              If the money was guaranteed to only be used to pay the costs, there probably wouldn't be any problems
              1. usr1106 · · focus · HN ↗
                No. What happened to their remote library scheme? They did not make if for profit, but still...
            2. usr1106 · · focus · HN ↗
              Interesting, in all the years I have never noticed that IP has 2 meanings (well probably more...) Yeah, I am an engineer and usually try to avoid the legal BS. Although I hate that AI has made stealing legal if you are big enough.
            3. bonestamp2 · · focus · HN ↗
              It wouldn't be behind their back, like I said, "Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it."
              1. mitxela · · focus · HN ↗
                You wouldn't need archive.org for that - you could negotiate with the actual website. I think they'd demand quite a lot of money.
          2. progval · · focus · HN ↗
            They already provide this service at <a href="https:&#x2F;&#x2F;archive-it.org&#x2F;archive-it&#x2F;" rel="nofollow">https:&#x2F;&#x2F;archive-it.org&#x2F;archive-it&#x2F; though for some reason they don&#x27;t seem to publicize it. Some info at <a href="https:&#x2F;&#x2F;help.archive.org&#x2F;help&#x2F;archive-it-information&#x2F;" rel="nofollow">https:&#x2F;&#x2F;help.archive.org&#x2F;help&#x2F;archive-it-information&#x2F; as well.
        3. bee_rider · · focus · HN ↗
          I wonder if there would be concern on their part about appearing to be a company that was basically offering paywall circumvention as a product.
          1. cloakley · · focus · HN ↗
            It wouldnt be a paywall, more like an option for companies to not pay scrappers. At least the payment deviates to the source.
        4. pantsforbirds · · focus · HN ↗
          I mean we already paid for the article from the source itself. I guess I&#x27;d expect a better &quot;diff&quot; source from them, but if they dont even update the article itself, i guess i wouldn&#x27;t expect a paid service to have those updates either?
          1. pantsforbirds · · focus · HN ↗
            ah, i think i misunderstood your original post. if you mean the wayback-machine&#x2F;arkive, then I suspect it&#x27;d be hard to justify? You are essentially paying a third-party source to validate that diffs didn&#x27;t go through on the source material.

            with llms, at some point it probably becomes easier to use your paid api connection to manage your own cached version yourself?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.