‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. toomuchtodo · · focus · HN ↗
      It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

      <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tragedy_of_the_commons" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tragedy_of_the_commons

      (no affiliation)

      1. ronsor · · focus · HN ↗
        Reddit has no excuses for the anonymous old.reddit.com removal; they&#x27;re simply greedy.

        On the other hand, the Internet Archive is a non-profit offering a free public resource.

        1. toomuchtodo · · focus · HN ↗
          Examples provided as technical examples, strong feelings are out of scope for this thread.
          1. itintheory · · focus · HN ↗
            As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we&#x27;d be looking at at least 250k&#x2F;yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it&#x27;s increasingly clear that this is a temporary bandaid.

            The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

            [0] <a href="https:&#x2F;&#x2F;people.kernel.org&#x2F;monsieuricon&#x2F;creepy-crawlies" rel="nofollow">https:&#x2F;&#x2F;people.kernel.org&#x2F;monsieuricon&#x2F;creepy-crawlies

            1. toomuchtodo · · focus · HN ↗
              No strong feelings here is what I meant. Certainly, that energy is best directed into aggressive countermeasures and defense in depth of public goods.

              <a href="https:&#x2F;&#x2F;hn.algolia.com&#x2F;?dateRange=all&amp;page=0&amp;prefix=true&amp;query=author%3Adang%20%E2%80%9Cstrong%20feelings%E2%80%9D&amp;sort=byDate&amp;type=comment" rel="nofollow">https:&#x2F;&#x2F;hn.algolia.com&#x2F;?dateRange=all&amp;page=0&amp;prefix=true&amp;que...

            2. userbinator · · focus · HN ↗
              I say put your data up in torrents, host a few KB of plain HTML linking to them, and let decentralisation do the rest.
              1. itintheory · · focus · HN ↗
                The data IS available. You can download it all from several sources in one big dump. And yet we&#x27;re still scraped.
                1. userbinator · · focus · HN ↗
                  Further evidence that they&#x27;re not actually going after your data, but just DDoS&#x27;ing.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.