‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. Kodiack · · focus · HN ↗
      I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests.

      However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.

      1. sippingabonedry · · focus · HN ↗
        How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks?

        Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

        1. fc417fc802 · · focus · HN ↗
          I thought at least google (and possibly others) provided a way to verify the user agent?
          1. usr1106 · · focus · HN ↗
            Sorry, not following. I thought the user agent is a string that the caller can set to anything. There is no immediate, reliable way to tell whether the string is correct.

            Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.

            What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.

            1. grumbelbart2 · · focus · HN ↗
              Google's scraper bot at least used to be behind IPs that you could identify via reverse-then-forward DNS. Not sure if that is still up to date, though.

              <a href="https:&#x2F;&#x2F;developers.google.com&#x2F;search&#x2F;blog&#x2F;2006&#x2F;09&#x2F;how-to-verify-googlebot?hl=en" rel="nofollow">https:&#x2F;&#x2F;developers.google.com&#x2F;search&#x2F;blog&#x2F;2006&#x2F;09&#x2F;how-to-ver...

              1. dewey · · focus · HN ↗
                Yep, verifying the IPs is still the way to go. You often see websites that do it wrong when you set your user agent to Google Bot and they give you a different version of the page without validating that.
        2. Kodiack · · focus · HN ↗
          They have their own ASN, which I’ve explicitly allowed requests from.

          <a href="https:&#x2F;&#x2F;www.peeringdb.com&#x2F;asn&#x2F;7941" rel="nofollow">https:&#x2F;&#x2F;www.peeringdb.com&#x2F;asn&#x2F;7941

          1. junon · · focus · HN ↗
            TIL they have their own ASN. This is helpful, thanks.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.