‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. packetslave · · focus · HN ↗
      This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
      1. bsimpson · · focus · HN ↗
        It's an open secret that you can often circumvent paywalls by searching Wayback.
        1. gambiting · · focus · HN ↗
          Every single paid article linked on HN has the way back machine link as the very first comment.
          1. ValentineC · · focus · HN ↗
            The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
            1. petcat · · focus · HN ↗
              ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
              1. sandcat_ · · focus · HN ↗
                That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
                1. petcat · · focus · HN ↗
                  It's bad form to scrape the scrapers?
                  1. sandcat_ · · focus · HN ↗
                    Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
                    1. petcat · · focus · HN ↗
                      You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason.

                      The end result is exactly the same.

                      1. fc417fc802 · · focus · HN ↗
                        It is different, precisely because the end result is not the same - one broadly benefits the public while the other doesn't.

                        Substitute almost any disruptive public service to see the issue with your line of reasoning. For example - you seem to think that [ bulldozing private property ] to "construct an emergency fire break" is somehow different than [ bulldozing private property ] for any other reason.

                        Never mind that the sort of scraping being objected to is actually harmful to service health while what the wayback machine does is almost entirely unnoticeable.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.