‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. packetslave · · focus · HN ↗
      This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
      1. bsimpson · · focus · HN ↗
        It's an open secret that you can often circumvent paywalls by searching Wayback.
        1. gambiting · · focus · HN ↗
          Every single paid article linked on HN has the way back machine link as the very first comment.
          1. ValentineC · · focus · HN ↗
            The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
            1. petcat · · focus · HN ↗
              ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
              1. organsnyder · · focus · HN ↗
                They're different sites, with different goals, run by different people.
                1. petcat · · focus · HN ↗
                  That provide the same functional service....

                  Hence, distinction without a difference.

                  1. fluffybucktsnek · · focus · HN ↗
                    Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
                    1. petcat · · focus · HN ↗
                      Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access.

                      So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

                      1. HDBaseT · · focus · HN ↗
                        I think you have the wrong impression of the Internet Archive.

                        The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".

                      2. fluffybucktsnek · · focus · HN ↗
                        Internet Archive's traffic may not matter to you, but that's the main topic of this discussion, regardless of what you care or use website archival tools for.
                      3. publlus_enigma · · focus · HN ↗
                        I suspect you may be conflating two different things.

                        Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.

                  2. celsoazevedo · · focus · HN ↗
                    They are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc.

                    I think it's a distinction worth making.

                    Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.

                  3. rpdillon · · focus · HN ↗
                    Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.
                    1. DaSHacka · · focus · HN ↗
                      Exactly this

                      archive.org is the more straight-laced archive that doesn't circumvent sites that try to block it, and removes content they deem 'problematic' even if not illegal or requested by the site owner.

                      Meanwhile archive.today/ph/is/etc is the guerrilla alternative run by a die-hard datahoarder that seeks to archive the information itself, bypassing whatever blockers/login pages/whathaveyou to achieve the result.

                      It's nice to have both options. When I archive a site, I usually use both for added resiliency.

                  4. mitxela · · focus · HN ↗
                    No they don't. Archive.org is co-operative, it respects robots.txt and allows deletion. It's also very slow. Archive.* is adversarial and archives sites that don't like it. That's why the FBI is trying to take it down.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.