‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. hedora · · focus · HN ↗
      My use of wayback has skyrocketed recently due to anti-bot measures.

      I often cannot get past captchas, and archive.org is one of the fallbacks I try.

      However, archive.is, etc are more reliable.

      I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.

      They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.

      1. account42 · · focus · HN ↗
        Ironically, archive.is itself has a captcha that doesn't like my home FF install.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.