> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
simonw · · focus · HN ↗
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
packetslave · · focus · HN ↗
bsimpson · · focus · HN ↗
gambiting · · focus · HN ↗
ValentineC · · focus · HN ↗
petcat · · focus · HN ↗
organsnyder · · focus · HN ↗
petcat · · focus · HN ↗
Hence, distinction without a difference.
fluffybucktsnek · · focus · HN ↗
petcat · · focus · HN ↗
So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.
publlus_enigma · · focus · HN ↗
Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.