> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
I'm guessing you use search engines, right? Those use scrapers and have to use scrapers. It's how they work.
A "scraper" is simply an automated process that fetches URLs intended for display to a human, and processes it. The act of scraping doesn't imply anything about:
1. The frequency of the fetches,
2. The way that the resulting page is processed.
Search engines scrape. Again, they have to. Same goes for archive.org.
Thing is, there aren't tens of thousands of search engines/archive.orgs that can overload a site at once.
simonw · · focus · HN ↗
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
packetslave · · focus · HN ↗
bsimpson · · focus · HN ↗
gambiting · · focus · HN ↗
ValentineC · · focus · HN ↗
petcat · · focus · HN ↗
sandcat_ · · focus · HN ↗
petcat · · focus · HN ↗
sandcat_ · · focus · HN ↗
petcat · · focus · HN ↗
The end result is exactly the same.
Sophira · · focus · HN ↗
A "scraper" is simply an automated process that fetches URLs intended for display to a human, and processes it. The act of scraping doesn't imply anything about:
1. The frequency of the fetches,
2. The way that the resulting page is processed.
Search engines scrape. Again, they have to. Same goes for archive.org.
Thing is, there aren't tens of thousands of search engines/archive.orgs that can overload a site at once.