‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. packetslave · · focus · HN ↗
      This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
      1. bsimpson · · focus · HN ↗
        It's an open secret that you can often circumvent paywalls by searching Wayback.
        1. gambiting · · focus · HN ↗
          Every single paid article linked on HN has the way back machine link as the very first comment.
          1. ValentineC · · focus · HN ↗
            The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
            1. eek2121 · · focus · HN ↗
              Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
              1. normie3000 · · focus · HN ↗
                > Why folks continue to use them confuses me.

                I use them. I haven't ever heard mention that the content is edited. Do you have a source?

                1. Mogzol · · focus · HN ↗
                  See the &quot;Background&quot; section of the Wikipedia RFC on banning archive.today links: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Requests_for_comment&#x2F;Archive.is_RFC_5" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Requests_for_comment...

                  They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

                  1. sam345 · · focus · HN ↗
                    Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it&#x27;s fine for gossip and speculation but it shouldn&#x27;t be in Ars Technica then.
                    1. fc417fc802 · · focus · HN ↗
                      Because a lot of us watched the drama unfold in real time.
                      1. avadodin · · focus · HN ↗
                        We have always been at war with Eastasia.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.