‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. packetslave · · focus · HN ↗
      This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
      1. bsimpson · · focus · HN ↗
        It's an open secret that you can often circumvent paywalls by searching Wayback.
        1. gambiting · · focus · HN ↗
          Every single paid article linked on HN has the way back machine link as the very first comment.
          1. ValentineC · · focus · HN ↗
            The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
            1. petcat · · focus · HN ↗
              ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
              1. organsnyder · · focus · HN ↗
                They're different sites, with different goals, run by different people.
                1. petcat · · focus · HN ↗
                  That provide the same functional service....

                  Hence, distinction without a difference.

                  1. fluffybucktsnek · · focus · HN ↗
                    Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
                    1. petcat · · focus · HN ↗
                      Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access.

                      So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

                      1. HDBaseT · · focus · HN ↗
                        I think you have the wrong impression of the Internet Archive.

                        The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".

                      2. fluffybucktsnek · · focus · HN ↗
                        Internet Archive's traffic may not matter to you, but that's the main topic of this discussion, regardless of what you care or use website archival tools for.
                      3. publlus_enigma · · focus · HN ↗
                        I suspect you may be conflating two different things.

                        Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.

                  2. celsoazevedo · · focus · HN ↗
                    They are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc.

                    I think it's a distinction worth making.

                    Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.

                  3. rpdillon · · focus · HN ↗
                    Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.
                    1. DaSHacka · · focus · HN ↗
                      Exactly this

                      archive.org is the more straight-laced archive that doesn't circumvent sites that try to block it, and removes content they deem 'problematic' even if not illegal or requested by the site owner.

                      Meanwhile archive.today/ph/is/etc is the guerrilla alternative run by a die-hard datahoarder that seeks to archive the information itself, bypassing whatever blockers/login pages/whathaveyou to achieve the result.

                      It's nice to have both options. When I archive a site, I usually use both for added resiliency.

                  4. mitxela · · focus · HN ↗
                    No they don't. Archive.org is co-operative, it respects robots.txt and allows deletion. It's also very slow. Archive.* is adversarial and archives sites that don't like it. That's why the FBI is trying to take it down.
              2. sandcat_ · · focus · HN ↗
                That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
                1. petcat · · focus · HN ↗
                  It's bad form to scrape the scrapers?
                  1. sandcat_ · · focus · HN ↗
                    Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
                    1. petcat · · focus · HN ↗
                      You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason.

                      The end result is exactly the same.

                      1. DaSHacka · · focus · HN ↗
                        The minuscule traffic generated by the wayback machine, which serves to preserve the content for years to come, is completely incomparable to the scrapers that hammer every single href linked on a website.
                      2. sippingabonedry · · focus · HN ↗

                        [dead]

                      3. fc417fc802 · · focus · HN ↗
                        It is different, precisely because the end result is not the same - one broadly benefits the public while the other doesn't.

                        Substitute almost any disruptive public service to see the issue with your line of reasoning. For example - you seem to think that [ bulldozing private property ] to "construct an emergency fire break" is somehow different than [ bulldozing private property ] for any other reason.

                        Never mind that the sort of scraping being objected to is actually harmful to service health while what the wayback machine does is almost entirely unnoticeable.

                      4. Sophira · · focus · HN ↗
                        I'm guessing you use search engines, right? Those use scrapers and have to use scrapers. It's how they work.

                        A "scraper" is simply an automated process that fetches URLs intended for display to a human, and processes it. The act of scraping doesn't imply anything about:

                        1. The frequency of the fetches,

                        2. The way that the resulting page is processed.

                        Search engines scrape. Again, they have to. Same goes for archive.org.

                        Thing is, there aren't tens of thousands of search engines/archive.orgs that can overload a site at once.

            2. eek2121 · · focus · HN ↗
              Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
              1. DaSHacka · · focus · HN ↗
                Ironically, your framing of the situation is infinitely more disingenuous to push a personal agenda versus anything the archive.today guy did.
              2. normie3000 · · focus · HN ↗
                > Why folks continue to use them confuses me.

                I use them. I haven't ever heard mention that the content is edited. Do you have a source?

                1. uqers · · focus · HN ↗
                  See

                  <a href="https:&#x2F;&#x2F;arstechnica.com&#x2F;tech-policy&#x2F;2026&#x2F;02&#x2F;wikipedia-might-blacklist-archive-today-after-site-maintainer-ddosed-a-blog&#x2F;" rel="nofollow">https:&#x2F;&#x2F;arstechnica.com&#x2F;tech-policy&#x2F;2026&#x2F;02&#x2F;wikipedia-might-...

                  <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Archive.today_guidance#Why_are_we_doing_this" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Archive.today_guidan...?

                  Besides tampering with content, the site was also using visitors to DDOS a blog that mentioned the owner of archive.today.

                  1. sam_lowry_ · · focus · HN ↗

                    [dead]

                2. Mogzol · · focus · HN ↗
                  See the &quot;Background&quot; section of the Wikipedia RFC on banning archive.today links: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Requests_for_comment&#x2F;Archive.is_RFC_5" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Requests_for_comment...

                  They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

                  1. sam345 · · focus · HN ↗
                    Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it&#x27;s fine for gossip and speculation but it shouldn&#x27;t be in Ars Technica then.
                    1. fc417fc802 · · focus · HN ↗
                      Because a lot of us watched the drama unfold in real time.
                      1. avadodin · · focus · HN ↗
                        We have always been at war with Eastasia.
                    2. Mogzol · · focus · HN ↗
                      There is a megalodon.jp archive of an archive.today page that shows the edits: <a href="https:&#x2F;&#x2F;megalodon.jp&#x2F;2026-0219-1654-07&#x2F;https:&#x2F;&#x2F;archive.ph:443&#x2F;LZymb" rel="nofollow">https:&#x2F;&#x2F;megalodon.jp&#x2F;2026-0219-1654-07&#x2F;https:&#x2F;&#x2F;archive.ph:44...

                      There&#x27;s also a bunch of previous hackernews discussions about it:

                      - <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47474255">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47474255

                      - <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46624740">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46624740

                      - <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47092006">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47092006

                      - <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46843805">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46843805

                      And you obviously have no reason to believe me, but I was following this when it was happening at the start of this year and can confirm that the DDoS script and archive text replacements really did happen.

              3. fn-mote · · focus · HN ↗
                &gt; Why folks continue to use them confuses me.

                Seems like the “confused” is disingenuous when not paying for content you read is a clear motivation.

                1. opello · · focus · HN ↗
                  Convenience as a higher order motivator than disgust at the bad behavior of archive.{today,ph,...} mentioned elsewhere, I think is the point of the comment to which you replied.
                  1. IAmBroom · · focus · HN ↗
                    Convenience is by definition one step.

                    Disgust requires research, analysis, and decision. The research alone is beyond most users&#x27; general practice.

                    Why would anyone walk in the front door when they could hop a fence, pry open a window, and crawl in?

                    1. thereforegrin · · focus · HN ↗
                      disgust is the simplest of emotions and so requires none of what you claim.

                      You&#x27;re mistaking it with substantiated criticism.

              4. Meneth · · focus · HN ↗
                &gt; Why folks continue to use them

                Because there&#x27;s no working alternative.

                1. gpvos · · focus · HN ↗
                  unwall.app works for at least some sites.
            3. Anonyneko · · focus · HN ↗
              Which sucks because these just don&#x27;t work for me for some reason (Finland, no luck with VPNs either).
        2. zymhan · · focus · HN ↗
          Only some of them, it is not universal.
          1. ghostly_s · · focus · HN ↗
            &quot;often&quot;
        3. koolala · · focus · HN ↗
          One site was doing that which archive in their name but wasn&#x27;t apart of archive.org
          1. sam_lowry_ · · focus · HN ↗
            archive.is or archive.today?

            Why being shy in the era of stealing AI?

            1. LoganDark · · focus · HN ↗
              archive.today uses clients to perform DDoS, I would not recommend using their site.
              1. schnebbau · · focus · HN ↗

                [dead]

                1. LoganDark · · focus · HN ↗
                  &lt;<a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Archive.today_guidance#Why_are_we_doing_this?" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Wikipedia:Archive.today_guidan...&gt; for those without a search engine.
                  1. DonHopkins · · focus · HN ↗

                    [dead]

                2. x______________ · · focus · HN ↗
                  Sure it is! This has been going on for years and global attention was gained at the beginning of this one.[0]

                  Wikipedia deprecates Archive.today, starts removing archive links (arstechnica.com) 616 points by nobody9999 6 months ago | hide | past | favorite | 368 comments

                  0 <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47092006">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47092006

                3. jimmydorry · · focus · HN ↗
                  It&#x27;s been posted multiple times of the last few months as they were removed from wikipedia and cloudflare. [1] [2] [3]

                  1. <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46843805">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46843805

                  2. <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47092006">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47092006

                  3. <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47474255">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47474255

                4. gpvos · · focus · HN ↗
                  This is well-known. If you are very sure it&#x27;s not happening anymore, please provide a reputable source.
                5. DonHopkins · · focus · HN ↗
                  You can&#x27;t make other people forget things by declaring yourself ignorant of the facts.
                  1. throw10920 · · focus · HN ↗
                    I literally had never heard of this before. I don&#x27;t check HN every single day.

                    It&#x27;s extremely reasonable to ask for a link, very easy to include one when making a claim, and attacking someone for asking for evidence is extremely anti-intellectual independent of the level of effort required.

                    (n.b. that doesn&#x27;t excuse the hostile way that they asked for proof - &quot;Let&#x27;s all just believe this baseless assertion shall we&quot;)

                    1. DonHopkins · · focus · HN ↗

                      [dead]

                      1. LoganDark · · focus · HN ↗
                        They were pointing out the lack of evidence on my part. I agree it was rude but I don&#x27;t think it&#x27;s helpful to start calling it misinformation with no evidence. They had a valid point that not everybody Just Knows already, hence why I did reply with a link. I don&#x27;t think it&#x27;s constructive to jab much more than I did in that reply.
                        1. throw10920 · · focus · HN ↗
                          And thank you for sending that link, because it was extremely useful and detail-packed.
            2. mitxela · · focus · HN ↗
              In some parts of the internet you can&#x27;t mention a pirate site (or left wing stuff, anything sexual, or Palestine) without being banned. HN isn&#x27;t one of them, but people have learned to be overly cautious.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.