‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. bradly · · focus · HN ↗
      Just yesterday from my one of my sessions with Sol:

      > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

      1. TeMPOraL · · focus · HN ↗
        As it should.

        Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

        1. bradly · · focus · HN ↗
          Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
          1. aaron_m04 · · focus · HN ↗
            robots.txt?
            1. bradly · · focus · HN ↗
              Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.
              1. dhx · · focus · HN ↗
                robots.txt was only intended to help search index crawlers not get stuck in endless crawl loops for badly designed websites.

                What you suggest is explicitly not a purpose of robots.txt per RFC9309[1]:

                "These rules are not a form of access authorization."

                HTTP 429 and HTTP 403 are what servers are meant to return to clients to slow them down or tell them to stop doing something without having first gained authorisation.

                [1] <a href="https:&#x2F;&#x2F;datatracker.ietf.org&#x2F;doc&#x2F;html&#x2F;rfc9309#section-1" rel="nofollow">https:&#x2F;&#x2F;datatracker.ietf.org&#x2F;doc&#x2F;html&#x2F;rfc9309#section-1

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.