‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. simonw · · focus · HN ↗
    > Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

    I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

    In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

    1. bradly · · focus · HN ↗
      Just yesterday from my one of my sessions with Sol:

      > Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

      1. TeMPOraL · · focus · HN ↗
        As it should.

        Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

        1. bradly · · focus · HN ↗
          Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
          1. Analemma_ · · focus · HN ↗
            I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.
            1. compiler-guy · · focus · HN ↗
              I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.
              1. cruffle_duffle · · focus · HN ↗
                Then make agent friendly content. Take the text and make a markdown version.
                1. fineIllregister · · focus · HN ↗
                  People doing this say it makes things worse because then the bots download both.
                  1. compiler-guy · · focus · HN ↗
                    Not to mention that it solves none of the rate issues. If the scrapers are hitting your site 10,000 times a day, adding markdown isn’t going to change that at all.
                  2. TeMPOraL · · focus · HN ↗
                    Because not enough people do this earnestly, and many more do it maliciously (bot endpoints that lie, or provide significantly less information than people endpoints) or put it behind a business contract (yes, APIs), so the bots or agents can't trust it in general.

                    Also let's not forget that innocent sites suffering from floods of scrapers are actually the minority here - this is just a special case; the main reason for the tension is simply that most websites and businesses on-line rely on users wasting their time, and cannot abide any form of end-user automation. Their business plans hinge on their ability to force themselves on you.

                    1. account42 · · focus · HN ↗
                      It's defensively, not maliciously. Malice would imply that the the site owner is morally obliged to serve the bots.
                      1. TeMPOraL · · focus · HN ↗
                        Morally, site owner should not be trying to discriminate between "people" and "bot" traffic in the first place.
              2. daveoc64 · · focus · HN ↗
                Is scale what we're discussing though?

                e.g. a prompt of "fetch <article URL> and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

                1. compiler-guy · · focus · HN ↗
                  It’s easy to write instructions that have the agent check once every fifteen minutes, or even once an hour, in perpetuity, which never sleeps. And people do write such instructions. A human can’t do that by hand for very long.

                  The problem is that it is hard to distinguish your one off (which seems perfectly fine) from the tidal wave of bad actors.

                2. kelnos · · focus · HN ↗
                  Sure, but all the time I'll ask Claude a question, and then I'll see it fetch 5-10 different URLs to come up with answer. I certainly would not be fetching those URLs at that rate if I were doing it myself. I would probably be visiting those pages, one by one, over the span of 10-20 minutes.

                  That's the scale argument.

                  1. TeMPOraL · · focus · HN ↗
                    As would I when researching anything myself. I'll do a web search, and if I see some highly relevant results, I'll middle-click them so they open in a new tab, and I'll easily do 5+ at a time, before then going to read the first one.

                    Same with browsing HN, btw. I have a row of 9 HN tabs open, all of them opened at the same time, as I scrolled the front page and middle-clicked on thread link to anything interesting.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.