‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. ilamont · · focus · HN ↗
    Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website?

    My blogs are getting slammed and there are issues with cloudflare or captchas.

    1. iamacyborg · · focus · HN ↗
      > Shouldn't the solution be to gate bulk access for automated services for a price?

      Fine in theory but determined scrapers will use residential proxies in bulk.

      1. LastTrain · · focus · HN ↗
        These should be illegal unless users sign off on every fucking byte.
      2. mitxela · · focus · HN ↗
        Make a user download and hash 100GB of junk data before being allowed in. Residential proxies cost a lot per GB.
    2. hamboomger · · focus · HN ↗
      This! But only wayback machine, other services I'm not sure.

      But maybe the problem is that they can't serve the data from the other websites like this, if they use it commercially. Right now they have non-commercial use, from what I understand.

    3. mrweasel · · focus · HN ↗
      > Shouldn't the solution be to gate bulk access for automated services for a price?

      The problem is that many of the people who are scraping this data doesn't want to pay. These are organisations who would rather not clone your git repo, and instead scrape every single page on your Forgejo installation. These are NOT nice people.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.