‹ BackHN Continuity

Thread

An update on Wayback Machine access

685 points · 362 comments · ChrisArchitect

  1. Onavo · · focus · HN ↗
    Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

    It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

    I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

    1. croes · · focus · HN ↗
      It’s one thing to archive other companies content, it’s another to sell the access to it
      1. faefox · · focus · HN ↗
        Yeah, who does the Internet Archive think it is, (insert literally any AI company here)?
        1. bonoboTP · · focus · HN ↗
          Which AI company is selling access to reliable verbatim copies of websites? I don't mean "it may regurgitate a paragraph", but as a reliable service where you can repeatably get website content snapshots to a reliability level that makes such a use case viable?

          Using the information for training purposes is not the same thing. Not legally the same and otherwise.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.