‹ BackHN Continuity

Thread

Stay discoverable in search while disallowing AI training

87 points · 50 comments · djfergus

  1. 1vuio0pswjnm7 · · focus · HN ↗
    "Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."

    Is that really true

    CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

    CF fingerprints www users

    As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused

    IA = Internet Archive

    1. devmor · · focus · HN ↗
      As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.
      1. RobotToaster · · focus · HN ↗
        The "good bots" also just happen to be from companies that pay cloudflare a lot of money, I imagine.
        1. actionfromafar · · focus · HN ↗
          Goodness Tokens
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.