Cloudflare lets sites block AI training without losing Google search traffic

Stay discoverable in search while disallowing AI training

Cloudflare lets sites block AI training without losing Google search traffic

Cloudflare's new Disallow AI Training setting publishes a no-training preference in robots.txt while keeping mixed-use crawlers like Googlebot and Bingbot indexing your site for search. Apple, Google, and Microsoft have committed to honor it. The move gives site owners granular control over search, training, and agent crawling, and sets the stage for controlling AI summaries next year.

Without proper controls, website owners have long faced a difficult tradeoff: allow your content to be used for AI training, or risk losing discoverability in search.
  1. DharmaPolice

    I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway.

    Ultimately this reminds me of those really early social media profiles (before people understood privacy settings if they even existed) which would say "If you're not my friend you're not allowed to read this page".

    If you don't want your content to end up in some database/archive don't publish it for the whole world to see.

  2. 1vuio0pswjnm7

    "Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."

    Is that really true

    CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

    CF fingerprints www users

    As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused

    IA = Internet Archive

  3. nirmeetimthebes

    "Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.

  4. skybrian

    I didn't know websites could opt out of providing data to Google's AI training. Looks Google added support for this via 'Google-Extended' in robots.txt back in 2023:

    https://blog.google/innovation-and-ai/products/an-update-on-...

  5. shark1

    "Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting."

    Specified Time Frame ;)

More from this day

2026-09-16