Cloudflare launches 'Disallow AI Training' so sites keep search indexing while refusing training crawls
- Cloudflare announced a Disallow AI Training setting that publishes a no-training preference in robots.txt while leaving the same crawler free to index the site for search, so site owners no longer have to choose between being found and refusing training.
- Fewer than 1% of Cloudflare sites block Search bots, while 17% already enable some mechanism to block AI training, which Cloudflare cites as the reason a single "Block AI" toggle was too blunt.
- Apple, Google, and Microsoft meet the new Accountable designation, which requires an opt-out for AI training, an opt-out for AI summaries, URL-level visibility into what was made available for training, and a guarantee that opting out does not affect search results.
- Because robots.txt cannot identify, classify, or stop a crawler that ignores it, Cloudflare enforces the preference at the network level and reports what each operator actually does on Radar; Amazon, Anthropic, Meta, and OpenAI separate their search and training crawlers, so their training-only crawlers stay blocked.
- The old Block and "Block on pages with ads" settings now apply to all training crawlers including mixed-use ones, and Cloudflare aims by early next year to let sites control how much of their content appears in AI summaries from a single Cloudflare setting.
Hacker News opinions
Cloudflare enabling the problem and the solution, in the same company. How long has it been now?
A bit too late honestly. So many people now just read the AI summary as the search result, and the sites that allow AI win by attrition. There's no going back from this.
Tearing down ads as an income source just demonizes independent publishers trying to make money, while the huge corporations get let off the hook because that's just what they do.
So what does this setting actually do? Does it block their IP ranges too? The idea that Meta will respect an Accountable label, by itself or through partners, is eyebrow raising at best.
I don't trust either Meta or OpenAI to use their separate search and training crawlers only for those purposes. Their pinky promises have no value when both are built on deceptive behavior.
Calling it Accountable with a capital A sounds deliberate, but all I picture is the renewal email asking if you want to renew your Accountable license by pinky promising again that you use your IPs the way you said.
Accountable is just a fancy word for a pinky promise with a label. Nothing stops the data from ending up in a training run once it has already been fetched.
The only way to actually stop that is DRM'ing everything, and that is a level of dystopia I don't want, Stallman never even imagined it.
Remember, you could already opt out of Google's AI training. Google-Extended in robots.txt has been supported since 2023.
What weirds me out is the analytics. There's no way my empty index.html is getting 10,000 hits a day.
They bullshit on those numbers. On a site with proven visitors, Cloudflare told me they saved me 90,000 visitors out of over a million. Big numbers make people feel good.
The part that bugs me is Cloudflare classifying anyone not on a popular browser with JavaScript as a bot. They discriminate on user agent instead of actual behavior.
After months of fighting DDoS from Anthropic and OpenAI across 50+ sites, Cloudflare still lets what it calls good bots through all your blocking rules, and there's no way to turn that off unless you pay.
I'm going to test this against my own cloud IP ranges database, but I have my doubts. A formal capitalized title doesn't stop anyone from lying about their traffic.
Seems useful for my site. I don't want competitors' AI systems training on our research data.
All these schemes to classify data as public but not really are doomed. Even if every AI company honored the terms, someone else can index it and sell it along. If you don't want it in a database, don't publish it for the whole world.