By Tech & Web Infrastructure Desk
Published: September 2024
Web infrastructure and security giant Cloudflare has rolled out a significant update to how website owners manage artificial intelligence crawlers. Following months of industry-wide anxiety over how to protect intellectual property without sacrificing search engine optimization (SEO), Cloudflare has introduced a new "Disallow AI Training" setting. This feature shifts how legacy blocking tools operate, aiming to protect content creators from having their data scraped for large language model (LLM) training while preserving vital search engine traffic.
The update marks a major pivot from previous iterations of Cloudflare’s bot management policy. Originally, the company had signaled a harsher crackdown that would have penalized sites blocking AI training by simultaneously cutting off standard search indexing from major tech players like Google, Apple, and Microsoft. Following extensive negotiations with crawler operators, Cloudflare has calibrated its approach, introducing a nuanced framework built around "Accountable" crawlers.
Main Facts: What the New "Disallow AI Training" Setting Does
The core of Cloudflare’s new update revolves around a refined separation of powers between search crawlers and AI training scrapers.
Previously, website administrators faced a blunt instrument: blocking AI training bots often meant blocking the primary indexers for Google, Apple, and Bing, because these tech giants utilize dual-purpose, "mixed-use" crawlers that serve both search ranking and AI model training functions. Refusing one meant losing both, presenting webmasters with a terrible dilemma: forfeit organic search traffic or allow their proprietary content to train commercial AI models for free.
The new "Disallow AI Training" setting, deployed as part of Cloudflare’s tripartite Training control suite (which also features Search and Agent controls), solves this dilemma for compliant operators.
- Selective Restriction: When enabled, the setting injects a no-training preference into a site’s
robots.txtfile. - Search Preservation: It explicitly allows trusted mixed-use crawlers (such as Googlebot, Applebot, and Bingbot) to continue indexing the site for traditional web search purposes.
- Aggressive Blocking Alternative: For site owners who still want a total blackout of these major tech crawlers, Cloudflare has clarified that selecting the stark "Block" option will entirely halt Googlebot, Applebot, and Bingbot, stripping the site of search visibility entirely.
Cloudflare notes that the vast majority of its customers do not need to take immediate action. Existing site configurations utilizing "Block" or "Block on pages with ads" for training are being automatically migrated to the new framework, while older toggles like "Block AI Bots" and Cloudflare’s Managed Robots.txt feature are being phased out.
Chronology: From Summer Threats to the September 15 Rollout
The path to Cloudflare’s current policy has been iterative, characterized by pushback from the SEO community and tense negotiations with major tech firms.
- July: Cloudflare first previewed a controversial strategy. The company announced that starting September 15, any website utilizing its tools to block AI training crawlers would automatically find their access to Google, Apple, and Bing search crawlers severed as well. The rationale was that these companies relied on unified infrastructure for both scraping and searching.
- August: Recognizing the immense pushback from publishers and webmasters who rely heavily on search engine traffic, Cloudflare shifted gears. The company published detailed briefings outlining a new concept: "Accountable" mixed-use crawlers. This laid the foundation for separate controls covering search, training, and AI agents.
- September 15: Cloudflare officially deployed the "Disallow AI Training" feature. Legacy block selections were migrated over to the new system, and mixed-use crawlers were granted a reprieve to index search results—provided their operators met strict accountability criteria.
Supporting Data & The Criteria for "Accountable" Crawlers
Cloudflare’s decision to spare Google, Apple, and Bing from blanket bans hinges on a newly minted "Accountable" designation. This label was born out of bilateral talks initiated in July between Cloudflare and major crawler operators. To earn and keep this classification, tech operators must meet—or commit to meeting—four foundational requirements:
- Opt-Out Mechanics: Operators must provide a functional way to opt out of AI training via standard mechanisms like
robots.txtor equivalent protocols. - AI Summaries Control: Operators must offer clear controls to opt out of AI-generated summaries (features currently rolling out with operators and expected through Cloudflare natively next year).
- Granular Transparency: Operators must provide URL-level visibility, detailing precisely which pages were accessed for training alongside performance metrics on how content appeared in search results.
- Search Independence: Operators must offer ironclad assurances that opting out of AI training will have zero negative impact on traditional search engine rankings and results visibility.
According to Cloudflare, tech heavyweights Apple, Google, and Microsoft currently meet these standards, having rolled out baseline features alongside firm commitments and deadlines to fulfill the remaining requirements. Meanwhile, dedicated training crawlers operated by companies like Amazon, Anthropic, Meta, and OpenAI are classified differently; because these firms run distinct, separate scrapers specifically for training (rather than mixed-use search bots), those dedicated scrapers remain blocked under the new setting.
Official Responses: How Major Search Engines Handle the Setting
The implementation of "Disallow AI Training" looks different depending on the specific search engine ecosystem:
At Google, Cloudflare’s setting interfaces directly with Google-Extended, the specific robots.txt token designated by Google to let webmasters opt their content out of Gemini and other generative AI training pipelines. Google’s official developer documentation explicitly notes that utilizing Google-Extended does not penalize a site’s inclusion or ranking in standard Google Search.
Additionally, Google manages generative AI visibility—such as appearance in AI Overviews, AI Mode, and generative Discover features—through a separate toggle inside Google Search Console. Google maintains that this Search Console preference operates independently of AI training rules.
Apple
For Apple, the setting leverages Applebot-Extended. Apple’s documentation states that Applebot-Extended does not crawl live web pages for standard indexing and is entirely decoupled from search ranking algorithms. To keep web content out of the broad knowledge summaries generated by Siri and Apple Search, webmasters must still rely on traditional nosnippet meta tags.
Microsoft (Bing)
Microsoft represents an outlier in the initial rollout. Choosing "Disallow AI Training" on Cloudflare does not currently transmit a no-training preference via robots.txt to Bing, because Microsoft has not yet built support for the protocol. Cloudflare has confirmed that Microsoft’s support for this specific robots.txt integration is slated to arrive much later, targeting early 2027.
In the interim, Bing users must continue relying on legacy methods, specifically the NOARCHIVE meta tag. According to Bing’s developer documentation, any content tagged with NOARCHIVE is excluded from training Microsoft’s generative AI models and is intentionally omitted from links generated via Copilot and Bing Chat.
Implications for Webmasters, Publishers, and the Future of AI
Cloudflare’s updated stance has profound implications for the digital publishing ecosystem.
1. Granular Control Without SEO Suicide
For years, site owners felt trapped. Allowing AI scrapers meant watching proprietary journalism, artwork, and data get ingested by tech giants to build commercial products without compensation or attribution. Conversely, blocking those scrapers risked catastrophic drops in organic search traffic. By leveraging the "Accountable" crawler framework, Cloudflare has successfully decoupled search visibility from AI scraping, giving publishers the technical means to say "no" to model training without vanishing from Google or Apple search results.
2. A Fragmented Technical Landscape
Despite Cloudflare’s streamlining efforts, managing visibility remains complex. Webmasters must still navigate a fragmented landscape where different tech giants rely on different protocols. Google uses Google-Extended, Apple uses Applebot-Extended, and Microsoft/Bing currently relies on the NOARCHIVE meta tag while lagging behind on robots.txt standardization.
3. Looking Ahead: The Battle Over AI Summaries
With the infrastructure for training opt-outs largely established, Cloudflare’s roadmap points toward the next major battleground for content creators: AI summaries.
As search engines increasingly replace traditional blue links with synthesized, zero-click AI answers, publishers are losing traffic even when their content is cited. Cloudflare has announced that its primary focus moving forward is tackling AI summaries. By early next year, the company aims to roll out a unified control panel allowing site owners to govern how much of their content is slurped into AI-generated summaries across all major platforms with a single toggle, bypassing the need to negotiate or configure disparate rules operator by operator.
For now, website administrators using Cloudflare can breathe a collective sigh of relief. The September 15 migration ensures their search visibility remains intact while providing a definitive, legally and technically sound mechanism to slam the door on unauthorized AI training.

