Without proper controls, website owners have long faced a difficult tradeoff: allow your content to be used for AI training, or risk losing discoverability in search. That tradeoff exists because some of the largest organizations on the Internet use mixed-use crawlers: a single crawler serving both search and AI training. Refuse one, and you refuse the other.
Today, Cloudflare is announcing a new Disallow AI Training setting that lets you easily stay indexed for search while refusing to let that same crawler train on your content. Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting. Mixed-use crawlers were the hard part of the training question.
AI Summaries are next. A site-wide yes or no is too blunt: how much of your content appears in a summary matters as much as whether it appears at all. An opt-out for AI summaries is already one of the requirements we've set for mixed-use crawler operators.
By early next year, our goal is to let you control how much of your content is included — set once on Cloudflare, rather than with each operator separately. Why asking isn’t enough Most site owners want to be found: by humans, agents, and (good) bots. But a significant portion of the open Internet is funded by advertising, subscriptions, or direct relationships with visitors, and those models only pay when someone actually arrives.
Almost every site owner considers Search beneficial: less than 1% of Cloudflare sites choose to block Search bots. Training, however, is a different story: 17% of sites choose to enable some mechanism to block training. This is exactly why we decided site owners needed more granular controls, rather than a one-size-fits-all “Block AI.”
A robots. txt directive alone cannot solve this problem. Anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it.
A network can solve it, however: we publish the preference, identify who is crawling, classify why they are crawling, and block the ones that ignore it – then report what each operator actually does on Radar . But blocking removes a crawler. It doesn't change how crawlers behave.
The better outcome is operators that don't make you choose at all. So since July, we've been talking to them directly. The response has been encouraging: almost all agreed that site owners should have control and transparency into how their content is used, and reassurance that their choices will be respected.
To help site owners understand that, we created a designation: Accountable. The Accountable designation recognizes both capabilities available today and concrete commitments to deliver them. To qualify, a bot operator must meet or commit to meeting the following requirements: A mechanism for site owners to opt out of AI training, through robots.
txt or a similar standard. A mechanism for site owners to opt out of AI summaries set with the operator directly, and next year through Cloudflare (see section below for more detail). URL-level visibility into which pages were made available for training, along with metrics showing how content appeared in search.
Assurance that opting out of AI training will not affect traditional search results. Apple, Google, and Microsoft all demonstrate that they meet the qualifications to be Accountable. Each combines capabilities available today with time-bound commitments for those still in development.
The details of each of these companies’ crawlers are shared below. New security setting options Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior. Three behaviors are available as controls: Search - crawling to build a search index.
Training - crawling to train or fine-tune a model. Agent - user-directed agents visiting a page on behalf of a human, such as chat fetch bots and browser-use agents. A mixed-use crawler is a single crawler doing both Search and Training.
Without controls, that combination creates the tradeoff described above: site owners cannot refuse one use without refusing the other. To avoid blocking Accountable mixed-use crawlers — the ones that don't force that tradeoff on website owners — we are introducing a new setting: Disallow AI Training. Disallow AI Training is named for the Disallow: directive it publishes in your robots.
txt. “Block” setting now means something different Block and “Block on pages with ads” previously did not apply to mixed-use crawlers because blocking them could also affect search discoverability. Now that we have the new Disallow AI Training setting, Block and “Block on pages with ads” apply to all training crawlers, including mixed-use crawlers.
Training, Search, and Agent controls are applied at the domain level. With the addition of Disallow AI Training, the available settings are: Allow : All crawlers are allowed, unless blocked by another setting or a WAF rule. Disallow AI Training : Bot Preference Sync publishes the applicable no-training preference in robots.
txt. Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search.
Disallow AI Training is only available as a setting for Training, not Search or Agent. Block on pages with ads : Crawlers, including mixed-use crawlers, are blocked only on pages detected to be serving an ad. Block : All crawlers, including mixed-use crawlers, are blocked.
Disallow AI Training works by publishing a preference in robots. txt. An ads-only preference cannot be expressed that way: Cloudflare can detect which pages serve ads, but that list is too large and changes too frequently to enumerate in robots.
txt. That's why there's no Disallow AI Training on pages with ads.
Originally published at blog.cloudflare.com


