Cloudflare Separates Search, Agent and Training

Cloudflare Separates Search, Agent and Training
Sponsored

Cloudflare is giving website owners separate controls for Search, Agent and Training traffic, with a newer “Disallow AI Training” option intended to preserve search access while expressing a no-training preference. The development makes crawling policy more specific, but its effectiveness depends on the crawler, the selected setting and the operator’s implementation.

The July 1 changelog describes three behaviours: Search builds an index, Agent acts on a person’s behalf, and Training collects material to train or fine-tune models. A bot can serve more than one purpose, so an operator’s name alone does not describe every use of the content it retrieves.

Disallow is different from blocking access

Cloudflare’s September 15 announcement says Disallow AI Training publishes applicable preferences through Bot Preference Sync. Accountable mixed-use crawlers remain allowed for search; other training crawlers are blocked. Selecting Block instead can prevent mixed-use crawlers from accessing a site for both search and training.

Cloudflare says Apple, Google and Microsoft meet or have made time-bound commitments to its accountability requirements. There is a current limitation for Bing: support for the domain-level robots.txt no-training preference is targeted for early 2027. Until then, the setting does not automatically convey that preference to Bing through robots.txt. The announcement identifies separate current Bing controls.

For publishers, this is a reason to review the effect of a setting before adopting it. Allowing search access does not guarantee indexing, and a no-training preference should not be described as universal prevention of content reuse.

Authentication answers another question

Web Bot Auth uses cryptographic signatures in HTTP messages to authenticate bot requests. Cloudflare documents it as a verification method for bots and agents, relying on public-key directories and signed requests.

Identity verification and content-use permission are separate concerns. An authenticated request can help a site establish who sent it; that alone does not prove compliance with every condition governing subsequent use. A useful governance model needs both an access decision and a clear understanding of what the permitted use means.

Robots.txt remains part of the system

The shift is better understood as adding layers to crawling governance. Robots.txt still carries preferences in the new mechanism. Classification, access controls and authentication address other parts of the problem. Describing the change as the end of robots.txt would obscure how the feature actually works.

In our analysis, SEO and security teams should define the outcome they want before choosing controls. A publisher may want discovery in search while restricting training; another business may value an agent retrieving current product information for a customer. Those intentions can require different treatment even when both requests are automated.

Check the policy against observed results

A practical review should connect each crawler decision to a business purpose, identify any mixed-use services affected and record rollout limitations. After a change, teams can examine access outcomes and search visibility alongside their stated policy. A toggle’s label is not enough to establish that the intended result occurred.

Cloudflare’s controls offer a more detailed vocabulary for those decisions. The opportunity for publishers is to make access deliberate and reviewable, while preserving the distinction between a declared preference, a technical block and an authenticated identity.

0%