FR
live

Cloudflare splits AI bots into Search, Agent, and Training and tightens the defaults on September 15

On July 1, 2026, Cloudflare replaced its one-click ’Block AI Bots’ switch with three distinct classes — Search, Agent, and Training. On September 15, 2026, Training and Agent bots will be blocked by default on pages that show ads: check your settings before that date.

A sorting conveyor that routes grey parcels into three separate chutes, a single amber tag placed on one of the chutes.

July 1, 2026. Cloudflare held its second “Content Independence Day” and replaced the blunt “Block AI Bots” switch with a three-way classification — Search, Agent, and Training. September 15, 2026. New domains onboarding to Cloudflare will receive fresh defaults: Training and Agent blocked on pages that display ads, Search left open. 2025. The first edition shipped a blunt block plus a Pay-Per-Crawl marketplace; the second hands publishers a real negotiating layer. Why it matters: in ten days, multi-purpose crawlers like Googlebot risk being blocked by default, and many site owners have not seen it coming.

The thirty-year bargain that AI broke

For three decades the web ran on an implicit deal: a crawler indexes your pages, and in return it sends referral traffic back to you. AI answer engines broke that bargain. They read your content, answer the user directly on the results page, and send nothing back to the publisher — a dilemma Cloudflare describes as a “Faustian bargain”: show up in search and get trained on for free, or block everything and vanish from discovery.

The blunt switch no longer matched reality. Google is simultaneously a search engine, an answer engine, and a model trainer, often through overlapping automation. Blocking “all AI bots” therefore shoots you in the foot while you defend yourself. The three-way classification is Cloudflare’s answer: stop asking “is this bot an AI?” and start asking “what is it doing on my site, what is it storing, and how will it reuse my content?”.

Three classes, three real use cases

The taxonomy, detailed by Jin-Hee Lee and Bryan Becker, splits automation into three independently manageable behaviors, available even on the Free plan:

  • Search — the crawler indexes your content so it can answer questions about it later. It is building a database of your site, with the expectation of referral traffic in return;
  • Agent — automation acting in real time on a person’s behalf: chat fetch bots like ChatGPT-User, or browser-use agents such as Gemini or Claude driving Chrome;
  • Training — a crawler taking your content to train or fine-tune a model, permanently absorbing your data into the model’s architecture.

The structural point is that a single bot can carry more than one label. Cloudflare now tracks all of a bot’s purposes instead of forcing one tag, and is pressuring bot operators to split their automation into separate, clearly named crawlers. Alongside the three AI uses, the taxonomy also covers Transact, Data Collection, Security Testing, SEO, Ads Verification, social previews, and monitoring.

The September 15 default, and the Googlebot trap

The heaviest change is the new default policy taking effect on September 15, 2026. For any new domain onboarding to Cloudflare, the Training and Agent categories will be blocked by default on pages that display ads, while Search stays allowed. The logic: an ad signals that a human was meant to land on that page, so the bots that prevent that human attention — Training and Agent — get locked out.

The subtlety that will hurt: defaults are enforced by the most restrictive applicable rule. A multi-purpose crawler like Googlebot, Applebot, or BingBot, which combines Search and Training, inherits the Training block wherever a customer blocks that category — through the new options or the legacy Block AI Bots service. Existing customers who want to keep things exactly as they are can opt out in their Security settings before September 15, and Cloudflare says it will keep notifying customers as the date approaches.

For a publisher that monetizes through ads, the net result is a silent flip: Googlebot can end up blocked on ad pages without any human touching robots.txt. It is exactly the kind of change you discover in Search Console a week later, as a drop in indexing.

A use= signal in robots.txt to negotiate reuse

Fine-grained classification only helps if the publisher can also express how much of their content a bot may keep and reshare. Cloudflare therefore adds a “content use” setting to Bot Management, with three levels, from least to most permissive:

text
# robots.txt — content-reuse signal (Cloudflare Content Signals)
User-agent: *
Allow: /
use=reference

The three values of the use= signal:

  • use=immediate — interact, but store and reuse nothing;
  • use=reference (default) — index, excerpt, and link back;
  • use=full — summarize and reproduce.

The use= signal extends the existing Content Signals and lets owners write rules grouped by behavior — “allow Search, SEO, and Ads Verification bots, but only up to the reference level” — instead of negotiating bot by bot. For anyone shipping an AI agent or answer engine, that signal becomes a first-class part of the crawl negotiation.

BotBase, and what it changes for both sides

Cloudflare is also launching BotBase, a visibility database reserved for Enterprise Bot Management: a searchable catalogue of all verified bots and agents, with their classification, and direct controls planned for later this year. It is the missing piece for auditing what actually crawls a fleet, beyond the User-Agent string alone.

The change draws the same line for both camps. For publishers, combined with Pay-Per-Crawl and the content-use tiers, the system turns content access into a negotiation layer — trade access for compensation rather than choosing between free exploitation and total invisibility. For bot operators, the consequence is a push toward transparency: split crawlers by purpose and respect use=, or see your automation blocked across a growing slice of the web.

The ten-day checklist

Before September 15, three checks are enough to cross the line without breakage. First, open your Cloudflare dashboard, SecurityBots, and read the current state of your three classes. If you are an existing customer, the default is not applied to you automatically: opting out is voluntary, but it only works if you know the option exists. Second, cross-check against Search Console: if Googlebot starts reporting crawl errors on your ad pages after the 15th, the Training block is the first suspect. Third, document the rule for your team — an SEO or content person discovering the change three weeks late costs more than ten minutes of checking today.

The subtlest trap remains the multi-purpose crawler. Googlebot is not a pure training bot: it indexes and feeds training corpora. Because the default applies the most restrictive rule, blocking Training means blocking Googlebot on ad pages. That is a defensible choice — you refuse to let your content train a model with no compensation — but it should be a choice, not an accidental consequence of a box left ticked.

Verdict

If you run an ad-monetized site, log into your Cloudflare dashboard before September 15 and explicitly set your Search, Agent, and Training policies — do not let a default decide for you, especially if you depend on Google indexing for traffic.

If you operate a crawler, an AI agent, or an answer engine, split your automation into separate, named crawlers now, and implement the use= signal in your robots.txt reading. That is the price of staying indexed past September 15.

If you are on the Free plan, the three-way classification is already available to you — only BotBase and the granular content-use controls are gated behind Enterprise.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

CloudFront flat-rate plans become manageable through the API and IaC

On September 3, 2026, AWS opened programmatic management of CloudFront flat-rate plans through the new PricingPlanManager API, the CLI, CloudFormation and the CDK. Teams can now codify subscribing, changing tiers and cancelling a no-overage monthly price, with a two-phase approval that prevents unintended billing.

AWS opens its first Saudi Arabia region and commits 50 MW of AI with HUMAIN

Announced at LEAP in Riyadh, AWS’s first infrastructure region in Saudi Arabia will go live in December 2026, bringing the global network to 40 regions on a planned investment of over $5.3 billion. For teams serving the Gulf, it is the long-awaited answer to data-residency requirements, doubled with a 50 MW AI Zone planned for 2028.

← Back to the feed

Type at least two characters.

navigate open esc dismiss