Cloudflare starts blocking AI crawlers that hide behind search bots by default

Cloudflare's new default crawler policy took effect September 15, automatically blocking AI training and agent bots on any page that carries advertising, unless the crawler clearly declares which of three purposes it's serving. The change closes a gap that let crawlers claim a search-indexing identity while simultaneously harvesting content for AI model training — a practice publishers have complained about since large language models began scraping the open web at scale.
Three buckets, one enforcement rule
Cloudflare now sorts AI bot traffic into three categories. Search covers crawling that indexes content to later answer questions about it, the kind of activity site owners have tolerated for decades in exchange for referral traffic. Agent covers automated behavior acting in real time on a person's behalf — tools like OpenAI's ChatGPT-User or Anthropic's Claude when it's driving a browser directly for a user. Training covers crawlers that ingest content specifically to train or fine-tune a model, where the data becomes permanently absorbed into the system rather than referenced on demand.
The enforcement logic is what makes this consequential: when a single crawler performs multiple functions — Google's own crawler, for instance, combines Search with Training — Cloudflare now applies whichever rule is most restrictive across the board. A bot that mixes search indexing with training data collection gets treated as a training crawler for blocking purposes, full stop, even on requests where it's nominally just indexing.
What site owners actually control
Publishers can now set content permissiveness independently of the bot-category defaults: immediate access lets a crawler interact without storing or reusing content, reference access allows indexing, excerpting, and linking back (the new default for ad-supported pages), and full access permits summarization and reproduction. The granular controls themselves went live back on July 1; September 15 is when the new default posture — training and agent bots blocked, search bots allowed — applies automatically to new customers, newly created sites, and all existing free-tier customers. Paid customers who already configured their own settings keep them unchanged.
Why this matters beyond one CDN's policy
Cloudflare sits in front of a large share of the web's ad-supported publisher content, which means this default shift functions as a de facto industry standard rather than one company's house rule — sites that never touch their Cloudflare settings will simply start blocking AI training crawlers today. For AI labs, it raises the practical cost of large-scale web training data collection: crawlers now have to honestly declare their purpose or risk being blocked entirely, rather than quietly riding along with search indexing traffic. It's also a bet that publishers will use the leverage this creates to negotiate payment — Cloudflare's broader Pay Per Crawl marketplace, which lets publishers charge for AI access to their content, was built for exactly this moment.
Originally reported by Cloudflare Blog. Read the original article for additional details.
View original source