Cloudflare’s new AI bot defaults land September 15. Here’s what to check first to protect your store’s organic and AI search visibility.
On July 1, Cloudflare announced changes to how it classifies and manages AI bot traffic across its network, roughly one in five websites globally.
On September 15, new defaults take effect, and Cloudflare’s own press release confirms that existing free-tier customers who haven’t touched their settings by then will be automatically moved onto them too.
This isn’t limited to new sites signing up. If your site sits behind Cloudflare and nobody checks this before the deadline, your settings could change without you doing anything.
You might not configure this yourself. It’s something your dev team, hosting provider, or agency likely needs to check on your behalf.
Here’s what’s changing, and the one action worth asking for directly: don’t let a setting meant to protect your content from AI training end up blocking Google too.
Search, Agent, and Training will no longer be one thing
Cloudflare used to offer a single “Block AI Bots” switch. Simple, but it lumped together three very different behaviours: a bot indexing your site for search results, a bot fetching a page in real time to answer a shopper’s question, and a bot harvesting your product pages to train a model.
Cloudflare now splits these into three categories that every customer can manage separately, Search, Agent, and Training.
Search is the behaviour most likely to send you traffic back. Training takes your content with nothing in return. Agent sits in between, a bot like ChatGPT or Gemini acting on a real shopper’s behalf.

The default can catch Google in the crossfire
From September 15, new sites onboarding to Cloudflare will have Training and Agent bots blocked by default on any page carrying ads. Search stays allowed. The logic is that an ad page is built for a human to see it, so bots that don’t send traffic back get shut out there. Existing free-tier customers who haven’t reviewed their settings by that date get moved onto the same defaults automatically, according to Cloudflare’s press release, so “we’re not a new customer” isn’t a reason to skip the check.
Separate from that default, there’s a broader rule taking effect the same day that applies to any Cloudflare customer, new or existing: Googlebot crawls for both Search and Training.
If you block Training traffic, whether through the new options or the older “Block AI Bots” toggle, you block Googlebot entirely, search ranking included, because Cloudflare enforces the most restrictive applicable rule when a bot serves multiple purposes. The same applies to Bingbot and Applebot.
This is the specific thing to get right. We obviously wouldn’t recommend blocking Training if the side effect is losing Google. For most retailers, disappearing out of Google search outweighs whatever protection you gain from keeping product content out of a model’s training data.
The action item for your dev team or agency isn’t “review the settings”, it’s a direct question with a specific answer: confirm Search is allowed and unaffected, and if Training is being blocked on your site, either turn that off or use Cloudflare’s opt-out (available in Security settings before September 15) so Search-only traffic from multi-purpose crawlers like Googlebot keeps flowing regardless of what’s set for Training.
A new signal for how content gets reused
Cloudflare is also adding a content-use signal to robots.txt, giving site owners a way to say not just whether a bot can crawl, but what it can do with what it finds.
Three levels: immediate (interact only, nothing stored), reference (index and link back, the default), and full (summarise and reproduce). Bots that reproduce content in full can’t hold Cloudflare’s “Verified” status at all, and any bot caught ignoring the content-use signal it agreed to loses that status too. Verified itself no longer means automatic access either way.
This won’t be something you set yourself either, but it’s a useful line item for a technical audit: is our robots.txt using this, and does it reflect how we actually want our content used by AI tools versus search engines?
Why this is a competitive question, not just a technical one
Search and AI-driven discovery are converging into the same channel. The infrastructure decisions made this year will determine whether your products surface when a shopper asks an AI assistant for a recommendation, or whether a competitor’s do instead.
You first need to make sure that no one accidentally trades away your Google visibility while trying to keep AI training bots out.
If you want a clear view of how your site actually performs in organic and AI search against the competition, you can request a free audit or arrange a quick call with our team.



