Cloudflare’s July 1, 2026 changelog set a date: September 15. From that day, domains newly joining Cloudflare get defaults that block Training and Agent bots on pages that display ads, while Search stays allowed. That part is mild. The part that already bites is older, because your Cloudflare AI crawler policy treats Googlebot as a training bot too. So if you ever ticked the box to block AI training, you may also be blocking the crawler your rankings depend on.
Cloudflare AI crawler traffic now moves in three lanes
Cloudflare used to give you one AI switch. Now it sorts bot behavior into three groups: Search, Agent and Training. Search means indexing a page so an engine can answer a question later, and it usually sends referral traffic back to you. Agent means a bot acting live for a real person, like ChatGPT-User. Training means harvesting text to build a model, with nothing sent back at all.
Those three lanes are not equal, so treating them as one setting is the mistake. In fact the split is the useful half of this change. You can now keep Search open, decide about Agent, and refuse Training, instead of choosing between all and nothing.
Each lane deserves a different answer. Search should stay open for almost every business, because that is the lane that sends people back. Training is the one you can close with the least regret. Agent is the genuinely hard call, and most teams have not thought about it yet.
The September 15 headline is wrong about who it hits
Plenty of posts say Cloudflare will block AI crawlers by default on September 15. That is not what the changelog says. The new defaults apply to domains onboarding to Cloudflare from that date, and existing customers keep the settings they already have. The default is narrower than it sounds, too, because it covers only pages that display ads. Search remains allowed either way.
So if your site already sits on Cloudflare, nothing changes for you on September 15 unless you change it. Still, every customer can opt out of the new defaults beforehand in the security dashboard. Read your own Cloudflare AI crawler settings rather than the headline about them.
Googlebot is a multi-purpose Cloudflare AI crawler
Here is the part worth ten minutes today. Cloudflare’s changelog says multi-purpose crawlers that combine Search and Training are affected by the new defaults to block Training. Googlebot, Applebot and BingBot all sit in that group, because Google, Apple and Microsoft use one crawl for ranking and for model work.
Practitioners have hit this already. Search Engine Journal reported on August 4, 2026 that a site owner who turned on Cloudflare’s AI training block saw Googlebot and BingBot get 403 errors on sitemap requests. Google’s John Mueller asked the reporter for details so he could investigate. One privacy setting, and your sitemap stops being readable by the engine you actually want.
The trap is structural rather than a bug. Cloudflare classifies these bots as multi-purpose because one crawler does both jobs, so any Cloudflare AI crawler rule aimed at training lands on search as well. That is deliberate pressure. TechCrunch reported on July 1, 2026 that the policy exists to push AI firms into running separate crawlers for each purpose. Until they do, the collateral damage stays yours to manage.
The IRS shows what an over-tight bot rule costs
Search Engine Roundtable documented the IRS website dropping out of Google on August 27, 2026 for terms like tax brackets and free tax filing. Lily Ray posted Sistrix charts showing the collapse. Glenn Gabe dug into the technical side. Meanwhile favicons broke, meta descriptions vanished, and PDFs fell out of the index at the same time.
Barry Schwartz pinned the likely cause on bot blocking set too aggressively, catching Google along with everything else. Rankings came back within about a day. Still, a federal agency lost its highest-value keywords because of a security rule, not a content decision.
Training bots take most of the crawl and send back almost nothing

Cloudflare’s own numbers explain why owners reach for the block. In an August 28, 2025 analysis, training accounted for roughly 80% of AI bot crawling, while search and live user actions together came to under 5%. Most of that traffic is machines reading your pages so a model can learn from them.
The referral math is worse. Over the first week of August 2025, Cloudflare measured Anthropic crawling about 50,000 pages for every referral it sent back, OpenAI at 887 to one, and Perplexity at 118 to one. So the instinct to shut the door is rational. Execution is where sites hurt themselves.
Blocking is still defensible, but do it on purpose
We have argued before that far fewer sites block AI crawlers than the headlines claim, and that most sites should not bother. That still holds. But if you do want to block, block Training only and leave Search alone.
Because Googlebot is multi-purpose, though, “Training only” is never as clean as the label suggests. So test it. Fetch your sitemap and a few page templates with a Googlebot user agent from outside your network, then read the status code instead of trusting the toggle.
There is a middle option worth knowing about. Cloudflare now runs pay-per-crawl experiments, so some publishers charge for access rather than refusing it. That suits large content libraries. For most service businesses, however, the honest answer is simpler: your problem is not that models read you too much, it is that answer engines never name you at all.
Your CDN is now an SEO setting
More than 20% of the web sits behind Cloudflare, and more than half of all internet traffic is now non-human, by Cloudflare’s own July 1, 2026 count. So edge rules decide what search engines and answer engines can reach, long before your content gets a vote.
That moves a config screen into the SEO stack. Most teams still file the CDN under infrastructure that somebody else owns. Meanwhile the person writing your meta descriptions has no idea a checkbox is returning 403 to the crawler they depend on.
The IRS case makes the ownership gap concrete. Nobody on that site wanted to lose tax brackets in August. Somebody tightened a bot rule, and the ranking loss showed up days later in a different team’s dashboard. So put the edge settings on the same review list as redirects and robots directives, because they fail the same way and just as quietly.
Audit your Cloudflare AI crawler settings this week

- Open your security settings and write down which of Search, Agent and Training you currently allow.
- Request your sitemap and three key templates with a Googlebot user agent from outside your network, then log every status code.
- Check Search Console crawl stats for a rise in failed fetches since your last edge change.
- Decide the Agent lane deliberately, because agent bots act for real people who may buy something.
- Note the settings in your change log, so the next person knows edge rules move rankings.
Do this before September 15 if you plan to launch a new domain on Cloudflare, because that domain lands on the new defaults. Otherwise do it anyway. The switch is already live on your account, and nobody sends you an email when it fires.
Want help auditing your Cloudflare AI crawler setup?
We audit edge rules, crawl access and AI visibility together, because splitting them apart is how sites go invisible to the engine they care about most. Get in touch and we will tell you what your Cloudflare AI crawler settings actually return today.