Blocking AI bots with one click can cost you Google: what changes on September 15

Blocking "AI bots" with a single click now means closing three different doors, and one of them leads to Google. Since July 1, 2026, Cloudflare has sorted the bots that visit your site into three categories: search, agent and training. In that same announcement, Cloudflare states that multi-purpose crawlers such as Googlebot, Applebot and Bingbot will be blocked for customers who have chosen to block training. From September 15, 2026, new domains onboarding to Cloudflare will block the training and agent categories by default on pages that display ads, while search stays allowed.

What exactly changes at Cloudflare on September 15, 2026?

Cloudflare applies a new default configuration to domains that join its network after September 15, 2026: training and agent blocked on ad-monetized pages, search allowed through. Sites already on Cloudflare keep their current settings, and Cloudflare says a site owner can opt out of the new defaults from their security settings. The date matters less than the principle it installs: AI bot access is no longer one checkbox, it is a three-way decision.

The three categories cover very different behaviors. A search bot collects and indexes your content so it can answer questions about it later. An agent acts in real time on behalf of a person, for instance when someone asks an AI assistant to go read a specific page. A training crawler takes your content to train or fine-tune a model. Refusing the third is a defensible publishing decision. Refusing the first is refusing to exist inside the answers.

Why can blocking AI cost you Google?

Because Googlebot does not sit on the AI side of the fence: it is treated as multi-purpose. Cloudflare states that multi-purpose crawlers such as Googlebot, Applebot and Bingbot are judged on all of their behaviors, so a site that blocks training blocks them as well. Google's own documentation says Googlebot is used to build Google's search indexes and affects Google Search, Discover and every Search feature. Losing Googlebot is not losing a slice of AI visibility. It is losing search.

A common belief holds that refusing Google-Extended is the safe move. Google's documentation says close to the opposite of what people credit it with: Google-Extended governs whether crawled content may be used to train future generations of Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI, and Google states that the token does not affect a site's inclusion in Google Search and is not a ranking signal. It is not the lever most site owners think they are pulling.

What does this change for a small or midsize business?

For a small business, the risk is not the September 15, 2026 change itself. The risk is discovering a decision someone else made two years ago. Almost no small company manages a firewall day to day. An agency, a contractor or a developer ticked a "block AI bots" box back when that box drew no line between search, agent and training. That one click, made to protect content, now covers bots that feed the answers your customers actually read.

The economics point one way. Ahrefs measured across 300,000 keywords, comparing December 2023 with December 2025, that the click-through rate of the top organic position is 58% lower when an AI Overview is displayed. Search engines still carry the overwhelming majority of commercial queries, and generative answers are eating into the clicks those queries used to produce. You are not choosing between Google and AI assistants. You have to hold both, and a training block bought at the price of Googlebot trades the busiest surface for protection on the smaller one.

Which bots deserve to pass, and which are a genuine judgment call?

The major AI companies have split their crawlers apart, and each of them documents the split. Refusing training no longer forces you to refuse search.

  • OAI-SearchBot powers ChatGPT search. OpenAI writes that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.
  • GPTBot is used to train OpenAI's models. OpenAI notes that each setting is independent of the others, so blocking GPTBot does not remove you from ChatGPT search.
  • Claude-SearchBot navigates the web to improve search result quality, while ClaudeBot collects web content that could contribute to training Anthropic's models. Claude-User handles requests started by a person asking Claude a question.
  • Googlebot feeds Google Search, Discover and every Search feature. It is the crawler a careless block catches first, and the most expensive one to lose.
  • Google-Extended covers training of future Gemini models and grounding in Gemini Apps. Google states it does not affect inclusion in Google Search.

Does that mean you should open the door to everyone?

No, and refusing training remains a legitimate call. A publisher, a database, a photography catalogue or any business whose product is the content itself has good reason to refuse training while still allowing search. The change is not that everything should be open. The change is that the decision is now made category by category, and a single toggle no longer expresses anything coherent.

Opening the door also creates no reason to be cited. A site with no content that answers a customer question gains nothing from letting every crawler through: it will be readable, and ignored. Access is a necessary condition, never a sufficient one. A site whose pages do not say what it sells, to whom, or where, has an editorial problem no firewall setting will fix.

What should you check before September 15, 2026?

Three questions are enough, and none of them requires technical skill to ask. First: who actually controls the door to your site, your platform, your host or your CDN? The answer shapes everything else, and it varies a great deal with the tool your site is built on, as our article "Shopify, Wix, Webflow, Framer: what your platform actually lets you do for AI" sets out.

Second: what does your robots.txt say today, and who last read it? It is the most requested file on your site and the least examined inside your company. Keep in mind that it declares an intention and constrains nobody, a point our article "Your robots.txt is not a lock: AI bots walk past it, and OpenAI says so" develops in full. Third: is your domain joining Cloudflare after September 15, 2026? If so, the new default applies to you without anyone ticking anything.

Frequently asked questions

Does blocking GPTBot remove me from ChatGPT?

No. OpenAI documents two separate crawlers, GPTBot for model training and OAI-SearchBot for ChatGPT search, and states that each setting is independent of the others. Blocking OAI-SearchBot is what removes you from ChatGPT search answers.

My site is already on Cloudflare. Do I need to do anything on September 15, 2026?

Your current settings do not change on their own, since the new default configuration applies to domains onboarding after that date. The useful question is what was ticked on your account, and when. A setting made before the three categories existed no longer means what it meant.

Does blocking AI bots actually protect my content?

Partly, and less than people assume. A network-level block bites harder than a robots.txt line, which is only a declaration of intent. Neither prevents content already republished elsewhere from circulating, nor stops a user from pasting your page into a conversation.

Is refusing model training a bad decision?

Not at all, when the content is the product. A publisher, a database or a photo library has a direct interest in refusing training. What changed is that they can now do it without cutting themselves off from the search crawlers that feed generative answers.

Topics
  • ai bots
  • cloudflare
  • googlebot
  • ai visibility
  • geo
  • crawlers
Blocking AI bots with one click can cost you Google: what changes on September 15