LYRENTH
DocsPricingBenchmarksIndex statsAboutBlogFor site ownersContact
July 30, 2026 · site owners · crawlers · policy

Cloudflare will block AI crawlers by default on September 15: what site owners should actually do

Cloudflare blocks Training and Agent crawlers by default on ad pages from September 15, 2026. What changes, who it affects, and a site-owner checklist.

Cloudflare will block AI crawlers by default on September 15: what site owners should actually do

On July 1, Cloudflare announced the biggest change to how the web treats automated readers since robots.txt: starting September 15, 2026, new defaults that block whole categories of AI crawlers unless a site owner decides otherwise. If you run a website, this change reaches you whether you use Cloudflare or not, because it resets what "normal" looks like for every bot operator on the web.

Here is what was actually announced, what it means for your site, and a short checklist to work through before the deadline.

What Cloudflare announced, precisely

The announcement (Cloudflare calls it their second Content Independence Day) has three parts.

1. A taxonomy: Search, Agent, Training. Instead of asking whether a bot "is AI," Cloudflare now classifies bots by what they do with your content:

  • Search: collects or indexes your content so it can answer questions about it later, and is expected to send referral traffic or other compensation back.
  • Agent: acts in real time on a person's behalf, like chat fetchers and browser-driving agents, usually with a human waiting on the other end.
  • Training: takes your content to train or fine-tune a model, absorbing it permanently into the model's weights.

All Cloudflare customers, including the free tier, can now allow or block each category separately.

2. New defaults on September 15. For new domains onboarding to Cloudflare, Training and Agent crawlers will be blocked by default on pages that display ads, while Search remains allowed by default. Cloudflare's reasoning: an ad is a signal that a page was meant for human attention, so bots that bypass that attention get blocked there, while search, the behavior that funnels visitors back, stays open.

3. Multi-purpose crawlers get judged by all of their behaviors. This is the part with the widest blast radius. A crawler that combines Search with Training is subject to the most restrictive applicable rule. Cloudflare names Googlebot, Applebot, and BingBot specifically: customers who block Training will block those crawlers too, because they crawl for both purposes with one bot. Site owners who do not want that outcome can opt out in their security settings before September 15.

What this means depending on who you are

If your site runs ads: you are the direct target of the new defaults. After September 15, new Cloudflare domains will serve your ad pages to human visitors and search crawlers, and refuse training scrapers and real-time agent fetchers by default. If you want agents to be able to read your pages (for example, because AI assistants recommending your product is how customers find you), that is now an explicit choice you have to make, not a default you inherit.

If you sell through AI recommendations: be careful with the blunt settings. Blocking the whole Agent category also blocks the fetcher that reads your pricing page when a customer asks their assistant "which of these tools should I buy?" The taxonomy exists so you can make a finer choice than "block all automation."

If you rely on Google: the multi-purpose rule is the one to understand. Blocking Training without reading the fine print can block the crawler your search traffic depends on. Check which of your rules apply to combined-purpose bots before the defaults land.

If you are not on Cloudflare at all: the defaults still matter, because they reset expectations. Bot operators now have a strong incentive to split their crawlers by purpose and identify themselves clearly, and site owners everywhere will start asking the same question Cloudflare's dashboard asks: what does this bot do with my content?

The question behind the question: who is actually reading you?

Category rules only work if bots tell the truth about their category, and anyone can put anything in a User-Agent string. That is why the other half of this shift is verification: cryptographic request signatures, published IP ranges, and reverse DNS that let you confirm a crawler is who it claims to be. Cloudflare maintains a directory of verified bots for exactly this reason.

This is also where we should say plainly where Lyrenth sits in this taxonomy. Lyrenth operates a single-purpose, search-class crawler: it fetches public pages once, indexes them, and serves every subsequent reader from the index, with attribution and a link back to the source. We do not train foundation models on the content we crawl, and every fetch honors robots.txt. Our crawler signs its requests, publishes its IP ranges, and documents both of its User-Agent strings in our crawler policy, so you can verify us instead of trusting a string.

If you want to see which AI bots are reading your site today, verified against the impostors, that is exactly what our free site-owner tools show: see which AI bots read your site.

The checklist before September 15

  1. Inventory your bot traffic now. Before defaults change, know your baseline: which crawlers visit, in which category, and which pages they read. Decisions made blind are the ones you reverse in October.
  2. Decide per category, not per vibe. Search: almost every site should keep it open, it is where discovery comes from. Training: your call, and it is a real trade with real leverage only if your content is distinctive. Agent: think about whether AI assistants acting for customers are a channel for you or a cost.
  3. Check the multi-purpose consequences. If you block Training, confirm you are comfortable with what that does to combined-purpose crawlers like Googlebot under the most-restrictive rule.
  4. Mind your ad pages specifically. The new defaults key off ad presence. Know which of your pages carry ads and whether the default treatment there matches what you want.
  5. Say your policy in robots.txt too. Cloudflare settings govern Cloudflare's network; robots.txt speaks to every well-behaved crawler everywhere. Keep the two consistent so honest bots outside Cloudflare hear the same policy.
  6. Prefer verified crawlers. Whatever you allow, allow it for bots that prove their identity. An allowlist built on User-Agent strings alone is an invitation.

The bigger picture

The 30-year-old deal ("crawl me, send me visitors") broke when reading stopped producing visits, and September 15 is the first time a major infrastructure provider changes the web's defaults to reflect that. The direction is clear: purpose-labeled, identity-verified, permission-based crawling. That is the world Lyrenth was built for from day one, and it is a better world for site owners than the anonymous free-for-all it replaces.

If you want your site to be represented accurately to every AI reader that comes through our index, verify your domain and you author the canonical version agents get, with a dashboard showing exactly who reads you.

All postsRead a URL in 5 minutes