Get a free website with any plan

See how
CLOUDFLARE

Controlling AI crawlers, AI assistants and model training on your site

Last updated

IN SHORT

The Cloudflare AI crawlers panel in Flashcloud lets you manage AI search crawlers, AI assistants, and model training bots independently. You set rules like Allow or Block for each behavior. This lets you block bots from scraping site content for model training while still allowing real-time AI assistants to answer user questions.

Go to the Cloudflare CDN page in the portal, pick your domain, and open the AI crawlers panel. You'll see three rows: Search crawlers, AI assistants, and Model training. Each has its own control, so you can let a chatbot fetch a page on a user's behalf while still refusing to let your content train the next model.

The three rows, and what each one covers

These are three different behaviors, not one "AI" switch:

  • Search crawlers - bots indexing your pages for AI-powered search and answer engines (the AI equivalent of a search engine indexer).
  • AI assistants - bots that fetch a specific page in real time because a user asked an assistant a question, like an assistant pulling up your pricing page to answer someone's question.
  • Model training - bots scraping content to include in a future model's training data. This one has an extra option beyond Allow and Block.

Setting each row

Search crawlers and AI assistants each take one of three settings:

  • Allow - the bot can fetch your pages freely.
  • Block on ad pages - blocked on pages you've marked as ad-supported, allowed elsewhere.
  • Block - blocked everywhere on the domain.

Model training has the same three settings, plus a fourth: Ask nicely (robots.txt), a lighter-touch option between Allow and Block.

Why the settings are split this way

A lot of site owners want AI assistants to be able to answer questions about their business (that's free visibility) while still refusing to hand over the entire site as training data. Splitting the panel into three independent rows makes that combination the default case, not a workaround.

What changed on 2026-09-15

Cloudflare updated the default posture for Training and AI assistants on ad-supported pages for newly proxied domains starting 2026-09-15. Sites created through Flashcloud stay fully open under this change. If your domain was added or re-proxied after that date, check those two rows rather than assuming older defaults still apply.

How this fits with the rest of your setup

If a change here doesn't seem to take effect, confirm proxy is on for the domain.

This is a separate mechanism from llms.txt. llms.txt is a voluntary summary file that tells an assistant what your site is about and which pages matter; the AI crawlers panel is an access control that decides whether a bot gets to fetch anything at all. A site can publish an llms.txt to make itself easy to summarize while still blocking model training in this panel. They don't conflict, they just answer different questions: llms.txt says "here's what we are," the AI crawlers panel says "here's what you're allowed to do about it."

If pages still seem to serve old behavior to bots after a change, confirm proxy is on for the domain, and see purging the Cloudflare cache if you're troubleshooting something that looks stale for other reasons.

Common questions

Why are my AI bot block rules not working?

Cloudflare proxy is likely disabled. Confirm proxy is turned on for the domain. If pages still serve old rules to bots, purge the Cloudflare cache.

Can I let chatbots view pages without allowing model training?

Yes. Set the AI assistants row to Allow and the Model training row to Block. The panel treats assistant lookups and model training as separate behaviors.

Does an llms.txt file conflict with Cloudflare AI crawler settings?

No. An llms.txt file provides a voluntary summary of your site, while the panel enforces actual access control. You can publish llms.txt and still block scrapers.

What does Ask nicely do for model training?

Ask nicely sets robots.txt rules to request that training bots skip your content. It provides a middle ground between Allow and Block exclusively on the Model training row.

CAN'T FIND IT?

Real humans answer fast.

Hosting with us? Open a ticket and a real person replies - no scripts, no upsells. Still choosing a host? The same team is included with every plan, from day one.