Guide

Is Cloudflare blocking AI crawlers on a site you look after?

Every guide to this starts with "log into your Cloudflare dashboard". Fine if it's your site. If you run a dozen client sites, half of them were put on Cloudflare by someone who left, and the login is in an inbox nobody checks.

You don't need the dashboard to find out whether AI crawlers are getting in. You need to ask for the page the way the crawler does and see what comes back.

Why it's worth checking client sites now

On 15 September 2026 Cloudflare changed its defaults for AI crawlers on the free plan. We measured the same 1,046 sites the day before and the day after. On Cloudflare-fronted sites, refusals of the two training crawlers went up (GPTBot 18.9% to 22.0%, ClaudeBot 19.9% to 22.8%). Refusals of search and agent crawlers dropped sharply, OAI-SearchBot from 16.9% to 3.0%. Sites not on Cloudflare barely moved.

So a site's access can change without anyone touching it. Which way it moved for any one client, you only find out by checking that site.

What to check, from outside

  • The page itself, fetched as a normal browser and as each AI crawler. If the browser gets it and GPTBot gets a 403 or a challenge page, there's a block.
  • robots.txt, read the way each crawler reads it. A blanket "User-agent: * Disallow: /" or a copied block list can shut out crawlers you wanted.
  • Whether the text is in the HTML. AI crawlers don't run JavaScript, so a page that builds itself in the browser can look empty to them even when nobody refuses anything.

Our free check does all three in about five seconds, for 19 AI crawlers and opt-out rules. You get a private link you can send to the client, and the bare address won't open it for anyone else.

If something is blocked, then you need the login. On Cloudflare the switches are under Security, then Bots. Search, Agent and Training are separate, so you can keep training crawlers out and still let ChatGPT search and user fetches in.

If you look after more than a few sites, checking once isn't enough, because defaults, plugins and hosts keep changing. That's what SeenSure's paid plans watch for, and they alert you when a site you manage stops letting a crawler in.

Related: robots.txt allows it but it's still blocked · Cloudflare's 15 September change