Guide

robots.txt says yes. The crawler still gets a 403.

You checked robots.txt. OAI-SearchBot is allowed, GPTBot is allowed, nothing says Disallow. And ChatGPT still can't read the page.

That happens a lot, and the reason is simple. robots.txt is a note on the door. It tells polite crawlers what you'd like. It doesn't open or lock anything. The thing that actually answers the request is your server, or whatever sits in front of it: a CDN, a firewall, a bot-protection rule, a security plugin. Any of those can refuse a crawler that robots.txt invited in.

How often does that happen?

We fetched 1,046 websites as an ordinary Chrome browser and as eight AI crawlers, in the same few seconds, from the same machine. The browser was refused by none of them. The AI crawlers were refused by roughly one site in six. On 2 September, OAI-SearchBot (the crawler that builds ChatGPT's search index) was refused by 17.0% of the sites that served a browser normally. ChatGPT-User, the fetch that happens when a person asks ChatGPT about a page, was refused by 16.6%.

Where those refusals happen is very lopsided. Sites behind Cloudflare were 71.3% of the sample and 94.6% of the refusals. On Cloudflare-fronted sites OAI-SearchBot was refused 22.6% of the time. Everywhere else it was 3.1%.

One caution. Our sample leans toward newly registered domains, which CDNs challenge harder, so the gap on a typical established business site is smaller (1.4x to 2.8x on the agency sites in the sample, against 4x to 12x overall). And a refusal that passes through Cloudflare isn't proof Cloudflare decided it. The site's own server or a WAF rule someone wrote can return the 403 and Cloudflare just carries it back. From outside, those look the same.

Where to look, in the order we'd check

  1. Your CDN's bot settings. On Cloudflare that's Security, then Bots. Look at Bot Fight Mode and the AI crawler controls (Search, Agent, Training are separate switches).
  2. WAF or firewall rules that match on user agent. A rule written to stop scrapers years ago ("block anything with bot in the name") catches every AI crawler too.
  3. Security plugins on the site itself. WordPress firewall plugins are a common source.
  4. Hosting-level protection. Some hosts filter bots before your site ever sees the request.

How to see it without logging into anything

Ask for the same page twice at the same moment, once as a browser and once as the crawler, and compare the answers. If the browser gets the page and the crawler gets a 403 or a challenge page, something in front of the site is refusing it, whatever robots.txt says.

That's exactly what our free check does, for 19 AI crawlers and opt-out rules, in about five seconds. No account, and the result stays private.

Related: Every AI crawler and what blocking it costs · Cloudflare's 15 September change · The full 1,046-site study