We are measuring Cloudflare’s 15 September change. Here is the method, before the event.
On 15 September 2026 Cloudflare changes how it treats AI crawlers. Afterwards nobody will be able to say what any given site looked like before — the answer stops being recoverable the moment it changes. So we measured 1,046 websites first, and we are publishing the method and the baseline now, while the result is still unknown to us. The measurement runs on 16 September and this page will show it either way.
The short version
- 1,046 websites, each fetched by ten identities — eight AI crawlers, an ordinary browser, and the same browser again.
- 746 sit behind Cloudflare, 300 do not. The 300 are the point: if access changes in both groups, the cause was not Cloudflare, and without them there would be no way to know that.
- The three search crawlers stay in the set as a control inside every host, precisely because Cloudflare says they are unaffected. Training and agent crawlers move while search does not ⇒ Cloudflare did what it announced. Search moves too ⇒ that is the headline.
- Baseline, taken 2026-09-02: sites behind Cloudflare already refused AI crawlers roughly four to ten times more often than sites that are not — 25.3% versus 5.8% for the most-refused identity — before anything changed.
- Refusals concentrate on Cloudflare-fronted sites far beyond their share of the sample. Those sites are 71.3% of the hosts and 94.6% of the refusals. We cannot say who decided — see below, it is the most important limit on this page.
- The after-measurement is outstanding. It runs on 16 September and lands here.
Why this is being published before the answer is known
A before/after study is only worth anything if the “before” was recorded before. That is not a claim a reader can check after the fact, and a measurement produced entirely after an event is exactly as trustworthy as the people who produced it — which, for a company nobody has heard of, is not very.
So the method, the host list construction, the controls and the full baseline are here now, dated, while we genuinely do not know what the answer will be. If the result turns out to be dull, this page will say so with the same numbers. Fixing a method after seeing the data is the failure this ordering exists to make impossible.
How it is measured
Each host is fetched by the same ten identities in a fixed order, 300 ms apart, six hosts in flight at a time, recording status code, response length, a hash of the body, the server header and whether a cf-ray was present. Every host experiences an identical sequence on every run.
The order brackets the crawlers between two identical browser requests — an ordinary browser first, then the eight AI crawlers, then that same browser again at the end. The closing fetch is the noise floor: it spans the whole sequence, so a host that changed halfway through is caught rather than blamed on whichever crawler happened to be asking at the time.
The method is pinned by a version number and the comparison tool refuses to compare two runs whose method versions differ, because a delta produced by our own change to the instrument would look exactly like a delta produced by the open web.
The runs are spaced deliberately, and this was corrected once. An anchor on 2 September, a drift run on 10 September, the before-run on 13 September, the after-run on 16 September. The 10th exists so that ordinary drift is measured over three days — the same interval as 13 → 16 — rather than over ten. Comparing a ten-day drift rate against a three-day change assumes drift accumulates linearly in time, which is the precise assumption a control group is there to avoid making.
The baseline, 2026-09-02
Refusal rate per identity, as a share of the hosts that answered an ordinary browser normally (1,028 of 1,046).
| Identity | Kind | All sites | Behind Cloudflare 735 | Not behind Cloudflare 293 |
|---|---|---|---|---|
ClaudeBot | Training | 19.7% | 25.3% | 5.8% |
GPTBot | Training | 19.4% | 24.4% | 6.8% |
OAI-SearchBot | Search | 17.2% | 22.7% | 3.4% |
ChatGPT-User | Agent | 16.8% | 22.4% | 2.7% |
PerplexityBot | Search | 16.3% | 21.5% | 3.4% |
Claude-SearchBot | Search | 15.6% | 20.5% | 3.1% |
Claude-User | Agent | 15.0% | 20.1% | 2.0% |
Perplexity-User | Agent | 15.0% | 20.1% | 2.0% |
An ordinary browser | Control | 0.0% | 0.0% | 0.0% |
The gap between the last two columns is the finding in this table. Before Cloudflare changed anything, an AI crawler was already several times more likely to be refused by a site behind Cloudflare than by one that is not. An ordinary browser was refused by neither, anywhere, which is what makes the comparison meaningful rather than a measure of how many sites were simply down.
We are naming no sites, here or on 16 September. Most of these owners have not been told, and a monitoring company that publishes other people’s configuration before telling them is not one anybody should trust with their own.
What we will not claim, and why
These are refusal rates, not blocking rates, and we cannot tell you who decided. 1,357 of 1,436 HTTP refusals (94.5%) carried a cf-ray. That header proves Cloudflare was in the path. It does not prove Cloudflare made the decision. A site’s own origin can return 403 and have that answer proxied back through Cloudflare, and from outside it is indistinguishable from a refusal Cloudflare issued itself. Nothing measurable from here separates Cloudflare’s own bot handling from a rule a customer wrote in their WAF.
What survives that objection is the concentration. Cloudflare-fronted hosts are 71.3% of this sample but account for 94.6% of the refusals, and not one of the 78 refusals from the non-Cloudflare control group carried a cf-ray at all. Where refusals happen is measurable from outside. Who chose them is not, and a monitoring product that blurred that line would be doing the thing it sells against.
This paragraph is here because an outside reader was right and we were wrong. An earlier draft said 94.5% of refusals “came from a CDN, not from the site”. That is not what the header shows, the objection came from someone who has measured this at far greater scale than we have, and the correction makes the finding narrower and the page better.
We lead on status codes, not on body changes. 361 of 1,026 hosts (35.2%) returned different bodies to two identical browser requests in the same run. Those pages rewrite themselves between fetches, so a body difference on them says nothing about crawlers. Excluding them cuts the body-change signal by about four fifths, and what remains is flat across all eight crawlers — which is the signature of content that changes on its own, not of anything crawler-specific. An earlier draft of this study would have reported that raw figure as “sites serve crawlers different content”. It would have been wrong, and the control is the only reason we know.
The announced scope is narrow, and that cuts against the drama. Of 3,535 hosts considered, 2,090 were reachable, 746 sat behind Cloudflare, and 52 were running ad tags. Both at once: 15. If the change stays inside the scope Cloudflare announced, it touches about 0.7% of these sites. Either it is far narrower than the reaction suggests, or it will not stay inside its stated scope — and only the controls can tell those apart. That number is measured on agency sites, which are structurally ad-free; it is not a statement about the web, and it is not a statement about the client sites those agencies look after.
What happens on 16 September
The same 1,046 hosts are fetched again by the same ten identities under the same pinned method, and the result is compared against the 13 September run. This page then shows the delta — including if the delta is nothing.
A null result gets published in the same place, at the same size. “Cloudflare’s change did not measurably alter AI crawler access to 1,046 sites” is a real finding and a useful one, and a study that would only report a change is not a study.
Why we are the ones doing this
SeenSure monitors whether AI crawlers can actually reach a website, and tells the owner when that changes. This measurement is the same thing our customers get, run across a thousand sites at once. If you want the answer for your own sites rather than for a sample, check one now — it takes a few seconds and needs no account.
Method, host-list construction and the verification that every run must pass before it can be compared are in the open: tools/capture-baseline.cjs. A run that fails its own verification cannot be published by us at all — the tool refuses to compare it, which is how a measurement made off a dropped connection stops becoming a finding about Cloudflare.