How it works

Three signals, ranked by how much we trust them

Most checkers give you a confident yes or no. Some of those answers are guesses, and a guess dressed as a fact is worse than no answer at all. Here is exactly what we can prove, what we can only infer, and where we will tell you we do not know.

How the paired request worksONE PAGE, TWO IDENTITIES, SAME SECONDIDENTITY AA person's browserUser-Agent: Mozilla/5.0 … ChromeIDENTITY BAn AI crawlerUser-Agent: GPTBot/1.2yourclient.comCDN / firewallhost defaults, bot rulesOrigin serverplugins, .htaccessWhere blocks actually live —and where nobody is looking.200 OK618 KB. The whole page.403 Forbidden0 KB. Nothing at all.Then we compare the two answersIdentical: the site is fine todayDifferent: a real blockCDN verifies bots: unverifiable

The whole method in one picture. Same page, same second, same machine — the only thing we change is who we say we are.

Signal 1 — what your robots.txt declares

Every AI crawler reads a file at yoursite.com/robots.txt before it visits, and obeys what it finds. If that file says a crawler is not welcome, the crawler stays away. There is no ambiguity here, so we treat this as fact.

This one check catches a surprising share of real problems: templates copied from another site, a line added years ago to stop scrapers, or a security plugin that rewrote the file during an update.

# A real example we find constantly
User-agent: GPTBot
Disallow: /

Two lines. Nobody remembers adding them. ChatGPT has not read the site since.

Signal 2 — what a live crawler request actually gets

robots.txt is only a request. A server can still refuse a crawler outright. So we ask for your page twice, at the same moment, from the same machine — once identifying as an ordinary browser, once as the crawler. Then we compare.

What we seeWhat it means
Both get the pageYou are fine.
Browser gets it, crawler is refusedA real block. We tell you where and how to lift it.
Both get a much smaller pageContent is hidden behind JavaScript.
Neither gets itYour site is down. Not an AI problem — we say so rather than raising a false alarm.

Signal 3 — verified crawler visits

Signals 1 and 2 stand outside your site and work out what a crawler would get. This one records what crawlers actually got, and it is the only signal here that produces evidence rather than a well-reasoned inference.

It exists because a user-agent header is a claim, not a fact. If your access log shows a thousand GPTBot requests, you do not know how many came from OpenAI — the string is just text, and anyone can send it. That matters more than it sounds: if you grant known AI crawlers access you would not give an unknown scraper, a spoofed name turns your own rule against you.

So on Pro plans your server reports each crawler visit to us, and we check the visitor's address against the range that crawler's operator publishes. An address inside OpenAI's own published range is not a claim anybody can make.

It has to run on your server rather than in the browser, and that is not a preference. AI crawlers do not execute JavaScript — that is half of what this product checks for — so a script tag would faithfully record every human visitor and not one crawler. It is a few lines in a Cloudflare Worker, in Node middleware, or in PHP.

Install it at the edge if you can, and here is the honest reason why. A full-page cache sits in front of your application. WordPress serves a cached page from advanced-cache.php, or your web server answers from disk, before PHP or your framework ever sees the request. A beacon running inside the application is skipped entirely on those hits — it does not slow anything down and it does not corrupt your cache, it simply never runs, and that crawler visit goes unrecorded.

That matters more here than it would in most tools, because we infer a block from a gap in visits, and an unrecorded visit looks exactly like an absent one. We would rather tell you this before you install than send you an alert about a crawler that never actually left. A Cloudflare Worker runs ahead of the cache and does not have the problem; if you install in PHP, treat it as a floor on what visited, not a complete record.

This is also the honest answer to the limitation above. Cloudflare, Fastly, Akamai and Imperva verify a crawler's real identity by IP range, and increasingly by cryptographic signature. No outside tool can imitate that — ours included. So from outside we say "cannot confirm from outside". A verified visit settles the same question from the inside, where the evidence actually is.

Absence detection. "GPTBot fetched your site every day for a month, then stopped 14 days ago." No test run today can produce that sentence, however thorough — it only exists if something kept the record. A crawler that had a habit and broke it is usually a block nobody knew they had. That is proof, not a guess.

What else we check

  • JavaScript dependency. Most AI crawlers do not run JavaScript. If your content only appears after the browser does work, they see an empty page.
  • llms.txt. Optional, but a broken one is worse than none — it hands crawlers a map to pages that no longer exist. We check every link.
  • Page basics. Title, description, headings, structured data, sitemap. These decide whether an AI describes you accurately or invents something.

How often

Every six hours on paid plans, daily on the free plan. We only email you when something changes — a break, or a recovery. Silence means everything is fine.