Three signals, ranked by how much we trust them
Most checkers give you a confident yes or no. Some of those answers are guesses, and a guess dressed as a fact is worse than no answer at all. Here is exactly what we can prove, what we can only infer, and where we will tell you we do not know.
The whole method in one picture. Same page, same second, same machine — the only thing we change is who we say we are.
Signal 1 — what your robots.txt declares
Every AI crawler reads a file at yoursite.com/robots.txt before it visits, and obeys what it finds. If that file says a crawler is not welcome, the crawler stays away. There is no ambiguity here, so we treat this as fact.
This one check catches a surprising share of real problems: templates copied from another site, a line added years ago to stop scrapers, or a security plugin that rewrote the file during an update.
User-agent: GPTBot
Disallow: /
Two lines. Nobody remembers adding them. ChatGPT has not read the site since.
Signal 2 — what a live crawler request actually gets
robots.txt is only a request. A server can still refuse a crawler outright. So we ask for your page twice, at the same moment, from the same machine — once identifying as an ordinary browser, once as the crawler. Then we compare.
| What we see | What it means |
|---|---|
| Both get the page | You are fine. |
| Browser gets it, crawler is refused | A real block. We tell you where and how to lift it. |
| Both get a much smaller page | Content is hidden behind JavaScript. |
| Neither gets it | Your site is down. Not an AI problem — we say so rather than raising a false alarm. |
Signal 3 — verified crawler visits
Signals 1 and 2 stand outside your site and work out what a crawler would get. This one records what crawlers actually got, and it is the only signal here that produces evidence rather than a well-reasoned inference.
It exists because a user-agent header is a claim, not a fact. If your access log shows a thousand GPTBot requests, you do not know how many came from OpenAI — the string is just text, and anyone can send it. That matters more than it sounds: if you grant known AI crawlers access you would not give an unknown scraper, a spoofed name turns your own rule against you.
So on Pro plans your server reports each crawler visit to us, and we check the visitor's address against the range that crawler's operator publishes. An address inside OpenAI's own published range is not a claim anybody can make.
It has to run on your server rather than in the browser, and that is not a preference. AI crawlers do not execute JavaScript — that is half of what this product checks for — so a script tag would faithfully record every human visitor and not one crawler. It is a few lines in a Cloudflare Worker, in Node middleware, or in PHP.
Install it at the edge if you can, and here is the honest reason why. A full-page cache sits in front of your application. WordPress serves a cached page from advanced-cache.php, or your web server answers from disk, before PHP or your framework ever sees the request. A beacon running inside the application is skipped entirely on those hits — it does not slow anything down and it does not corrupt your cache, it simply never runs, and that crawler visit goes unrecorded.
That matters more here than it would in most tools, because we infer a block from a gap in visits, and an unrecorded visit looks exactly like an absent one. We would rather tell you this before you install than send you an alert about a crawler that never actually left. A Cloudflare Worker runs ahead of the cache and does not have the problem; if you install in PHP, treat it as a floor on what visited, not a complete record.
This is also the honest answer to the limitation above. Cloudflare, Fastly, Akamai and Imperva verify a crawler's real identity by IP range, and increasingly by cryptographic signature. No outside tool can imitate that — ours included. So from outside we say "cannot confirm from outside". A verified visit settles the same question from the inside, where the evidence actually is.
What else we check
- JavaScript dependency. Most AI crawlers do not run JavaScript. If your content only appears after the browser does work, they see an empty page.
- llms.txt. Optional, but a broken one is worse than none — it hands crawlers a map to pages that no longer exist. We check every link.
- Page basics. Title, description, headings, structured data, sitemap. These decide whether an AI describes you accurately or invents something.
How often
Every six hours on paid plans, daily on the free plan. We only email you when something changes — a break, or a recovery. Silence means everything is fine.