Is your website blocking AI crawlers?

Enter a web address. The checker reads your robots.txt, then visits the page as a browser, as OpenAI's GPTBot and as Common Crawl's CCBot, and shows you where the stories differ.

Free, no sign-up. Four requests to your site per check. Results are kept for 10 minutes, then deleted.

How the check works

Most AI crawler checkers read your robots.txt and stop there. That file only records what you asked for. What matters is what your server does when a crawler turns up, and on a lot of sites the two don't match, because a CDN, firewall or security plugin makes its own decisions.

So this tool makes four requests. It fetches your robots.txt and works out, crawler by crawler, what the rules allow on the page you entered. It loads the page as an ordinary browser. Then it loads the same page identifying as GPTBot, and again as CCBot. If the browser gets your page and a crawler gets a 403 or a challenge screen, something between you and the internet is treating that crawler differently.

What it looks for that most checkers don't

  • Host-level CCBot blocks. Hosting companies often block Common Crawl's CCBot at server level because it crawls heavily. Nothing appears in robots.txt, so you'd never know. It happened to this site. CCBot doesn't feed live AI answers, but Common Crawl's dataset is a major source of AI training data.
  • Rules you didn't write. Cloudflare's managed robots.txt setting can add AI blocks to the file crawlers see, without them appearing in your CMS or SEO plugin.
  • Live-fetch agents robots.txt can't stop. OpenAI's documentation has said since December 2025 that robots.txt rules may not apply to ChatGPT-User, because a person started the request. Perplexity-User is in the same position.
  • The Google-Extended mix-up. Blocking Google-Extended keeps you out of Gemini training. It does nothing to AI Overviews or AI Mode, which use ordinary Googlebot.
  • Copilot. Microsoft Copilot answers draw on Bing, so a site that blocks Bingbot drops out of Copilot too. Most AI checkers leave Bingbot off their list.
  • Group precedence. A crawler with its own group in robots.txt ignores your User-agent: * rules entirely, so paths you thought were closed to everyone may be open to it.
  • Page-level switches. noindex, nosnippet and max-snippet:0 on the page, and pages that show almost no text until JavaScript runs.

What it can't see

Our requests come from our server, not from OpenAI's. Some security tools reject anything claiming to be GPTBot unless it comes from OpenAI's published IP addresses, which is correct behaviour and looks identical to a block from the outside. If you get a GPTBot result you didn't expect, ask your host to confirm whether the real crawler gets through.

It can't see Google's "Search generative AI" setting in Search Console either. That setting, added in June 2026 after the Competition and Markets Authority required it, can remove a whole site from AI Overviews and AI Mode while leaving normal rankings alone. If you or a previous agency have access to Search Console, it's worth a look.

For a live test of every major AI crawler one by one, the 365i AI Crawler Checker sends a request as each of 14 bots. The two tools cover different ground, and running both takes under a minute.

The checker covers three of the five places a block can hide: robots.txt, your server or CDN, and the page itself. For the other two, Google's Search generative AI setting and what your server logs show, see the full guide: Is your website blocking AI? How to check and fix AI crawler access.

Questions

Can I block ChatGPT-User with robots.txt?

Not reliably. Since December 2025 OpenAI's documentation says robots.txt rules may not apply to ChatGPT-User, because a person started the request. Blocking it takes a firewall or server rule. GPTBot and OAI-SearchBot still follow robots.txt.

Does blocking Google-Extended remove my site from AI Overviews?

No. AI Overviews and AI Mode use ordinary Googlebot and Google's normal search index. Google-Extended only controls Gemini training and grounding. The control for AI Overviews and AI Mode is the Search generative AI setting in Search Console.

Why does the checker say GPTBot was blocked when my robots.txt allows it?

robots.txt only states your rules. A firewall, CDN or security plugin can still reject a request that identifies as GPTBot. Some of them only reject traffic pretending to be GPTBot and let the real crawler through, so ask your host which is happening before changing anything.

Access looks fine but AI still leaves you out?

That's the usual finding. Once crawlers can reach a site, the gap is almost always in the content. The AI visibility report tests the questions your buyers ask across ChatGPT, Perplexity and Google's AI Overviews, and shows where competitors get named instead of you.

To cite this tool: Wood, D. (2026). AI crawler checker. aivisibilitygap.com. https://aivisibilitygap.com/ai-crawler-checker.html