Is your website blocking AI? How to check and fix AI crawler access in 2026
Most websites that are blocking AI never made a conscious decision to do so. Someone ticked a box in a security plugin a couple of years ago, or Cloudflare added a managed rule, or the host runs a server-level filter nobody mentioned during setup. The robots.txt file looks fine. The crawler still gets a 403.
So a proper check means looking in more than one place. Pick a page your customers actually land on, a service page or a well-read article, and work through the checks below in order. For one URL it takes about ten minutes.
For a quick first read, our free AI crawler checker shows the robots.txt rules served for the page you enter, side by side with live responses when our server requests it as an ordinary browser, as GPTBot and as CCBot. Those live requests come from our server. If GPTBot gets refused there, treat it as a lead and confirm it against OpenAI's published IP ranges before calling it a block.
How to check if your website is blocking AI crawlers: five places a block can hide
Each AI company runs several bots, and each layer of your stack can treat them differently. A training crawler, a search crawler and a fetch triggered by someone typing a question into ChatGPT can all get different answers from the same website.
- robots.txt as the public receives it, including anything your CDN adds on the way out.Check 1
- Your CDN, host or firewall. A CDN (content delivery network) such as Cloudflare sits in front of your site and can return a
403, a429or a challenge page whatever robots.txt says.Check 2 - The page itself: login walls, redirect chains, JavaScript-only content, noindex and snippet controls.Check 3
- Google's Search Console setting for AI Overviews and AI Mode, which has nothing to do with robots.txt.Check 4
- What actually reached your server, in your logs or bot reports.Check 5
For Google, a page has to be indexed and eligible to show a snippet in ordinary Search before it can appear as a supporting link in AI Overviews or AI Mode. Google's guidance also names CDN and hosting blocks specifically as something to rule out.1 Eligible just means eligible to be chosen as a link, though. Getting chosen as a supporting link is a separate question.
How many websites block AI crawlers?
Fewer than you'd think. And even fewer deliberately.
The ai-visibility.org.uk quarterly robots.txt study, run by Mark McNeece at 365i, crawled 1,744 websites. In that sample:2
GPTBot also gets blocked roughly twice as often as ChatGPT-User. That one is interesting. Plenty of site owners blocking OpenAI have shut out the training crawler and left the user-triggered fetcher alone, which may or may not be what they intended.
These figures count declared rules in robots.txt. They can't show whether a firewall let the unnamed bots through, or whether any crawler obeyed. The number I'd pay most attention to is the 84.2%. Most sites are running on whatever defaults their platform, plugin or CDN shipped with, and have never checked.
Which AI crawler do you need to allow for ChatGPT, Claude, Perplexity and Google?
Be specific here. Almost every bad AI crawler decision I come across starts with blocking the wrong bot.
OpenAI runs three main agents. GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot surfaces sites in ChatGPT search, and OpenAI recommends allowing it if you want to appear there. Sites that block it can still turn up as plain navigational links. ChatGPT-User fetches a page when someone asks ChatGPT to look at it, and OpenAI says robots.txt rules may not apply because a person started the request.3,4 A robots.txt change takes about 24 hours to reach ChatGPT search.3
So blocking GPTBot keeps you out of training data. ChatGPT search is governed by OAI-SearchBot.
| Agent | Operator | What it does | Does robots.txt control it? | How to check |
|---|---|---|---|---|
GPTBot | OpenAI | Collects content for model training | Yes | Named group, OpenAI IP list3 |
OAI-SearchBot | OpenAI | ChatGPT search results | Yes, this is the search opt-out | Named group, host and CDN access |
ChatGPT-User | OpenAI | Fetches a page when a user asks | OpenAI says rules may not apply | WAF response, source IP |
ClaudeBot | Anthropic | Collects content for model training | Yes5 | Named group |
Claude-SearchBot | Anthropic | Claude search indexing | Yes | Named group |
Claude-User | Anthropic | Fetches a page when a user asks | Yes, per Anthropic's policy | Named group, logs |
PerplexityBot | Perplexity | Perplexity search index | Yes, and Perplexity recommends allowing it6 | Named group, Perplexity IP list |
Perplexity-User | Perplexity | Fetches a page during a user's question | Perplexity says it generally ignores robots.txt6 | WAF response, source IP |
Googlebot | Search crawl, including AI Overviews and AI Mode eligibility | Yes | URL Inspection | |
Google-Extended | A control token for certain Google AI uses, including model training | Yes, but no crawler sends requests under this name7 | Rule only. It will never show in your logs | |
bingbot | Microsoft | Bing index, which Copilot draws on | Yes | Bing Webmaster Tools |
CCBot | Common Crawl | Open web dataset, widely used to train AI models | Yes8 | Rule plus a live test |
Note: Anthropic says Claude-User respects robots.txt, the reverse of OpenAI's position on ChatGPT-User. Anthropic also warns that blocking its IP addresses can stop it reading your robots.txt at all, and asks you to add rules to every subdomain you want covered.5
And on CCBot: blocking it keeps you out of Common Crawl, which is widely used to train AI models.
Check 1: Is your robots.txt blocking GPTBot, OAI-SearchBot or other AI crawlers?
Fetch https://yourdomain.com/robots.txt from a browser or the command line, from outside your CMS, and test the exact page URL you care about against each agent. The rules for which group a crawler follows and how paths match are set out in RFC 9309.9 Three mistakes come up again and again.
The blanket disallow
User-agent: *
Disallow: /
Every compliant crawler, Googlebot and bingbot included, is told to stay out. If your file opens like this and you want to be found in search at all, fix it before thinking about AI-specific rules.
Named groups ignore the wildcard
User-agent: *
Disallow: /private/
User-agent: GPTBot
Allow: /
GPTBot finds its own group and follows that group only. The Disallow: /private/ line under * doesn't apply to it, so /private/ is open. To keep it closed, repeat the rule inside the GPTBot group:
User-agent: GPTBot
Disallow: /private/
robots.txt was never security, either. Anything private needs a login or a server-level block.
The file doesn't load properly
curl -i https://yourdomain.com/robots.txt
curl -I https://yourdomain.com/important-page/
Check the status code, any redirects and the body that comes back. A 403 or 5xx on robots.txt is a problem in its own right, and crawlers handle it in different ways. One more trap: if robots.txt blocks a page, crawlers never get to read that page's noindex tag, so they can't act on it.10
Check 2: Is Cloudflare or your web host blocking AI bots?
This is where most accidental blocks live.
Cloudflare's managed robots.txt setting, available on every plan, adds AI crawler rules to the file your visitors receive. If you already have a robots.txt, Cloudflare puts its rules in front of yours and serves the two as one file. If you don't have one, Cloudflare serves its own. Its documented example disallows GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended, Amazonbot, Bytespider and meta-externalagent.11 Cloudflare also says that on the Free plan, a domain with no robots.txt of its own and managed robots.txt switched off will show crawlers its Content Signals Policy instead.11
In practice that means WordPress or your SEO plugin can show you one robots.txt while the public gets another. Always check the live URL.
Where to look on each platform:
- Cloudflare: the served robots.txt, the managed robots.txt toggle, AI Crawl Control and Security Events. AI Crawl Control is the enforcement side and can block a named AI crawler outright, whatever robots.txt says.11
- Squarespace: Settings → Crawlers. The AI crawler option edits robots.txt for you, and you can't edit the file directly.12
- WordPress: the live file, your SEO plugin, any security plugin, and whatever the host or CDN adds on top.
- Wix, Shopify or Webflow: compare the platform's robots settings with the file actually served, then check any app firewall in front of it.
What if your host is blocking AI crawlers and robots.txt says allow?
I found this one the hard way. The robots.txt on this site allowed CCBot, and CCBot was still being turned away. The block sat at server level on HostGator, somewhere I couldn't see or edit, and it took a fair amount of back-and-forth with their support before it was lifted. It's the reason the checker on this site runs a live CCBot request next to the robots.txt read.
So if your robots.txt says allow and a live request gets a 403, ask your host directly about server-level crawler rules. Be specific about the genuine bot. A checker sends its request from its own IP address, and a host can reasonably say it blocks that IP and leave it there. What you need to know is how real Common Crawl, OpenAI or Anthropic traffic gets treated.
Check 3: Can AI crawlers actually read your page once they get in?
A 200 response is a good start. Now look at what came back with it.
Google says its generative AI features use publicly accessible, crawlable content,13 and the same holds for every other AI search product. Pages that return 200 but hand a crawler a JavaScript shell with no text in it, a cookie or login wall, a bot challenge dressed up as a normal page, or a redirect that dumps everything on the home page all fail quietly. Nothing in your analytics will flag it. Snippet controls matter here too. nosnippet keeps a page out of AI Overviews, but it strips your ordinary search snippet at the same time.10
View the source of the page, or fetch it with curl, and check the words a buyer would need are in the HTML. Test a representative article and a service page. The home page is usually the best-built page on the site, so it tells you the least.
Check 4: Why isn't your site showing in Google AI Overviews or AI Mode?
Blocking or allowing Google-Extended changes nothing here, whatever you've read.
AI Overviews and AI Mode run on Googlebot access and ordinary Search eligibility. On top of that, Search Console now has its own control under Settings → Search generative AI, with three options: include (the default), exclude, or inherit from a parent property. Excluding a site removes its links and content from AI Overviews, AI Mode and generative AI features in Discover. Google rolled the control out to all websites worldwide as of 31 August 2026. It doesn't affect AI training, and Google points to Google-Extended for that. Changes generally take a few days to apply.14 Google's optimisation guide now states plainly that a site must be set to include to be eligible at all.13
The setting exists because of the UK. On 3 June 2026 the Competition and Markets Authority imposed a conduct requirement on Google under the Digital Markets, Competition and Consumers Act, giving publishers a way to opt out of AI search features without losing their place in ordinary results.15 The requirement also covers opting out of AI fine-tuning, clearer attribution with direct links, and directory and page-level controls that haven't arrived yet.15,16 The current control works at property level.
Note: If you run subdomains or country folders as separate Search Console properties, check each of them, because an exclude on the parent flows down to any child set to inherit.
| If you want to… | Use | What happens |
|---|---|---|
| Appear in ordinary Search and AI Overviews | Googlebot access, indexable pages, snippets allowed, Include in Search Console | You're eligible. Inclusion isn't guaranteed |
| Opt out of AI Overviews without losing Google rankings | Exclude in the Search generative AI control | No links or content in the covered AI features. Ordinary rankings unaffected, per Google |
| Limit Google's use of your content for model training | Google-Extended | AI Overviews and AI Mode links stay as they are |
| Hide one passage from all snippets | data-nosnippet | Also hides it from ordinary search previews |
| Leave Google entirely | noindex on accessible pages | You lose ordinary Search too |
Once you're set to include, open the Generative AI performance report in Search Console. It shows impressions for your pages inside AI Overviews and AI Mode, which is the nearest thing Google offers to a direct answer on whether you're appearing in AI Overviews.17
Check 5: How to see if GPTBot, ClaudeBot and other AI bots are visiting your site
Server logs show the requests that reached your server and the response it sent. Don't expect Google Analytics to help here. It only counts visits that run its JavaScript tag, and crawlers don't, so GA4 will show no AI bots even when they're visiting every day. You need something that records requests on the server side, and what's available depends on your setup:
- Behind a CDN such as Cloudflare: anything the CDN stopped never reaches your server, so check its security events as well as your server logs.
- Shared hosting with cPanel: raw access logs, plus AWStats, which lists named bots in its robots and spiders section. No command line needed.
- Self-hosted WordPress: a bot-logging plugin records crawler visits on the server side, so you get a bot log even if your host doesn't provide one.
- Wix: built-in bot reports under Analytics & Reports → Reports → Marketing & SEO, including bot traffic by page and response status over time.18
- Shopify or Squarespace: you won't get bot logs, so rely on the live checks and your platform's crawler settings.
Whatever you're on, referral visits from chatgpt.com or perplexity.ai in your analytics are a useful indirect sign. They can't show whether every crawler gets in, but they do prove some of your pages are being fetched and cited.
If you do have logs, remember that user-agent strings are trivial to fake. Match any suspicious IP against the operator's published list before you draw conclusions. OpenAI and Perplexity both publish JSON IP lists from their crawler pages.3,6
# Adjust the log path and format for your server
grep -Ei 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|PerplexityBot|CCBot|bingbot' /path/to/access.log | tail -50
What the response codes mean for the page you're testing
| Response | What it usually means |
|---|---|
200 with real HTML | Access is fine. Look at indexing, snippet settings and the content. |
403 | A host, CDN, WAF or security plugin rule. |
429 | Rate limiting or bot protection. |
200 with a challenge page or empty shell | Back to Check 3. |
| Nothing logged | Check CDN logs and a longer time window. The crawler may just not have come yet. |
A missing entry isn't proof of a block. Neither is one 403 from a tool pretending to be GPTBot.
Free AI crawler checker vs other tools: which should you use?
Each tool answers a different question. A robots tester reads your stated rules. A live checker sees how your server responds to its own network. Only your logs and the operators' own tools get close to genuine bot traffic. Google makes the point itself: no third-party tool has access to its internal ranking or AI systems,13 and that includes ours.
| Tool | Good for | Can't tell you on its own |
|---|---|---|
| AI crawler checker | Free check of one URL: robots.txt by agent, plus live browser, GPTBot and CCBot responses side by side. Flags likely Cloudflare-managed rules | How OpenAI's or Common Crawl's real IPs are treated. It doesn't live-test every agent or see your logs |
| 365i AI Crawler Checker | Wider coverage: 14 agents checked against robots.txt and the live server, plus noindex and nosnippet checks | Same limit. Its requests come from its own IPs |
| Google Search Console: robots.txt report, URL Inspection, Generative AI performance report | Google's own view of robots availability, indexing, fetch problems and AI feature impressions | Anything about other AI companies |
| Bing Webmaster Tools robots.txt tester | How Bing reads a URL against your robots.txt | OpenAI or Anthropic access |
| Cloudflare AI Crawl Control | AI crawler requests and decisions at Cloudflare, if you use it | Your host's separate rules |
curl plus server and CDN logs | The exact public response and every request that arrived | Whether a request really came from the operator, without IP checks |
I built our checker as a deliberately small tool. It tests one URL at a time and live-requests three agents: an ordinary browser as the baseline, GPTBot because it's the AI crawler sites name most often, and CCBot because of the HostGator episode above. A tool that fires every known AI user agent at any URL someone types in can also be pointed at other people's servers, and I'd rather keep it narrow. If you need broader coverage, run the 365i checker next to it. A Screaming Frog or Sitebulb crawl under a chosen user agent is also useful for finding rendering and internal link problems across the whole site.
Does llms.txt help AI crawlers access your site?
No. Not for access.
Google says AI Overviews and AI Mode need no new AI text file or special markup.13 An llms.txt file can't undo a 403, an exclude setting in Search Console or a login wall. If you already keep one for an agent workflow, keep it accurate. Otherwise spend the time on access first, then on pages that answer your buyers' questions directly. The guide to getting cited by ChatGPT covers that second part.
What to do when AI bots ignore robots.txt
robots.txt is a request. Cloudflare's own documentation says compliance is voluntary and some operators ignore it.11 If you need a block that holds, set a server or WAF rule tied to a verified identity or IP list. Then test it, because an over-broad rule is exactly how sites end up blocking the crawlers they wanted.
The Perplexity episode shows why evidence matters. In August 2025 Cloudflare published an analysis attributing undeclared crawling to Perplexity on sites that had blocked its declared bots.19 Perplexity rejected Cloudflare's account.20 A user-agent string couldn't settle that argument, and it won't settle one on your server either. Compare request patterns, source IPs and response logs before blaming a named company for a strange fetch.
How to block AI training crawlers but stay visible in ChatGPT and AI search
This is the setup many businesses actually want: keep content out of model training, stay findable in AI search. A narrow robots.txt policy does most of it. Treat this as a starting point. It doesn't cover every AI operator.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Googlebot
Allow: /
User-agent: bingbot
Allow: /
Keep private paths protected at the server. Decide on Google-Extended separately, since it covers Google's training uses and leaves AI Overviews alone. And a ChatGPT-User line won't reliably stop user-initiated fetches, given OpenAI's own caveat.
Newer signals are worth knowing about. Cloudflare's Content Signals let a robots.txt state separate preferences for search, real-time AI input and model training.21 The Really Simple Licensing (RSL) standard and the IETF's AI Preferences working group are both working on clearer ways to state permissions.22,23 These describe what you'd like to happen. Before relying on one, check whether the operators you care about honour it.
The 10-minute AI crawler access check
- Open your public
/robots.txtand confirm it loads with a200. - Test one high-value URL against
Googlebot,bingbot,OAI-SearchBot,GPTBot,ClaudeBotandPerplexityBot. - Compare the public file with your CMS, host and CDN settings. Look for rules you didn't write.
- Run the AI crawler checker and compare the browser, GPTBot and CCBot responses.
- View the page source and confirm the key content is in the HTML.
- In Search Console, check Settings → Search generative AI for the property, and for any parent it inherits from.
- Use URL Inspection to confirm indexing and snippet eligibility.
- Look for
403,429and challenges in CDN events, server logs or your platform's bot reports, and verify crawler IPs where you can. - Fix the layer that failed, re-run the check on the same URL, and note the date and result.
Frequently asked questions
Why isn't my website showing up in ChatGPT?
Start with access. Check that OAI-SearchBot is allowed in robots.txt and that your host or CDN doesn't refuse it. If access is fine, the gap is usually content: ChatGPT cites pages that answer the specific question clearly, and pages that other sources already corroborate.
Does blocking GPTBot stop ChatGPT citing my site?
Not directly. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and OpenAI advises allowing it if you want to appear there. Allowing access doesn't guarantee a citation.
Can robots.txt block ChatGPT-User?
OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User visits. If you need to enforce a block, use a server or WAF rule and verify the request really came from OpenAI.
Does blocking Google-Extended remove my site from AI Overviews?
No. AI Overviews depend on Googlebot access, indexing, snippet eligibility and the Search generative AI setting in Search Console.
How do I opt out of AI Overviews without losing Google rankings?
Set Exclude under Settings → Search generative AI in Search Console. Google says the setting isn't used as a ranking signal elsewhere in Search, but you give up any links, impressions and traffic from AI Overviews, AI Mode and generative features in Discover.
Is Cloudflare blocking AI crawlers on my site?
It might be. Its managed robots.txt setting can add AI crawler disallows to your public file, and AI Crawl Control or other bot rules can block requests outright. Fetch your live robots.txt and check Security Events for the agent and URL you care about.
How do I unblock AI crawlers on Cloudflare?
Turn off managed robots.txt, or edit your own file so the crawlers you want are allowed, then set those crawlers to allow in AI Crawl Control. Re-test the live URL afterwards, because other bot protection settings can still challenge them.
What's the best free way to check AI crawler access?
Run the AI crawler checker on a high-value page to compare declared rules with live responses, then confirm anything odd in your host or CDN logs. Use Search Console for Google's indexing and AI settings.
Every check came back clean?
Then access isn't your problem. If your business still doesn't show up when buyers ask AI tools for recommendations, the AI visibility gap analysis explains how to see which competitors get named instead and why.
And if you'd rather have someone run it for you, the AI visibility report is free and starts with a short conversation about the questions your buyers actually ask.
To cite this page: Wood, D. (2026). "Is your website blocking AI? How to check and fix AI crawler access in 2026." AI Visibility Gap. https://aivisibilitygap.com/is-your-website-blocking-ai.html
Sources
- Google Search Central, "AI features and your website." developers.google.com/search/docs/appearance/ai-features
- AI Visibility (365i), "Is Your Website Blocking AI Crawlers? How to Check." ai-visibility.org.uk/blog/is-your-website-blocking-ai/
- OpenAI, "Overview of OpenAI Crawlers." developers.openai.com/api/docs/bots
- OpenAI Help Center, "Publishers and Developers FAQ." help.openai.com/en/articles/12627856-publishers-and-developers-faq
- Anthropic, "Does Anthropic crawl data from the web, and how can site owners block the crawler?" support.claude.com/en/articles/8896518
- Perplexity, "Perplexity Crawlers." docs.perplexity.ai/docs/resources/perplexity-crawlers
- Google Search Central, "Google's common crawlers" (Google-Extended). developers.google.com/search/docs/crawling-indexing/google-common-crawlers
- Common Crawl, "CCBot." commoncrawl.org/ccbot
- IETF, "RFC 9309: Robots Exclusion Protocol." datatracker.ietf.org/doc/html/rfc9309
- Google Search Central, "Robots meta tag, data-nosnippet, and X-Robots-Tag specifications." developers.google.com/search/docs/crawling-indexing/robots-meta-tag
- Cloudflare Docs, "robots.txt setting" (managed robots.txt). developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/
- Squarespace Help Center, "Request that AI models exclude your site." support.squarespace.com/hc/en-us/articles/360022347072
- Google Search Central, "Google's guide to optimizing for generative AI features on Google Search." developers.google.com/search/docs/fundamentals/ai-optimization-guide
- Google Search Console Help, "Search generative AI control." support.google.com/webmasters/answer/16908024
- Competition and Markets Authority, "CMA secures fairer deal for publishers and improves Google search services in UK," 3 June 2026. gov.uk/government/news/cma-secures-fairer-deal-for-publishers-and-improves-google-search-services-in-uk
- Bronwyn Howell, "Google Opt-Outs: Greater or Less Transparency for Consumers?" AEI, 12 June 2026. ctse.aei.org/google-opt-outs-greater-or-less-transparency-for-consumers/
- Google Search Console Help, "Generative AI performance report (Search)." support.google.com/webmasters/answer/16984139
- Wix SEO Hub, guide to Wix bot log reports. wix.com/seo/learn/post/wix-bot-log-reports
- Cloudflare Blog, "Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives," 4 August 2025. blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/
- Perplexity, "Agents or Bots? Making Sense of AI on the Open Web," 4 August 2025. perplexity.ai/hub/blog/agents-or-bots-making-sense-of-ai-on-the-open-web
- Cloudflare, Content Signals Policy. contentsignals.org
- RSL Collective, "Really Simple Licensing (RSL) standard." rslstandard.org
- IETF, AI Preferences (aipref) Working Group. datatracker.ietf.org/wg/aipref/about/