AI Crawler Checker
The AI Crawler Checker tests whether ChatGPT, Claude, Perplexity, Google, Bing, Apple and Meta can reach a page on your site. It reads your robots.txt the way each crawler does, then requests the page with each crawler’s user agent to catch firewall and CDN blocks that robots.txt doesn’t show.
Free, no sign-up. We check one page and its robots.txt. Nothing about you is stored.
Three kinds of AI crawler
Not every AI crawler does the same job, and blocking the wrong one is the most common mistake we see on education websites. Only the first two groups affect whether AI answers can use and cite your pages.
Search and AI answers
- Googlebot Google. Indexes pages for Google Search, including AI Overviews and AI Mode.
- Bingbot Microsoft. Indexes pages for Bing, which also grounds Copilot answers.
- OAI-SearchBot OpenAI. Indexes pages so ChatGPT can show and link them in search answers. Not used for training.
- Claude-SearchBot Anthropic. Indexes pages to improve the search results Claude gives users.
- PerplexityBot Perplexity. Indexes pages so Perplexity can surface and link them. Not used for training.
- Applebot Apple. Indexes pages for Apple's search features.
- Meta-WebIndexer Meta. Indexes pages to improve search results in Meta AI.
Live visits for a user's question
- ChatGPT-User OpenAI. Opens a page when a ChatGPT user's question needs it. Not used for training.
- Claude-User Anthropic. Opens a page when a Claude user's question needs it.
- Perplexity-User Perplexity. Opens a page to answer a Perplexity user's question.
- Meta-ExternalFetcher Meta. Opens links at a user's request in Meta AI.
Model training
- GPTBot OpenAI. Collects pages that may be used to train OpenAI's models.
- ClaudeBot Anthropic. Collects pages that may be used to train Anthropic's models.
- Google-Extended Google. A robots.txt control only. Decides whether Google may use your pages to train Gemini.
- Applebot-Extended Apple. A robots.txt control only. Decides whether Apple may train its models on pages Applebot collected.
- meta-externalagent Meta. Collects pages for training Meta's AI models and improving its products.
- CCBot Common Crawl. Builds the open Common Crawl archive, a common source of AI training data.
- Bytespider ByteDance. ByteDance's crawler, associated with training its AI models.
What the colours mean
- Open
- robots.txt allows it and the live test reached the page.
- Check
- robots.txt allows it, but the server refused or changed the page for that user agent, or the test couldn’t confirm access.
- Blocked
- A search or answer crawler is blocked by robots.txt or kept out of the index by noindex.
- Policy
- A training crawler. Allowing or blocking it is your choice and doesn’t change whether AI answers cite you.
What this test can’t tell you
It checks access, not outcomes. A page every crawler can reach may still never be cited, because AI answers choose sources on relevance, clarity and trust.
The live test also can’t use the vendors’ real IP addresses, so confirm any firewall finding in your server or CDN logs.
To see what AI assistants actually say about your institution, and which sources they cite, that is what an audit is for.
Frequently asked questions
- If I block GPTBot, will ChatGPT stop citing my site?
- No. GPTBot collects pages for model training. ChatGPT search uses OAI-SearchBot, and live page visits use ChatGPT-User. You can block GPTBot and still be cited, as long as the other two are allowed.
- Does blocking Google-Extended remove my site from AI Overviews?
- No. Google-Extended only controls whether Google may use your pages to train Gemini. Google says it does not affect inclusion in Google Search, and AI Overviews and AI Mode are part of Search, which follows Googlebot.
- Why would my site block AI crawlers when nobody decided to?
- Usually a firewall or CDN setting. Several providers offer AI-bot blocking as a one-click or default option, and security plugins often block unfamiliar user agents. robots.txt then looks open while the server refuses the crawler. That is what the live test in this tool is for.

