What the AI crawler checker tests
robots.txt per AI bot
We read your robots.txt the way the bots do (RFC 9309): their own group before “*”, the longest rule wins. For eight bots you see whether they may read the checked page and which line decides it.
Firewall cross-check
We fetch the page once as a browser and once with the user agent of GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. If only the bot request is rejected, a firewall or CDN is blocking, even if your robots.txt allows everything.
Text without JavaScript
We count the words directly in the HTML. From 150 words the page counts as readable, under 50 as empty.
Indexing
We look for noindex and nosnippet in the “robots” meta tag and in the X-Robots-Tag HTTP header.
llms.txt
We check whether /llms.txt returns a text file with the heading “# Name” and count its links. An HTML page at this address doesn't count.
Structured data
We read all JSON-LD blocks on the page, show their types (such as Organization or LocalBusiness) and report broken JSON.
AI crawlers and their user agents: GPTBot, OAI-SearchBot, ClaudeBot & co.
Not every bot does the same job. Search bots build the index that ChatGPT search, Claude or Perplexity answer from. Training crawlers collect text used to train future models.
| Bot | Provider | Job according to the provider |
|---|---|---|
OAI-SearchBot | OpenAI | Search index for ChatGPT search. If you block it, you no longer appear in its answers. |
ChatGPT-User | OpenAI | Fetches on behalf of a user. According to OpenAI, robots.txt rules may not apply. |
GPTBot | OpenAI | Collects content for training OpenAI models. |
Claude-SearchBot | Anthropic | Search index for Claude's answers. |
ClaudeBot | Anthropic | Collects content for training Claude. |
PerplexityBot | Perplexity | Perplexity's search index, according to the provider not used for training. |
Google-Extended | Not a crawler of its own, just a robots.txt entry: controls Gemini training and grounding, not Google Search. | |
CCBot | Common Crawl | Common Crawl's crawler. The open dataset is a training source for many models. |
GPTBot in robots.txt: block AI training, allow AI search?
Many companies want to appear in AI answers but don't want their texts used for training. You can separate the two: allow search bots, block training crawlers. According to OpenAI, blocking GPTBot doesn't remove you from ChatGPT search; OAI-SearchBot handles that. And according to Google, Google-Extended has no influence on Google Search.
# Allow AI search
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
# Block AI training (optional)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /Giving training crawlers a group of their own only blocks those. All other bots keep following the “User-agent: *” group.
Why JavaScript can hide your content from AI crawlers
Many modern websites only build their content in the browser with JavaScript. For people it looks the same, for AI crawlers it doesn't: according to an analysis by Vercel (December 2024), the crawlers of OpenAI and Anthropic do fetch JavaScript files but don't execute them. Anything built only in the browser stays invisible to them.
Frequently asked questions
Should I block GPTBot?
It's a trade-off. GPTBot collects training data for OpenAI models. According to OpenAI, blocking it doesn't remove you from ChatGPT search; OAI-SearchBot handles that. Without training, though, a future model may know your brand less well.
Why does the check report a firewall block although my robots.txt allows everything?
The robots.txt is a request to the bots, a firewall is a door. Many CDNs and security plugins block AI bots by their user agent, regardless of the robots.txt. Since July 2025, Cloudflare asks new domains during setup whether to block AI crawlers. Our request doesn't come from the providers' IP addresses, so a block may also target bots that merely pretend to be GPTBot.
Does every bot follow the robots.txt?
According to their providers, the crawlers in this check do. The providers name exceptions themselves: according to OpenAI, the rules may not apply to ChatGPT-User, and according to Perplexity, Perplexity-User generally ignores the robots.txt. Both only fetch pages on behalf of a user. If you want to lock them out reliably, you need a firewall rule.
How quickly does a robots.txt change take effect?
OpenAI and Perplexity state about 24 hours in their docs. You can rerun the check right away: it reads your robots.txt fresh every time.
What is Google-Extended?
An entry for the robots.txt, not a crawler of its own. Google keeps reading your pages with Googlebot. Google-Extended only decides whether that content may be used to train Gemini and for grounding in Gemini apps and Vertex AI. According to Google, it has no influence on inclusion or ranking in Google Search.
Which page does the check test?
The address you enter. Without a path that's your homepage, with a path for example a blog post or product page. On top come the files in the root directory, i.e. robots.txt and llms.txt. If your robots.txt blocks individual areas, we show them per bot below the table.
Do you store my data?
No. We fetch your website, analyse it and send you the result. Neither the address nor the result is stored. To limit requests we briefly count in memory, without storing your IP address.
Sources
- OpenAI: Overview of OpenAI Crawlers
- Anthropic: Anthropic's crawlers
- Perplexity: Perplexity Crawlers
- Google: Common crawlers (Google-Extended)
- Google: AI features and your website
- Common Crawl: CCBot
- Vercel: The rise of the AI crawler
- RFC 9309: Robots Exclusion Protocol
As of October 2026
