GEO / AI search

AI crawler

Also known as: LLM crawler, AI bot, GPTBot, ClaudeBot, Google-Extended, PerplexityBot

AI crawlers are specialised bots that AI providers operate in order to crawl websites for two purposes: training future LLM models and live retrieval for answer generation. Unlike classic search engine bots, each provider often has several user agents with different purposes — anyone who wants to maintain visibility in AI answers has to allow the live retrieval bots, but can block the training bots if required.

The most important AI crawlers

Strategy: which to allow, which to block

Three pragmatic recommendations: (1) Full visibility — allow all AI crawlers. Maximum GEO leverage, and at the same time consent to training. (2) Live yes, training no — allow live crawlers (OAI-SearchBot, Claude-SearchBot, Perplexity) and block training crawlers (GPTBot, ClaudeBot, Google-Extended). This preserves GEO visibility without agreeing to the training use. (3) Complete block — all AI crawlers out. This protects content from any AI use, but the domain is invisible in AI answers — the equivalent of ”does not exist” in the multi-engine future.

A concrete robots.txt example (variant 2)

# Block training crawlers
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

# Allow live crawlers (no entry = allowed)
# OAI-SearchBot, Claude-User, Perplexity are not listed = free access

Example from practice

Example: A magazine blocked all AI crawlers completely in 2023 — strategy: ”We do not hand over our content for AI training.” The consequence in 2025: in 47 of 50 test queries in ChatGPT Search and Perplexity the magazine is not cited, even though its classic Google rankings are good. Competitors with similar content who allow live crawlers win a growing share of the AI referral traffic. Correction in early 2026: variant 2 (training blocked, live crawlers free) — visibility in AI answers recovers after 8 weeks to 35 of 50 queries.

Frequently asked questions

What is an AI crawler?
An AI crawler is a specialised bot from AI providers (OpenAI, Anthropic, Google, Perplexity) that visits websites for two purposes: training future LLM models and live retrieval for current answers. Examples: GPTBot (OpenAI training), OAI-SearchBot (ChatGPT live), ClaudeBot (Anthropic training), PerplexityBot (Perplexity index).
Should you block AI crawlers?
It depends on the objective. Anyone who wants to be cited in AI answers has to allow live crawlers — otherwise the domain does not exist for ChatGPT, Claude and Perplexity. Training crawlers can be blocked separately if content should not feed into future models. A complete block costs GEO visibility entirely.
How do you block an AI crawler?
Via robots.txt:
User-agent: GPTBot
Disallow: /

For live crawlers, add Claude-User, OAI-SearchBot and PerplexityBot. Some providers additionally observe meta robots tags such as noai or noimageai.
Do AI crawlers obey robots.txt?
The established providers do — OpenAI, Anthropic, Google and Perplexity document their user agents and respect robots.txt. Exceptions: individual scrapers and aggregators ignore the rules. For critical data, server-side bot recognition with user-agent checking and rate limiting is advisable in addition.
Which AI crawlers exist in 2026?
More than 20 known user agents. The most important: GPTBot, OAI-SearchBot, ChatGPT-User (OpenAI), ClaudeBot, Claude-User, Claude-SearchBot (Anthropic), Google-Extended (Gemini), PerplexityBot, Perplexity-User (Perplexity), Bytespider (TikTok/Doubao), CCBot (Common Crawl), Meta-ExternalAgent (Meta), Amazonbot (Alexa). New crawlers are added every few months.

Used in Rankmio for

AI bot configuration in the robots.txt audit

Go to the feature →

Last updated: 2026-06-17  ·  Browse all glossary entries

Free SEO & GEO Check

SEO score, AI visibility and citability of your website in 30 seconds — no registration required.

Check for free now

Ready to optimize your website?

Register for free, get 10 credits and start right away.

Register now