block-ai-crawlers-robots-txt-guide-2026.png

Should You Block AI Crawlers? robots.txt Guide for 2026

Should you block AI crawlers? The short answer is: block the ones that train on your content without attribution, and allow the ones that cite you in AI search results. Those are different bots, run by the same companies, and you can control them independently in robots.txt. The businesses getting this wrong in September 2026 fall into two camps: those who block everything and become invisible in AI search, and those who allow everything and donate their content to model training with nothing in return.

The distinction that makes a smart policy possible is this: OpenAI, Anthropic, and Google now run separate bots for training and for search. GPTBot trains ChatGPT’s model. OAI-SearchBot indexes pages for ChatGPT search citations. ClaudeBot trains Claude. Claude-SearchBot indexes for Claude’s search answers. Google-Extended feeds Gemini training. Googlebot feeds Google Search, including AI Overviews. You can block one without blocking the other.

Training crawlers now account for 67.5% of AI-driven traffic by volume (Pixis, 2026). GPTBot alone grew from 5% to 30% of all AI crawler traffic between May 2024 and May 2025. Most websites built or last audited before 2023 are blocking AI crawlers by default, often unintentionally, through aggressive Cloudflare, CDN, or security plugin configurations that treat OAI-SearchBot and PerplexityBot the same as malicious scrapers. If you haven’t checked your robots.txt and CDN settings specifically for AI bots, you’re almost certainly either over-exposed or over-blocked.

Key Takeaways

  • Training bots and search bots are now separate. You can block AI training (GPTBot, ClaudeBot, Google-Extended) while staying eligible for AI search citations (OAI-SearchBot, Claude-SearchBot, PerplexityBot).
  • The defensible 2026 default: block training crawlers, allow retrieval crawlers. This protects your content from unattributed model training while keeping you visible in AI-generated answers.
  • Most sites built before 2023 are accidentally blocking AI search bots through CDN or security configurations that treat all AI crawlers identically. Check your CDN settings, not just your robots.txt.
  • robots.txt is voluntary, not a guarantee. Compliance is opt-in. Some crawlers (Bytespider) have historically ignored it. Real-time user-directed fetches (ChatGPT-User, Perplexity-User) may bypass robots.txt as “user-directed access.”
  • Blocking OAI-SearchBot removes you from ChatGPT search entirely. OpenAI’s documentation states explicitly: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.”
  • Blocking everything is the most common mistake. Businesses that want AI visibility but block all AI bots are creating the exact invisibility they’re trying to prevent.

Why Businesses Want to Block AI Crawlers in 2026

The concern is legitimate. AI companies crawl websites to build training datasets, and the resulting AI models generate responses that compete with the original content for user attention. A business that spent years creating authoritative guides may see an AI answer that paraphrases their content without attribution, reducing traffic to the source while benefiting the AI platform.

Three categories of businesses have the strongest case for blocking training crawlers:

Publishers who monetize content directly. News organizations, research publishers, and paywalled content providers whose business model depends on users visiting their site to consume content. AI training that lets models reproduce the substance of that content without a visit or a payment directly undermines the revenue model. Many major publishers block GPTBot and ClaudeBot.

Businesses with proprietary data or IP. Original research, proprietary frameworks, unique datasets, and competitive intelligence that would lose value if incorporated into a publicly available model. Training crawlers don’t attribute; they absorb.

Any business that wants control. Even without a clear commercial harm, some owners want to decide how their content is used. Blocking training crawlers is a legitimate exercise of that control, and robots.txt makes it easy.

Businesses that should NOT block retrieval crawlers: Any business that wants to appear in AI-generated answers. If you sell to people who ask ChatGPT for recommendations, who search Perplexity for comparisons, or who see AI Overviews above Google results, blocking retrieval crawlers removes you from those answers entirely.

The nuance is that most businesses are in both camps: they want control over training, AND they want AI citation visibility. The training-vs-retrieval split is the mechanism that makes both possible simultaneously.

The Critical Distinction: Training Bots vs Search Bots

This is the single most important section of this guide. Every robots.txt decision follows from understanding which bot does what.

raining-bots-vs-retrieval-bots-ai-crawlers.png

OpenAI (ChatGPT)

BotPurposeBlock =
GPTBotCrawls content for model trainingYour content won’t train future ChatGPT models
OAI-SearchBotIndexes pages for ChatGPT searchYou won’t appear in ChatGPT search answers at all
ChatGPT-UserFetches pages in real-time when a user asks ChatGPT to read a URLMay bypass robots.txt (treated as user-directed)

OpenAI’s own documentation: “Disallowing GPTBot indicates your site’s content should not be used in training generative AI foundation models.” And: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.”

Anthropic (Claude)

BotPurposeBlock =
ClaudeBotCrawls for model trainingYour content won’t train future Claude models
Claude-SearchBotIndexes for Claude’s search answersYou won’t appear in Claude’s search citations
Claude-UserReal-time user-initiated fetchesMay bypass robots.txt (user-directed)

Anthropic documents three distinct bots (updated February 20, 2026), each independently controllable.

Google (Gemini)

BotPurposeBlock =
GooglebotCrawls for Google Search (including AI Overviews)You disappear from Google entirely — never do this
Google-ExtendedControls non-Search Google uses (Gemini training, Vertex AI)Your content won’t feed Gemini’s training pipeline

Critical: Blocking Googlebot removes you from Google Search entirely, including organic rankings and AI Overviews. Never block Googlebot. Google-Extended is the bot that controls Gemini training specifically.

Perplexity

BotPurposeBlock =
PerplexityBotIndexes for Perplexity’s search indexYou won’t appear in Perplexity answers
Perplexity-UserReal-time user-initiated fetchesOften bypasses robots.txt (documented behavior)

Other Training Bots Worth Blocking

BotCompanyPurpose
CCBotCommon CrawlOpen training dataset used by many AI companies
BytespiderByteDanceTraining data for TikTok AI; historically ignores robots.txt
meta-externalagentMetaTraining data for Meta AI
AmazonbotAmazonTraining data for Alexa AI
Applebot-ExtendedAppleTraining data for Siri/Apple Intelligence

Our technical SEO guide covers robots.txt fundamentals, and our llms.txt guide covers the companion file that provides context alongside robots.txt access rules.

How to Block AI Crawlers in robots.txt: The Recommended Configuration

robots-txt-ai-crawler-configuration-2026.png

Here is the robots.txt configuration we use on our own site and recommend for businesses that want AI search visibility while protecting content from training:

# === Standard Search Engines (ALWAYS ALLOW) ===

User-agent: Googlebot

Allow: /

User-agent: Bingbot

Allow: /

# === AI SEARCH/RETRIEVAL BOTS (ALLOW — these power AI citations) ===

User-agent: OAI-SearchBot

Allow: /

User-agent: ChatGPT-User

Allow: /

User-agent: Claude-SearchBot

Allow: /

User-agent: Claude-User

Allow: /

User-agent: PerplexityBot

Allow: /

# === AI TRAINING BOTS (BLOCK — these train models without attribution) ===

User-agent: GPTBot

Disallow: /

User-agent: ClaudeBot

Disallow: /

User-agent: Google-Extended

Disallow: /

User-agent: CCBot

Disallow: /

User-agent: Bytespider

Disallow: /

User-agent: meta-externalagent

Disallow: /

User-agent: Amazonbot

Disallow: /

User-agent: Applebot-Extended

Disallow: /

# === llms.txt (allow all AI bots to read your site guide) ===

User-agent: GPTBot

Allow: /llms.txt

User-agent: ClaudeBot

Allow: /llms.txt

# === Default ===

User-agent: *

Allow: /

# === Sitemap ===

Sitemap: https://yoursite.com/sitemap.xml

Why this configuration works:

  • Search engines (Googlebot, Bingbot) are explicitly allowed, protecting all traditional SEO.
  • AI search/retrieval bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are allowed, preserving your eligibility to be cited in ChatGPT, Claude, and Perplexity answers.
  • AI training bots (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, meta-externalagent) are blocked, preventing your content from being used in model training.
  • llms.txt is allowed even for blocked training bots, so they can still read your site guide during inference.
  • The default User-agent: * with Allow: / ensures any unlisted crawler can access your site.

The Decision Framework: Who Should Block What

Not every business should use the same configuration. Here’s the framework:

ai-crawler-blocking-decision-framework

If you want maximum AI visibility and don’t care about training: Allow everything. This is the simplest configuration and maximizes your exposure across every AI surface. Appropriate for businesses whose primary goal is AI citations and recommendations.

If you want AI visibility but not training (most businesses): Use the configuration above. Block training bots, allow retrieval bots. This is the defensible 2026 default and what we recommend for most AutiMark clients.

If you’re a publisher who monetizes content directly: Block both training AND retrieval bots, but know the trade-off: you protect your content completely but become invisible in AI-generated answers. Some publishers consider this acceptable because their revenue model depends on direct visits, not AI citations.

If you have proprietary data: Block all AI bots on pages containing proprietary information, but allow retrieval bots on your public-facing marketing and content pages. Use path-level directives:

User-agent: GPTBot

Disallow: /proprietary/

Allow: /blog/

Allow: /services/

Our AI visibility guide covers how to measure whether your crawler configuration is producing the AI citation results you want.

Common Mistakes When Blocking AI Crawlers

Blocking everything because you’re angry about AI scraping. Understandable, but if you also want AI search visibility, blanket blocking is self-defeating. Block training bots specifically; allow retrieval bots.

Not checking your CDN settings. Cloudflare’s “Managed Content” bot-management rules, Sucuri, and other security services often block AI crawlers at the CDN level, overriding your robots.txt entirely. Your robots.txt can say Allow: / for OAI-SearchBot, but if Cloudflare blocks it before it reaches your server, the Allow directive never applies. Check your CDN’s bot management settings independently.

cdn-overriding-robots-txt-ai-crawlers-warning.

Treating all AI bots identically. GPTBot and OAI-SearchBot are run by the same company but serve completely different purposes. Blocking both because “they’re both OpenAI” is like blocking Googlebot because you don’t want to appear in Google Ads. The bots are separate; treat them separately.

Forgetting that robots.txt is voluntary. Compliant bots honor it. Non-compliant bots (Bytespider has been documented ignoring it) don’t. For non-compliant crawlers, the only real defense is at the server or WAF level, not in robots.txt. Verify with reverse DNS lookups, not by trusting user-agent strings alone.

Not re-checking after platform updates. OpenAI added OAI-SearchBot after GPTBot was already established. Anthropic added Claude-SearchBot and Claude-User in 2026. If your robots.txt was written in 2024, it doesn’t account for bots that didn’t exist yet. Review quarterly.

Blocking Google-Extended and wondering why you’re invisible in Gemini. Google-Extended controls Gemini training. Blocking it is a legitimate choice, but know that it removes your content from Gemini’s training pipeline while Googlebot (which you should never block) still feeds AI Overviews through Google Search.

Our on-page SEO service includes robots.txt auditing and AI-crawler configuration as standard technical SEO work.

How to Verify Your Configuration

After updating robots.txt, verify that your directives are working:

Test robots.txt syntax. Use Google’s robots.txt Tester in Search Console or any online validator to check for syntax errors.

Check your CDN separately. Visit your site from a clean browser with your CDN’s bot-management dashboard open. Verify that the bots you want to allow aren’t being blocked at the CDN layer. In Cloudflare, check Security → Bots → Configure Super Bot Fight Mode.

Monitor server logs. Check for AI crawler visits. If OAI-SearchBot and PerplexityBot are hitting your pages, your Allow directives are working. If they’re absent, something (CDN, WAF, or misconfiguration) is blocking them.

Test AI citations. Search your target queries in ChatGPT (with web search enabled), Perplexity, and Google AI Mode. If you appear in citations, your retrieval bots are accessing your content successfully. If you don’t, check your configuration.

Our how to measure SEO success guide covers the full AI citation monitoring methodology.

The Honest Limitations of robots.txt for AI Crawlers

robots.txt is a norm, not a law. It’s a voluntary protocol. Compliant crawlers honor it. Non-compliant ones don’t. There is no legal enforcement mechanism built into robots.txt itself, though unauthorized access may have legal implications depending on jurisdiction.

User-directed fetches often bypass it. When a person asks ChatGPT to “read this URL,” ChatGPT-User fetches the page on the user’s behalf. Providers treat this as user-directed access, not crawling, so robots.txt training directives may not apply.

It doesn’t prevent content already used. If GPTBot crawled your site before you added a block, that content may already be in the training dataset. Blocking now prevents future crawling, not retroactive removal.

AI answers can use your content without crawling. AI models may have learned your content during training or found it cached in third-party datasets. robots.txt controls direct crawling, not indirect exposure.

These limitations don’t make robots.txt useless. They mean it should be part of a layered approach: robots.txt for compliant bots, CDN/WAF rules for non-compliant ones, and llms.txt to guide the bots you do allow.

Where robots.txt Fits in Your AI Strategy

robots.txt is the access layer. It sits alongside four other AI-readiness components:

robots.txt → controls which AI bots can access your content llms.txt → guides allowed bots to your most important pages Schema markup → declares entities and relationships in machine-readable format Direct-answer content → structures pages for AI extraction Entity consensus → builds consistent brand signals across platforms

Together, these layers create the controlled, structured AI presence that earns citations without giving away content for free.

Our schema markup guide covers the structured data layer. Our AEO guide and GEO guide cover the content and entity strategies. Our ChatGPT search guide and AI Overviews guide cover platform-specific optimization.

What This Costs

Updating robots.txt is free and takes 15 minutes. The real cost is getting it wrong: blocking retrieval bots accidentally means losing AI search visibility that takes months to rebuild.

Professional AI-crawler auditing, including CDN configuration review, llms.txt setup, and ongoing monitoring, is typically included in our AI SEO service and Smart SEO managed plans. Our pricing page shows what each plan covers.

Frequently Asked Questions

Should I block AI crawlers? Block training crawlers (GPTBot, ClaudeBot, Google-Extended) if you want to prevent your content from being used in model training. Allow retrieval crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) if you want to appear in AI search answers. The two are now separate bots, so you can do both simultaneously.

How do I block AI bots in robots.txt? Add a User-agent: line for each bot followed by Disallow: / to block or Allow: / to permit. Each bot needs its own directive. Block training bots (GPTBot, ClaudeBot, CCBot, Bytespider, meta-externalagent) and allow retrieval bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) in separate directives.

Will blocking AI crawlers hurt my SEO? Blocking training crawlers (GPTBot, ClaudeBot) does not affect Google organic rankings or AI Overview appearance because those are powered by Googlebot, which you should never block. Blocking retrieval crawlers (OAI-SearchBot, PerplexityBot) removes you from those platforms’ AI-generated answers, which hurts your AI search visibility but not your Google rankings directly.

What is the difference between GPTBot and OAI-SearchBot? GPTBot crawls content for training ChatGPT’s underlying model. OAI-SearchBot indexes pages for ChatGPT’s web search feature. They are separate bots with separate purposes. Blocking GPTBot stops training. Blocking OAI-SearchBot stops you from appearing in ChatGPT search answers. You can block one without affecting the other.

Does blocking AI bots stop my content being used for training? For compliant crawlers (GPTBot, ClaudeBot), yes, going forward. It does not retroactively remove content already used in training. For non-compliant crawlers (Bytespider), robots.txt alone is insufficient; use CDN-level or WAF-level blocking. User-directed fetches (ChatGPT-User) may bypass robots.txt as “user-directed access.”

Which AI crawlers should I allow? OAI-SearchBot (ChatGPT search), ChatGPT-User (real-time fetches), Claude-SearchBot (Claude search), Claude-User (Claude real-time), and PerplexityBot (Perplexity index). These are the retrieval and search bots that power AI-generated answers. Allowing them keeps your content eligible for AI citations and recommendations.

The Bottom Line

The question isn’t whether to block AI crawlers. It’s which ones to block and which to allow, because they now serve fundamentally different purposes. Training bots absorb your content into a model with no attribution. Retrieval bots cite your content in AI answers with a link back to your site. The same company runs both, and you control them independently.

The defensible 2026 default: block training crawlers, allow retrieval crawlers, verify your CDN isn’t overriding your robots.txt, and review quarterly as new bots are introduced. That configuration protects your content while keeping you visible in the fastest-growing discovery channel in digital marketing.

If you want help auditing your current crawler configuration and building the full AI-readiness stack, book a free strategy call and we’ll check your robots.txt, CDN settings, and llms.txt in 15 minutes.

Popular post

Google Business Profile Optimization: The Complete 2026 Guide

How to Do an SEO Audit in 2026: Step-by-Step Guide + Checklist 

Leave a Reply

Your email address will not be published. Required fields are marked *