AI Crawler List: Every Bot, What It Wants, and How to Block It

An AI crawler list of all 62 user-agent strings, reconciled against operator docs and robots.txt telemetry, plus the bots and tokens no directory covers.

Published 28 min read

AI Crawler List: Every Bot, What It Wants, and How to Block It
Share:

Cloudflare Radar's AI-bot directory holds 87 AI bots, and 62 of them send an identifiable user-agent string. This AI crawler list gives all 62: the exact string, the operator, what each one collects, and the robots.txt directive that controls it. It also covers the bots and control tokens that directory leaves out.

Here's the problem nobody writing these lists says out loud: no single source lists AI crawlers correctly, because the sources disagree with each other. Cloudflare's directory is a verified-bot registry, not a census of the web. PerplexityBot, CCBot, Bytespider, cohere-ai, Diffbot and omgili aren't in it at all. Yet Cloudflare's own traffic data, over the 28 days to August 3, 2026, puts Bytespider at 4.73% of AI-bot requests. A bot can be absent from the registry and crawling you right now.

So I built this list by reconciling four sources instead of trusting one: Cloudflare Radar's bot directory (what's verified), Radar's traffic data (what's actually requesting pages), each operator's own documentation (the only authority on robots.txt behaviour and current names), and Radar's robots.txt telemetry (what site owners actually write rules about, including rules that no longer do anything). Where those four conflict, I say so rather than picking the tidiest answer.

I spent five years on Google's Search team working on large-scale crawling and indexing, and I now run the detection systems at TechnologyChecker that scan more than 50 million domains a month. That means I sit on the other side of this request. I know what a crawler's user-agent header costs to set, how easily it changes, and why a list assembled from third-party directories goes stale within a quarter. TechnologyChecker detects WordPress on 6,049,999 live domains as of July 26, 2026, a 64.06% category share and rank #1 in our WordPress usage data. For most of those sites, a text file is the only crawl control they have.

Two related reports cover ground this one deliberately skips. If you want to know which crawlers get blocked most, that's the robots.txt blocking report. If you want operator league tables and the crawl-to-refer economics, read the bot traffic statistics report. This page is the directory.

The directory, as a corridor. A twelve-second illustrative loop rather than real footage: each lit alcove stands for a bot with a published user-agent, and the dark ones for the entries no registry lists. Press play to start it, nothing downloads before that.

What is an AI crawler?

An AI crawler is an automated client that fetches web pages on behalf of an AI system rather than a human reader. That covers three separate jobs, and conflating them is the single most expensive mistake in this whole topic.

A training crawler collects content in bulk to build or improve a model. An AI search crawler indexes pages so an assistant can retrieve and cite them in an answer. A user-triggered fetcher grabs one page because a person just asked a question that needs it. Same protocol, same HTTP request, three completely different consequences for you.

Cloudflare Radar's AI category splits its 87 entries the same way: 36 crawlers, 39 assistants and 12 search bots. Traffic concentrates hard at the top. Googlebot accounts for 24.66% of AI-bot requests over the 28 days to August 3, 2026, with ClaudeBot second at 15.59%.

The three types of AI crawler: 33 bulk training crawlers, 11 AI search indexers, 18 user-triggered agents
The three jobs an AI crawler can be doing. Counts are the 62 AI bots that publish a user-agent string, grouped by Cloudflare Radar's own AI categories. — Source: Cloudflare Radar bot directory, fetched August 4, 2026.
📊
By the Numbers: HTML is 72.29% of what AI bots fetch, against 6.90% JSON and 5.21% images (Cloudflare Radar, 28 days to 2026-08-03, share of AI-bot requests). These bots want your prose, not your assets.

One more distinction matters before the tables. An AI scraper is usually the same thing as an AI crawler with a different tone attached. The technical difference that actually changes your defence isn't scraper-versus-crawler, it's whether the client sends a name you can match on. Twenty-five of the 87 AI bots in Radar's directory send no distinguishing user-agent at all. They get their own section below.

The complete AI crawler list: 62 user-agent strings

Cloudflare Radar tracks 87 AI bots in 2026, but only 62 of them send an identifiable user-agent string
87 AI bots tracked, 62 that you can actually name in a rule. The remaining 25 are identified by IP range and signed metadata, so no robots.txt directive can match them. — Source: Cloudflare Radar bot directory, fetched August 4, 2026.

Below are all 62 AI bots in Cloudflare Radar's directory that publish an identifiable user-agent string, grouped by the job each one does. Match the user-agent as a substring, not an exact string. Most operators append a version number that changes without notice, so a rule keyed to GPTBot/1.4 breaks the day OpenAI ships 1.5.

The directive column shows the User-agent: line only. Pair it with Disallow: / to block or Allow: / to permit, as in the full template further down.

Bulk crawlers (33)

These are Radar's AI_CRAWLER entries. Most collect content for model training, though the category is broader than that and includes monitoring and enrichment crawlers.

Bot User-agent string Operator What it wants robots.txt directive
GPTBot GPTBot OpenAI Content that may train OpenAI's foundation models User-agent: GPTBot
Claude ClaudeBot Anthropic Web content that may contribute to Claude's training User-agent: ClaudeBot
Meta-ExternalAgent meta-externalagent Meta Content for AI training and direct product indexing User-agent: meta-externalagent
Amazonbot Amazonbot Amazon Pages that help Alexa answer more questions User-agent: Amazonbot
PetalBot PetalBot Huawei An index for Petal Search and content recommendations User-agent: PetalBot
GoogleOther GoogleOther Google One-off crawls for internal research and development User-agent: GoogleOther
Google-CloudVertexBot CloudVertexBot Google Site-owner-requested crawls for targeted AI training User-agent: CloudVertexBot
Amazon Kendra amazon-kendra- Amazon Unstructured data for a customer's search index User-agent: amazon-kendra-
KimiBot KimiBot Moonshot AI Content used to train Kimi's foundation models User-agent: KimiBot
ICC Crawler ICC-Crawler NICT Web pages for a Japanese public research corpus User-agent: ICC-Crawler
Cotoyogi Cotoyogi Research Organization of Information and Systems Content for a managed AI research corpus User-agent: Cotoyogi
Cloudflare Crawler CloudflareBrowserRenderingCrawler/1.0 Cloudflare Rendered page content for a customer's crawl job User-agent: CloudflareBrowserRenderingCrawler/1.0
LINER Bot LinerBot Liner Source pages that can answer Liner users' questions User-agent: LinerBot
FishBot FishBot FishBot Pages for open-source AI training User-agent: FishBot
Novellum AI Crawl Novellum Novellum Site content fetched by customer-built agents User-agent: Novellum
Big Sur AI bigsur.ai Big Sur AI Customer site content for AI-powered experiences User-agent: bigsur.ai
Navu NavuBot HivePoint Customer and prospect sites used to train their AI User-agent: NavuBot
QualifiedBot QualifiedBot Qualified Customer site content that feeds hosted chatbots User-agent: QualifiedBot
atlassian-bot atlassian-bot Atlassian Third-party site data indexed for Rovo search User-agent: atlassian-bot
Anchor Browser Anchor Browser Anchor Page content for AI agents using its browser User-agent: Anchor Browser
Make.com make.com Make.com Data pulled from customer endpoints in automations User-agent: make.com
Brandwatch magpie-crawler Brandwatch Content indexed for social media monitoring User-agent: magpie-crawler
AwarioSmartBot Awario Awario New and updated web data for marketing monitoring User-agent: Awario
Echobot Bot w4mwnpbXf3MFAbxOkJRw Echobox Full article text for publisher distribution automation User-agent: w4mwnpbXf3MFAbxOkJRw
SemrushBot-OCOB SemrushBot-OCOB Semrush Content for Semrush's visibility tooling User-agent: SemrushBot-OCOB
SemrushBotSwa SemrushBot-SWA Semrush URL accessibility checks for the SEO Writing Assistant User-agent: SemrushBot-SWA
BorderxBot BorderxBot Borderxlab E-commerce product data User-agent: BorderxBot
Selectika AI SelectikaScraper Selectika Fashion imagery for computer-vision enrichment User-agent: SelectikaScraper
CitibotSiteCrawler CitibotSiteCrawler Citibot Public government site data for civic AI tools User-agent: CitibotSiteCrawler
netEstate Imprint Crawler netEstate NE Crawler netEstate Public contact details from imprint pages User-agent: netEstate NE Crawler
payroll-bot AdpResearchBot ADP Public legal and payroll documentation User-agent: AdpResearchBot
WARDBot WARDBot WEBSPARK URL status codes for uptime monitoring User-agent: WARDBot
YGS Group Falconer Scraper ygs-scraper-bot Not published Partner sites that granted scraping permission User-agent: ygs-scraper-bot

AI search crawlers (11)

These index your pages so an assistant can retrieve and cite them. Blocking one removes you from that assistant's answers.

Bot User-agent string Operator What it wants robots.txt directive
Claude-SearchBot Claude-SearchBot Anthropic Pages to index for Claude's search answers User-agent: Claude-SearchBot
Applebot Applebot Apple Content powering Spotlight, Siri and Safari User-agent: Applebot
Amzn-SearchBot Amzn-SearchBot Amazon Pages that improve Alexa and Rufus results User-agent: Amzn-SearchBot
Bravebot Bravebot Brave Software New pages to index for Brave Search User-agent: Bravebot
AI Search Cloudflare-AI-Search Cloudflare A customer's own connected data User-agent: Cloudflare-AI-Search
AI Search External Cloudflare-AI-Search-External Cloudflare External sites feeding a customer's AI search User-agent: Cloudflare-AI-Search-External
Kernel Search KernelSearchBot Kernel Technologies Pages for AI-powered search and retrieval User-agent: KernelSearchBot
ShapBot ShapBot Parallel Sites indexed for Parallel's web APIs User-agent: ShapBot
Alphalens Bot alphalens-bot Alphalens Companies and their offerings, indexed for B2B search User-agent: alphalens-bot
Direqt Anomura Anomura Direqt Its customers' own websites User-agent: Anomura
Element451Bot Element451Bot Element451 Pages for a higher-education knowledge hub User-agent: Element451Bot
⚠️
Common Mistake: Writing a rule against the full user-agent string. OpenAI's documentation says the version number in GPTBot/1.4 may change, so match the substring GPTBot instead. The same applies to every versioned bot in these tables.

User-triggered fetchers (18)

These arrive because a person asked a question. Radar files them under AI_ASSISTANT. Read the next section before you block any of them, because robots.txt is not a reliable control here.

Bot User-agent string Operator What it wants robots.txt directive
ChatGPT-User ChatGPT-User OpenAI A page a ChatGPT user's question needs User-agent: ChatGPT-User
Claude-User Claude-User Anthropic A page needed to answer a Claude user User-agent: Claude-User
Meta-ExternalFetcher meta-externalfetcher Meta Individual links opened at a user's initiative User-agent: meta-externalfetcher
MistralAI-User MistralAI-User Mistral AI Pages opened on request inside Le Chat User-agent: MistralAI-User
DuckAssistbot DuckAssistBot DuckDuckGo Pages for DuckDuckGo's assistant User-agent: DuckAssistBot
Google-Agent Google-Agent Google Pages agents navigate and act on for a user User-agent: Google-Agent
Devin Devin Devin AI Pages an engineering agent needs for a task User-agent: Devin
Apify Website Content Crawler ApifyWebsiteContentCrawler Apify Page content converted to feed a customer's AI app User-agent: ApifyWebsiteContentCrawler
Instapaper Instapaper Instant Paper An article a user saved to read later User-agent: Instapaper
Retool Retool Retool Data for a customer's internal app User-agent: Retool
QAtechBot QATechBot QA.tech Pages under automated QA test User-agent: QATechBot
TwinAgent TwinAgent Twin Pages in an end-to-end automated operation User-agent: TwinAgent
Cledara SaaS Management Agent CledaraBot Cledara Invoices and SaaS admin pages User-agent: CledaraBot
HarkBot HarkBot Hark Pages for a personal intelligence assistant User-agent: HarkBot
HIFIBot HIFIBot HIFI Royalty statements for music clients User-agent: HIFIBot
Chathive crawler ChathiveCrawler Inteso Group A customer's own site, to power their assistant User-agent: ChathiveCrawler
EasyScan EasyScan codire GmbH Content reviewed for potential legal issues User-agent: EasyScan
Nava Labs ASP Nava Nava Labs Benefits sites navigated for social workers User-agent: Nava

Training, AI search, or user-triggered: what does each one actually want?

Why AI crawlers visit: training 43.9%, mixed purpose 40.3%, search 11.8% and user action 2.7% of AI-bot requests
Why AI crawlers visit, by the purpose each operator declares. Every share is a percentage of AI-bot requests, not of all web traffic. — Source: Cloudflare Radar, 28 days to August 3, 2026.

Training crawlers want volume, AI search crawlers want coverage, and user-triggered fetchers want one specific page right now. Over the 28 days to August 3, 2026, Cloudflare Radar attributes 43.86% of AI-bot requests to training, 11.82% to search and 2.67% to user action, with 40.33% mixed and 1.32% undeclared. Those percentages are shares of AI-bot requests, not of your traffic.

The practical difference is what each one gives back. A training crawler takes content and returns nothing directly; the payoff, if any, arrives later as model knowledge. An AI search crawler is the closest thing to the old bargain, since it indexes you so an assistant can cite and link you. A user-triggered fetcher is a person, one step removed, and blocking it means the person gets a worse answer about you.

Anthropic's documentation is worth reading here because it draws the line cleanly across all three of its bots: ClaudeBot for training, Claude-SearchBot for search indexing, Claude-User for user questions. Anthropic states plainly that its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt," with no carve-out for user-initiated fetches.

🔑
Key Takeaway: Blocking a training crawler does not block the citation crawler. GPTBot and OAI-SearchBot are separate tokens with separate rules, and so are ClaudeBot and Claude-SearchBot. Block the first pair, keep the second, and you keep the citations.

Two operators break the pattern, and they do it in their own docs. Perplexity writes that for Perplexity-User, "since a user requested the fetch, this fetcher generally ignores robots.txt rules." OpenAI's wording for ChatGPT-User is nearly identical: "because these actions are initiated by a user, robots.txt rules may not apply." Neither statement is an accusation. Both are published design decisions, and both mean the same thing for you: for this class of agent, robots.txt is not the control surface. Network-layer rules are.

Which AI crawlers are missing from the directories?

Nine AI bots named in robots.txt but absent from Cloudflare's directory, led by CCBot on 635 domains
Nine crawlers site owners write rules about that no directory lists. Values are counts of domains naming each token in robots.txt, not percentages and not request volumes. — Source: Cloudflare Radar robots.txt telemetry, August 3, 2026.

The AI crawlers most site owners actually write rules about are missing from Cloudflare Radar's AI-bot directory. Radar's directory is a verified-bot registry, built from operators who came forward and got their identity confirmed. It was never a census of what crawls the web. Nine of the most-named tokens in robots.txt files have no entry in it, including PerplexityBot, CCBot, Diffbot and OpenAI's own OAI-SearchBot.

The gap is easy to miss and expensive to inherit. Bytespider has no directory entry, yet Radar's traffic data ranks it eighth among AI bots at 4.73% of AI-bot requests. Anyone who builds a blocklist by exporting a registry ends up with a file that omits the bot sitting in their access logs.

Here are the absent tokens, with what their operators do and don't document, and how many domains name each one in robots.txt as of August 3, 2026.

Token Operator Purpose robots.txt behaviour, per operator Domains naming it
CCBot Common Crawl Open dataset used upstream of many models Documented: a Disallow for CCBot is honoured 635
Bytespider ByteDance Not documented No official documentation published 552
PerplexityBot Perplexity Search and citation, not model training Documented as controllable via robots.txt 442
OAI-SearchBot OpenAI Surfaces sites in ChatGPT search Documented as the search opt-out token 334
cohere-ai Cohere Not documented No official documentation published 239
Diffbot Diffbot Knowledge-graph and search index, not training Adheres by default, overridable by agreement 208
omgili Webz.io Legacy dataset crawler Successor crawlers documented as honouring robots.txt 203
Perplexity-User Perplexity Fetches a page a user asked about Operator states it generally ignores robots.txt 195
YouBot You.com Not documented No official documentation published 188
🚩
Red Flag: Any AI crawler list generated purely from a verified-bot registry will silently omit Bytespider, CCBot and PerplexityBot. Cross-check every generated blocklist against your own access logs before you trust it.

Three of those rows deserve a caveat I won't paper over. ByteDance, Cohere and You.com publish no crawler documentation I could find. The user-agent strings and behaviours circulating for their bots come from third-party directories, and in YouBot's case those directories disagree with each other on the exact string. I've listed the robots.txt tokens, because those are what site owners write and what Radar counts, and left the rest blank.

Common Crawl is the opposite case, and worth a specific note: CCBot isn't an AI company's crawler at all, but its corpus is the most common upstream training dataset, so a Disallow for CCBot is an indirect training control. The official CCBot documentation publishes the directive alongside a warning that other crawlers falsely identify as CCBot, so verify by reverse DNS before you act on a log entry.

One more limit on the registry, phrased precisely because it gets misquoted constantly: Cloudflare affirmatively verifies robots.txt compliance for just 10 of the 87 AI bots in its directory. That is verification coverage, not a compliance rate. ClaudeBot, Claude-SearchBot and GoogleOther all sit in the unverified group while their operators document compliance in writing.

Two entries on every AI crawler list that aren't crawlers

Google-Extended and Applebot-Extended are permission tokens that never crawl, blocked on 655 and 477 domains
1,132 domains writing rules for something that never sends a request. Both tokens govern how already-crawled data may be used; neither appears as a user-agent in any access log. — Source: Cloudflare Radar robots.txt telemetry, August 3, 2026, with behaviour verified against Google and Apple documentation.

Google-Extended and Applebot-Extended appear on nearly every AI crawler list, and neither one is a crawler. Both are robots.txt control tokens: strings you write to govern how already-crawled data may be used. Neither sends an HTTP request, and neither will ever appear in your access logs.

Google is unambiguous about it. According to Google's crawler documentation, "Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity." The token governs whether crawled content may train and ground Google's generative models. Googlebot keeps crawling under its own name either way.

Apple says the same thing in fewer words. According to Apple's Applebot support page, "Applebot-Extended does not crawl webpages" and "is only used to determine how to use the data crawled by the Applebot user agent." Apple adds the reassurance that matters most: "webpages that disallow Applebot-Extended can still be included in search results."

⚠️
Common Mistake: Grepping access logs for Google-Extended or Applebot-Extended and reading zero hits as "the block is working." Zero hits is what a correctly behaving crawler produces, because these tokens never send a request in the first place.

Now the number that makes this worth a section. Google-Extended is the third most-named token in robots.txt files at 655 domains, and Applebot-Extended is eighth at 477. That's 1,132 domains writing rules for things that never send a request. Those rules aren't wasted, since they're the correct way to opt out of generative training. But blocking Google-Extended does not reduce your crawl traffic by a single byte, and it does not remove you from Google Search. If you disallowed it hoping to cut server load, you changed a usage right, not a request.

Legacy tokens are the mirror image of the same confusion. anthropic-ai sits in 272 robots.txt files and Claude-Web in 216, and neither string appears anywhere in Anthropic's current crawler documentation, which describes exactly three bots. Anthropic has published no retirement statement for either, so I'll put it no stronger than this: they are absent from the current docs. Leaving a stale Disallow for anthropic-ai costs nothing. Treating it as protection from ClaudeBot costs you the block you thought you had.

The 25 AI agents that have no user-agent at all

Twenty-five of the 87 AI bots in Radar's directory send no distinguishing user-agent string. Cloudflare identifies them by IP range and signed-agent metadata instead. For this group, robots.txt is not weak, it's structurally unreachable: a directive needs a name to match, and there is no name.

Iceberg diagram: 62 AI bots send a user-agent above the waterline, 25 nameless agents sit below it
The part of the AI crawler population robots.txt cannot see. A directive needs a name to match, and 25 of the 87 AI bots Cloudflare tracks publish none. — Source: Cloudflare Radar bot directory, fetched August 4, 2026.

The list is mostly agents that act rather than read. Nine are regional endpoints of Amazon Bedrock AgentCore Browser (US East 1 and 2, US West 2, EU West 1, EU Central 1, and four Asia-Pacific regions), each a cloud browser that AI agents drive. Alongside them sit OpenAI's ChatGPT agent, Cloudflare Browser Run, Browserbase, Kernel, Manus Bot and AGI Agent.

Then there's a commerce cluster that will matter more every quarter: FirmlyAI Bot, Henry Shopping Agent, RyeBot, Strivve Automation, Payhawk's invoice-fetching agent and the Visually.io Shopify editor. These check out, pay, and pull invoices on a user's behalf. Three Cloudflare AI Crawl Control bots (Daric2, Daric3, Daric4) and Klaviyo's KlaviyoAIBot round out the 25.

Kernel is the detail I'd point at if you only read one line of this section. The same company runs KernelSearchBot, which announces itself in the tables above, and a browser platform that doesn't. Identifying isn't a company-level trait. It's a per-product decision, and the same vendor can go both ways.

🚩
Red Flag: If your entire AI-access policy lives in robots.txt, you have zero coverage of these 25 agents, including every agentic checkout bot. That gap only closes at the network layer.

Which raises the uncomfortable arithmetic. TechnologyChecker detects Cloudflare Turnstile on 48,488 live domains, a 3.61% category share and third in its category, and Akamai Bot Manager on 177,580 domains, both as of July 26, 2026. Set that beside 6.05 million live WordPress domains. Dedicated bot management is a rounding error against the population of sites that can edit a text file. For most of the web, the only available lever is exactly the one that a whole class of agent openly ignores.

Should you allow or block AI crawlers?

AI crawler traffic share: Googlebot 24.7%, ClaudeBot 15.6%, Meta-ExternalAgent 12.9%, GPTBot 9.8%, Bingbot 9.0%
Which AI bots actually generate the requests. Shares are of AI-bot requests observed by Cloudflare. Note Bytespider at 4.73% despite having no entry in the bot directory at all. — Source: Cloudflare Radar, 28 days to August 3, 2026.

Allow the AI search crawlers, decide on training deliberately, and handle agents at the network layer. That's the framework, and it maps to the three groups in the tables above rather than to operator brands.

Allow AI search crawlers. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, Amzn-SearchBot and Bravebot exist to put you in an answer with a link. Blocking them is the AI-era equivalent of noindexing yourself. Perplexity's documentation states that PerplexityBot "is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models," which is about as direct as an operator gets.

Decide on training crawlers on your own terms. There's no universally right answer here, and anyone who tells you otherwise is selling something. Blocking GPTBot, ClaudeBot, meta-externalagent, CCBot and the rest is a content-licensing position, not a technical fix. The economics behind that choice, including how many pages each operator crawls per referral it sends back, are in our crawl-to-refer analysis. Some operators crawl thousands of pages per visitor returned.

Handle agents at the network layer. For the user-triggered class and the 25 nameless agents, write the robots.txt rule anyway as a statement of intent, then enforce with rate limits, bot detection and CAPTCHA tools, or challenges if enforcement actually matters to you. Cloudflare's AI Crawl Control is one option for that layer, and it's the product behind the Daric bots above.

📌
Pro Tip: Decide by purpose, never by operator. OpenAI runs a training bot, a search bot, an ads bot and a user fetcher under four different tokens. "Block OpenAI" is not a policy, it's four separate decisions collapsed into one.

Here is that whole framework on one screen, with the bots that matter most in each bucket.

AI crawler cheat sheet: allow OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot and Amzn-SearchBot; decide on GPTBot, ClaudeBot, meta-externalagent, CCBot, Amazonbot and Bytespider; robots.txt will not reach ChatGPT-User, Perplexity-User or the 25 agents with no user-agent
The allow, decide and cannot-reach buckets on one screen. The third column is not a list of bad actors: ChatGPT-User and Perplexity-User are placed there because their operators publish a robots.txt bypass in writing, and the 25 nameless agents because no directive can match a bot that sends no name. — Source: Cloudflare Radar bot directory plus operator documentation, August 2026.

If you're checking which of these are already reaching your pages, or auditing a portfolio of sites, that's the sort of question our technology detection plans are built for.

What robots.txt rules should you actually write?

Streams of blue particles passing through a narrow lit gateway while others deflect away into darkness

Start from a selective template rather than a blanket block. The one below allows the crawlers that can cite you, blocks the bulk collectors, and sets both usage-control tokens. Per-bot directives for all 62 named crawlers are in the tables above; paste only the lines that reflect a decision you've actually made.

# AI search crawlers: allow, so assistants can cite and link you
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Applebot
Allow: /

User-agent: Amzn-SearchBot
Allow: /

User-agent: Bravebot
Allow: /

# Bulk training crawlers: block if you want content out of training sets
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: PetalBot
Disallow: /

User-agent: KimiBot
Disallow: /

User-agent: GoogleOther
Disallow: /

# Usage-control tokens: no crawl traffic, no search-visibility cost
User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# User-triggered fetchers: advisory only for some operators
User-agent: ChatGPT-User
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Perplexity-User
Allow: /

Two honest caveats about that file. An Allow: / block is technically a no-op, since anything not disallowed is already permitted, so treat those blocks as documentation of a decision rather than as a mechanism. And GPTBot is the most-named token in robots.txt at 781 domains, with ClaudeBot second at 686, which tells you what other site owners chose, not what's right for you.

📌
Pro Tip: Both OpenAI and Perplexity document a lag of up to 24 hours between a robots.txt edit and their systems reflecting it. Don't judge a new rule by the next hour of logs.

For the full implementation walkthrough, including rule-ordering pitfalls, wildcard interactions and when to escalate to firewall rules, our robots.txt blocking guide covers the mechanics in depth.

Our own bot

A single robot standing lit in a spotlight while identical figures stay unlit in the darkness behind it

I'd be writing a dishonest article if I catalogued everyone else's crawler and left ours out. TechnologyChecker runs a crawler too. It reads roughly 50 million domains a month to work out what each site is built with, and that data is what powers the detection numbers cited throughout this post. If our crawler reaches your site, you're entitled to the same information I've demanded of every operator above.

TCBot

User-agent string

Mozilla/5.0 (compatible; TCBot/1.0; +https://technologychecker.io/bot)

Robots.txt

User-agent token in robots.txt: TCBot

Obeys robots.txt: Yes. TCBot strictly fetches and parses robots.txt to honor site owner preferences. We hold ourselves to standard web crawling protocols, ensuring that any allow or disallow directives targeting our user-agent token are fully respected.

Obeys crawl delay: Yes. TCBot parses and honors crawl-delay directives. While we generally only crawl each domain roughly once a month—meaning the practical load on any single site is already minimal by design—we will always respect the specific pacing requested by site administrators.

Purpose

Builds the technology detection dataset behind TechnologyChecker, a technographic intelligence platform. TCBot fetches publicly available pages, HTTP headers, DNS records, and script tags to identify which technologies a site runs. It does not collect content for AI model training, and it does not resell page text. What it produces is a technology profile: this domain runs WordPress, that one runs Shopify.

To exclude your site

Name the token and disallow it, exactly as you would for any bot in the tables above:

User-agent: TCBot
Disallow: /

If you'd rather not wait for the next crawl to pick up the change, block the TCBot user-agent at your CDN, WAF or server config, or write to us through the contact page and we'll exclude your domain at our end.

⚠️
Common Mistake: Publishing a bot page and honouring robots.txt are two different commitments, and plenty of crawler documentation blurs them. A user-agent string only tells you who came. It says nothing about whether that operator reads your rules. Check both, including for us.

Frequently asked questions

What is an AI crawler bot?

An AI crawler bot is an automated client that requests web pages for an AI system rather than a human reader. It does one of three jobs: collecting content for model training, indexing pages so an assistant can cite them, or fetching a single page because a user asked a question. Cloudflare Radar's AI directory tracks 87 of them.

Should I allow AI crawlers?

Allow the AI search crawlers and decide separately on training crawlers. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot index you so assistants can cite and link you, so blocking them removes you from those answers. Training crawlers give nothing back directly, which makes blocking them a content-licensing decision rather than a technical one.

Is ChatGPT a web crawler?

ChatGPT itself is not a crawler, but OpenAI operates four distinct bots behind it. GPTBot collects training content, OAI-SearchBot indexes pages for ChatGPT search, ChatGPT-User fetches pages when a person asks a question, and OAI-AdsBot checks landing pages submitted as ads. Each takes its own robots.txt token, and OpenAI names only GPTBot and OAI-SearchBot as the tags webmasters use to manage crawling.

What user-agent strings do major AI crawlers use?

The major training tokens are GPTBot, ClaudeBot, meta-externalagent, Amazonbot, CCBot and Bytespider. The major search tokens are OAI-SearchBot, Claude-SearchBot, PerplexityBot and Applebot. The major user-triggered tokens are ChatGPT-User, Claude-User, Perplexity-User and MistralAI-User. All 62 strings with an identifiable user-agent are in the tables above.

How do I block AI crawlers with robots.txt?

Add a User-agent: line naming the bot's token, followed by Disallow: / on the next line, then repeat per bot. robots.txt has no wildcard that reliably targets AI bots as a class, so you name them individually. The rule works only for crawlers that choose to honour it, which excludes the user-triggered fetchers whose operators document a bypass and the 25 agents with no user-agent to match.

How does Cloudflare block AI crawlers, and what is AI Crawl Control?

Cloudflare blocks AI crawlers at the network layer by matching verified bot identities and IP ranges rather than trusting the user-agent header. AI Crawl Control is its managed interface for that, letting site owners allow, block or meter individual AI bots. Because it operates in front of your origin, it reaches the IP-identified agents that robots.txt cannot address.

How do AI scrapers differ from traditional crawlers?

An AI scraper wants the text of your page as training or answer material; a traditional search crawler wants to index it so it can send you a visitor. The technical request is identical. The difference that actually changes your defence is whether the client identifies itself, since 25 of the 87 AI bots in Radar's directory send no distinguishing user-agent at all.

How can I analyse AI crawler traffic on my site?

Filter your server or CDN logs on the user-agent tokens in the tables above, then group them by purpose rather than by operator so you can see training volume separately from search and user fetches. Verify the big operators by reverse DNS or their published IP files, since user-agent headers are trivially spoofed. For network-wide context on how those volumes compare, see our AI crawl purpose data.

Which AI crawlers should SaaS companies block in 2026?

Most B2B SaaS teams land on the same shape: block the bulk training crawlers (GPTBot, ClaudeBot, CCBot, meta-externalagent, Bytespider), allow every AI search crawler, and enforce rate limits on agents that can't be named. Documentation and pricing pages are the exception worth thinking hard about, since those are exactly the pages buyers ask assistants about.

Is there a free AI web crawler available?

Yes, though it's a different question from blocking one. Common Crawl publishes its corpus free for anyone to use, and several vendors offer free tiers for content extraction, including the Apify crawler listed in the tables above. Running your own crawler puts you on the other side of this list, where the courtesy is to publish a named user-agent and honour robots.txt.

Methodology and sources

This AI crawler list reconciles four datasets, all pulled or verified in the first days of August 2026.

Cloudflare Radar bot directory. The 87 AI-category entries, their categories, operators, descriptions and user-agent patterns came from Radar's bots endpoints, fetched August 4, 2026, out of 689 bots in the directory overall. The split of 62 with a user-agent string and 25 identified by IP is derived from that pull. Verified-robots.txt counts come from the same records.

Cloudflare Radar traffic and robots.txt telemetry. Bot traffic shares by user-agent, crawl-purpose shares and content-type shares cover the 28 days from July 6 to August 3, 2026, and are expressed as shares of AI-bot requests. Domain counts for robots.txt tokens are a snapshot dated August 3, 2026, and count domains naming a token, not requests. Source: Cloudflare Radar.

Operator documentation. Every robots.txt behaviour, quote and current bot name was verified against the operator's own docs: OpenAI, Anthropic, Google, Apple, Perplexity, Common Crawl, Diffbot and Webz.io. Aggregators and third-party bot directories were excluded by design. Where an operator publishes nothing, the entry says so rather than borrowing a claim.

TechnologyChecker detection data. WordPress, Cloudflare Turnstile and Akamai Bot Manager figures are live-domain counts from TechnologyChecker's own detection platform as of July 26, 2026, drawn from the scans that cover more than 50 million domains a month.

Two limits worth stating. Radar's followsRobotsTxt flag records affirmative verification, so a false value means unverified, never non-compliant; I've used it only as a coverage measure. And user-agent strings change without announcement, which is why every string above is meant to be matched as a substring and re-checked quarterly against the operator docs linked here.

Related research: