AI Crawler List: Every Bot, What It Wants, and How to Block It
An AI crawler list of all 62 user-agent strings, reconciled against operator docs and robots.txt telemetry, plus the bots and tokens no directory covers.
Published •28 min read

Cloudflare Radar's AI-bot directory holds 87 AI bots, and 62 of them send an identifiable user-agent string. This AI crawler list gives all 62: the exact string, the operator, what each one collects, and the robots.txt directive that controls it. It also covers the bots and control tokens that directory leaves out.
Here's the problem nobody writing these lists says out loud: no single source lists AI crawlers correctly, because the sources disagree with each other. Cloudflare's directory is a verified-bot registry, not a census of the web. PerplexityBot, CCBot, Bytespider, cohere-ai, Diffbot and omgili aren't in it at all. Yet Cloudflare's own traffic data, over the 28 days to August 3, 2026, puts Bytespider at 4.73% of AI-bot requests. A bot can be absent from the registry and crawling you right now.
So I built this list by reconciling four sources instead of trusting one: Cloudflare Radar's bot directory (what's verified), Radar's traffic data (what's actually requesting pages), each operator's own documentation (the only authority on robots.txt behaviour and current names), and Radar's robots.txt telemetry (what site owners actually write rules about, including rules that no longer do anything). Where those four conflict, I say so rather than picking the tidiest answer.
I spent five years on Google's Search team working on large-scale crawling and indexing, and I now run the detection systems at TechnologyChecker that scan more than 50 million domains a month. That means I sit on the other side of this request. I know what a crawler's user-agent header costs to set, how easily it changes, and why a list assembled from third-party directories goes stale within a quarter. TechnologyChecker detects WordPress on 6,049,999 live domains as of July 26, 2026, a 64.06% category share and rank #1 in our WordPress usage data. For most of those sites, a text file is the only crawl control they have.
Two related reports cover ground this one deliberately skips. If you want to know which crawlers get blocked most, that's the robots.txt blocking report. If you want operator league tables and the crawl-to-refer economics, read the bot traffic statistics report. This page is the directory.
What is an AI crawler?
An AI crawler is an automated client that fetches web pages on behalf of an AI system rather than a human reader. That covers three separate jobs, and conflating them is the single most expensive mistake in this whole topic.
A training crawler collects content in bulk to build or improve a model. An AI search crawler indexes pages so an assistant can retrieve and cite them in an answer. A user-triggered fetcher grabs one page because a person just asked a question that needs it. Same protocol, same HTTP request, three completely different consequences for you.
Cloudflare Radar's AI category splits its 87 entries the same way: 36 crawlers, 39 assistants and 12 search bots. Traffic concentrates hard at the top. Googlebot accounts for 24.66% of AI-bot requests over the 28 days to August 3, 2026, with ClaudeBot second at 15.59%.
One more distinction matters before the tables. An AI scraper is usually the same thing as an AI crawler with a different tone attached. The technical difference that actually changes your defence isn't scraper-versus-crawler, it's whether the client sends a name you can match on. Twenty-five of the 87 AI bots in Radar's directory send no distinguishing user-agent at all. They get their own section below.
The complete AI crawler list: 62 user-agent strings
Below are all 62 AI bots in Cloudflare Radar's directory that publish an identifiable user-agent string, grouped by the job each one does. Match the user-agent as a substring, not an exact string. Most operators append a version number that changes without notice, so a rule keyed to GPTBot/1.4 breaks the day OpenAI ships 1.5.
The directive column shows the User-agent: line only. Pair it with Disallow: / to block or Allow: / to permit, as in the full template further down.
Bulk crawlers (33)
These are Radar's AI_CRAWLER entries. Most collect content for model training, though the category is broader than that and includes monitoring and enrichment crawlers.
| Bot | User-agent string | Operator | What it wants | robots.txt directive |
|---|---|---|---|---|
| GPTBot | GPTBot |
OpenAI | Content that may train OpenAI's foundation models | User-agent: GPTBot |
| Claude | ClaudeBot |
Anthropic | Web content that may contribute to Claude's training | User-agent: ClaudeBot |
| Meta-ExternalAgent | meta-externalagent |
Meta | Content for AI training and direct product indexing | User-agent: meta-externalagent |
| Amazonbot | Amazonbot |
Amazon | Pages that help Alexa answer more questions | User-agent: Amazonbot |
| PetalBot | PetalBot |
Huawei | An index for Petal Search and content recommendations | User-agent: PetalBot |
| GoogleOther | GoogleOther |
One-off crawls for internal research and development | User-agent: GoogleOther |
|
| Google-CloudVertexBot | CloudVertexBot |
Site-owner-requested crawls for targeted AI training | User-agent: CloudVertexBot |
|
| Amazon Kendra | amazon-kendra- |
Amazon | Unstructured data for a customer's search index | User-agent: amazon-kendra- |
| KimiBot | KimiBot |
Moonshot AI | Content used to train Kimi's foundation models | User-agent: KimiBot |
| ICC Crawler | ICC-Crawler |
NICT | Web pages for a Japanese public research corpus | User-agent: ICC-Crawler |
| Cotoyogi | Cotoyogi |
Research Organization of Information and Systems | Content for a managed AI research corpus | User-agent: Cotoyogi |
| Cloudflare Crawler | CloudflareBrowserRenderingCrawler/1.0 |
Cloudflare | Rendered page content for a customer's crawl job | User-agent: CloudflareBrowserRenderingCrawler/1.0 |
| LINER Bot | LinerBot |
Liner | Source pages that can answer Liner users' questions | User-agent: LinerBot |
| FishBot | FishBot |
FishBot | Pages for open-source AI training | User-agent: FishBot |
| Novellum AI Crawl | Novellum |
Novellum | Site content fetched by customer-built agents | User-agent: Novellum |
| Big Sur AI | bigsur.ai |
Big Sur AI | Customer site content for AI-powered experiences | User-agent: bigsur.ai |
| Navu | NavuBot |
HivePoint | Customer and prospect sites used to train their AI | User-agent: NavuBot |
| QualifiedBot | QualifiedBot |
Qualified | Customer site content that feeds hosted chatbots | User-agent: QualifiedBot |
| atlassian-bot | atlassian-bot |
Atlassian | Third-party site data indexed for Rovo search | User-agent: atlassian-bot |
| Anchor Browser | Anchor Browser |
Anchor | Page content for AI agents using its browser | User-agent: Anchor Browser |
| Make.com | make.com |
Make.com | Data pulled from customer endpoints in automations | User-agent: make.com |
| Brandwatch | magpie-crawler |
Brandwatch | Content indexed for social media monitoring | User-agent: magpie-crawler |
| AwarioSmartBot | Awario |
Awario | New and updated web data for marketing monitoring | User-agent: Awario |
| Echobot Bot | w4mwnpbXf3MFAbxOkJRw |
Echobox | Full article text for publisher distribution automation | User-agent: w4mwnpbXf3MFAbxOkJRw |
| SemrushBot-OCOB | SemrushBot-OCOB |
Semrush | Content for Semrush's visibility tooling | User-agent: SemrushBot-OCOB |
| SemrushBotSwa | SemrushBot-SWA |
Semrush | URL accessibility checks for the SEO Writing Assistant | User-agent: SemrushBot-SWA |
| BorderxBot | BorderxBot |
Borderxlab | E-commerce product data | User-agent: BorderxBot |
| Selectika AI | SelectikaScraper |
Selectika | Fashion imagery for computer-vision enrichment | User-agent: SelectikaScraper |
| CitibotSiteCrawler | CitibotSiteCrawler |
Citibot | Public government site data for civic AI tools | User-agent: CitibotSiteCrawler |
| netEstate Imprint Crawler | netEstate NE Crawler |
netEstate | Public contact details from imprint pages | User-agent: netEstate NE Crawler |
| payroll-bot | AdpResearchBot |
ADP | Public legal and payroll documentation | User-agent: AdpResearchBot |
| WARDBot | WARDBot |
WEBSPARK | URL status codes for uptime monitoring | User-agent: WARDBot |
| YGS Group Falconer Scraper | ygs-scraper-bot |
Not published | Partner sites that granted scraping permission | User-agent: ygs-scraper-bot |
AI search crawlers (11)
These index your pages so an assistant can retrieve and cite them. Blocking one removes you from that assistant's answers.
| Bot | User-agent string | Operator | What it wants | robots.txt directive |
|---|---|---|---|---|
| Claude-SearchBot | Claude-SearchBot |
Anthropic | Pages to index for Claude's search answers | User-agent: Claude-SearchBot |
| Applebot | Applebot |
Apple | Content powering Spotlight, Siri and Safari | User-agent: Applebot |
| Amzn-SearchBot | Amzn-SearchBot |
Amazon | Pages that improve Alexa and Rufus results | User-agent: Amzn-SearchBot |
| Bravebot | Bravebot |
Brave Software | New pages to index for Brave Search | User-agent: Bravebot |
| AI Search | Cloudflare-AI-Search |
Cloudflare | A customer's own connected data | User-agent: Cloudflare-AI-Search |
| AI Search External | Cloudflare-AI-Search-External |
Cloudflare | External sites feeding a customer's AI search | User-agent: Cloudflare-AI-Search-External |
| Kernel Search | KernelSearchBot |
Kernel Technologies | Pages for AI-powered search and retrieval | User-agent: KernelSearchBot |
| ShapBot | ShapBot |
Parallel | Sites indexed for Parallel's web APIs | User-agent: ShapBot |
| Alphalens Bot | alphalens-bot |
Alphalens | Companies and their offerings, indexed for B2B search | User-agent: alphalens-bot |
| Direqt Anomura | Anomura |
Direqt | Its customers' own websites | User-agent: Anomura |
| Element451Bot | Element451Bot |
Element451 | Pages for a higher-education knowledge hub | User-agent: Element451Bot |
User-triggered fetchers (18)
These arrive because a person asked a question. Radar files them under AI_ASSISTANT. Read the next section before you block any of them, because robots.txt is not a reliable control here.
| Bot | User-agent string | Operator | What it wants | robots.txt directive |
|---|---|---|---|---|
| ChatGPT-User | ChatGPT-User |
OpenAI | A page a ChatGPT user's question needs | User-agent: ChatGPT-User |
| Claude-User | Claude-User |
Anthropic | A page needed to answer a Claude user | User-agent: Claude-User |
| Meta-ExternalFetcher | meta-externalfetcher |
Meta | Individual links opened at a user's initiative | User-agent: meta-externalfetcher |
| MistralAI-User | MistralAI-User |
Mistral AI | Pages opened on request inside Le Chat | User-agent: MistralAI-User |
| DuckAssistbot | DuckAssistBot |
DuckDuckGo | Pages for DuckDuckGo's assistant | User-agent: DuckAssistBot |
| Google-Agent | Google-Agent |
Pages agents navigate and act on for a user | User-agent: Google-Agent |
|
| Devin | Devin |
Devin AI | Pages an engineering agent needs for a task | User-agent: Devin |
| Apify Website Content Crawler | ApifyWebsiteContentCrawler |
Apify | Page content converted to feed a customer's AI app | User-agent: ApifyWebsiteContentCrawler |
| Instapaper | Instapaper |
Instant Paper | An article a user saved to read later | User-agent: Instapaper |
| Retool | Retool |
Retool | Data for a customer's internal app | User-agent: Retool |
| QAtechBot | QATechBot |
QA.tech | Pages under automated QA test | User-agent: QATechBot |
| TwinAgent | TwinAgent |
Twin | Pages in an end-to-end automated operation | User-agent: TwinAgent |
| Cledara SaaS Management Agent | CledaraBot |
Cledara | Invoices and SaaS admin pages | User-agent: CledaraBot |
| HarkBot | HarkBot |
Hark | Pages for a personal intelligence assistant | User-agent: HarkBot |
| HIFIBot | HIFIBot |
HIFI | Royalty statements for music clients | User-agent: HIFIBot |
| Chathive crawler | ChathiveCrawler |
Inteso Group | A customer's own site, to power their assistant | User-agent: ChathiveCrawler |
| EasyScan | EasyScan |
codire GmbH | Content reviewed for potential legal issues | User-agent: EasyScan |
| Nava Labs ASP | Nava |
Nava Labs | Benefits sites navigated for social workers | User-agent: Nava |
Training, AI search, or user-triggered: what does each one actually want?
Training crawlers want volume, AI search crawlers want coverage, and user-triggered fetchers want one specific page right now. Over the 28 days to August 3, 2026, Cloudflare Radar attributes 43.86% of AI-bot requests to training, 11.82% to search and 2.67% to user action, with 40.33% mixed and 1.32% undeclared. Those percentages are shares of AI-bot requests, not of your traffic.
The practical difference is what each one gives back. A training crawler takes content and returns nothing directly; the payoff, if any, arrives later as model knowledge. An AI search crawler is the closest thing to the old bargain, since it indexes you so an assistant can cite and link you. A user-triggered fetcher is a person, one step removed, and blocking it means the person gets a worse answer about you.
Anthropic's documentation is worth reading here because it draws the line cleanly across all three of its bots: ClaudeBot for training, Claude-SearchBot for search indexing, Claude-User for user questions. Anthropic states plainly that its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt," with no carve-out for user-initiated fetches.
Two operators break the pattern, and they do it in their own docs. Perplexity writes that for Perplexity-User, "since a user requested the fetch, this fetcher generally ignores robots.txt rules." OpenAI's wording for ChatGPT-User is nearly identical: "because these actions are initiated by a user, robots.txt rules may not apply." Neither statement is an accusation. Both are published design decisions, and both mean the same thing for you: for this class of agent, robots.txt is not the control surface. Network-layer rules are.
Which AI crawlers are missing from the directories?
The AI crawlers most site owners actually write rules about are missing from Cloudflare Radar's AI-bot directory. Radar's directory is a verified-bot registry, built from operators who came forward and got their identity confirmed. It was never a census of what crawls the web. Nine of the most-named tokens in robots.txt files have no entry in it, including PerplexityBot, CCBot, Diffbot and OpenAI's own OAI-SearchBot.
The gap is easy to miss and expensive to inherit. Bytespider has no directory entry, yet Radar's traffic data ranks it eighth among AI bots at 4.73% of AI-bot requests. Anyone who builds a blocklist by exporting a registry ends up with a file that omits the bot sitting in their access logs.
Here are the absent tokens, with what their operators do and don't document, and how many domains name each one in robots.txt as of August 3, 2026.
| Token | Operator | Purpose | robots.txt behaviour, per operator | Domains naming it |
|---|---|---|---|---|
CCBot |
Common Crawl | Open dataset used upstream of many models | Documented: a Disallow for CCBot is honoured |
635 |
Bytespider |
ByteDance | Not documented | No official documentation published | 552 |
PerplexityBot |
Perplexity | Search and citation, not model training | Documented as controllable via robots.txt | 442 |
OAI-SearchBot |
OpenAI | Surfaces sites in ChatGPT search | Documented as the search opt-out token | 334 |
cohere-ai |
Cohere | Not documented | No official documentation published | 239 |
Diffbot |
Diffbot | Knowledge-graph and search index, not training | Adheres by default, overridable by agreement | 208 |
omgili |
Webz.io | Legacy dataset crawler | Successor crawlers documented as honouring robots.txt | 203 |
Perplexity-User |
Perplexity | Fetches a page a user asked about | Operator states it generally ignores robots.txt | 195 |
YouBot |
You.com | Not documented | No official documentation published | 188 |
Three of those rows deserve a caveat I won't paper over. ByteDance, Cohere and You.com publish no crawler documentation I could find. The user-agent strings and behaviours circulating for their bots come from third-party directories, and in YouBot's case those directories disagree with each other on the exact string. I've listed the robots.txt tokens, because those are what site owners write and what Radar counts, and left the rest blank.
Common Crawl is the opposite case, and worth a specific note: CCBot isn't an AI company's crawler at all, but its corpus is the most common upstream training dataset, so a Disallow for CCBot is an indirect training control. The official CCBot documentation publishes the directive alongside a warning that other crawlers falsely identify as CCBot, so verify by reverse DNS before you act on a log entry.
One more limit on the registry, phrased precisely because it gets misquoted constantly: Cloudflare affirmatively verifies robots.txt compliance for just 10 of the 87 AI bots in its directory. That is verification coverage, not a compliance rate. ClaudeBot, Claude-SearchBot and GoogleOther all sit in the unverified group while their operators document compliance in writing.
Two entries on every AI crawler list that aren't crawlers
Google-Extended and Applebot-Extended appear on nearly every AI crawler list, and neither one is a crawler. Both are robots.txt control tokens: strings you write to govern how already-crawled data may be used. Neither sends an HTTP request, and neither will ever appear in your access logs.
Google is unambiguous about it. According to Google's crawler documentation, "Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity." The token governs whether crawled content may train and ground Google's generative models. Googlebot keeps crawling under its own name either way.
Apple says the same thing in fewer words. According to Apple's Applebot support page, "Applebot-Extended does not crawl webpages" and "is only used to determine how to use the data crawled by the Applebot user agent." Apple adds the reassurance that matters most: "webpages that disallow Applebot-Extended can still be included in search results."
Now the number that makes this worth a section. Google-Extended is the third most-named token in robots.txt files at 655 domains, and Applebot-Extended is eighth at 477. That's 1,132 domains writing rules for things that never send a request. Those rules aren't wasted, since they're the correct way to opt out of generative training. But blocking Google-Extended does not reduce your crawl traffic by a single byte, and it does not remove you from Google Search. If you disallowed it hoping to cut server load, you changed a usage right, not a request.
Legacy tokens are the mirror image of the same confusion. anthropic-ai sits in 272 robots.txt files and Claude-Web in 216, and neither string appears anywhere in Anthropic's current crawler documentation, which describes exactly three bots. Anthropic has published no retirement statement for either, so I'll put it no stronger than this: they are absent from the current docs. Leaving a stale Disallow for anthropic-ai costs nothing. Treating it as protection from ClaudeBot costs you the block you thought you had.
The 25 AI agents that have no user-agent at all
Twenty-five of the 87 AI bots in Radar's directory send no distinguishing user-agent string. Cloudflare identifies them by IP range and signed-agent metadata instead. For this group, robots.txt is not weak, it's structurally unreachable: a directive needs a name to match, and there is no name.
The list is mostly agents that act rather than read. Nine are regional endpoints of Amazon Bedrock AgentCore Browser (US East 1 and 2, US West 2, EU West 1, EU Central 1, and four Asia-Pacific regions), each a cloud browser that AI agents drive. Alongside them sit OpenAI's ChatGPT agent, Cloudflare Browser Run, Browserbase, Kernel, Manus Bot and AGI Agent.
Then there's a commerce cluster that will matter more every quarter: FirmlyAI Bot, Henry Shopping Agent, RyeBot, Strivve Automation, Payhawk's invoice-fetching agent and the Visually.io Shopify editor. These check out, pay, and pull invoices on a user's behalf. Three Cloudflare AI Crawl Control bots (Daric2, Daric3, Daric4) and Klaviyo's KlaviyoAIBot round out the 25.
Kernel is the detail I'd point at if you only read one line of this section. The same company runs KernelSearchBot, which announces itself in the tables above, and a browser platform that doesn't. Identifying isn't a company-level trait. It's a per-product decision, and the same vendor can go both ways.
Which raises the uncomfortable arithmetic. TechnologyChecker detects Cloudflare Turnstile on 48,488 live domains, a 3.61% category share and third in its category, and Akamai Bot Manager on 177,580 domains, both as of July 26, 2026. Set that beside 6.05 million live WordPress domains. Dedicated bot management is a rounding error against the population of sites that can edit a text file. For most of the web, the only available lever is exactly the one that a whole class of agent openly ignores.
Should you allow or block AI crawlers?
Allow the AI search crawlers, decide on training deliberately, and handle agents at the network layer. That's the framework, and it maps to the three groups in the tables above rather than to operator brands.
Allow AI search crawlers. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, Amzn-SearchBot and Bravebot exist to put you in an answer with a link. Blocking them is the AI-era equivalent of noindexing yourself. Perplexity's documentation states that PerplexityBot "is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models," which is about as direct as an operator gets.
Decide on training crawlers on your own terms. There's no universally right answer here, and anyone who tells you otherwise is selling something. Blocking GPTBot, ClaudeBot, meta-externalagent, CCBot and the rest is a content-licensing position, not a technical fix. The economics behind that choice, including how many pages each operator crawls per referral it sends back, are in our crawl-to-refer analysis. Some operators crawl thousands of pages per visitor returned.
Handle agents at the network layer. For the user-triggered class and the 25 nameless agents, write the robots.txt rule anyway as a statement of intent, then enforce with rate limits, bot detection and CAPTCHA tools, or challenges if enforcement actually matters to you. Cloudflare's AI Crawl Control is one option for that layer, and it's the product behind the Daric bots above.
Here is that whole framework on one screen, with the bots that matter most in each bucket.
If you're checking which of these are already reaching your pages, or auditing a portfolio of sites, that's the sort of question our technology detection plans are built for.
What robots.txt rules should you actually write?

Start from a selective template rather than a blanket block. The one below allows the crawlers that can cite you, blocks the bulk collectors, and sets both usage-control tokens. Per-bot directives for all 62 named crawlers are in the tables above; paste only the lines that reflect a decision you've actually made.
# AI search crawlers: allow, so assistants can cite and link you
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Applebot
Allow: /
User-agent: Amzn-SearchBot
Allow: /
User-agent: Bravebot
Allow: /
# Bulk training crawlers: block if you want content out of training sets
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: PetalBot
Disallow: /
User-agent: KimiBot
Disallow: /
User-agent: GoogleOther
Disallow: /
# Usage-control tokens: no crawl traffic, no search-visibility cost
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
# User-triggered fetchers: advisory only for some operators
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-User
Allow: /
User-agent: Perplexity-User
Allow: /
Two honest caveats about that file. An Allow: / block is technically a no-op, since anything not disallowed is already permitted, so treat those blocks as documentation of a decision rather than as a mechanism. And GPTBot is the most-named token in robots.txt at 781 domains, with ClaudeBot second at 686, which tells you what other site owners chose, not what's right for you.
For the full implementation walkthrough, including rule-ordering pitfalls, wildcard interactions and when to escalate to firewall rules, our robots.txt blocking guide covers the mechanics in depth.
Our own bot

I'd be writing a dishonest article if I catalogued everyone else's crawler and left ours out. TechnologyChecker runs a crawler too. It reads roughly 50 million domains a month to work out what each site is built with, and that data is what powers the detection numbers cited throughout this post. If our crawler reaches your site, you're entitled to the same information I've demanded of every operator above.
TCBot
User-agent string
Mozilla/5.0 (compatible; TCBot/1.0; +https://technologychecker.io/bot)
Robots.txt
User-agent token in robots.txt: TCBot
Obeys robots.txt: Yes. TCBot strictly fetches and parses robots.txt to honor site owner preferences. We hold ourselves to standard web crawling protocols, ensuring that any allow or disallow directives targeting our user-agent token are fully respected.
Obeys crawl delay: Yes. TCBot parses and honors crawl-delay directives. While we generally only crawl each domain roughly once a month—meaning the practical load on any single site is already minimal by design—we will always respect the specific pacing requested by site administrators.
Purpose
Builds the technology detection dataset behind TechnologyChecker, a technographic intelligence platform. TCBot fetches publicly available pages, HTTP headers, DNS records, and script tags to identify which technologies a site runs. It does not collect content for AI model training, and it does not resell page text. What it produces is a technology profile: this domain runs WordPress, that one runs Shopify.
To exclude your site
Name the token and disallow it, exactly as you would for any bot in the tables above:
User-agent: TCBot
Disallow: /
If you'd rather not wait for the next crawl to pick up the change, block the TCBot user-agent at your CDN, WAF or server config, or write to us through the contact page and we'll exclude your domain at our end.
Frequently asked questions
What is an AI crawler bot?
An AI crawler bot is an automated client that requests web pages for an AI system rather than a human reader. It does one of three jobs: collecting content for model training, indexing pages so an assistant can cite them, or fetching a single page because a user asked a question. Cloudflare Radar's AI directory tracks 87 of them.
Should I allow AI crawlers?
Allow the AI search crawlers and decide separately on training crawlers. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot index you so assistants can cite and link you, so blocking them removes you from those answers. Training crawlers give nothing back directly, which makes blocking them a content-licensing decision rather than a technical one.
Is ChatGPT a web crawler?
ChatGPT itself is not a crawler, but OpenAI operates four distinct bots behind it. GPTBot collects training content, OAI-SearchBot indexes pages for ChatGPT search, ChatGPT-User fetches pages when a person asks a question, and OAI-AdsBot checks landing pages submitted as ads. Each takes its own robots.txt token, and OpenAI names only GPTBot and OAI-SearchBot as the tags webmasters use to manage crawling.
What user-agent strings do major AI crawlers use?
The major training tokens are GPTBot, ClaudeBot, meta-externalagent, Amazonbot, CCBot and Bytespider. The major search tokens are OAI-SearchBot, Claude-SearchBot, PerplexityBot and Applebot. The major user-triggered tokens are ChatGPT-User, Claude-User, Perplexity-User and MistralAI-User. All 62 strings with an identifiable user-agent are in the tables above.
How do I block AI crawlers with robots.txt?
Add a User-agent: line naming the bot's token, followed by Disallow: / on the next line, then repeat per bot. robots.txt has no wildcard that reliably targets AI bots as a class, so you name them individually. The rule works only for crawlers that choose to honour it, which excludes the user-triggered fetchers whose operators document a bypass and the 25 agents with no user-agent to match.
How does Cloudflare block AI crawlers, and what is AI Crawl Control?
Cloudflare blocks AI crawlers at the network layer by matching verified bot identities and IP ranges rather than trusting the user-agent header. AI Crawl Control is its managed interface for that, letting site owners allow, block or meter individual AI bots. Because it operates in front of your origin, it reaches the IP-identified agents that robots.txt cannot address.
How do AI scrapers differ from traditional crawlers?
An AI scraper wants the text of your page as training or answer material; a traditional search crawler wants to index it so it can send you a visitor. The technical request is identical. The difference that actually changes your defence is whether the client identifies itself, since 25 of the 87 AI bots in Radar's directory send no distinguishing user-agent at all.
How can I analyse AI crawler traffic on my site?
Filter your server or CDN logs on the user-agent tokens in the tables above, then group them by purpose rather than by operator so you can see training volume separately from search and user fetches. Verify the big operators by reverse DNS or their published IP files, since user-agent headers are trivially spoofed. For network-wide context on how those volumes compare, see our AI crawl purpose data.
Which AI crawlers should SaaS companies block in 2026?
Most B2B SaaS teams land on the same shape: block the bulk training crawlers (GPTBot, ClaudeBot, CCBot, meta-externalagent, Bytespider), allow every AI search crawler, and enforce rate limits on agents that can't be named. Documentation and pricing pages are the exception worth thinking hard about, since those are exactly the pages buyers ask assistants about.
Is there a free AI web crawler available?
Yes, though it's a different question from blocking one. Common Crawl publishes its corpus free for anyone to use, and several vendors offer free tiers for content extraction, including the Apify crawler listed in the tables above. Running your own crawler puts you on the other side of this list, where the courtesy is to publish a named user-agent and honour robots.txt.
Methodology and sources
This AI crawler list reconciles four datasets, all pulled or verified in the first days of August 2026.
Cloudflare Radar bot directory. The 87 AI-category entries, their categories, operators, descriptions and user-agent patterns came from Radar's bots endpoints, fetched August 4, 2026, out of 689 bots in the directory overall. The split of 62 with a user-agent string and 25 identified by IP is derived from that pull. Verified-robots.txt counts come from the same records.
Cloudflare Radar traffic and robots.txt telemetry. Bot traffic shares by user-agent, crawl-purpose shares and content-type shares cover the 28 days from July 6 to August 3, 2026, and are expressed as shares of AI-bot requests. Domain counts for robots.txt tokens are a snapshot dated August 3, 2026, and count domains naming a token, not requests. Source: Cloudflare Radar.
Operator documentation. Every robots.txt behaviour, quote and current bot name was verified against the operator's own docs: OpenAI, Anthropic, Google, Apple, Perplexity, Common Crawl, Diffbot and Webz.io. Aggregators and third-party bot directories were excluded by design. Where an operator publishes nothing, the entry says so rather than borrowing a claim.
TechnologyChecker detection data. WordPress, Cloudflare Turnstile and Akamai Bot Manager figures are live-domain counts from TechnologyChecker's own detection platform as of July 26, 2026, drawn from the scans that cover more than 50 million domains a month.
Two limits worth stating. Radar's followsRobotsTxt flag records affirmative verification, so a false value means unverified, never non-compliant; I've used it only as a coverage measure. And user-agent strings change without announcement, which is why every string above is meant to be matched as a substring and re-checked quarterly against the operator docs linked here.
Related research:


