AI Crawler Directory

CCBot Common Crawl (nonprofit)'s web archive crawler

CCBot is the crawler of Common Crawl, a nonprofit that publishes a free public archive of the web. The twist: that archive has become a foundational training corpus for the AI industry — GPT-3 and many other models trained heavily on Common Crawl data. Blocking CCBot is therefore the single robots.txt line with the widest AI-training reach: it opts you out of a dataset used by countless model builders at once, including ones that don't run their own crawlers.

Web archive crawlerReviewed August 2026
Quick facts
Operator
Common Crawl (nonprofit)
Feeds
The Common Crawl public web archive — a foundation of many AI training datasets
Type
Web archive crawler
robots.txt token
CCBot
Respects robots.txt
Yes
Official docs
commoncrawl.org

What is CCBot?

CCBot itself is exemplary: it honors robots.txt and crawl-delay, identifies itself clearly, and crawls at modest rates for a genuinely open archive.

The policy question is downstream use. Once your content is in a Common Crawl snapshot, anyone — AI labs included — can train on it. Blocking CCBot only affects future snapshots; existing ones already circulate.

How to identify CCBot

CCBot identifies itself with the following user-agent string:

CCBot/2.0 (https://commoncrawl.org/faq/)

Never trust the user-agent string alone. Scrapers routinely impersonate well-known crawlers to inherit their access. Common Crawl documents CCBot in its FAQ; it crawls from Amazon-hosted infrastructure and identifies itself consistently. It is a well-behaved, verifiable crawler.

Note that most AI crawlers and fetchers, CCBot included, do not execute JavaScript — so this traffic is invisible to GA4 and every script-based analytics tool. Server logs, CDN analytics, or a dedicated bot-analytics layer are the only places you will see it. For the full picture of measuring AI-driven visits, see our guide on how to track AI traffic.

Controlling CCBot with robots.txt

To refuse CCBot access to your entire site, add this to your robots.txt:

User-agent: CCBot
Disallow: /

To restrict it from specific sections only (for example, premium content) while leaving the rest open:

User-agent: CCBot
Disallow: /premium/
Disallow: /members/

Should you block or monetize CCBot?

The case for blocking: One disallow line withdraws your future content from the default training corpus of the entire AI industry — by far the highest leverage-per-line in robots.txt. The collateral cost is that Common Crawl also serves academic research and archiving.

The case for allowing or monetizing: There is nothing to charge Common Crawl itself — it's a nonprofit archive. The monetization logic is indirect: content absent from free corpora is content AI companies must come get on your terms.

Bottom line
Content businesses should block it: it is the widest unpaid training funnel that exists. If you value contributing to open research archives, that's a legitimate reason to differ — just make it a decision rather than a default.
Where Oasy fits

See exactly what CCBot does on your site — then decide what that access is worth.

Oasy detects and fingerprints 50+ AI crawlers with per-URL analytics, blocks the ones you exclude at the edge, and turns the rest into revenue — licensed RAG access and sponsored placement inside AI answers, settled weekly. Analytics scripts can't see this traffic; your server logs can, and so can we.

Join the waitlist

Frequently asked questions

Why does blocking CCBot matter more than blocking individual AI crawlers?+

Because Common Crawl's archive feeds many AI companies' training pipelines at once — including companies that run no crawler of their own. One disallow line covers all of that future collection.

Does blocking CCBot remove my existing content from AI training data?+

No. Published Common Crawl snapshots remain available, and models already trained on them keep that knowledge. Blocking only keeps future content out of future snapshots.

Is CCBot itself an AI company's crawler?+

No — Common Crawl is a nonprofit with an open archive that predates the LLM era. It became AI-relevant because its archive is free, huge, and therefore the default training corpus for much of the industry.

Stay informed on life in Europe as an expat. The Local delivers daily news, guides and essential info across 9 European countries. Sign up for the free newsletter at thelocal.com/free-newsletter