AI Crawler Directory

Amazonbot Amazon's ai training crawler

Amazonbot crawls the web to improve Alexa's question answering and other Amazon AI services. Amazon states it respects robots.txt, though publishers have repeatedly reported it among the heaviest AI-era crawlers by request volume — worth watching in your logs even if you allow it.

AI training crawlerReviewed August 2026
Quick facts
Operator
Amazon
Feeds
Alexa question answering and Amazon AI services
Type
AI training crawler
robots.txt token
Amazonbot
Respects robots.txt
Yes

What is Amazonbot?

Amazonbot crawls broadly and persistently; multiple publishers and infrastructure providers have flagged its aggregate volume as disproportionate to any visible traffic return.

The visibility payoff is opaque: Alexa answers rarely cite or link sources in a way that returns measurable traffic, which distinguishes Amazonbot from search indexers whose citations you can at least measure.

How to identify Amazonbot

Amazonbot identifies itself with the following user-agent string:

Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)

Never trust the user-agent string alone. Scrapers routinely impersonate well-known crawlers to inherit their access. Amazon documents Amazonbot and supports reverse-DNS verification of its crawler hosts; validate suspect traffic before whitelisting the user-agent string.

Note that most AI crawlers and fetchers, Amazonbot included, do not execute JavaScript — so this traffic is invisible to GA4 and every script-based analytics tool. Server logs, CDN analytics, or a dedicated bot-analytics layer are the only places you will see it. For the full picture of measuring AI-driven visits, see our guide on how to track AI traffic.

Controlling Amazonbot with robots.txt

To refuse Amazonbot access to your entire site, add this to your robots.txt:

User-agent: Amazonbot
Disallow: /

To restrict it from specific sections only (for example, premium content) while leaving the rest open:

User-agent: Amazonbot
Disallow: /premium/
Disallow: /members/

Should you block or monetize Amazonbot?

The case for blocking: High crawl cost, negligible referral return, and training use for a commercial assistant: the ledger reads like a training crawler's, and many publishers treat it as one and block it.

The case for allowing or monetizing: As with other training-type access, the coherent path is licensed access rather than free crawling — Amazon licenses plenty of data commercially and can license yours.

Bottom line
Block by default unless Alexa-surface visibility genuinely matters to your audience. If you allow it, watch its crawl volume — it has a reputation for enthusiasm.
Where Oasy fits

See exactly what Amazonbot does on your site — then decide what that access is worth.

Oasy detects and fingerprints 50+ AI crawlers with per-URL analytics, blocks the ones you exclude at the edge, and turns the rest into revenue — licensed RAG access and sponsored placement inside AI answers, settled weekly. Analytics scripts can't see this traffic; your server logs can, and so can we.

Join the waitlist

Frequently asked questions

Does Amazonbot respect robots.txt?+

Amazon says yes, and disallow rules do stop the declared crawler. Its reputation issue is volume rather than compliance — it can crawl heavily within the rules.

What do I gain from allowing Amazonbot?+

Potential presence in Alexa's spoken answers, which are largely uncited and unlinked. For most publishers that is the weakest visibility return among the major AI crawlers.

Stay informed on life in Europe as an expat. The Local delivers daily news, guides and essential info across 9 European countries. Sign up for the free newsletter at thelocal.com/free-newsletter