# Which AI crawlers exist, and how should you set robots.txt? | Winin



- Canonical: https://winin.ai/en/answers/ai-crawlers-robots-txt/

- Published: 2026-09-25

- Updated: 2026-09-25



[← Content & Evidence](/en/blog/topics/content-evidence/)
Answers / Practical methods
Which AI crawlers exist, and how should you set robots.txt?
On this page
[Quick answer](#quick-answer)
[Details](#details)
[Common questions](#common-questions)
[Related facts](#related-facts)
Quick answer
AI companies' crawlers serve three purposes: collecting training data, indexing for AI search, and fetching a page when a user asks the AI to read it in a conversation. The three can be allowed or blocked separately. To be found and cited in AI answers, allow at least "search indexing" and "user-triggered fetching"; whether to allow training is a separate decision. Before blocking anything, find out which crawlers actually visit your site today.
Details
Three kinds of AI crawlers
Purpose
What it does
Common official identities (User-Agent)
Training collection
Collects public web pages for model training
GPTBot (OpenAI), ClaudeBot (Anthropic), KimiBot (Moonshot AI), CCBot (Common Crawl)
AI search indexing
Builds an index for AI search or web-search answers
OAI-SearchBot (OpenAI), Claude-SearchBot (Anthropic), PerplexityBot, Kimi-SearchBot
User-triggered fetching
Opens a link when a user asks the AI to
ChatGPT-User, Claude-User, Perplexity-User, Kimi-User
These identities follow each company's published crawler documentation. Some companies also publish official IP ranges, which can be used to verify that a request really comes from them.
Two more names are often confused:
Google-Extended and Applebot-Extended
are not separate crawlers but policy tokens written in robots.txt, declaring whether content may be used for those companies' AI training. They do not affect normal search indexing.
Bytespider
is ByteDance's crawler. A visit from it does not mean Doubao is citing your page; the two cannot be equated.
Crawlers of Chinese models
As of publication,
Kimi
publishes three crawler identities (KimiBot, Kimi-SearchBot, Kimi-User) with corresponding IP lists.
DeepSeek, Qwen
and other platforms have not published dedicated crawler identities. Their web-search answers usually rely on search engines or the platform's own retrieval, so a crawler name cannot be used to infer that "DeepSeek visited". Baidu (Baiduspider), 360, Huawei Petal and others are search engine crawlers and likewise cannot be mapped to a specific AI product.
How to write robots.txt
Decide three things first: whether to allow training, whether to allow AI search indexing, and whether to allow user-triggered fetching. An example that "allows AI search and fetching but not training":
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Google-Extended
Disallow: /
Notes:
A crawler matches only its most specific group of rules; do not write contradictory rules in two places.
robots.txt is a public declaration, not access control. Well-behaved crawlers follow it; requests faking an identity do not.
Do not rely on robots.txt to protect private pages, admin areas or report links; use login and permissions.
Does blocking affect recommendations?
It does, but consider each case separately:
Blocking search indexing and user-triggered fetching
: web-search answers find it harder to read your site, more citations come from third-party pages, and you lose control over how you are described.
Blocking training collection
: affects the "memory" of future model versions, with less effect on current web-search answers.
No setting guarantees a recommendation. robots.txt only decides "whether it can be read"; whether the content is then used depends on how clear, credible and relevant it is.
How to know whether AI crawlers have visited
Check server access logs
: filter by the User-Agent identities in the table above.
Verify IPs
: User-Agents can be faked. OpenAI, Anthropic, Perplexity and Kimi publish official IP lists; only matching requests are credible.
Separate "served successfully" from "used"
: a successful GET in the logs only shows the page was read, not that it was indexed, trained on or cited.
If you use Cloudflare or similar protection
Many CDNs and firewalls block or challenge AI crawlers by default. If you want AI search to read you, allow the corresponding search and user-fetching crawlers explicitly in your protection settings, and use logs to confirm they receive 200 responses rather than a challenge page or 403.
Common questions
Q: Should I block GPTBot?
A: GPTBot is OpenAI's training collection crawler. Blocking it does not stop ChatGPT's web search from reading your pages, which uses OAI-SearchBot and ChatGPT-User. Whether to allow training depends on your stance on licensing your content.
Q: Which AI crawlers are there in China?
A: The main one with published dedicated identities is Kimi. ByteDance's Bytespider, Baidu's Baiduspider and similar crawlers cannot be mapped to specific AI products. When other platforms have not published a dedicated identity, do not guess from names.
Q: Will blocking AI crawlers affect AI recommendations?
A: Blocking search indexing and user fetching makes it harder for web-search answers to read your site; blocking training collection mainly affects future model versions. No setting guarantees a recommendation.
Related facts
[Models Winin monitors](/en/facts/models-monitored/)
[Reach / Shortlist / Recommended](/en/facts/stages-reach-shortlist-recommended/)
Contact:
contact@winin.ai
Keep reading
[Content & Evidence](/en/blog/topics/content-evidence/)
[How to build brand fact pages: from conflicting information to verifiable facts](/en/guides/fact-pages/)
[What is AI Recommendation Intelligence?](/en/answers/ai-recommendation-intelligence/)
[What is fact governance for AI answers?](/en/answers/fact-governance-ai-answers/)
FROM READING TO ACTION
Start with website readiness. Find what you can improve.
[Check my website](/en/free-site-audit/)
