What is Bytespider?
According to ByteDance's Toutiao Search webmaster docs, Bytespider is the Toutiao Search crawler that fetches web pages to build the search index.
- Operator
- ByteDance
- Purpose
- AI search index
- Feeds
- Toutiao Search
- robots.txt token
Bytespider- Follows robots.txt
- Not stated
- IP ranges
- Published list
- Source
- ByteDance documentation
Bytespider user agent
Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36
PC, Android and iOS examples from the official Chinese-language page, last saved 2022-08-17, which may be out of date.
Match it in robots.txt with the token Bytespider. User agents can be faked, so check requests against ByteDance's published IP ranges before trusting them. Reverse DNS (host / nslookup / dig -x) should return a *.bytedance.com hostname (e.g. bytespider-...crawl.bytedance.com); ByteDance says anything else is an impostor. The IP list is inline on the doc page, with no JSON.
Should you block Bytespider?
ByteDance doesn't say whether Bytespider follows robots.txt. Translated: Toutiao's 'Robots matching' page says path matching is “consistent with Google's matching method”; it never says Bytespider complies.
Block Bytespider
User-agent: Bytespider Disallow: /
Allow Bytespider
User-agent: Bytespider Allow: /
What others report about Bytespider
These claims come from third parties, not ByteDance, and ByteDance hasn't confirmed them.
- Research cited by Fortune found Bytespider scraping at many times the rate of the LLM-training bots of Google, Meta, Amazon, OpenAI and Anthropic, and refers to 'training data scraped by Bytespider'; Fortune notes ByteDance's Doubao LLM predates that data. (fortune.com)
- The same research reported that Bytespider 'does not respect robots.txt'. (fortune.com)
- Third-party bot directories describe Bytespider as allegedly downloading training data for ByteDance LLMs, including Doubao. (knownagents.com)
Bytespider FAQ
What is Bytespider?
According to ByteDance's Toutiao Search webmaster docs, Bytespider is the Toutiao Search crawler that fetches web pages to build the search index.
What is the Bytespider user agent?
Bytespider identifies itself as: Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36. In robots.txt, match it with the token "Bytespider".
Does Bytespider respect robots.txt?
ByteDance doesn't say whether Bytespider follows robots.txt.
How do I block Bytespider?
Add "User-agent: Bytespider" followed by "Disallow: /" to the robots.txt file at the root of your site.
Crawlable isn't the same as cited
Letting AI crawlers in is step one. Citedify checks whether ChatGPT, Claude, Perplexity and Google AI Overviews actually recommend your brand, and what to fix when they don't.
Facts on this page come from ByteDance's documentation, checked 2026-09-28, except the section marked as third-party reports. Sources: zhanzhang.toutiao.com/page/outer/docs/26899, zhanzhang.toutiao.com/page/outer/docs/520, zhanzhang.toutiao.com/page/outer/docs, fortune.com/2024/10/03/bytedance-tiktok-bytespider-scraper-bot/, knownagents.com/agents/bytespider