What is Bytespider?

According to ByteDance's Toutiao Search webmaster docs, Bytespider is the Toutiao Search crawler that fetches web pages to build the search index.

Operator
ByteDance
Purpose
AI search index
Feeds
Toutiao Search
robots.txt token
Bytespider
Follows robots.txt
Not stated

Bytespider user agent

Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36

PC, Android and iOS examples from the official Chinese-language page, last saved 2022-08-17, which may be out of date.

Match it in robots.txt with the token Bytespider. User agents can be faked, so check requests against ByteDance's published IP ranges before trusting them. Reverse DNS (host / nslookup / dig -x) should return a *.bytedance.com hostname (e.g. bytespider-...crawl.bytedance.com); ByteDance says anything else is an impostor. The IP list is inline on the doc page, with no JSON.

Should you block Bytespider?

ByteDance doesn't say whether Bytespider follows robots.txt. Translated: Toutiao's 'Robots matching' page says path matching is “consistent with Google's matching method”; it never says Bytespider complies.

Block Bytespider

User-agent: Bytespider
Disallow: /

Allow Bytespider

User-agent: Bytespider
Allow: /

What others report about Bytespider

These claims come from third parties, not ByteDance, and ByteDance hasn't confirmed them.

  • Research cited by Fortune found Bytespider scraping at many times the rate of the LLM-training bots of Google, Meta, Amazon, OpenAI and Anthropic, and refers to 'training data scraped by Bytespider'; Fortune notes ByteDance's Doubao LLM predates that data. (fortune.com)
  • The same research reported that Bytespider 'does not respect robots.txt'. (fortune.com)
  • Third-party bot directories describe Bytespider as allegedly downloading training data for ByteDance LLMs, including Doubao. (knownagents.com)

Bytespider FAQ

What is Bytespider?

According to ByteDance's Toutiao Search webmaster docs, Bytespider is the Toutiao Search crawler that fetches web pages to build the search index.

What is the Bytespider user agent?

Bytespider identifies itself as: Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36. In robots.txt, match it with the token "Bytespider".

Does Bytespider respect robots.txt?

ByteDance doesn't say whether Bytespider follows robots.txt.

How do I block Bytespider?

Add "User-agent: Bytespider" followed by "Disallow: /" to the robots.txt file at the root of your site.

Crawlable isn't the same as cited

Letting AI crawlers in is step one. Citedify checks whether ChatGPT, Claude, Perplexity and Google AI Overviews actually recommend your brand, and what to fix when they don't.

Facts on this page come from ByteDance's documentation, checked 2026-09-28, except the section marked as third-party reports. Sources: zhanzhang.toutiao.com/page/outer/docs/26899, zhanzhang.toutiao.com/page/outer/docs/520, zhanzhang.toutiao.com/page/outer/docs, fortune.com/2024/10/03/bytedance-tiktok-bytespider-scraper-bot/, knownagents.com/agents/bytespider