A mixed-use crawler is an AI crawler that blends several jobs — search indexing, AI-agent fetching, and model training — behind a single user-agent. Because one token does more than one thing, a site owner cannot cleanly permit the use they want (say, live retrieval for citations) without also permitting the ones they may not want (training). It is the practical reason "just block the AI bots" is rarely simple.
The term moved from concept to consequence in mid-2026. From 15 September 2026, Cloudflare's new defaults block AI training and agent crawlers on ad-bearing pages while leaving pure search crawlers allowed — and a crawler that mixes those functions gets blocked on ad pages unless its operator lets site owners separate the jobs. Purpose, not just identity, now decides access. Attributed to Cloudflare's July 2026 policy; single-operator and dated, since defaults are still moving.
Mixed use is why verified identity and clear per-purpose signalling matter: when a bot's job is ambiguous, the safe default drifts toward deny. For a brand that wants AI visibility, the takeaway is to decide access per crawler purpose, and to confirm that the crawlers feeding live answers can still reach the pages you want cited, rather than being caught in a mixed-use block meant for training bots.