
Key takeaways
- Cloudflare says 52% of crawler requests were for AI training as of June 2026, up from 22% in Spring 2025.
- Mixed-use crawlers, which blend search, agent use and training, account for over 36% of crawler activity.
- Cloudflare describes pure search crawling as a small and declining share, despite remaining critical for publisher visibility.
- More than one-third of crawler activity still comes from mixed-use bots, so purpose labels cannot fully separate training from search.
- Cloudflare defines its crawl-to-refer ratio as HTML requests from a platform's crawlers divided by HTML requests referred from that platform, normalized to one referral.
Most crawling headlines count bots. Cloudflare’s July 2026 report does something more useful for an SEO: it splits crawler requests by what the crawler is for. On that split, AI search visibility is the smaller half of the story, because most requests are not asking to send anyone back.
What share of crawler requests are for AI training?
Cloudflare writes that “52% of crawler requests are now for AI training as of June 2026, up from 22% in Spring 2025.” It also says mixed-use crawlers, those “blending search, agent use, and training”, represent over 36% of activity, and that pure search crawling is now “a small and declining share of overall crawler activity, despite remaining critical for publisher visibility.”
Read together, the two named categories cover at least 88% of identified crawler requests. Cloudflare does not publish the remaining split on the page, so this report does not estimate it.
Why do mixed-use crawlers muddy the number?
Cloudflare says more than one-third of crawler activity on its network still comes from mixed-use bots. A crawler that fetches a page for both training and answer retrieval cannot be counted as either, so the 52% training figure is a floor for training-related fetching, not a clean ceiling. For technical SEO work, that means a single robots.txt rule for such a bot blocks or allows training and retrieval together.
How does Cloudflare measure crawl-to-refer ratio?
The ratio comes from a separate Radar feature. Cloudflare divides “the total number of requests from relevant user agents associated with a given search or AI platform where the response was of Content-type: text/html” by the number of HTML requests whose Referer header names that platform, normalized to a single referral. Cloudflare notes that referrals from Google’s native apps carry no Referer header, so ratios for some platforms can read higher than reality.
What this report does not show
The 2026 post does not give a crawl-to-refer ratio, an absolute request count or a sample size for the purpose split, so none appear here. Cloudflare also reports that some of the most heavily crawled categories have seen human traffic decline as much as 40% in less than one year, but the page does not tie that figure to the purpose split, and this report does not either.
The evidence
- Period
- Spring 2025 to June 2026
- Sample
- Cloudflare network-wide crawler requests; not disclosed
Method: Cloudflare classifies crawlers by declared purpose (training, search, agent or mixed use) and compares the share of crawler requests per purpose at two points, Spring 2025 and June 2026. Its report states that the data is compiled from Cloudflare Radar and the Cloudflare Investor Day 2026 presentation, and gives no sampling detail beyond that. The crawl-to-refer ratio is defined in Cloudflare's earlier Radar post as HTML-response requests from a platform's user agents divided by HTML requests whose Referer header names that platform. Those ratios exclude referrals from Google's native apps, which send no Referer header.
Sources
- 1.Content Independence Day, one year on: building the business model for the agentic Internet - Cloudflare Blog, July 1, 2026Primary
- 2.The crawl before the fall... of referrals: understanding AI's impact on content providers - Cloudflare Blog, July 1, 2025Primary
About the author

Founder
Founder, UpgradIQ, Inc.
Adam Hafez works on technical SEO and search measurement: how pages get crawled, indexed, ranked and now quoted by answer engines. He founded UpgradIQ, which reads Google Search Console and GA4 to tie ranking movement back to the changes that caused it. He publishes what the data supports and states the limits of it.
- Technical SEO
- Search Console and GA4 measurement
- Answer engine optimization
- Structured data
Related reading
How three CTR trackers reported three different position-1 rates
Three trackers measured 2026 organic CTR by position and reported position 1 anywhere from 27.6% to 39.8%, because none used the same method.
How five llms.txt trackers reported five different adoption rates
Five independent trackers measured llms.txt adoption in 2026 and reported numbers from 7.4% to 28%, because none counted the same thing.
Who blocks AI crawlers: 60 robots.txt files, read in full
We read the robots.txt of 60 publishers and SEO vendors. News sites block AI crawlers almost universally. The SEO industry does not.
llms.txt is a vendor standard, not a publisher one
Only 14 of 60 major sites serve an llms.txt. Half of them sell SEO or hosting software. One is a news publisher. The file has an audience problem.


