Skip to content
SEO Madmanby Adam Hafez
ReportsAI search

Which AI crawlers do news sites block most? CCBot tops 75%

A robots.txt scan of 100 top UK/US news sites finds CCBot blocked by 75%, ahead of Anthropic-ai, ClaudeBot and GPTBot. Google-Extended is blocked least.

Published · 2 min read

Written byAdam Hafez
Share
Markdown
Pile of stacked newspapers

Key takeaways

  • CCBot, Common Crawl's bot, is the single most-blocked AI crawler among the 100 news sites checked, disallowed by 75% of them.
  • GPTBot is blocked by 62% of the sites studied, well behind CCBot, Anthropic-ai and ClaudeBot but still ahead of Google-Extended.
  • Google-Extended is the least-blocked training bot at 46%, and US publishers block it far more often than UK publishers, 58% versus 29%.
  • Live retrieval bots are blocked less than training bots overall, with Perplexity-User disallowed by only 17% of sites versus 67% for PerplexityBot.
  • Just 14% of the news sites studied block every AI bot the researchers checked, while 18% block none of them.

Our own report on live 403 rates tracks how often an AI crawler request is actually refused at the edge, and cites a robots.txt scan that ranks crawlers by how often they are named in a disallow rule. BuzzStream ran a separate scan in the same spirit but narrowed to one publisher category: the top 100 UK and US news sites, individually checked against 11 named AI bots rather than a broad web-wide sample. That narrower lens surfaces a per-crawler ranking and a UK-versus-US split that a general sample would average away.

Which crawlers do news sites block most?

Across the 100 sites, CCBot is blocked by 75%, the highest rate BuzzStream recorded. Anthropic-ai follows at 72%, then ClaudeBot at 69%, PerplexityBot at 67%, Claude-Web at 66%, GPTBot at 62% and Applebot-Extended at 61%. Live-retrieval bots trail behind: OAI-SearchBot is blocked by 49% of sites, Google-Extended by 46%, ChatGPT-User by 40%, and Perplexity-User by just 17%, the lowest figure in the study.

Why is Google-Extended blocked less than GPTBot?

Google-Extended’s 46% overall rate splits sharply by country: US publishers block it at 58% against 29% for UK publishers, a wider regional gap than BuzzStream found for any other crawler in the study. The publication does not explain the split, but Google-Extended controls only whether a site’s content feeds Gemini and AI Overviews training, not whether Google can still crawl and rank the page through Googlebot, which may make UK publishers more willing to leave it open.

How much do training bots and search bots differ?

BuzzStream groups the 11 crawlers into training bots and live-retrieval or user-triggered bots, and the gap between the two groups is consistent: 79% of the 100 sites block at least one training bot, against 71% for live-retrieval bots. Within retrieval bots, the split is starker still, PerplexityBot at 67% against Perplexity-User at 17%, meaning many sites disallow Perplexity’s background crawl while still letting a live, user-initiated Perplexity fetch through.

Why we care

A crawler-by-crawler ranking, not just an aggregate block rate, tells a publisher which specific bot to check first. If CCBot is blocked by three in four of the largest news sites, a smaller publisher weighing the same decision is looking at what has become close to a category norm rather than a fringe move, while Google-Extended’s 46% overall rate shows the opposite is true for it.

The UK-US gap on Google-Extended is the most actionable single number in the study. A US publisher blocking that crawler is following the regional majority at 58%; a UK publisher doing the same is in the 29% minority, and either one should know which group they are joining before setting the rule, since our earlier robots.txt research found that publishers frequently copy a competitor’s policy without checking what it actually blocks.

The evidence

Period
April 2026 snapshot
Sample
100 top news sites (50 UK, 50 US)

Method: BuzzStream identified the top 50 news sites in the UK and the top 50 in the US by SimilarWeb traffic share, for a combined sample of 100 sites after deduplication, and fetched each site's robots.txt file to check for disallow directives against 11 AI-related crawlers, grouped into training bots, live-retrieval or search bots, and user-triggered fetch bots. A site counts as blocking a given crawler if its robots.txt contains a disallow rule naming that crawler's user agent, which records the published policy rather than whether the crawler honors it or whether any request was actually refused at the server.

Share of top news sites blocking each AI crawler via robots.txt
CCBot75%
Anthropic-ai72%
ClaudeBot69%
PerplexityBot67%
GPTBot62%
Google-Extended46%

Sources

  1. 1.Which News Sites Block AI Crawlers in 2026? [New Data] - BuzzStream, April 8, 2026Primary

Frequently asked questions

Which AI crawler do the most news sites block?

CCBot, the crawler behind the Common Crawl dataset that many AI labs train on, is blocked by 75% of the 100 top UK and US news sites BuzzStream checked in April 2026, the highest rate of any crawler in the study.

Do news sites block GPTBot more than Google-Extended?

Yes. GPTBot is disallowed by 62% of the sites studied, compared with 46% for Google-Extended, which was the least-blocked training bot in BuzzStream's sample. US publishers drove most of that gap, blocking Google-Extended at 58% against 29% for UK publishers.

About the author

Adam Hafez
Adam Hafez

Founder, UpgradIQ FZC LLC

Adam Hafez works on technical SEO and search measurement: how pages get crawled, indexed, ranked and now quoted by answer engines. He founded UpgradIQ, which reads Google Search Console and GA4 to tie ranking movement back to the changes that caused it. He publishes what the data supports and states the limits of it.

  • Technical SEO
  • Search Console and GA4 measurement
  • Answer engine optimization
  • Structured data

The briefing

Only what actually changed in search, delivered in full by RSS, Atom or JSON feed.

Follow