---
title: "Which AI crawlers do news sites block most? CCBot tops 75%"
url: https://seomadman.com/reports/news-site-ai-crawler-blocking-by-bot-2026
section: reports
published: 2026-09-20T00:00:00.000Z
modified: 2026-09-20T00:00:00.000Z
author: Adam Hafez
topics: ["AI search"]
---

# Which AI crawlers do news sites block most? CCBot tops 75%

## The short answer

A robots.txt scan of the top 50 UK and top 50 US news sites finds CCBot blocked by 75%, the highest of any crawler checked, ahead of Anthropic-ai at 72%, ClaudeBot at 69% and GPTBot at 62%. Google-Extended is blocked least among training bots, at 46%. Overall, 79% of sites block at least one AI training bot.

## Key takeaways

- CCBot, Common Crawl's bot, is the single most-blocked AI crawler among the 100 news sites checked, disallowed by 75% of them.
- GPTBot is blocked by 62% of the sites studied, well behind CCBot, Anthropic-ai and ClaudeBot but still ahead of Google-Extended.
- Google-Extended is the least-blocked training bot at 46%, and US publishers block it far more often than UK publishers, 58% versus 29%.
- Live retrieval bots are blocked less than training bots overall, with Perplexity-User disallowed by only 17% of sites versus 67% for PerplexityBot.
- Just 14% of the news sites studied block every AI bot the researchers checked, while 18% block none of them.

Our own [report on live 403 rates](/reports/ai-crawler-blocking-rate-trend-2026) tracks how often
an AI crawler request is actually refused at the edge, and cites a robots.txt scan that ranks
crawlers by how often they are named in a disallow rule. BuzzStream ran a separate scan in the
same spirit but narrowed to one publisher category: the top 100 UK and US news sites, individually
checked against 11 named AI bots rather than a broad web-wide sample. That narrower lens surfaces a
per-crawler ranking and a UK-versus-US split that a general sample would average away.

## Which crawlers do news sites block most?

Across the 100 sites, CCBot is blocked by 75%, the highest rate BuzzStream recorded. Anthropic-ai
follows at 72%, then ClaudeBot at 69%, PerplexityBot at 67%, Claude-Web at 66%, GPTBot at 62% and
Applebot-Extended at 61%. Live-retrieval bots trail behind: OAI-SearchBot is blocked by 49% of
sites, Google-Extended by 46%, ChatGPT-User by 40%, and Perplexity-User by just 17%, the lowest
figure in the study.

## Why is Google-Extended blocked less than GPTBot?

Google-Extended's 46% overall rate splits sharply by country: US publishers block it at 58%
against 29% for UK publishers, a wider regional gap than BuzzStream found for any other crawler in
the study. The publication does not explain the split, but Google-Extended controls only whether a
site's content feeds Gemini and AI Overviews training, not whether Google can still crawl and rank
the page through Googlebot, which may make UK publishers more willing to leave it open.

## How much do training bots and search bots differ?

BuzzStream groups the 11 crawlers into training bots and live-retrieval or user-triggered bots, and
the gap between the two groups is consistent: 79% of the 100 sites block at least one training bot,
against 71% for live-retrieval bots. Within retrieval bots, the split is starker still, PerplexityBot
at 67% against Perplexity-User at 17%, meaning many sites disallow Perplexity's background crawl
while still letting a live, user-initiated Perplexity fetch through.

## Why we care

**A crawler-by-crawler ranking, not just an aggregate block rate, tells a publisher which specific
bot to check first.** If CCBot is blocked by three in four of the largest news sites, a smaller
publisher weighing the same decision is looking at what has become close to a category norm rather
than a fringe move, while Google-Extended's 46% overall rate shows the opposite is true for it.

**The UK-US gap on Google-Extended is the most actionable single number in the study.** A US
publisher blocking that crawler is following the regional majority at 58%; a UK publisher doing the
same is in the 29% minority, and either one should know which group they are joining before setting
the rule, since [our earlier robots.txt research](/research/ai-crawler-policies-news-sites) found
that publishers frequently copy a competitor's policy without checking what it actually blocks.

## Frequently asked questions

### Which AI crawler do the most news sites block?

CCBot, the crawler behind the Common Crawl dataset that many AI labs train on, is blocked by 75% of the 100 top UK and US news sites BuzzStream checked in April 2026, the highest rate of any crawler in the study.

### Do news sites block GPTBot more than Google-Extended?

Yes. GPTBot is disallowed by 62% of the sites studied, compared with 46% for Google-Extended, which was the least-blocked training bot in BuzzStream's sample. US publishers drove most of that gap, blocking Google-Extended at 58% against 29% for UK publishers.

## Sources

1. [Which News Sites Block AI Crawlers in 2026? [New Data]](https://www.buzzstream.com/blog/publishers-block-ai-study/) - BuzzStream (primary)