Statistics · AI Crawler Monitor

How many German websites block AI crawlers?

A daily robots.txt measurement across Germany’s top 1000 websites: which AI bots get locked out, how often — and what the 20 biggest sites do. All numbers are free to cite.

As of 3 July 2026 · Panel: Deutschland (CrUX Top-1000), 891 sites reachable

32.2 %of the top websites lock out GPTBot (robots.txt or server-side)
14 / 20of the 20 best-known sites block GPTBot
891websites in the daily measurement panel
12AI crawlers tracked

Block rate per AI crawler

Share of reachable panel sites that effectively lock out each bot (robots.txt incl. server-side blocks where measured):

CCBot · Common Crawl · Datensatz33.4 %
GPTBot · OpenAI · Training32.2 %
Bytespider · ByteDance · Training30.4 %
ClaudeBot · Anthropic · Training28.4 %
Google-Extended · Google · Gemini-Training26.3 %
meta-externalagent · Meta · Training26.2 %
Applebot-Extended · Apple · Training25.3 %
Amazonbot · Amazon · Assistent22.4 %
anthropic-ai · Anthropic · Training (alt)20.8 %
PerplexityBot · Perplexity · Suche20.3 %
ChatGPT-User · OpenAI · On-Demand16.4 %
OAI-SearchBot · OpenAI · Suche11.9 %

Method & context

Every day the robots.txt of Germany’s top 1000 websites (CrUX panel) is measured; for 891 reachable sites the rules for 12 known AI crawlers are evaluated. A subset is additionally probed for server-side blocks (e.g. 403 for bot user agents).

“Effectively blocked” means: an explicit robots.txt rule against the bot or a verified server-side block. Sites without a rule count as not blocking — and robots.txt is a request, not a technical barrier.

Tracked since 3 July 2026. The page is recomputed from the current measurement with every site build.