← All dispatches
Dispatches · #intelligence · ScrapeOps

AI Crawler Extraction Ratio Hits 1,917 Pages Per Link

August 3, 2026 · Abhishek Gupta
Infographic: AI crawler extraction ratio 2026 — Anthropic crawls 1,917 pages per referral sent back, GPTBot blocked by 633 domains, AI bot 403 rate at 9.64 percent

Anthropic's crawler read 1,917 pages in July 2026 for every single referral link it sent back to a publisher. Google's ratio, for comparison, is 4.7 pages per link. That AI crawler extraction ratio is the actual number behind a question every data team is asking this year: if the deal with the open web was "we crawl you, you get discovered," is that deal still real?

The short version

  • Anthropic's crawl-to-referral ratio hit 1,917:1 in July 2026 — nearly two thousand pages read for every one referral link sent back to the site it took from. (Cloudflare Radar, via TechnologyChecker.io)
  • GPTBot is the most-blocked AI crawler in robots.txt files: 633 domains disallow it in a July 27, 2026 snapshot of 4,223 sites, ahead of CCBot (567) and ClaudeBot (563).
  • 403 Forbidden responses served to AI bots climbed to 9.64% of requests in July 2026, up from 5.67% a year earlier, peaking at 12.92% in the final week of the month.
  • 89.4% of AI crawler activity in July 2026 was training or "mixed purpose" traffic — only 10.2% of crawls were tied to an action capable of sending a reader back to the source.
  • Bots overall made up 34.94% of measured web traffic in the 28 days to August 1, 2026, up from 30.76% a year earlier.

Why Is Anthropic's Crawl-to-Referral Ratio So High?

Anthropic's crawlers synthesize answers inside Claude instead of sending readers to a results page, so far fewer crawls end in a click-through. The ratio is actually improving — it was 3,386:1 in June 2026, so July's 1,917:1 is roughly a 43% drop in pages taken per referral sent.

That improvement doesn't erase the gap with search-style crawling. Here is how the major AI operators compared in July 2026, pages crawled for every referral returned:

OperatorRatio (crawl : referral)Trend vs. prior month
Anthropic1,917 : 1Improved from 3,386:1
Perplexity289 : 1Worsened from 207:1
OpenAI251 : 1Improved from 647:1
Google4.7 : 1Baseline

Google's ratio is roughly 400 times smaller than Anthropic's. That gap is the entire argument publishers are making when they say AI crawlers take without giving back — and it is why so many are now writing GPTBot, ClaudeBot, and CCBot directly into their robots.txt disallow rules.

The Sites Are Fighting Back

A July 27, 2026 snapshot of 4,223 robots.txt files found GPTBot named in 633 disallow rules — more than any other AI crawler, ahead of Common Crawl's CCBot at 567 and Anthropic's ClaudeBot at 563. Google-Extended (522) and ByteDance's Bytespider (518) round out the top five.

Blocking a crawler in robots.txt is a request, not a wall, so enforcement is moving to the HTTP layer. The share of requests from AI bots that get an outright 403 Forbidden rose from 5.67% in July 2025 to 9.64% in July 2026, spiking to 12.92% in the last week of July. Retail absorbs 28.71% of all AI crawler traffic — more than double any other sector — which tracks with how aggressively e-commerce sites have started rate-limiting and blocking.

What Does This Mean for Anyone Building on Web Data?

If your RAG pipeline or training set depends on crawling the open web directly, you are now competing with rising 403 rates, a growing disallow list, and sites that increasingly treat any automated request as hostile by default. Training and mixed-purpose crawling — 89.4% of all AI crawler activity in July — gets the least sympathy from site owners, because it returns nothing to the page it read.

This is the exact problem ScrapeOps exists to route around. One query into ScrapeOps returns hundreds of deduplicated, comprehension-ready sources instead of a homegrown crawler that breaks the moment a target site adds a new block rule or a fresh anti-bot check.

Is Blocking AI Crawlers Actually Working?

Blocking is reducing exposure at the margins, not reversing the trend. Bot traffic still grew from 30.76% to 34.94% of the measured web year over year, even as 403 rates to AI bots nearly doubled — sites are filtering harder, but crawler volume keeps climbing faster than enforcement can catch up.

None of this is a reason to stop building on web data — it's a reason to stop assuming raw crawling will keep working the way it did in 2024. The sites that mattered to your dataset last quarter may return a 403 this quarter, and the ones that don't are getting crawled by five different AI operators at once, each with its own idea of what "fair use" means. Track the extraction ratio, not just the block list — it tells you where the exchange still makes sense and where it's quietly become one-directional. We cover more of these shifts at dekryptlabs.com/dispatches, and the underlying research behind our own data infrastructure work lives at dekryptlabs.com/research.

Frequently Asked Questions

What is an AI crawler extraction ratio? It's the number of pages an AI crawler reads for every referral link it sends back to the site it crawled. Anthropic's ratio was 1,917:1 in July 2026, versus Google's 4.7:1 — meaning Anthropic's crawler took far more than it sent back in return traffic.

Which AI crawler is blocked by the most websites? GPTBot, operated by OpenAI, is the most-blocked AI crawler. A July 27, 2026 snapshot of 4,223 robots.txt files found 633 domains explicitly disallowing it, ahead of CCBot (567 domains) and ClaudeBot (563 domains).

Why are 403 error rates to AI bots rising? Sites are enforcing blocks at the HTTP layer instead of just robots.txt, since disallow rules are only a request. The share of AI bot requests receiving a 403 Forbidden rose from 5.67% in July 2025 to 9.64% in July 2026, per Cloudflare Radar data.

How much of the internet's traffic is bots now? Bots made up 34.94% of measured web traffic in the 28 days ending August 1, 2026, up from 30.76% in the same period a year earlier, according to Cloudflare Radar.

Abhishek Gupta is Co-Founder at Dekrypt Labs, building ScrapeOps — the data acquisition engine that turns any question into clean, deduplicated, comprehension-ready sources. dekryptlabs.com