Anthropic's crawler read 1,917 pages in July 2026 for every single referral link it sent back to a publisher. Google's ratio, for comparison, is 4.7 pages per link. That AI crawler extraction ratio is the actual number behind a question every data team is asking this year: if the deal with the open web was "we crawl you, you get discovered," is that deal still real?
The short version
Anthropic's crawlers synthesize answers inside Claude instead of sending readers to a results page, so far fewer crawls end in a click-through. The ratio is actually improving — it was 3,386:1 in June 2026, so July's 1,917:1 is roughly a 43% drop in pages taken per referral sent.
That improvement doesn't erase the gap with search-style crawling. Here is how the major AI operators compared in July 2026, pages crawled for every referral returned:
| Operator | Ratio (crawl : referral) | Trend vs. prior month |
|---|---|---|
| Anthropic | 1,917 : 1 | Improved from 3,386:1 |
| Perplexity | 289 : 1 | Worsened from 207:1 |
| OpenAI | 251 : 1 | Improved from 647:1 |
| 4.7 : 1 | Baseline |
Google's ratio is roughly 400 times smaller than Anthropic's. That gap is the entire argument publishers are making when they say AI crawlers take without giving back — and it is why so many are now writing GPTBot, ClaudeBot, and CCBot directly into their robots.txt disallow rules.
A July 27, 2026 snapshot of 4,223 robots.txt files found GPTBot named in 633 disallow rules — more than any other AI crawler, ahead of Common Crawl's CCBot at 567 and Anthropic's ClaudeBot at 563. Google-Extended (522) and ByteDance's Bytespider (518) round out the top five.
Blocking a crawler in robots.txt is a request, not a wall, so enforcement is moving to the HTTP layer. The share of requests from AI bots that get an outright 403 Forbidden rose from 5.67% in July 2025 to 9.64% in July 2026, spiking to 12.92% in the last week of July. Retail absorbs 28.71% of all AI crawler traffic — more than double any other sector — which tracks with how aggressively e-commerce sites have started rate-limiting and blocking.
If your RAG pipeline or training set depends on crawling the open web directly, you are now competing with rising 403 rates, a growing disallow list, and sites that increasingly treat any automated request as hostile by default. Training and mixed-purpose crawling — 89.4% of all AI crawler activity in July — gets the least sympathy from site owners, because it returns nothing to the page it read.
This is the exact problem ScrapeOps exists to route around. One query into ScrapeOps returns hundreds of deduplicated, comprehension-ready sources instead of a homegrown crawler that breaks the moment a target site adds a new block rule or a fresh anti-bot check.
Blocking is reducing exposure at the margins, not reversing the trend. Bot traffic still grew from 30.76% to 34.94% of the measured web year over year, even as 403 rates to AI bots nearly doubled — sites are filtering harder, but crawler volume keeps climbing faster than enforcement can catch up.
None of this is a reason to stop building on web data — it's a reason to stop assuming raw crawling will keep working the way it did in 2024. The sites that mattered to your dataset last quarter may return a 403 this quarter, and the ones that don't are getting crawled by five different AI operators at once, each with its own idea of what "fair use" means. Track the extraction ratio, not just the block list — it tells you where the exchange still makes sense and where it's quietly become one-directional. We cover more of these shifts at dekryptlabs.com/dispatches, and the underlying research behind our own data infrastructure work lives at dekryptlabs.com/research.
What is an AI crawler extraction ratio? It's the number of pages an AI crawler reads for every referral link it sends back to the site it crawled. Anthropic's ratio was 1,917:1 in July 2026, versus Google's 4.7:1 — meaning Anthropic's crawler took far more than it sent back in return traffic.
Which AI crawler is blocked by the most websites? GPTBot, operated by OpenAI, is the most-blocked AI crawler. A July 27, 2026 snapshot of 4,223 robots.txt files found 633 domains explicitly disallowing it, ahead of CCBot (567 domains) and ClaudeBot (563 domains).
Why are 403 error rates to AI bots rising? Sites are enforcing blocks at the HTTP layer instead of just robots.txt, since disallow rules are only a request. The share of AI bot requests receiving a 403 Forbidden rose from 5.67% in July 2025 to 9.64% in July 2026, per Cloudflare Radar data.
How much of the internet's traffic is bots now? Bots made up 34.94% of measured web traffic in the 28 days ending August 1, 2026, up from 30.76% in the same period a year earlier, according to Cloudflare Radar.
Abhishek Gupta is Co-Founder at Dekrypt Labs, building ScrapeOps — the data acquisition engine that turns any question into clean, deduplicated, comprehension-ready sources. dekryptlabs.com