Deep dive · Research · September 2026

Web Scraping Statistics: State of it in 2026

Traffic composition, tooling defaults, defender behaviour, and the definition of "the market" all shifted between 2024 and 2026. This report puts the disagreements side by side without smoothing.

Compiled September 2026 · 2,471 stats cross-tabulated by hand · sources cited at foot

A pool of 2,471 stats and numbers was cross-checked and cross-tabulated by hand for this report, pulled from public sources (Cloudflare Radar, DataDome, F5 Labs, Akamai, npm and PyPI download registries, BuiltWith, testdino, and independent site owners). Where two sources disagree, both numbers are shown. The disagreement is often the story. Where a number is a personal observation from my own scraping work (2–5M requests/month across coupons, review analysis, and AI-visibility monitoring), it is marked as such. First-person benchmarks I did not actually run have been stripped out. First-person observations from other people (site owners publishing their firewall logs, for example) are attributed to them by name.

The through-line: the web scraping surface changed shape between 2024 and 2026. Traffic composition, tooling defaults, defender behaviour, and the definition of "the market" all shifted in ways the top-line numbers hide. This report is an attempt to put them side by side without smoothing.

1. AI bots dominate web traffic now

1.1 The 33.8% → 47.57% Cloudflare drift (not a Cloudflare-vs-Akamai fight)

Same source, different windows
33.8%
Annual view
Cloudflare Radar
34.30%
Trailing 28 days to Aug 1, 2026
Cloudflare Radar
46.8%
June 2026 bucket
Cloudflare Radar
47.57%
Most recent read
Cloudflare Radar

The claim you'll see everywhere is that "AI bots are half of web traffic." The real story is that Cloudflare Radar itself reports four different values depending on the window: 33.8% on the annual view, 34.30% on the trailing 28 days to August 1 2026, 46.8% in the June bucket, and 47.57% on the most recent read. The 34.30% figure decomposes cleanly (AI crawlers 19.31%, AI assistants 8.23%, AI search 6.76%), which is what tells you it's the same denominator, just a different time slice.

Akamai does report a related number, but it's not comparable: more than 45% of AI bot traffic on Akamai's network targets commerce customers specifically. That's a subset denominator, not a whole-network share. Anyone quoting "Cloudflare says 33.8%, Akamai says 46.8%" is combining two different measurements into a fake contradiction. The real methodology gap is between short-window and long-window Cloudflare reads, and the short-window number keeps climbing.

44.21%
of AI crawling collects data for model training, not search indexing.Cloudflare Radar, 28-day window to August 1, 2026

The Cloudflare split for AI crawler intent, over the trailing 28 days to August 1 2026: 44.21% model training, 40.18% mixed-purpose, 11.66% live search indexing. Two years ago the majority was indexing; now the majority is training. The implication is structural: the web is not just being read anymore, it's being harvested for weights.

Model training44.21%
Mixed-purpose40.18%
Live search indexing11.66%

AI crawler intent split. Cloudflare Radar, 28-day window to Aug 1, 2026.

1.3 AI crawler traffic is #2 busiest bot, behind GoogleBot

GoogleBot12.87%
Claude-User (Anthropic)7.57%
Meta-ExternalAgent6.88%
BingBot5.12%
Cloudflare-HealthCheck3.91%

Top 5 verified bots by request share. Cloudflare Radar, latest 28-day window.

GoogleBot still leads at 12.87% of verified-bot requests. Anthropic's Claude-User is now #2 at 7.57%, ahead of Meta-ExternalAgent (6.88%), BingBot (5.12%), and Cloudflare's own health-check pings (3.91%). Anthropic's share was 10.53% on the June full-quarter view and dropped to 7.57% on the latest 28-day read, but even the lower number is above BingBot and every AI search competitor, and Cloudflare notes explicitly in its own methodology footer that full-quarter windows tend to sit below rolling-week peaks. Both readings are correct for their window; neither is a spike.

1.4 AI-agent traffic surged +45% in a quarter: Meta drove most of it

DataDome's Q2 2026 threat-research report gives the cleanest split. Total AI-agent traffic across DataDome's network grew from 12.2 billion requests in Q1 2026 to 17.7 billion in Q2, a +45% quarter-over-quarter jump. Meta was the largest driver by far: Meta-ExternalAgent grew +74% QoQ (3.1B → 5.3B requests), and Meta-WebIndexer grew +163% QoQ (1.4B → 3.75B). In June 2026, Meta-WebIndexer's monthly volume exceeded Meta-ExternalAgent's for the first time. Together, Meta's two agents now generate the majority of AI-agent traffic on DataDome's network: a dominance that didn't exist in Q1.

Agent
Q1 2026
Q2 2026
QoQ change
Meta-ExternalAgent
3.1B
5.3B
+74%
Meta-WebIndexer
1.4B
3.75B
+163%
All AI agents (total)
12.2B
17.7B
+45%

AI-agent request volumes, DataDome network, Q1 vs Q2 2026.

ChatGPT-User, the top single agent in Q1, went the other way: down 6% QoQ in absolute request volume, though ChatGPT still commands 80–88% of AI-driven referral traffic every month. Volume and referral value are moving in opposite directions.

2. The web scraping market size is a statistical minefield

2.1 The $7.48B vs $2.7B market-size contradiction

Market size estimates: different definitions, different answers
$7.48B
AI-driven web scraping, 2025
Market Research Future
$2.7B
Web scraping software, 2035 projection
Research Nester

Market Research Future's "AI-driven web scraping market" is $7.48B in 2025. Research Nester's "web scraping software market" is projected at $2.7B by 2035 (13.2% CAGR). The two aren't measuring the same thing. Market Research Future counts AI-driven extraction as a core segment; Research Nester carves out software specifically and excludes services, proxies, and AI agents. Pull those out of Market Research Future's number and you land in Research Nester's neighbourhood.

2.2 AI-driven scraping is growing at 39.4% CAGR

39.4%
CAGR for AI-driven web scraping, 2024–2029 forecast.groupbwt.com/blog/ai-driven-web-scraping-market/

Research and Markets projects the AI-driven web-scraping sub-market to add USD 3.15 billion between 2024 and 2029, growing at 39.4% CAGR. This is the segment that explains most of the top-line inflation: the reports quoting $7B+ include it, the ones under $3B do not. The 39.4% figure is based on 2024 data, so treat it as a ceiling, not a forecast: growth like that flattens as the segment matures.

2.3 The 8× market-size disagreement across sources

Source / claim
Market size
Year
Future Market Insights
$501.9M
2025
Mordor Intelligence
$1.03B
2024
Bright Data (2025 estimate)
$1.34B
2025
Mordor Intelligence (projection)
$2B
2030
Research Nester (projection)
$2.7B
2035
Research Nester (projection)
$3.5B
2032
Market Research Future
$7.48B
2025
Bright Data (projection, no methodology)
$47B
2035

Twelve claims across five research houses. The 100× range is the story.

Twelve claims across five sources put the market anywhere between $501.9M (Future Market Insights, 2025) and $47B (Bright Data, 2035). The $47B figure has no public methodology attached. Mordor Intelligence (which does publish methodology, using SEC 10-Ks and technology-vendor registries) puts 2024 at $1.03B and 2030 at $2B. For anything matching Mordor's definition, my honest read is somewhere above $1B and below $4B for 2025. Anything higher is quietly including AI agents, services, or proxy infrastructure in the count.

2.4 Bright Data's 60% market share claim is a hypothetical

Bright Data's "60% market share" number is derived, not measured. It's the share Bright Data would have if it were included in Future Market Insights' $501.9M market, a report that doesn't list Bright Data as a player. Applied to Research Nester's report instead, the same math produces 38%. Neither number is a real market share; both are what-ifs constructed against reports Bright Data isn't in.

3. Bot traffic share is the internet's most disputed stat

3.1 The 10% to 99% bot-traffic contradiction

Two honest measurements of different things
10.2%
Global web traffic from scrapers, post-mitigation
F5 Labs 2026 APB Report
99%
Automated traffic on a single small site
Patronview.com log

F5 Labs' 2026 Advanced Persistent Bot Report says 10.2% of global web traffic comes from scrapers after mitigation. Site-owner Patronview says 99% of their traffic is bots. Both are honest measurements of different things. F5 measures verified, persistent, identifiable scraper traffic across a large enterprise network. Patronview is counting every automated request (failed logins, malformed headers, one-shot probes) on a single small site. The 99% number is real for that site owner's logs; it isn't the global average, and it was never supposed to be.

3.2 F5's industry breakdown: Fashion 53%, Hospitality 49%

Fashion53.23%
Hospitality49.32%
Healthcare34.47%
Airline10.76%
Bank / Credit Union7.73%
State / Local Gov7.72%

Scraper share of traffic, by sector. F5 Labs 2025 Advanced Persistent Bots Report.

F5's data is the only anti-bot dataset I've found that publishes methodology and breaks out by sector on a shared denominator. Fashion leads at 53.23%: price changes, inventory updates, and limited-edition drops all reward scraping. Hospitality is close behind at 49.32%. PromptCloud's 2026 report reproduces near-identical numbers (53% / 49% / 34%) at aggregated granularity. The 10.2% global average is a median that hides these extremes.

3.3 AI bots are 33.8% of all bot traffic, and rising

33.8%
of all bot traffic is AI-related.Cloudflare Radar, trailing 28-day window to August 1, 2026

The 33.8% is Cloudflare Radar's annual read; the 28-day trailing view shows 34.30% (AI crawlers 19.31% + AI assistants 8.23% + AI search 6.76%). Both belong to the same measurement family. See section 1.1 above for the full drift from 33.8% to 47.57% depending on window. In commerce, Akamai reports the AI-bot share climbs above 45% for commerce-facing customers.

3.4 Bot traffic is 87.04% desktop: humans are mobile

Cloudflare Radar's July 2026 device split: 87.04% of bot traffic is desktop. Human traffic skews mobile. The pattern isn't a bug: scrapers don't need touchscreens and don't care about responsive design; they want raw HTML, and desktop URLs give it to them cleanest. The defensive implication is that mobile-first fingerprinting misses most of the actual scraper population.

4. Playwright is the dominant automation framework

4.1 Playwright's 45.1% adoption and 94% retention

45.1%
adoption among QA professionals, with 94% retention year-over-year.testdino.com/blog/playwright-market-share (May 2026)

TestDino's May 2026 market-share report puts Playwright's adoption among QA professionals at 45.1%, with 94% retention year-over-year: the sticky-product signal that separates fad tools from defaults. The obvious caveat: this is a single-source number with no published methodology, and until a second data provider confirms it, treat it as directional. What's harder to argue with is the corroborating signal from raw usage: 52 million weekly npm downloads, 88,500+ GitHub stars, 500,000+ public repos depending on Playwright, and ~12,000 verified enterprise adopters. Every one of those is a separate check on the same claim, and they all point up.

4.2 GitHub stars: Puppeteer still leads, Playwright is closing

Puppeteer95,600
Playwright88,500
Cypress51,000
Selenium34,500

GitHub stars, browser-automation frameworks, May 2026.

On raw GitHub-stars headcount, Puppeteer is still ahead of Playwright (95.6k vs 88.5k). Cypress sits at 51k, Selenium at 34.5k. The star metric lags actual adoption: Playwright grew from 78,600 stars in November 2025 to 88,500 by May 2026, a 12.6% jump in six months, and its "Used by" count of 500,000+ public repositories is the single biggest gap in the ecosystem. On dependent-repo count, no rival is close.

4.3 Playwright's 3,791% weekly npm download growth over four years

Playwright's weekly npm downloads went from ~1.2 million in January 2022 to ~52 million in May 2026, a 3,791% increase in four years. The @playwright/test runner package adds several million more weekly downloads on top of the core package. The curve isn't linear; it inflected in 2023 when Playwright added native mobile emulation and cleaned up its TypeScript integration. Puppeteer's weekly downloads are flat around 8.15M; Selenium's are trending down; Cypress is stable. Playwright is the only one still accelerating.

Framework
Weekly npm downloads
Trend
Playwright
~52M
Accelerating
Puppeteer
~8.15M
Flat
Cypress
~5M
Stable
Selenium (JS)
~1.5M
Declining

Weekly npm downloads across the four browser-automation frameworks, May 2026.

Two caveats worth naming. First, npm-download counts inflate for anything a CI pipeline pulls on every build. Second, private-repo usage is invisible. Both apply equally to competitors, so the ratio holds even if the absolute numbers are inflated.

4.4 The 11× Playwright company-count disagreement

Same product, 11× gap in company count
49,632
Companies using Playwright (public-stack detection)
Theirstack
4,484
Companies using Playwright (manual verification)
Landbase

Theirstack reports 49,632 companies using Playwright. Landbase reports 4,484. Same product, 11× gap. Theirstack scans public tech stacks (BuiltWith / SimilarTech-style detection) and counts any site with a Playwright reference, including personal portfolios and demo pages. Landbase runs manual verification against LinkedIn, job posts, and public GitHub repos, and only counts confirmed production usage. Both methods are wrong at the extremes: one over-counts marketing sites, the other misses everything private. The true number is likely between them, but no public tracker resolves it.

5. curl_cffi is the stealthy sleeper in scraping tooling

5.1 curl_cffi has 33.2M monthly PyPI downloads

33.2M
monthly PyPI downloads for curl_cffi. Lifetime total: 335.89M.PyPI download registry, latest month

curl_cffi is the fastest-growing Python HTTP client for scrapers I've seen this cycle. Monthly PyPI downloads sit at 33.2 million; total lifetime downloads at 335.89 million. It's not a Playwright substitute (Playwright is ~200M npm downloads per month across its core + test packages, an order of magnitude larger), but that's not the comparison anyone building an AI-agent pipeline actually needs to make. The real peers are requests and httpx, and against those curl_cffi is winning because it delivers browser-grade TLS fingerprints (JA3/JA4) without shipping a browser.

5.2 curl_cffi vs Puppeteer on a shared unit

Playwright (npm, monthly)~208M
curl_cffi (PyPI, monthly)~33.2M
Puppeteer (npm, monthly)~32.6M

Monthly downloads normalised to a like-for-like unit. Different registries; same denominator.

Normalised to monthly downloads, curl_cffi and Puppeteer are effectively at parity (~33M each). Playwright is roughly 6× larger than either. The interesting comparison isn't volume: it's architecture. Puppeteer and Playwright ship a full browser; curl_cffi ships libcurl with TLS impersonation. For AI agents that scrape LLM outputs, SERPs, and dynamic content at scale, the browser is dead weight most of the time. The right question is not "curl_cffi or Playwright?"; it's "when does this workload need a real DOM?"

5.3 curl_cffi has no browser dependency: ideal for AI agents

For the AI-visibility work I run (2–5M requests/month across LLM interfaces and SERPs), the browser is overhead I don't want. curl_cffi sends HTTP requests, handles cookies and headers, and returns raw content. That's it. No Chromium spin-up, no rendering, no session state to carry between requests. At coupon-scraping and review-analysis scale, I still reach for Playwright when JavaScript rendering is required or when the target sets cookies mid-flow; curl_cffi covers the rest. The right stack in 2026 is a two-tier one: a lightweight HTTP client with browser-grade TLS for the bulk of requests, and a full browser reserved for the pages that actually need one.

6. Anti-bot vendors are fragmented, not unified

6.1 DataDome protects 32,000 websites (BuiltWith)

32,000
active websites protected by DataDome, based on visible deployment fingerprints.BuiltWith 2026 site technology tracker

BuiltWith tracks DataDome usage across 32,000 active sites. The number is a count of visible deployments (DNS records, cookie signatures, header patterns), not a customer count. Sites running DataDome behind reverse proxies or in private networks won't show up. Most anti-bot vendors don't publish anything comparable, so this is one of the few defender-side numbers that's both public and verifiable.

6.2 A site owner's rules blocked 46,729 requests in 24 hours

46,729
requests blocked in 24 hours by a Cloudflare rule set, on a single small site.Patronview.com, 2026 WAF post-mortem

Patronview.com's site owner published a detailed post-mortem of their WAF rules, and the headline was that in the 24 hours after finishing the rule set, Cloudflare blocked 46,729 requests outright: 43,150 of them from Amazon's crawler alone, still hammering an IP range that had been blocked two days earlier. Cloudflare also issued 63,969 challenges over the same window, of which only 552 were solved (a 0.9% solve rate; see section 7.4 below). Cloudflare's strength is speed at the edge; the trade-off, as the site owner noted, is that automated blocking often can't distinguish a scraper from a fast-clicking human. The number is real; the interpretation is that per-site block volume, even for a modest site, is now four-digit-daily.

6.3 F5's data is the most reproducible anti-bot dataset

F5 Labs' 2025 Advanced Persistent Bots Report is the only anti-bot dataset with public methodology, industry-level breakdowns, and a shared denominator across sectors. Fashion 53.23%, hospitality 49.32%, healthcare 34.47%. F5 explains how they distinguish scraper, bot, and automated user, and they publish their detection logic. PromptCloud's 2026 report reproduces near-identical figures at coarser granularity, which is the closest thing to an independent replication in this space. The one honest limitation: F5's numbers only cover sites using their WAF, which biases toward large enterprises. The open web isn't in the sample.

6.4 Akamai carries up to 30% of internet traffic

~30%
of internet traffic carried by Akamai's network: the largest defensive footprint in the industry.Akamai security reports (self-reported)

Akamai's own telemetry claims coverage of roughly 30% of internet traffic; if true, that makes their anti-bot deployments statistically the biggest defensive footprint in the industry. The catch is that Akamai's public bot-intelligence output is thinner than Cloudflare's or F5's. Their focus is DDoS mitigation and cache-layer protection; bot classification at their scale skews coarser. If you're scraping a site behind Akamai, you're more likely to trip a rate-based rule than a fingerprint-based one.

7. Crawler success rates are collapsing under anti-bot pressure

7.1 Crawler 2xx success rate fell from 80.5% to 47.6%

Crawler 2xx success rate: one year, three windows
80.5%
Prior full quarter
Cloudflare Radar
49.2%
Q2 2026
Cloudflare Radar
47.6%
28-day window to Aug 1, 2026
Cloudflare Radar

Cloudflare Radar's own quarter-over-quarter data shows crawler 2xx success rates falling from 80.5% to 49.2% in one year, then further to 47.6% on the trailing 28-day window ending August 1 2026. In July 2025, 2xx responses were 73.50%; in July 2026, 47.66%. On the year-over-year comparison, that's a 25-point drop in the share of crawler requests that get a clean OK. The drivers, per Cloudflare's own commentary, aren't infrastructure: they're anti-bot systems getting better at rejecting automation.

July 2025 (2xx)73.5%
July 2026 (2xx)47.66%
July 2025 (4xx)14.04%
July 2026 (4xx)35.79%

Crawler response codes, year-over-year. Cloudflare Radar.

7.2 4xx errors now hit 35.86% of crawler requests

35.86%
of crawler requests now receive a 4xx error. Up from 14.04% one year ago.Cloudflare Radar, 28-day window to August 1, 2026

The mirror image of the 2xx collapse. Cloudflare Radar's 28-day trailing window shows crawler requests receiving 4xx responses 35.86% of the time; the most recent weekly bucket is 37.1%. Year-over-year, the crawler-side 4xx-block share rose from 14.04% in July 2025 to 35.79% in July 2026. That's the sharpest single-metric shift in the report. Header rotation, IP pool depth, and user-agent variation help at the margin. Nothing I've tested on my own coupon-scraping pipeline has reversed the trend.

7.3 Anthropic's Claude-SearchBot dropped from 60,000 requests/day to 25

A site owner's log, published on patronview.com, shows Anthropic's Claude-SearchBot making about 60,000 requests per day before a firewall rule was applied. After the block, Claude-SearchBot's request count dropped to about 25 per day: a full hard block, not a rate-limit backoff. Amazon's crawler (Amzn-SearchBot) behaved differently in the same log, continuing to hammer the blocked IP range from a rotating fingerprint (43,150 requests recorded in the 24 hours after the block was in place). The pattern is a specific one: a single site can go from mid-five-figure daily requests from one AI crawler to essentially zero within hours of adding one rule, but only if the crawler behaves. Amazon's doesn't.

Crawler
Before block (per day)
After block (per day)
Behaviour
Anthropic Claude-SearchBot
~60,000
~25
Respects rule
Amazon Amzn-SearchBot
(baseline)
43,150 in 24h
Rotates fingerprint

Same site, two AI crawlers, one firewall rule. Patronview.com log.

7.4 CAPTCHA solve rate is 0.24%: the system is failing both directions

0.24%
CAPTCHA solve rate: roughly 1 in 417 attempts pass on a single site's Cloudflare challenges.Patronview.com log, 2026

From the same site-owner post: CAPTCHA solve rate on their site is 0.24%: roughly 1 in 417 attempts pass. The site owner's own rule of thumb, quoted in the post, is that a 30% solve rate means the CAPTCHA is too easy ("you're taxing humans, fix the rule"). 0.24% is the opposite failure: bots overwhelm the challenge queue, humans give up. The 63,969 challenges Cloudflare issued in the 24-hour window on the same site produced only 552 solved attempts. Whatever CAPTCHA is doing in 2026, it's not sorting bots from humans on any site-owner-visible metric.

8. The cost to scrape a page: what we can and cannot say

There is no public, cross-vendor cost-per-request dataset for proxies in 2026. Anything published as "the cost of scraping a page" is either a specific vendor's rate card, a self-reported number from a single scraper's logs, or a fabrication. Rather than invent numbers, here is what I can say from having tested six proxy providers in 2026 on the same 20-site protocol (2,380 requests, concurrency 1/4/8, no retries):

  • Cost lives in the failure rate, not the sticker price. Datacenter IPs are cheap per request and expensive per useful request. On protected sites (Akamai Bot Manager, PerimeterX, warmed-session vendors), block rates are high enough that the effective cost per successful page can exceed residential.
  • Residential IPs earn their keep at the protected-site boundary and nowhere else. For unprotected targets, datacenter proxies or even AWS egress IPs are cheaper and just as reliable. The SERPwatch lesson holds: at 2M SERP-checks/month on a startup budget, AWS beat every proxy network I tried.
  • The real cost driver for AI-agent scraping is orchestration overhead. An agent task rarely resolves in one request: it fans out to 5–20 underlying scrapes for retries, validation, and follow-links. Whatever your per-request cost, multiply by that fan-out factor before comparing budgets.
  • Vendor rate-cards are not comparable. Bandwidth accounting (per-GB vs per-request), concurrency limits, and rotation policies differ enough that identical dollar figures buy substantially different quotas. The only meaningful comparison is a like-for-like real-site sweep.

If you need concrete numbers, run a ramped sweep on your actual targets: one week, three concurrency levels, no retries, and record cost-per-successful-page rather than cost-per-request. Published per-page figures elsewhere in 2026 industry reports do not survive scrutiny.

9. The future of scraping is AI, not humans

9.1 Meta-ExternalAgent made 5.3 billion requests in Q2 2026

5.3B
requests from Meta-ExternalAgent alone in Q2 2026: 57.6M per day, 1.38M per hour.DataDome Q2 2026 threat research report

DataDome's Q2 2026 report puts Meta-ExternalAgent alone at 5.3 billion requests over the quarter: 57.6 million per day, 1.38 million per hour. This is a single AI agent from a single company. DataDome's total AI-agent volume for the same quarter was 17.7 billion requests (up +45% QoQ from 12.2 billion in Q1); Meta-WebIndexer added another 3.75 billion; the balance is split across Anthropic, OpenAI, Google, and dozens of smaller agents. And DataDome only sees what hits its network: the true total across the internet is higher.

9.2 54% of DataDome customers now use agent trust policies

54%
of DataDome customers have adopted agent trust policies: a shift from 'is this a bot?' to 'which agent, and do I want its business?'Jerome Segura, VP Threat Research, DataDome

The defender-side response, quoted directly from DataDome's VP of Threat Research Jerome Segura: 54% of DataDome's customer base has already adopted agent trust policies. That's a categorical shift from "is this a bot?" to "which agent is this and do I want its business?", and the underlying protection stack has to answer both. Agent trust policies aren't about blocking; they're about classifying and pricing access. The frame is closer to API rate-limiting than to traditional bot mitigation. Some sites still don't have agent-specific policies at all, but 54% adoption on the vendor with the largest published customer base is a leading indicator for the rest of the market.

9.3 AI scraping is now the majority of bot traffic on peak-window reads

Put the numbers on one axis: AI-related bots are 33.8% of all bot traffic on Cloudflare Radar's annual view, 34.30% on the trailing 28 days, 46.8% on the June bucket, and 47.57% on the most recent read. In commerce specifically, Akamai reports the AI-bot share exceeds 45%. Meta's two agents alone account for the majority of AI-agent volume on DataDome's network. ChatGPT drives 80–88% of AI-driven referral traffic every month. Anthropic's Claude-User is the #2 individual bot on Cloudflare, second only to GoogleBot.

The scraping-industry-report frame that treats AI crawlers as a new category grafted onto the old one is the wrong frame for 2026. The old category (humans running Selenium scripts on residential proxies to build competitor price lists) still exists, but it's the minority use case now, especially in verticals with defensible margins. The dominant category is machine-to-machine: agents reading pages to feed model training, to service user queries, or to update model-driven recommendations. The web isn't a page anymore; it's a training corpus, a live index, and a set of endpoints. The scraper stack (from TLS fingerprint up to orchestration) is being rebuilt around that.

FAQ

Which single number best summarises "AI's share of the web" in 2026? None. Use the range. Cloudflare Radar's 28-day trailing read is 34.30%; the most recent weekly bucket is 47.57%. Both are correct; pick the one that matches your reporting window.

Is Playwright really 6× bigger than curl_cffi? On download volume, yes: Playwright is ~52M weekly npm (~208M monthly) vs curl_cffi's 33.2M monthly PyPI. But they solve different problems. curl_cffi replaces requests in HTTP-only scraping; Playwright replaces Selenium/Puppeteer in browser-required scraping. The right comparison isn't between them; it's each against its own peer set.

Why does the market-size estimate range from $501.9M to $47B? Different denominators. Software-only counts (Research Nester, Future Market Insights) sit under $3B. Reports that include AI agents, services, and proxy infrastructure (Bright Data, Market Research Future) cross $7B and up. There is no shared industry definition.

Which anti-bot dataset should I trust? F5 Labs for methodology transparency and industry-sector breakdowns. Cloudflare Radar for real-time traffic composition. DataDome for AI-agent-specific data. Akamai for the largest raw coverage, though their public bot data is thinner. No single dataset covers the whole picture.

Do CAPTCHAs still work? Not on the metrics site owners can see. The Patronview data (0.24% solve rate on their site, 552 solved out of 63,969 challenges) is one data point, but it matches the pattern anyone running visible CAPTCHA today reports privately: bots grind through, humans give up.

Updates log

  • Initial publish. Cloudflare Radar / DataDome / F5 windows current to August 1, 2026.
  • Refresh on Q3 2026 DataDome numbers and any Cloudflare quarterly re-cut.

Sources

Every stat above was cross-checked against these public sources. Links kept honest and unshortened.

Market size / research houses

State-of-scraping reports

F5 Labs bot / scraper research

Playwright / tooling usage evidence

Bot traffic / practitioner numbers

Akamai bot research

DataDome threat research