Guide · Scraping · September 2026

curl_cffi didn't help me scrape Google SERPs

Scraping Google in 2026, the opening story. httpx got nowhere. curl_cffi × 17 TLS profiles: 30/30 blocked. Playwright + a 6-line stealth init script finally pulled a real SERP (597 KB, 8 h3s, status 200) on my home IP. I thought I had a working scraper. I did not.

Published September 2026 · one of my SERP-scraping writeups · updated as new findings land

1. Why I'm doing this

I'm building an open-source rank tracker for SEOs. The pitch is simple: why pay Serper, DataForSEO, or Bright Data $2-$5 per 1,000 Google queries when a residential proxy plan costs a fraction of that per GB? The whole product exists to prove the cost math and give SEOs an alternative to renting somebody else's scraper.

That thesis has one very load-bearing assumption: you can actually self-scrape Google.

Turns out that assumption needed a lot more testing than I expected. This is what I learned, in order, including all the wrong turns.

2. The starting stack

Nothing exotic. Python 3.14, httpx for the HTTP layer, BeautifulSoup for parsing. One fetch(query, ...) function per search engine, one parse(html, ...) per engine. Bing was working fine (still is). Google was returning 100% blocks, the previous session left it that way.

The Google fetcher was already sending Chrome-shaped headers (Sec-Ch-Ua, Sec-Fetch-*, all the right smells), Chrome-shaped URL params, everything a well-configured HTTP client would send. And Google's response was a 200 OK, ~92 KB of HTML, zero organic results. That's a "soft block": Google is answering, just not with anything useful.

The prior working note said "TLS fingerprint is probably the ceiling, need to swap httpx for curl_cffi." That felt right. I was ready to just do it.

Instead I ran probes first. Best decision of the day.

3. Probe 1, does the URL shape matter?

The first thing I wanted to know: are we being blocked because our request looks like a scraper, or because Google is fingerprinting the underlying HTTP client?

I wrote a small script that hits Google with six different URL/param/header combinations, five queries each, all from my home IP:

  • Bare ?q= only, with a curl-style User-Agent
  • Bare ?q= only, with the full Chrome header bundle
  • Current "approach A" (q, oq, sourceid=chrome, source=chrome.ob, ie=UTF-8)
  • Classic scraper (q, hl=en, gl=us, num=10), the shape we know is botlike
  • Approach A with only a User-Agent (no Sec-Ch-Ua, no Sec-Fetch)
  • Minimal Chrome-ish (q + ie=UTF-8)

Result: 30/30 blocked. Identical ~92 KB responses. Every variant. Doesn't matter if you send Chrome's whole handshake or just ?q=, Google answers the same way.

That killed one whole class of "fixes." No header tweak, no param tweak, no URL rewrite was going to get us anywhere. Google wasn't reading our request contents at all before deciding to block us. The decision was happening earlier.

I then looked at what "92 KB of blocked HTML" actually contains. This turned out to be the most important five minutes of the day:

  • Title: "Google Search"
  • Body text (visible): "Ве молиме кликнете овде ако не сте пренасочени по неколку секунди", Macedonian for "Please click here if you are not redirected in a few seconds"
  • No captcha, no "unusual traffic" warning, no /sorry/ redirect
  • A <meta http-equiv="refresh" content="0;url=/httpservice/retry/enablejs?sei=..."> in the noscript block

So this wasn't a "soft block." This was the /httpservice/retry/enablejs interstitial, Google's "please enable JavaScript" bounce page, deliberately localized to Macedonian to make the challenge page harder to grep. My prior code had been miscategorizing this for weeks.

I followed Google's own "click here" fallback link. It returned a smaller 13 KB page titled (translated) "Enable JavaScript to use search." No non-JS bypass path exists. That door is welded shut.

Lesson: before you spend a day fixing what you think is broken, spend ten minutes proving what is actually happening. My whole prior plan was based on calling this a "soft block" when it was really a JS challenge. Different problem, different solution space.

4. Probe 2, the TLS fingerprint hypothesis

Even after Probe 1 shifted the picture, TLS was still the leading suspect. Here's the reasoning: real Chrome and Python's httpx produce very different TLS ClientHellos. Chrome uses BoringSSL with a specific cipher-suite order, extension order, and elliptic-curve preference. Python uses OpenSSL. These get hashed into JA3 and JA4 fingerprints, and every serious anti-bot system in 2026 (Google, Cloudflare, Akamai, DataDome, PerimeterX) reads them.

The standard industry fix is curl_cffi, a Python wrapper around curl-impersonate, a patched libcurl that ships with byte-identical TLS stacks for Chrome, Firefox, Safari. You pick a target (chrome120, chrome124, etc.) and get a matching ClientHello. It's what most successful self-hosted scrapers use.

I installed it and wrote a comparison script that:

  1. Hits tls.browserleaks.com/json with plain httpx, records JA3/JA4/Akamai HTTP2 fingerprints
  2. Hits the same URL via curl_cffi(impersonate="chrome120"), records the same
  3. Then hits real Google /search with both, compares h3 count

The fingerprint diff was objective and clean:

Signal httpx curl_cffi chrome120
JA3 hash 478843a4cce1e9955bfcf160342b6e11 38524d20151eaee2d3c6abedefe61bc4
JA4 t13d1712h1_ab0a1bf427ad_... t13d1516h2_8daaf6152771_...
Akamai HTTP2 hash empty (HTTP/1.1) 52d84b11737d980aef856699f885ca86

The JA4 delta is especially telling: h1 vs h2 means httpx fell back to HTTP/1.1 while curl_cffi correctly negotiated HTTP/2, exactly like real Chrome. Everything the industry says would work, worked at the fingerprint layer.

Google's response to curl_cffi:

The exact same 92 KB enablejs interstitial.

That was a jaw-drop moment. I re-ran with 17 different impersonation profiles: every Chrome version curl_cffi ships (99, 110, 116, 120, 123, 124, 131, 133a), every Safari (15.5, 17.0, 17.2iOS, 18.0, 18.4), Firefox 133/135, Edge 99/101. All 17 returned the enablejs interstitial. Google doesn't care what TLS you present. That's not the discriminator.

This was the moment I had to accept: there is no HTTP-only way to scrape Google today. Not with better headers, not with better TLS, not with better anything at that layer.

5. The curl sanity check I almost skipped

At this point my user pushed back. "Did you actually just try curl 'https://www.google.com/search?q=best+proxy' from the command line? Because that URL loads fine in my incognito Chrome from the same IP."

Fair question. I hadn't. curl has yet another TLS stack (macOS SecureTransport) different from both httpx (OpenSSL) and curl_cffi (BoringSSL-cloned). Different fingerprint, different result maybe.

Three variants:

curl 'https://www.google.com/search?q=best+proxy+services'
curl -A '...Chrome/120...' 'https://www.google.com/search?q=best+proxy+services'
curl -L --http2 -A '...Chrome/120...' 'https://www.google.com/search?q=best+proxy+services'

All three: same enablejs interstitial, same ~92 KB.

That's now five independent HTTP clients (httpx, curl_cffi × 17 profiles, macOS curl HTTP/1.1, macOS curl HTTP/2) all blocked identically from the same IP where real Chrome incognito works.

The only thing Chrome does that none of them do: run JavaScript.

What incognito Chrome experienced but never showed the user: it received the same enablejs page. Its JS engine executed the challenge silently in under 100ms, got a session cookie, and rendered the real SERP under the same URL. The URL bar showing /search?q=best+proxy isn't proof the URL worked, that's Chrome's history.replaceState behavior. The URL is the destination. The JS challenge is the entry gate.

Proof, if you want it: open one of the 92 KB HTML files in a browser as a local file. Chrome will execute the JS and land you on real Google. Same file that shows blank when you cat it.

Lesson: your user's simplest possible test is usually the one you should have run first. Even if you're pretty sure you already covered that case.

6. Pivoting to Playwright

Once you know the answer is "must run JavaScript," the tool changes. You need a real browser. Playwright, Puppeteer, or Selenium.

I picked Playwright. It's the most mature Python browser-automation package, actively maintained by Microsoft, first-class async support, and it handles proxies per browser context (which was going to matter later). Install:

pip install playwright
playwright install chromium

The chromium download is chromium-headless-shell, ~95 MB, a stripped-down headless-only build.

First test, no cleverness:

b = await pw.chromium.launch(headless=True)
ctx = await b.new_context(user_agent="...Chrome/120...")
p = await ctx.new_page()
await p.goto("https://www.google.com/search?q=best+proxy+services")

Response: status 200, redirected to /sorry/index?continue=..., 6.6 KB, "unusual traffic from your computer network".

Playwright by default gets further than httpx (past the JS gate, since it actually has JS) but Google then detects headless-mode fingerprint and escalates to full reCAPTCHA. Common tells for automated Chromium:

  • navigator.webdriver === true (Playwright sets this)
  • Empty navigator.plugins
  • Empty or wrong navigator.languages
  • Missing window.chrome.runtime
  • permissions.query({name:'notifications'}) returns default instead of matching Notification.permission
  • The chromium-headless-shell binary has specific stripped-down properties full Chrome doesn't

I tried four variants in one shot:

  1. Default headless-shell, no stealth → captcha
  2. Headless-shell + a small stealth init script + --disable-blink-features=AutomationControlled → REAL SERP, 597 KB, 8 h3s
  3. Real Chrome (channel="chrome") + stealth, headless → captcha
  4. Real Chrome + stealth, headed (visible window) → captcha

Variant 2 worked. Variants 3 and 4 probably got captchaed as collateral damage from Google seeing three captcha requests from my IP in quick succession, but the point stands: stealth patches on headless-shell were enough.

The stealth init script I ended up with is tiny:

Object.defineProperty(navigator, 'webdriver', { get: () => undefined });
window.chrome = window.chrome || { runtime: {} };
Object.defineProperty(navigator, 'plugins',   { get: () => [1,2,3,4,5] });
Object.defineProperty(navigator, 'languages', { get: () => ['en-US','en'] });
// patch permissions.query('notifications') to return real state

Combined with --disable-blink-features=AutomationControlled at browser launch, that was enough to look like a real user's Chrome to Google's client-side detector.

7. The gotcha that cost me an hour: don't block stylesheets

Playwright lets you intercept network requests. For a scraper, you obviously want to block images, media, fonts: they're bandwidth you don't need, since you're only parsing the HTML.

My first pass also blocked stylesheets. Seemed harmless. I'm not looking at layout, I'm scraping DOM.

Immediate redirect to /sorry/.

Google appears to detect the specific pattern of "browser loads HTML but never fetches CSS" and treats it as a red flag. Blocking image, media, font is safe. Blocking stylesheet is a captcha trigger. Cost delta of leaving CSS on: some KB per query, which is nothing.

The final resource filter:

async def _route(route):
    if route.request.resource_type in {"image", "media", "font"}:
        await route.abort()
    else:
        await route.continue_()

8. Warming the context

One more thing to make it reliable. On a cold browser context, the very first search request captchas roughly 1 in 5 times. Google's classifier is stricter for "this session has zero history."

Fix: navigate to google.com/ first, wait for domcontentloaded (~1s), then do the search in the same context. The homepage visit sets whatever session cookie / cross-tab-ish state Google uses to soften subsequent scoring. Passing all 5 queries after warm-up on my next test.

await page.goto("https://www.google.com/", wait_until="domcontentloaded")
await page.goto(f"https://www.google.com/search?q={query}", wait_until="domcontentloaded")

That's the entire architecture change. One extra request per browser context, dramatically better first-request success.

9. Session pinning: the proxy problem I got completely wrong

At this point Google worked from my home IP. But my product ships to users who want to scrape at scale, from behind residential proxies (Decodo, Oxylabs, etc.). Time to test the proxy path.

Residential proxy providers offer two flavors:

  • Rotating: every TCP connection you open gets a new exit IP from the pool. Great for looking like a swarm of unrelated visitors.
  • Sticky: consecutive connections from you route to the same exit IP for some TTL (from 1 min on Decodo up to 30 min elsewhere).

I had a theory. Because I added a warm-up, session continuity should matter. If the warm-up GET / lands on IP 1.2.3.4 and the follow-up GET /search lands on IP 5.6.7.8, Google sees two different first-time visitors and the warm-up cookie is worthless. So obviously sticky should win, right?

Wrong. But I had to run the test wrong twice to learn that.

9.1 First attempt, where the dashboard revelations began

I ran a probe comparing what I called "pure rotating" (Decodo port :10001) vs "sticky session" (:10000 with a -session-<id> chunk added to the username). For provable correctness, each browser context hit https://api.ipify.org at startup to log its actual exit IP.

Both modes came back 10/10 pass, same IP throughout. Which was baffling: the rotating mode was supposed to give different IPs per request.

Then I looked at Decodo's actual dashboard.

9.2 Correction after checking the Decodo dashboard

Ports were inverted from what I had assumed:

  • gate.decodo.com:10000 = rotating endpoint. Bare username. New exit IP per TCP connection.
  • gate.decodo.com:10001, :10002, :10003, ... = sticky endpoints. Each port is one sticky session with its own exit IP. Username format: user-<base>-sessionduration-<minutes> (Decodo's minimum TTL is 1 minute, which is unusually short and quite useful).

Two things about this worth pointing out for anyone about to make the same mistake:

Decodo distinguishes concurrent sticky sessions by port number, not by a session-ID string. Oxylabs and most other providers put a sessid value in the username and use a single port. Decodo gives you 10+ sticky ports and you pick different ones for parallel sticky sessions. Different mental model, easy to trip over.

The 1-minute sticky TTL is a genuinely useful lever. Most providers cap at 10 or 30 minutes. Decodo's 1 minute means you can rotate off a flagged IP within a minute of it turning bad, which is much tighter control than the alternatives.

My first probe used :10001 (which I thought was rotating but was actually sticky) with a bare username. That's why I saw one IP for 10 requests: I was on a sticky port the whole time. It also used :10000 (actually rotating) with a -session-<random> username fragment that isn't Decodo's syntax, which Decodo appears to have just ignored. Then Playwright reused connections within one browser context, so even on the rotating port I got one IP. Two totally different wrong reasons producing the same misleading result.

9.3 The corrected test

New probe, three modes, correct Decodo semantics this time:

  • A, true rotating: :10000, bare username, fresh browser context per keyword. Each keyword's warm+search share an IP within their context (Playwright's connection pool), but each keyword gets a new IP.
  • B, sticky batch: :10001 with user-<base>-sessionduration-1, one shared browser context for all 10 keywords. Same exit IP for the whole run.
  • C, sticky per keyword: ports :10001 through :10010, one per keyword, fresh browser context each. Every keyword gets a unique-but-stable IP. This was my "best of both worlds" hypothesis: warm+search share IP AND IPs vary across keywords.

Same 10 SEO-flavored queries as before (best proxy, top proxy, decodo review, oxylabs review, smartproxy review, residential proxy, rotating proxy service, seo rank tracker, seo software review, best serp api). Each context hit ipify at startup to log its actual exit IP.

Results:

Mode Pass rate Unique exit IPs Avg latency
A, true rotating 6/10 10 7712ms
B, sticky batch (one IP for all) 1/10 1 6386ms
C, sticky per keyword (unique-but-stable IPs) 3/10 10 10338ms

9.4 What actually happened, per mode

Sticky batch was a disaster. Exit IP 81.56.101.92 served the first request, then Google returned HTTP 429 (rate limited) on every subsequent request for about 60 seconds. Nine consecutive blocks. Only the 10th query, after enough time had elapsed, broke through. My "one user browsing consistently" theory was exactly backwards: for Google specifically, one IP doing 10 searches in 90 seconds trips rate limiting instantly.

True rotating was the winner. Failures came in clusters: 4 pass, 3 fail, 2 pass, 1 fail. Google will 429 you on some residential IPs regardless of your fingerprint, likely because those IPs have recent captcha history from other scrapers sharing the same residential pool. Not something you can fix from your end. But 6/10 first-try is a real usable baseline.

Sticky-per-keyword performed worse than pure rotating despite offering the same IP diversity. My guess: sticky-port pools serve "colder" IPs (less traffic → looks more suspicious), plus the sticky assignment mechanism adds latency (~30% slower than rotating) with no upside.

9.5 The reversal, plainly

My original recommendation was use sticky-per-context, rotate contexts every N queries. Data says the opposite:

For Google SERP scraping, IP diversity beats IP continuity every time. Same-IP bursts trip rate limiting; fresh IPs look like a natural distribution of visitors (a residential ISP with many people happening to search Google, not one user hammering it).

The warm-up cookie continuity I thought mattered? Turns out it only matters within a single keyword's warm+search pair, which Playwright's connection pool preserves for free. It does not need to persist across keywords. Google seems perfectly happy to see a "new visitor" for each search.

Sticky mode is still useful: logged-in scraping, cart flows, pagination against a session cookie, anything with real per-user state. Just not for Google SERPs.

9.6 The lucky part

The engine I'd already written creates one fresh browser context per fetch() call. And my DB happened to have Decodo stored with port :10000 and a bare username. That combination is already true rotating mode. The winning configuration was what was already running. I just didn't know until the data proved it.

9.7 Next optimization

60% first-try is usable but not great. The obvious next step is retry-on-block in the orchestrator: if parse() reports a block reason, close the context and retry once with a fresh browser context (which gets a fresh IP on rotating). Compounding math: 6/10 first try + 6/10 of the failures = ~84% cumulative pass rate. Save for a dedicated session; ship the baseline first.

10. Where I am now

  • Bing: still on plain httpx + optional proxy. Works fine. No JS gate.
  • Google: rewritten around Playwright + stealth patches + homepage warm-up + resource blocking (images/media/fonts only). Runs headless. Proxy per browser context via session-pinned username.
  • Perf: ~2-5s per Google query (vs 200ms for httpx, when httpx was still failing). Fine for a rank tracker. A 500-keyword project finishes in ~15 minutes on 2-3 concurrent browser contexts.
  • Bandwidth: ~450-600 KB per rendered SERP with images/media/fonts blocked. Still hugely cheaper per keyword than managed API pricing.
  • Product thesis intact: the "proxies not APIs" pitch holds, just with a headless browser instead of raw HTTP for Google.

11. What I got wrong, listed plainly

Because failures are the whole point of this post.

  1. I trusted my own prior notes. They called the block a "soft block" (empty SERP). It was actually a JS challenge (enablejs interstitial). Different problem, different fix. Cost: probably a couple of extra sessions of chasing the wrong hypothesis.
  2. I put TLS fingerprint at the top of my suspect list. It was the industry-canonical answer for "why is my scraper blocked despite Chrome headers." Turns out, for Google specifically in 2026, TLS spoofing doesn't help. Fingerprint changed cleanly, response didn't budge.
  3. I nearly skipped the plain-curl test. User had to push me to run it. Simplest possible check, would've saved 45 minutes if I'd done it first.
  4. I blocked stylesheets. Instant /sorry/ redirect. One tiny detail with a huge cost.
  5. I overclaimed early results. My first proxy test was 2/5. I noted it. Same config later hit 10/10. Small samples are noisy; one 5-query test isn't a signal.
  6. I got Decodo's port convention exactly backwards. Assumed :10000 was sticky and :10001 was rotating. It's the reverse. Their dashboard made this obvious in 30 seconds. Cost: an entire probe run testing two misconfigured modes and producing a misleading "both worked" result.
  7. I recommended sticky-per-context and it turned out to be the worst option for Google. My reasoning was "warm-up cookies need IP continuity." Sticky-batch got 1/10 pass because same-IP bursts trigger 429 rate limiting in about one request. Pure rotating got 6/10. IP diversity, not continuity, is what Google wants for SERP requests.

12. Open questions I still need answered

For Decodo support (talking to them next):

  • Confirm the port scheme long-term: is :10000 always going to be rotating and :10001+ always sticky, or is that subject to change?
  • Do sticky ports higher than :10010 exist? What's the concurrent-sticky-session cap for a standard account?
  • Any per-country or per-region session pool differences? Is a -country-us username fragment supported, and how does it interact with sessionduration?
  • Any hint of which residential subnets are "hotter" (more captcha history), or is that just noise we live with?

For myself, next testing session:

  • Run the 3-mode probe 3× at different times of day, plus at least one weekend run. Today's 6/10 might not be stable.
  • Add retry-on-block to the orchestrator, measure how much it lifts cumulative pass rate.
  • Measure actual per-keyword bandwidth over a real 500-keyword project.
  • Calculate real $/1,000 keywords via Decodo, put it beside Serper's / DataForSEO's public pricing: that comparison is the whole product pitch.

For the reader, things worth reading up on (living reading list, I'll add real links as I revisit each one):

  • JA3 and JA4 TLS fingerprints: what they are, why every anti-bot system looks at them.
  • curl-impersonate and curl_cffi: the standard toolkit for Chrome-shaped TLS in Python, even though it didn't save me here.
  • Playwright's stealth patterns: the specific navigator.webdriver etc. tells that get patched.
  • patchright: a Playwright fork with more anti-detection patches baked in. Didn't need it here; would try it if vanilla Playwright started failing.
  • Residential proxy session syntax: Decodo and Oxylabs docs on their username formats.

For SEOs reading this: you have real leverage over scraper-API vendors. The tech to self-scrape exists, it works, and the cost math is on your side. The gap is that most self-hosted scrapers stopped at "HTTP client + proxy" and never made it to "headless browser + proxy," which is where you have to be for Google in 2026.

For devs reading this: the industry-canonical answer ("just use curl_cffi") is no longer sufficient for the biggest target. Update your mental model. The frontier has moved to JS execution + browser fingerprint stealth. httpx-only is still fine for a lot of the web, but not for Google SERPs.

For me: I'll keep updating this post as I learn more. Especially after Decodo support and I sort out what their rotation and session TTL semantics actually are. If you're an SEO or a dev doing similar work, tell me what I'm getting wrong.

Updates log

  • Initial publish. Home-IP breakthrough, the five rungs that got me there, and the proxy session reversal.
  • Light edits. Preserved the premature conclusion (that later writeups take apart) rather than rewriting it after the fact.