Guide · Scraping · October 2026

Scraping Google SERPs with Camoufox: 20/20 at ~$0.0007

The mechanism change that beat all the previous failures. Camoufox + one minted identity + parked tab + in-page fetch(). 20/20 pass on the same Decodo rotating pool that had been serving the old loop 6-7/10. ~$0.0007 per SERP in steady state, at parity with Serper.

Published October 2026 · one of my SERP-scraping writeups · updated as new findings land

Picking up from the pool-burn writeup: the baseline from the Decodo baseline test collapsed from 4/9 to 0/9 on a burned pool with nothing changed on my end, and no renderer-side patch would fix it because the problem wasn't inside the browser. This post is the mechanism change that beat all the previous failures by not touching the browser at all.

1. Where this picks up

By the end of the pool-burn writeup, the diagnosis was uncomfortably clean: pass rate is a function of pool reputation at a point in time, and nothing I can do on the client side changes pool reputation. The old loop (open a browser, navigate to /search, parse, close) was consuming pool reputation at full price per query. On a half-burned pool day that cost me 60-70% of my queries. On a wholly-burned day, 100%. The retry-on-block logic I'd been planning was going to make this worse, not better: more queries against a bad pool is just more /sorry/ pages at full bandwidth.

What I actually needed was a different mechanism. Not a better browser, not a better fingerprint, not a better retry policy. A way of using the pool that consumed less reputation per useful SERP.

The thing that worked wasn't something I thought of. It was in a scraper repo a friend pointed me at: a third-party Google-feed scraper that was doing something I initially thought was going to make Google clock it in half a second. It didn't. On the Decodo rotating pool that had been serving my old loop 6-7/10, this one went 10/10 twice, then 20/20, held one exit IP for the entire session, and cost ~$0.0007 per SERP in steady-state wire bytes. On the same pool, the same day.

I spent two days trying to prove it was wrong. It wasn't wrong.

2. TL;DR

  • Open a Camoufox browser through a rotating residential proxy. Load google.com. Do one real /search navigation: this is the "mint" where Google's JS challenge runs and the x5sec session cookie gets stamped onto the IP that earned it.
  • Navigate back to google.com/ and leave the tab open. The tab holds one live HTTP/2 connection, which on a per-connection-rotating provider (Decodo :10000) pins the exit IP for as long as the connection stays alive.
  • Every subsequent query runs as page.evaluate(fetch('/search?q=...', { credentials: 'include' })) from inside that parked tab. Same connection, same IP, same cookie, same approved identity. Google sees a session it already greenlit.
  • Measured on Decodo rotating :10000: 20/20 OK, ~1.0s per query, 137.9 KB wire per query post-gzip (what the proxy actually bills), ~$0.0007 per SERP in steady state at $5/GB, ~$0.0015 all-in with the mint amortized over 20 queries. Serper is $0.001. We're at parity, and I think there's headroom.
  • The things this does not fix: wholly-burned-pool days (the mint itself fails and the run dies), cross-provider portability (I've only measured this on Decodo so far), captcha-solver minting (I'm using warm-up-only; a real CapSolver mint would survive worse pool conditions). All three are open sessions.

So, what is this mechanism actually for? Rotating residential pools on Google /search in 2026, where you want to amortize one JS-challenge pass over many queries without earning a re-check. Not sticky-session providers, not datacenter-only pools without JS-challenge survival, not non-Google targets where the mechanism is overkill.

3. The repo that embarrassed me

A friend pointed me at a Google-feed scraper repo. I was looking at the code to understand how it handled pagination, which was the problem I thought I was working on that week. What caught my eye was that the scraper opened a browser, did one navigation to Google, and then never navigated again. Every subsequent query, and there were hundreds of them, ran as a page.evaluate(fetch('/search?q=...')) from inside the one open tab.

My reaction was that this was obviously going to get flagged. A single browser tab making a hundred /search requests from inside its own JavaScript context? That's not a browsing pattern a human would produce. Google's going to look at that and go "that's clearly a bot, flag the session, serve captcha." I honestly thought the repo was a toy that had never seen production load.

So I ran it. On my own Decodo credentials. Through :10000 rotating. Just to prove it would die.

It didn't die. Twenty queries, twenty passes, one exit IP, one coherent session. On a pool that had been giving my old loop 6-7/10 the same week. I ran it again. Same result. I interleaved it with the old loop so pool drift couldn't take the credit: alternating queries between arm A (fresh navigation per query, old shape) and arm B (minted identity, repo's shape) within the same probe run. On 2026-10-01 the two arms came back:

run arm pass avg body (decompressed) avg latency unique exit IPs
1 A (fresh nav) 6/10 156.5 KB (passing only) 8.3s 10 (one per query)
1 B (minted + parked) 10/10 543.0 KB 1.6s 1 IP for the whole session
2 A 7/10 248.7 KB 5.7s 10
2 B 10/10 442.6 KB 1.5s 1 IP (87.1.116.225) for all 10

The IP column is where the whole mechanism lives. Arm B held one exit IP across 20 queries. Arm A churned IPs on every connection. The arm-A failures weren't random; they were the queries that drew a residential IP Google had recently seen misbehave. Arm B drew one IP, Google approved it, and then every subsequent query rode that approval instead of being independently judged.

The point I am making is this: I had been working on the wrong problem for weeks. The thing I needed to minimize was not "how obviously am I a bot" but "how often does Google re-check me." Those are different problems with different fixes.

4. The two mechanisms I'd been missing

So why does this work? Here is my reconstruction, in the order I understood it. The first part I'm reasonably sure about from direct evidence; the second part I'm reasonably sure about from provider behavior; the composition of the two is the whole trick.

The cookie-IP binding. When Google's /search endpoint stamps a session cookie (the x5sec one, and probably a few others), the cookie is only valid when presented from the IP that earned it. I know this from an earlier session where I tried the obvious shortcut: run a browser to mint the cookie, extract the cookie, replay it via httpx or curl_cffi from a different process. 30/30 blocked. Two independent reasons: (a) a different process opens a new connection → new residential exit IP → cookie now presented from the wrong IP → invalid; (b) httpx/curl_cffi can't pass the JS challenge the cookie is supposed to attest to, so even if the IP binding were loose, the browser-check-equivalence isn't. The browser has to own the whole identity (fingerprint + cookie + connection + IP) coherently. Nothing less works.

Per-connection rotation. Decodo rotating :10000 assigns a new exit IP per TCP connection, not per request. If your scraper opens a new connection for every query (which is what goto does: the browser pool recycles idle connections aggressively on navigation to a new URL), you get a new IP per query. If your scraper holds one connection open and multiplexes queries through it, you get one IP for as many queries as the connection stays alive. HTTP/2 is designed for exactly this: one TCP connection, many streams. Google's /search is served over HTTP/2. A parked tab on google.com keeps the HTTP/2 connection open to Google's edge, and page.evaluate(fetch('/search')) from inside that tab reuses the same connection, which means the same exit IP.

Compose those two mechanisms and the trick is: one connection → one IP → one cookie → one approved session → N queries that each ride a check Google already ran at mint. The old loop was doing one check per query. The new loop does one check per session. If the session is long enough, you amortize the challenge across many queries and the per-query cost of "being approved by Google" collapses toward zero.

Dead-giveaway in retrospect: the old loop's 6-7/10 ceiling was never about my fingerprint. It was about Google re-rolling the dice on every connection, and some fraction of the dice came up "burned residential IP from your pool." The new loop rolls the dice once.

5. Why Camoufox specifically (and not Firefox-stealth)

The residential-wall test ended on Firefox-with-a-stealth-init-script 3/3 /sorry/ on Decodo. If the Firefox TLS stack is the thing that gets past layer 1, why did I switch to Camoufox for layer 2?

Because stacking stealth patches on Firefox was making it worse, not better. The problem with "Firefox + my own navigator.webdriver patch + my own navigator.plugins patch + my own prefs overrides" is that the fingerprint stops being coherent: I'm forcing a Firefox TLS handshake but a lightly-mangled Firefox renderer, and the composite doesn't look like any real Firefox install any user actually runs. Something downstream reads the mismatch and treats it as more suspicious than either a clean stock Firefox or a clean stock Chromium would be.

Camoufox is a patched Firefox fork maintained specifically for anti-detect automation. It ships one coherent fingerprint bundle: TLS, navigator properties, screen size, locale, UA, timing, canvas, WebGL, all consistent, all picked from a believable real-world distribution via browserforge. You don't bolt on stealth patches. You let it pick a bundle and leave it alone. The whole point is that mismatches between layers are the signal; Camoufox's job is to not leak mismatches.

from camoufox.async_api import AsyncCamoufox

async with AsyncCamoufox(
    headless=True,
    proxy={"server": "http://gate.decodo.com:10000",
           "username": "...", "password": "..."},
    humanize=True,         # built-in mouse jitter, ~200-600ms per navigation
    geoip=True,            # spoof geo from the proxy exit IP
    locale="en-US",
) as browser:
    context = await browser.new_context()
    parked = await context.new_page()
    # MINT: homepage, then one real /search to run the JS challenge
    await parked.goto("https://www.google.com/", wait_until="domcontentloaded")
    await parked.goto(f"https://www.google.com/search?q={warmup_q}",
                      wait_until="domcontentloaded")
    # PARK: navigate back to homepage, leave tab open
    await parked.goto("https://www.google.com/", wait_until="domcontentloaded")

    # REPLAY: every query is an in-page fetch from the parked tab
    for q in queries:
        result = await parked.evaluate(_FETCH_JS, {
            "url": f"https://www.google.com/search?q={q}",
            "timeoutMs": 20_000,
        })

No navigator.webdriver patch. No --disable-blink-features arg. No homepage warm-up as a separate step: the warm-up is the mint. The whole shape is ~30 lines. The _FETCH_JS is a one-page JavaScript function that calls fetch(url, {credentials: 'include'}) with an abort-signal timeout and returns {status, body} across the Playwright bridge.

So I added it as part of the engine. One wrinkle worth flagging: humanize=True costs ~200-600ms of latency per navigation. That is fine. A probe is not a benchmark, and in the engine the mint happens once per identity. The in-page fetches aren't navigations and don't trigger humanization, so the per-query latency stays at ~1.0s.

6. The numbers from the clean solo run

The interleaved A/B above is the proof-of-mechanism. The clean solo run is the number you'd actually bring to a cost conversation. Same probe, arm B only, N=20 queries, Decodo rotating :10000, with proper wire-byte accounting via Playwright's request.sizes() on the parked page:

  • 20/20 OK. No /sorry/, no enablejs, no 429. Twenty queries, twenty SERPs with parseable h3 results.
  • Avg decompressed body: 464.3 KB. This is the HTML the scraper actually parses.
  • Avg wire bytes: 137.9 KB. This is the post-gzip number: what travels across the proxy, what the proxy bills.
  • Avg latency: 1.0s. Down from ~5-8s in the old fresh-nav loop. Most of the latency savings come from not re-running the JS challenge per query.
  • Decodo dashboard delta for the run: +6.27 MB total, 19-20 billed requests, $0.03 billed. Dashboard-reported cost matches the wire measurement to within rounding. So the wire number is billing-grade, not an estimate.

The body-vs-wire gap is 3.4×. I didn't appreciate how big this gap was until this probe. Anyone computing scraper economics from len(response.text) (which is the decompressed body) is overpaying their own estimate by that same 3-4×. Use wire bytes for billing math. Use body bytes for "how much HTML am I actually parsing."

Breaking the cost out properly, at Decodo's $5/GB residential pricing:

per query $/SERP at $5/GB
Steady-state wire (arm B, post-mint) 138 KB $0.0007
All-in with mint amortized over 20 queries 313 KB $0.0015
Serper API (external, for comparison) $0.001

The mint itself costs ~3.5 MB (homepage + one real /search navigation = full page load with CSS/JS execution) which is ~$0.017 per identity. That's a fixed cost. The lever is how many queries you serve from one identity before you recycle. 20 queries → $0.0015. 100 queries → ~$0.0009. 500 queries, if an identity could survive that long → ~$0.00074 all-in.

I don't know how long one identity actually survives in production. The probe tops out at 20 queries and nothing started to drift. I think the ceiling is "somewhere above 25" because that's where I set the proactive-recycle counter in the engine wiring (Session 3.8) and I haven't observed a mid-identity failure yet. The honest version: this is a measured lower bound, not a measured ceiling. Identity-lifetime measurement is a separate session.

7. What this reframes about the earlier tests

Four things this changes about the mental model from everything earlier:

  1. The "rotating beats sticky" rule from how I got past the blocks needs an asterisk. Rotating still beats sticky when "sticky" means "pin the IP via the provider's sticky port, which draws from a smaller more-burned sub-pool." But rotating-with-a-held-connection is a third thing, and it beats both. Fresh IP that was never in the sticky-pool draw, pinned for the duration of one session via the HTTP/2 connection. You get the IP-freshness benefit of rotating and the IP-stability benefit of sticky in the same shape.
  2. The "layer 2 is pool reputation" finding from the pool-burn writeup stands, but with a workaround. Layer 2 is still pool reputation. I still can't fix pool reputation from the client side. What I can do is consume less pool reputation per useful SERP: one JS-challenge pass per 20+ queries instead of per query. On a half-burned pool day, the old loop would eat ~30-40% of its queries getting re-checked and failing; the new loop pays the pool-reputation tax once at mint and then is immune for the rest of the session.
  3. Fingerprint still matters, but it's one layer in a stack of mechanisms. The residential-wall test's finding (TLS fingerprint at layer 1) is correct. The pool-burn writeup's finding (IP reputation at layer 2) is correct. This article's finding is that connection lifecycle at layer 3 is the lever I'd been missing. Fix all three and the mechanism holds. Fix only one or two and the thing collapses in exactly the way it did for me across the earlier tests.
  4. The cost comparison from how I got past the blocks was accidentally approximately right. The pool-burn writeup added the "pool burn makes single-day numbers unreliable" asterisk, which still applies, but on working-pool days, this mechanism's $0.0007-$0.0015/SERP range is actually below Serper's $0.001. The product thesis survives these tests intact, though for completely different reasons than my first writeup claimed.

8. The caveats, laid out plainly

I don't want this post to read as "I solved Google." It is one working mechanism measured on one provider over a handful of runs. Four honest limits:

  • One provider so far. Decodo rotating :10000 only. The whole mechanism depends on per-connection (not per-request) rotation. If Oxylabs rotates per request, a held connection won't hold an IP, and the mechanism collapses into the old loop. If Oxylabs rotates per connection the way Decodo does, it should port cleanly. Cross-provider validation is Session 3.9 and it is not shipped yet. Call this a Decodo-shaped win until proven otherwise. If someone with Oxylabs rotating credentials runs the probe (scripts/probes/minted_identity_inpage_fetch.py --provider oxylabs) and reports back, that would change the shape of the recommendation substantially.
  • Warm-up-only mint, no CapSolver. The original repo mints through a captcha-solving extension; I'm minting by homepage warm-up + one real /search. On a day where the pool is burned enough that even the mint eats a /sorry/, the whole session dies at step one. The probe retries the mint 3× and gives up if all three fail: which is the right cheap-fail behavior, but it means some bad-pool days still produce 0/20. A real CapSolver mint would survive those days by solving the captcha instead of waiting for a clean IP draw.
  • 20 queries is the measured ceiling, not the actual ceiling. I don't know when one identity stops working. 50? 500? The engine wiring uses N=25 as a conservative default with proactive recycle, but that number is a guess. The right answer is probably "measure it with a --recycle 100 run and watch for the first drift," which I haven't done.
  • Pool burn still kills bad-pool days. The mechanism survives a bad draw at mint (via retry). It does not survive a pool-wide burn where every single mint attempt fails. On a 0/20 day the right answer is "stop running and schedule for later," not "retry harder." The engine's circuit-breaker (3 consecutive mint failures → raise PoolBurned, mark the project's runs as error with reason=pool_burned) is the correct behavior but it does not make the burn go away. On a burned pool, no mechanism saves you.

I think the mechanism is real. I don't think the mechanism is universal. The distance between those two statements is the whole open-question list below.

Note: the next thing on the list is whether a datacenter pool (Webshare's $2.99/mo static IPs) can host a minted identity. If the JS challenge survives on a datacenter IP (which I'm not sure it will, Google is quite good at distinguishing residential from datacenter) the cost per SERP drops roughly 8× because datacenter bandwidth is roughly 8× cheaper per GB. If Google's residential-vs-datacenter discriminator fires at mint time, the mechanism is residential-only and the cost ceiling is Decodo's $5/GB. One of those two is true and I don't know which yet.

9. What's still open

In rough order of what I'd measure next:

  • Cross-provider validation (Session 3.9). The per-connection rotation model holding on Oxylabs, Webshare, Shifter. Single most important measurement before the recommendation here is portable.
  • Identity-lifetime ceiling. How many queries one parked tab survives before Google forces a reconnect or revokes the cookie. Currently I run with 25 and recycle; the actual ceiling could be an order of magnitude higher.
  • CapSolver mint path. Survives worse pool conditions than warm-up-only. Worth integrating once the cross-provider matrix is clean enough that the mint is the remaining bottleneck.
  • Datacenter pool viability. The ~8× cost reduction if a datacenter IP can host a minted identity. The outcome shifts the product's unit economics materially either way.
  • Pool-burn correlation across providers (deferred from the pool-burn writeup). If Decodo and Oxylabs burn on the same hour, multi-provider failover doesn't help. If they burn independently, it does. This matters more for operations than for the mechanism itself.

10. The thing I'd tell past-me, if I could

Every test I ran before this one was me trying to make a bot look more like a human. Every lever I pulled was a fingerprint lever: stealth init scripts, TLS impersonation, headless-shell patches, Firefox-vs-Chromium. All of them lived in a layer that pool reputation sits above and connection lifecycle sits beside. I was tuning the wrong axis.

The fix wasn't a better bot. The fix was using the pool less. One mint instead of one-per-query. One connection instead of fresh-per-query. One approved identity instead of a fresh stranger per request. The mechanism change saved more pass rate than every fingerprint patch combined.

If you're building a self-hosted Google scraper in 2026 and you're a few weeks into fingerprint work and nothing is clicking, this is probably the pivot. Stop looking at your handshake. Start looking at your connection lifecycle. The question isn't "how do I look more human?" The question is "how do I earn an approved session once and keep it alive?"

The engine wiring is in src/core/engines/google_identity.py in the open-source repo. The probe is scripts/probes/minted_identity_inpage_fetch.py. If you run it on a provider I haven't measured and the per-connection rotation model holds or breaks, tell me which. If you measure an identity surviving past 100 queries without drift, tell me how many. If someone from a residential-pool provider reads this and wants to argue that pool burn isn't what I think it is, I would genuinely love to be wrong about parts of this.

I'm stopping here because this is the first mechanism that survived its own measurement. Not because it is finished.

Updates log

  • Initial publish. The mechanism, the interleaved A/B, the clean solo run, and the four honest limits.