Picking up from the residential-wall test: the Firefox TLS probe cleared layer 1 but ran straight into layer 2 (3/3 /sorry/), and no renderer-side patch I tried moved layer 2. This post is what I did instead: stop trying to climb layer 2 from inside the browser and just measure what my existing stealth-Chromium engine achieves on Decodo residential at N=9 real-SEO keywords. The number is 4/9. This is why that is actually the floor and not a step on the way to something better, at least not from inside the browser.
1. TL;DR
- The engine that ships right now (Chromium stealth +
--disable-blink-features=AutomationControlled+ 4-line init script + block image/media/font + homepage warm-up + fresh browser per fetch) hits 4/9 ≈ 44% on Decodo rotating:10000through a dashboard N=9 real-keyword test on 2026-09-06. Oxylabs residential on the same engine, same day, hit roughly the same 5/10 ≈ 50%. - Every "obvious optimization" I tried on top of this baseline either didn't move the number or made it worse. Persistent browser context (a 7× bandwidth win on my home IP) collapses to 0% on Decodo. Humanization added 68% bandwidth and 3-6× latency for zero pass-rate gain: that's the humanize A/B. Firefox stealth patches didn't change layer 2: that was the residential-wall test. Blocking stylesheets triggered
/sorry/. - Per-query bandwidth on this engine is ~1.95 MB on the proxy wire. Not 450 KB like the
len(response.text)number my internal logging reported: the Decodo and Oxylabs dashboards both read 4× that. The proxy bills you for what actually traverses the proxy, which is a lot more than what Python sees. - At $5/GB on Decodo and 50% pass rate, that's ~$0.019 per attempted keyword and ~$0.038 per successful SERP. Serper is $0.001. On this baseline alone, self-scrape is 15-30× more expensive than the managed API. The pitch of the whole product does not survive this number.
- The practical takeaway from the baseline: the ceiling is not inside the browser. Every hour spent on stealth patches, behavior sim, TLS hacks, header tweaks is an hour not spent on the two things that would actually move the number: different IP categories, or a different mechanism for using the IPs you have. That tee-up is what the humanize A/B, the pool-burn writeup, and the parked-tab writeup each take a swing at. Spoiler: only one of them worked, and it was the mechanism change.
2. Why I'm publishing the floor and not the ceiling
Most scraping writeups I've read have a shape I don't trust. "Here's the trick that made my scraper work." The trick is usually either (a) a single stealth patch presented without any A/B against the no-patch baseline, (b) a provider recommendation with no measured pass rate, or (c) a code snippet that compiles but has never been pointed at a hostile target under load. The implicit claim is "do this and you're fine." The implicit cost is that readers walk away thinking the problem is solved, build infrastructure assuming it works, and then quietly find out at scale that most of their queries are returning captcha walls.
I'd rather publish the opposite shape. Here is the recipe. Here is the measured pass rate on the hostile target. Here is the bandwidth bill. Here is every optimization I tried on top and why none of them helped. If someone reads this and knows a lever I haven't tried, that's a conversation. If someone reads it and matches the number from their own setup, that's a confirmation. If someone was about to launch a product on top of "I got 10/10 on my home IP" and sees 4/9 on residential: that's a cost saved before it got expensive.
The point I am making is this: a floor you measured honestly is more useful than a ceiling you built from a cherry-picked best run. My first writeup ended on a 10/10 home-IP result that I correctly read as "the mechanism works at all." It was not a production number. This post is the production number.
3. The recipe, exactly as it ships
This is the engine state at git tag chromium-decodo-working-2026-09-06 (commit 6e91594). Nothing fancy. The point of showing you the recipe is not that it's clever (it isn't) but that the baseline is this ordinary. Anyone with Playwright and a residential proxy credential can reproduce it in half a day.
DEFAULT_USER_AGENT = (
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
)
_STEALTH_ARGS = [
"--disable-blink-features=AutomationControlled",
"--disable-features=IsolateOrigins,site-per-process",
"--no-default-browser-check",
"--no-first-run",
]
_STEALTH_INIT_JS = """
Object.defineProperty(navigator, 'webdriver', { get: () => undefined });
window.chrome = window.chrome || { runtime: {} };
Object.defineProperty(navigator, 'plugins', { get: () => [1,2,3,4,5] });
Object.defineProperty(navigator, 'languages', { get: () => ['en-US','en'] });
"""
_BLOCKED_RESOURCE_TYPES = {"image", "media", "font"}
Per fetch:
pw.chromium.launch(headless=True, args=_STEALTH_ARGS): fresh browser per query, no persistent context.browser.new_context(proxy=..., user_agent=..., locale="en-US"): fresh context.context.add_init_script(_STEALTH_INIT_JS): hidenavigator.webdriver, add the chrome shim.context.route("**/*", abort_if_resource_type_in={image, media, font}): do not also block stylesheets. Blocking stylesheets triggered/sorry/in testing. The images/media/fonts block is safe; the stylesheets block is not.page.goto("https://www.google.com/"): warmup homepage. This step costs ~1s and ~0.5 MB of bandwidth, and skipping it noticeably bumps the/sorry/rate on fresh contexts.page.goto("https://www.google.com/search?q=..."): the actual search.page.wait_for_selector("h3", timeout=5s): organic result anchors.- Parse HTML.
A few things worth flagging because they tempted me to change them and I was wrong both times. Fresh browser per fetch reads like obvious wastefulness: why not reuse a persistent context across keywords? On my home IP, persistent context dropped bandwidth from 1,370 KB cold to 177 KB warm per query. On Decodo residential, the persistent-context flavor silently hung at the TLS handshake, and I spent most of a day thinking it was a Playwright bug. See the residential-wall test for the TLS story. The thing I needed to internalize is that persistent context is a home-IP optimization. On residential it stops being an optimization and starts being a liability.
Blocking stylesheets also reads obvious: I'm scraping DOM, not rendering a page, why do I need CSS? Google's edge appears to detect "browser loaded HTML but never fetched CSS" as a shape it treats as suspicious, and the response flips to /sorry/. Blocking image/media/font is fine. The incremental bandwidth saved by also blocking stylesheets is probably ~50 KB per query; the pass-rate cost of blocking stylesheets is roughly all of it. Not worth it.
The recipe is this ordinary because every attempt to make it less ordinary cost me pass rate.
4. The pass-rate matrix (and the day-to-day drift it hides)
Live-dashboard test, N=9 real commercial SEO keywords, 2026-09-06, cold runs through the dashboard:
| Provider | Pool type | Port | Live pass rate | Sample | Notes |
|---|---|---|---|---|---|
| Decodo residential | Backconnect rotating | :10000 |
4/9 ≈ 44% | Dashboard, 2026-09-06 | Non-US exit IPs dominant |
| Decodo residential | Backconnect rotating | :10000 |
4/5 ≈ 80% | Probe arm-A slice, N=5 | Same day, small-sample lottery |
| Oxylabs residential | Backconnect | default | ~5/10 ≈ 50% | Dashboard, 2026-09-06 | Comparable behavior to Decodo |
| Decodo residential (prior day) | Rotating :10000 |
6/10 ≈ 60% | Probe, 2026-09-05 | Different day, different exit IPs | |
| Decodo residential (prior day) | Sticky batch :10001 |
1/10 ≈ 10% | Probe, 2026-09-05 | HTTP 429 for 60s after request 1 |
Two things to read off this table. First, rotating lands 40-60% on any given day, with large same-day variance driven by which exit IPs you happen to draw. The 4/9 and 4/5 rows are the same engine on the same provider on the same day, differing only in which keywords ran when. I don't think that's statistically meaningful: small-sample lottery with the pass/fail decided largely by whether a given query drew a currently-flagged residential IP or not.
Second, sticky lands worse than rotating on Google. The 1/10 sticky result from 2026-09-05 is the specific pattern of "one IP serves one query successfully, then Google 429-rate-limits every subsequent request from that IP for about 60 seconds." For most non-Google scraping, sticky is the better default: logged-in flows, cart sessions, pagination against a session cookie. For Google specifically, sticky trips an aggressive rate limiter on first repeat and the whole premise collapses. How I got past the blocks has the sticky-vs-rotating reversal in more detail; this post just surfaces it in the matrix.
The honest disclaimer: a 9-keyword or 10-keyword test is a sample of a distribution, not a measurement of a parameter. 44% on Decodo on 2026-09-06 does not mean 44% is Decodo's "pass rate." It means on this day, with these IPs that happened to be in the pool, 4 of 9 queries got through. The pool-burn writeup is what happened when I re-ran the exact same probe six days later and got 0/9.
So this baseline number has a shelf life. It is accurate for the day it was measured. It is not a fixed property of my scraper or of Decodo's service.
5. The bandwidth reality (the number my own logging was getting wrong)
Here is a thing I had wrong for the first several weeks of this project. My Python engine was logging response_size = len(response.text) and I was computing cost from that. The number came out to roughly 450 KB per query. At $5/GB that's $0.0022 per attempted query, which comfortably beats Serper's $0.001 per query if you're hitting a pass rate above 50%. The product thesis survived that math.
Then I looked at the proxy dashboard.
Live dashboard readout from a 7-keyword test on Oxylabs residential, 2026-09-06:
- Total bandwidth: 13.64 MB over 51 requests
- Per keyword: ~1.95 MB
- Per-host breakdown:
google.com: 1.17 MB per keyword (2 requests: warmup + search)gstatic.com: 756 KB per keyword (~3.3 requests, mostly JS bundles the browser fetches to render the SERP)play.google.com,ogads-pa.clients6.google.com, etc.: ~23 KB combined
That matches Session 2.7.5's independently-measured Decodo baseline of 1.96 MB per query to within rounding. So the dashboard number is the real number. My len(response.text) number was decompressed HTML of one of the two or three requests the browser made per query: missing the warmup, missing the gstatic bundle fetches, missing everything the browser does on its own to actually render a SERP. I was underreporting cost by about 4×.
The gstatic bundles are the thing I can't block without breaking the JS challenge. The challenge Google serves to prove you're not a bot is itself JavaScript, hosted on gstatic.com. Blocking that domain kills the challenge-pass and the request returns an enablejs interstitial. So of the 1.95 MB per query, roughly 1.17 MB is directly my fault (two google.com navigations) and roughly 756 KB is the cost of being allowed to run the JS the challenge needs: I pay it whether I want to or not.
The practical consequence: if you are computing scraper economics from Python's view of response size, you're probably off by 3-4× low. The proxy bills what traverses the proxy. Python sees the decompressed body of some fraction of the responses. These are different numbers and the gap is systematic, not noise. Use the dashboard, not the client.
6. The cost math, the version where the product thesis loses to Serper
With the corrected 1.95 MB/query number, 50% pass rate, and provider pricing roughly as follows:
| Rate | $/attempted keyword | $/successful SERP (at 50% pass) |
|---|---|---|
| $5/GB (Decodo residential) | $0.0097 | $0.0195 |
| $8/GB (Oxylabs Advanced) | $0.0156 | $0.0312 |
| $10/GB (mid-tier residential) | $0.0195 | $0.0391 |
| $15/GB (small plans, pay-as-you-go) | $0.0293 | $0.0586 |
At 1000 queries per day, 30 days: roughly 58.5 GB per month. $300-900 depending on tier. Compare that to Serper at $30 per 1000 queries = $900 per month for 1000 successful queries per day: and remember Serper only charges for successful queries while residential bills you for every byte including the failed ones. So at 50% pass the comparison is $30/mo per 1000 successful queries on Serper vs $300-900/mo for the same 1000 successful queries on residential. The gap is 10-30× in Serper's favor on this baseline.
That is a brutal number for a product whose entire pitch is "self-scrape beats managed APIs on cost." I want to be careful not to spin it. The pitch does not survive the baseline. If I shipped the product at this engine's current pass rate and current bandwidth, every SEO team evaluating it would do the back-of-envelope and pick Serper.
What the pitch could survive is one of three things changing:
- Pass rate climbs to 90%+: same bandwidth, more successful queries per MB spent. The parked-tab writeup is the data point that moves this number.
- Bandwidth drops by 3-5×: same pass rate, less MB per success. The parked-tab writeup also does this, as a side effect of the mechanism change.
- Different IP class entirely: datacenter IPs at ~$0.10/GB instead of residential at $5/GB would collapse the cost gap by 50×, if datacenter IPs can even survive Google's
/searchlayer-2 check. I have not measured this; it is one of the open questions.
So the baseline cost is embarrassing. It is also the control against which the humanize A/B, pool-burn writeup, and parked-tab writeup each measure a change. If you only look at the parked-tab writeup's $0.0007 per SERP in isolation, it looks like a modestly good number. Compared to this baseline's $0.019-0.059 per success, it is a 25-80× improvement from the same provider through a mechanism change. The baseline is what makes that comparison legible.
7. What I tried on top of this baseline that didn't help
Five optimizations tested on top of the engine above, in roughly the order I tried them. Each either failed to move the pass rate or actively made the economics worse. I'm listing them with cross-links where the full A/B lives in its own writeup so this list doesn't try to be exhaustive.
-
Persistent browser context. Session 2.7.5 on my home IP: 1,370 KB cold per query → 177 KB warm per query, a 7.7× bandwidth win. Pointed at Decodo residential: the whole thing hangs silently at the TLS handshake. The residential-wall test has the TLS-at-the-handshake story. The practical answer: on residential, I ship fresh-browser-per-fetch and eat the cold-start bandwidth. The persistent-context win is real but only on IPs Google doesn't suspect in the first place.
-
Firefox instead of Chromium. On the same Decodo pool and same day as the Chromium silent-hang, Firefox passed the TLS handshake 3/3. Then returned 3/3
/sorry/on actual/searchrequests. Two-layer wall, the residential-wall test. I kept Chromium because Firefox didn't actually win; it just failed at a different layer. -
Firefox with stealth patches.
navigator.webdriver === falseverified in the browser's own self-check. Still 3/3/sorry/. The patches took effect; Google's/sorry/wall doesn't care about renderer-level stealth. The residential-wall test is the autopsy. -
Humanization (scroll + dwell + fake mouse moves). Interleaved A/B against the baseline. Zero pass-rate improvement, 68% more bandwidth, 3-6× latency. This one gets its own writeup because the mechanism-level reason it can't work is worth more than a bullet point: the humanize A/B. Short version: the thing deciding pass/fail happens before any of the gestures get a chance to run, and the gestures themselves trigger lazy-load content that costs money.
-
Aggressive resource blocking (adding
stylesheetto the block list). Instant/sorry/redirect, roughly 1/1 across the small test. "Browser loaded HTML but didn't fetch CSS" appears to be a specific signal Google's edge treats as a bot shape. The incremental bandwidth saved was maybe 50 KB per query; the cost was approximately all my pass rate. I removed stylesheet blocking and never retried.
The pattern across all five: client-side changes don't move the pass rate on a Decodo residential pool whose exit IPs are already on Google's proxy-intel feed. The thing deciding the outcome of each request is the IP, not my browser.
8. Why the ceiling is what it is, the one load-bearing paragraph
Here is my best read of why no amount of in-browser work moves the baseline above ~50% on Decodo residential. I'm going to flag the parts I'm sure of and the parts I'm guessing.
Pretty sure: Google scores incoming /search requests primarily on source-IP reputation, and the residential pools you can rent are not clean inputs to that scorer. Decodo's rotating pool draws from the same residential IPs other scrapers are also using. When those scrapers do things that light up Google's abuse signals, the IPs they touched go onto a provisional abuse list. Your next connection draws from that same pool and inherits the status of whichever IP you were handed. The pass-rate variance you see (44% today, 60% last week, 10% next week) tracks pool composition, not your code.
Pretty sure: the TLS-layer gate and the /sorry/-layer gate are different mechanisms. The residential-wall test walked through this. TLS gates on the specific ClientHello shape coming out of your proxy; /sorry/ gates on IP reputation. Patching the one does not fix the other.
Guessing: the exact signals Google uses to score an IP's reputation, how long a flagged IP stays flagged, how it recovers, whether providers rotate flagged IPs out of their active pool. I think it's some composite of recent /search traffic volume, pattern consistency, header-UA coherence, and whether the IP shows up on third-party proxy-intel feeds (Spur.us, IPQualityScore, DataDome). I'm inferring this from behavior; I don't have Google's side of the decision.
The practical corollary: the marginal return on more browser tuning is near zero. Every hour spent on stealth patches, behavior simulation, TLS impersonation, header reordering is an hour not spent on the two things that would actually move the number: a different IP category (static residential, mobile, datacenter) or a different mechanism for using the IPs you have (fewer connections to the pool, more work per connection, less re-sampling of Google's abuse scorer per useful SERP). Those are the pool-burn writeup and 6. The baseline is what sits here when neither of those levers has been pulled yet.
9. The checkpoint pattern, because every "improvement" above is a potential regression
One habit worth stealing from this project if you don't already have it. The engine state above is tagged in git as chromium-decodo-working-2026-09-06 on commit 6e91594. Rollback is one command:
git checkout chromium-decodo-working-2026-09-06 -- src/core/engines/google.py
The reason this matters: every "optimization" I tried on top of this baseline passed its own small-N test before I noticed it had regressed the real number. Persistent context looked like a 7× bandwidth win in a 3-query home-IP probe. On a 9-keyword Decodo dashboard run, it was 0%. Humanization looked plausibly like it was helping in the first 3 queries of its A/B. By query 10 it was clearly 68% more expensive for the same pass rate.
A tag you can roll back to is the difference between "an experiment failed" and "I accidentally shipped a regression and will find out at the next scrape." For a scraper that is this sensitive to small changes, I think a disciplined rollback anchor is almost mandatory. The verification recipe before you tag anything as an "improvement":
- Restart the server from the venv (not system Python, Playwright version mismatches silently break things).
- Pick ~10 real commercial keywords, not test strings: promo codes, coupons, tech reviews.
- Run them through the dashboard sequentially, note pass/fail per keyword.
- Any improvement over 44% on Decodo
:10000for a comparable keyword set = real signal. Below that = regression, revert.
The disciplined version of this project treats every optimization as guilty until it has beaten the tagged baseline on the same N and same keyword mix. The undisciplined version is how I spent two weeks thinking persistent-context on residential was working before I ran the dashboard check.
10. What the baseline opens up
So where does this leave things. The humanize A/B is one of the five "didn't help" items above, written up in full as a measured A/B, because humanization is the single most-recommended "trick" in the online scraping discourse and the mechanism-level reason it can't work is worth more than a bullet in a list. The pool-burn writeup is what happened when I ran the exact probe on this page six days after the 44% measurement and got 0/9. The parked-tab writeup is the mechanism change that beat all five of the "didn't help" items by not touching any of them.
The baseline is the reference point against which each of those three measures a change. I'd rather have published this post as a stub for weeks than skipped it: "the honest floor" is the piece of scaffolding each of those three leans on.
If you're building the same thing and your scraper is somewhere in the 40-60% pass range on residential proxies against Google: I think that's where you should be. Not because it's good. Because it is where the current browser-side state of the art lives, and anyone telling you they've got a stealth script that reliably clears 80% on residential is either on a very good pool day or hasn't published the matched failure days. If you've measured a stable higher number on residential against Google with a specific recipe you'd be willing to share, I would genuinely like to see your matrix. The thing that would be useful is not another "my scraper got 10/10 once" result; it is a side-by-side with the same keyword mix on the same day as a known baseline. That is the data point I'm still looking for.
Updates log
- Initial publish. The 4/9 baseline on Decodo rotating, the bandwidth correction (1.95 MB per query vs 450 KB my logging reported), and the five optimizations that didn't help.
- Light edits. Added cross-links to the parked-tab mechanism writeup.