Picking up from the humanize A/B: humanize didn't help on top of the baseline, so I shelved it and left the engine shipping at the 4/9 ≈ 44% floor from the Decodo baseline test. This post is what happened when I came back six days later to run the regression check with nothing changed on my end, and the reframe that forced on me.
1. Where this picks up
In the residential-wall test I had the honest version of the win: stealth-Chromium through Decodo residential 0/3 (silent-hang at TLS), Firefox 3/3 at the handshake but then 3/3 /sorry/ on the actual SERP, Firefox-stealth init script (navigator.webdriver === false verified) still 3/3 /sorry/. The block isn't one wall, it's two. The second wall was something I called "IP reputation" without really knowing what that means in practice.
I was not satisfied with 0/3 but I had a hypothesis: run the proper dashboard test, N=9 real keywords instead of 3, see if the small-N result was unlucky. Residential proxies rotate per connection on Decodo rotating :10000, so N=9 should pull nine different exit IPs, and if even some fraction of the pool is clean, I should see a non-zero pass rate.
I ran the test. 4/9 pass. 2026-09-06.
Not great. Usable, though. If I bolt on retry-on-block I can push 4/9 into maybe 7/9 cumulative, and I have a shippable floor. I checkpointed the state (git tag chromium-decodo-working-2026-09-06), moved on to pipeline work for a few days, and came back to run the regression check before writing it up.
0/9. Same code. Same keywords. Same provider. Same port. Six days later.
2. TL;DR
- The 4/9 → 0/9 collapse was not a code change. Git log clean. Same script, same keywords, same
gate.decodo.com:10000, six days apart. - Running the same probe through three different browser configurations on the same day all came back 0/anything. Not a script problem.
- The common denominator was the pool, not my code. Residential IP reputation moves on a timescale of days, driven by what other people on the same pool were doing to Google: none of which I control, see, or get warned about.
- The reframe this forced on me: pass rate is a function of pool reputation at a point in time, not a function of your script's quality. Any single-day "my scraper got X% pass rate" number is a sample of a moving distribution, not a measurement of a fixed property. Reading them like fixed properties is how you end up with the kind of claim my first writeup on this made and this post's whole existence is correcting.
- The practical consequence: "pick the best proxy provider" is probably not the right question. "Pick a mechanism that survives a bad pool day" is the real one.
What this post does not do: fix the problem. The fix is the parked-tab writeup. This one is about the shape of the problem, because if you don't understand the shape, the fix won't make sense.
3. The regression check that wasn't
Here is what I did on 2026-09-12. I came back to the project after a few days doing something else, pulled the branch, and ran the matrix check: same N=9 real-SEO keywords, same stealth-Chromium through Decodo rotating :10000, same engine code tagged at chromium-decodo-working-2026-09-06. Pacing 35s between queries, warm context, resource-blocking rules unchanged.
preset=warm_cookies pass=0/9 avg_body=92KB avg_lat=4.1s final_urls=9x/sorry/
preset=warm_full pass=0/9 avg_body=92KB avg_lat=4.4s final_urls=9x/sorry/
preset=cold pass=0/9 avg_body=92KB avg_lat=4.6s final_urls=9x/sorry/
Raw pass rate 0.0 across every preset. Not a timeout, not a silent hang, not a weird intermittent. Clean /sorry/ redirects on every single query. The scraper was doing its job perfectly, it was just that Google's answer to every one of its queries was "no."
My first reflex was obviously to check what I had changed on my side. git log --since="2026-09-06" on the engine files: nothing. git diff chromium-decodo-working-2026-09-06..HEAD -- src/core/engines/ src/core/scraper.py: nothing. The code was byte-identical to the version that got 4/9 six days earlier.
Second reflex: maybe the probe is broken, not the pool. So I ran a completely different probe (scripts/probes/camoufox_google.py, a Camoufox-based probe I had wired up for a different test track) through the same Decodo port. 3/3 /sorry/. Then I ran the Firefox stealth probe from the residential-wall test. 3/3 /sorry/.
Three different browser configurations. Three different fingerprint stacks. One shared pool. All zero.
That is the point where the question stops being "what did I change?" and starts being "did my view of the problem change, or did Google's view of this pool change?"
4. What pool reputation actually is (and why you don't get to see it)
Here is the version of the mechanism I'm reasonably confident about, with the hand-wavy parts flagged.
A residential proxy provider does not own IPs the way a datacenter provider does. They pay real residential users (via an SDK bundled into some free app) for the right to route traffic through those users' home internet connections. The "pool" is a rolling set of residential IPs, probably somewhere between 100k and several million IPs active at any given moment, depending on the provider, that anyone with a Decodo (or Oxylabs, or Smartproxy, or Shifter) subscription can route through.
When you buy residential proxy access, you are not buying "your own IPs." You are buying time-shared access to a pool that other scrapers are also using. The IPs you get served when you connect are IPs the provider's load balancer picked for you, probably weighted by availability. You do not know what other scrapers used that IP yesterday. You do not know what sites they scraped, how aggressively, with what pattern. You do not know if the IP showed up on an abuse feed somewhere last week.
Google, meanwhile, is probably running some version of this: track which IP ranges are producing anomalous /search traffic in the last N hours (volume, patterns, timing, header shapes) and down-weight their reputation accordingly. If a Decodo pool full of residential US IPs got hammered by somebody else running a 500-concurrency Google scraper for 48 hours, the IPs you get served the next time you connect carry that history. Google serves them /sorry/ on arrival not because your script is bad but because those specific IPs are now on a provisional abuse list.
That is my best read. The parts I'm hedging:
- I don't actually know Google's abuse-detection mechanics. I'm inferring from behavior. The specific signals, time windows, and thresholds are internal and I will never see them.
- I don't know whether providers quietly rotate "burned" IPs out of the active pool, or whether the pool composition is driven purely by residential-user activity (an SDK installer going online/offline). Probably some mix.
- I don't know whether the 2026-09-12 collapse was pool-wide or just unlucky draws against a mostly-burned pool. The fact that four different browser configs got the same 0/9 result on the same day suggests pool-wide, but I can't rule out a very bad draw.
- I don't know the recovery timescale. The pool was 4/9 clean on 2026-09-06, 0/9 burned on 2026-09-12. Does it recover in 3 days? 10? Never, until the provider explicitly rotates IPs out? I haven't run the long-duration probe (planned for a separate session) that would answer this. I think the answer is "days, driven by other scrapers' activity patterns, not predictable" but I don't have data.
The point I am making is this: pool reputation is a shared resource that you don't control, don't observe directly, and are billed for whether it's working or not. When it's good, you get 4/9 or 7/10. When it's bad, you get 0/9. The transition between those two states can happen overnight for reasons that have nothing to do with your code.
5. Oxylabs, briefly, because I know the question is coming
The obvious next question is: does Oxylabs do this? Shifter? Webshare?
Partial answer: yes, I expect the same pattern on any residential pool, because the economics that cause pool burn on Decodo are the same economics that fund every residential pool. All of them pay residential users for bandwidth, all of them share pool IPs across customers, all of them are therefore susceptible to pool-wide reputation collapse driven by other customers' activity. The question is not whether it happens on Oxylabs: I'm pretty sure it does. The question is how correlated the burns are across providers. If Decodo is burned, is Oxylabs burned? Are they burned at the same time?
That is a measurable question and I have not measured it. The cross-provider matrix is a planned session (3.4) and the pool-recovery-over-24h probe is another (3.6). Both have been sitting on the to-do list behind "actually fix the problem" (which is the parked-tab writeup). I'd rather have the fix shipped and the measurement pending than the other way around.
So I'll keep this honest: Oxylabs probably does the same thing. I don't have the clean cross-day A/B yet. Treat the "Decodo pool burns" claim as a measured fact and the "all residential pools burn" claim as a reasoned expectation.
Note: the thing I'd actually like to measure is cross-provider burn correlation on the same hour. If Decodo and Oxylabs burn simultaneously, that suggests Google's abuse detection is operating on something higher than per-provider pools, maybe residential-IP-range-wide. If they burn independently, that suggests multi-provider redundancy is actually a workable mitigation. These are different worlds and I don't know which one I live in.
6. What this does to the cost math
How I got past the blocks made a cost comparison between self-scraping and Serper. The number I quoted was based on a working pass rate on a working day. That whole comparison needs an asterisk.
Here is the honest version of the cost math under pool-burn conditions:
- On a 4/9 day at 500 KB/query wire, you need ~2.25 queries per successful SERP, so your effective bandwidth per success is ~1.1 MB, or ~$0.0056 per success at $5/GB.
- On a 7/10 day (my earlier measurement) with retry-on-block pushing cumulative to ~84%, you need ~1.2 queries per success, so ~0.6 MB per success, or ~$0.003 per success.
- On a 0/9 day, the cost per success is undefined. The denominator is zero. Every byte you spent that day was wasted. The proxy bill lands anyway.
If pool-burn days happen (and from a sample size of two measurements, six days apart, I observed one usable day and one zero day, so plausibly 50% zero days), the real cost per successful SERP is the usable-day cost divided by the fraction of usable days. If 50% of days are zero, your effective cost is 2× the usable-day cost. If 20% of days are zero, it's 1.25×.
I don't have enough data to put a number on the zero-day fraction. The 24-hour pool-recovery probe would answer that; it hasn't been run. For now, the honest version of the cost comparison in my first writeup is "self-scrape is competitive with managed APIs on days when the pool is working, with an unknown drag factor for pool-burn days."
That is less satisfying than that clean comparison. It is also more true.
7. What this reframes
Four things I would change about how I'd describe this project if I were starting over:
-
"Which proxy provider should I use?" is almost the wrong question. Pool quality from any single provider moves on a daily-to-weekly timescale that is larger than the differences between providers on a single day. On any given day, the "best" provider might be the one whose pool happened to not get hammered last week. You cannot pick that provider in advance.
-
Pass rate is a distribution, not a number. Any article (including, embarrassingly, my first writeup on getting past the blocks) that reports "X/Y pass rate" from a single-day test is reporting a sample of a distribution and labeling it as a parameter. The correct version is "X/Y on 2026-09-06, pool state unknown, re-run three days later gave Z/Y, so pool volatility is at least this much."
-
Retry-on-block is a half-measure. Retrying a failed query on a fresh connection = fresh exit IP works when the pool is partially burned: some IPs are flagged, others aren't, you can find a clean one in a few tries. When the pool is wholly burned, retrying just burns more bandwidth getting more
/sorry/pages. The retry logic needs a circuit breaker, which I didn't have at this point in the project. -
The mechanism matters more than the fingerprint. This is the biggest reframe. The whole first two articles were about getting the browser's fingerprint right: stealth init scripts, TLS impersonation, headless-shell patches. All of that lives in a layer that pool reputation sits above. If the pool is burned, no fingerprint will save you. If the pool is clean, a surprisingly lax fingerprint will survive. The thing I actually want is not a better fingerprint, it's a way of using the pool that minimizes how much pool reputation I'm consuming per query. That is the thing the parked-tab writeup figures out.
8. The concession, because this article has a hole in it
I want to be careful about one thing. "Pool reputation is the main variable" is not the same as "fingerprint doesn't matter." The residential-wall test's finding stands: on a bad TLS fingerprint through a residential proxy, you cannot even complete a handshake. The TLS layer matters when the pool is working; it stops mattering when the pool is burned, because the pool-burn wall sits after the TLS wall and kicks you out at /sorry/ instead of at the handshake.
The honest hierarchy is probably: TLS fingerprint (layer 1) → IP reputation (layer 2) → everything else. You have to clear layer 1 to even get the chance to fail at layer 2. The residential-wall test found a configuration that clears layer 1 (Firefox TLS through residential). This article found that layer 2 is a wall I don't know how to climb from inside the browser. The parked-tab writeup figures out how to climb it from outside the browser.
If someone has data that contradicts this hierarchy, if there's a provider where TLS fingerprint doesn't matter because they terminate TLS at the proxy, or if there's an abuse mechanism I'm missing that sits between layer 1 and layer 2, I'd love to see it. The hierarchy I've drawn here is the most parsimonious explanation of my data, not a proof.
9. What the parked-tab writeup does with this
The mechanism in the parked-tab writeup does not fix pool reputation. Nothing I can do on the client side fixes pool reputation. What it does is use the pool differently: mint one identity with Google (one JS-challenge pass, one x5sec cookie), hold the connection open, and replay queries through the same session so that Google is checking an already-approved identity instead of running a fresh background check on every query. That consumes dramatically less pool reputation per useful SERP. It still fails on wholly-burned-pool days (the mint itself fails and the whole run is dead) but it survives partially-burned-pool days where fresh-navigation-per-query would turn into a /sorry/ cascade.
That is the parked-tab writeup. If you've read this far and you're building the same thing, the TL;DR of everything I've measured so far is: your scraper is doing less work than you think, and the thing it's actually competing with is not Google's bot detector but the aggregate behavior of every other scraper on the same residential pool. The fix is a mechanism change, not a fingerprint change.
If you have cross-provider pool-burn data, especially if you've measured whether burns are correlated across Decodo / Oxylabs / Smartproxy on the same hour, I want to see it before I write session 3.4. That is the single piece of data most likely to change the final recommendation here. Reach out.
Updates log
- Initial publish. The 4/9 → 0/9 collapse, the pool-reputation reframe, and the cost-math asterisk.