Skip to content

About

Handling 403, 407, 429 and Cloudflare 1020 in Python scrapers: classification, backoff and IP rotation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Web Scraping Error Handling — 403, 407, 429 & Cloudflare 1020 in Python

Production scrapers don't fail on parsing — they fail on HTTP errors. This repo is a compact, battle-tested error-handling layer for requests-based scrapers: classify the status code, decide rotate IP / back off / re-render, and retry with exponential backoff.

The decision table

Status Meaning Correct reaction
403 IP or fingerprint blocked rotate IP, fix headers — full 403 guide
407 Proxy auth failed fix credentials, not retries — 407 explained
429 Rate limited exponential backoff + rotation — 429 playbook
1020 Cloudflare access denied residential exit + browser TLS — error 1020 fix
200 + empty JS-rendered page switch to rendering, not retries

The retry layer

import time, requests

USER, PASS = "YOUR_USERNAME", "YOUR_PASSWORD"
GATEWAY = f"http://{USER}:{PASS}@residentialboson.quantumproxies.io:9000"  # new IP per request

RETRYABLE = {403, 429, 502, 503}

def fetch(url: str, tries: int = 4) -> requests.Response:
    for attempt in range(tries):
        try:
            r = requests.get(url, proxies={"http": GATEWAY, "https": GATEWAY}, timeout=30)
        except requests.exceptions.ProxyError as e:
            # credentials / gateway issue — retrying won't help
            raise SystemExit(f"proxy error (check auth): {e}")
        if r.status_code == 200:
            return r
        if r.status_code not in RETRYABLE:
            r.raise_for_status()
        time.sleep(2 ** attempt)          # 1s, 2s, 4s, 8s
        # rotating port: next attempt exits from a different IP automatically
    raise RuntimeError(f"gave up on {url} after {tries} tries")

Because the gateway rotates the exit IP on every request, a retry is a rotation — no pool bookkeeping. Full runnable version with per-status handlers: errors.py.

Common traps

  • curl gets 403 but the browser works → it's TLS/header fingerprinting, not the IP. Why that happens.
  • ProxyError in requests → almost always credentials or scheme (http:// vs socks5h://). Debug checklist.
  • Retrying 429 on the same IP → makes the ban longer. Rotate first, then back off. How to avoid IP bans.

Get access

Examples run on QuantumProxies rotating residential proxies — per-request rotation means retry == new IP. Credentials: app.quantumproxies.io/register.

About

Handling 403, 407, 429 and Cloudflare 1020 in Python scrapers: classification, backoff and IP rotation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages