Tutorial · 7 min read

Async scraping in Python with proxies: httpx and aiohttp

Hundreds of concurrent requests through a proxy with httpx or aiohttp: connection limits, per-request rotation, retries with backoff and a running byte count.

A synchronous scraper spends almost all its time waiting for the network. Async scraping in Python fixes that: one process keeps dozens of requests in flight and handles each reply as it lands. Both httpx and aiohttp do this well, and both take a proxy in one argument.

The one argument is the easy part. This guide covers the rest: capping concurrency, getting a new IP when you want one (and keeping it when you do not), retrying without hammering, and counting the bytes you move, because a residential proxy bills every one of them.

How async scraping with a proxy works

You create one client, point it at the proxy, and start many requests at once with asyncio.gather. A semaphore caps how many are in flight, and each request opens or reuses a connection through the proxy. That reuse is the detail most guides skip, and it decides which IP each request leaves from.

httpx proxy: AsyncClient(proxy=...)

pip install httpx
import asyncio
import httpx

PROXY = "http://USER:[email protected]:8000"


async def main():
    async with httpx.AsyncClient(proxy=PROXY, timeout=20) as client:
        for _ in range(3):
            r = await client.get("https://httpbin.org/ip")
            print(r.json())


asyncio.run(main())

The argument is proxy=, singular; the httpx setup page covers the older spelling, per-pattern mounts and SOCKS.

Run that loop and look at the three addresses. There is a good chance they are all the same, even though the residential gateway rotates per connection. More on why below.

aiohttp proxy: per request, with credentials

import asyncio
import aiohttp

PROXY = "http://USER:[email protected]:8000"


async def main():
    async with aiohttp.ClientSession() as session:
        async with session.get("https://httpbin.org/ip", proxy=PROXY) as resp:
            print(await resp.json())


asyncio.run(main())

aiohttp takes the proxy per request, which makes mixing proxies in one session easy; the aiohttp setup page covers proxy authentication, environment variables and SOCKS.

Limit concurrency: a Semaphore plus connection limits

Two separate limits, and you want both. A semaphore caps how many requests your code has in flight. The client’s connection pool caps how many sockets it opens.

limits = httpx.Limits(max_connections=20, max_keepalive_connections=0)
client = httpx.AsyncClient(proxy=PROXY, limits=limits)

connector = aiohttp.TCPConnector(limit=20, limit_per_host=5, force_close=True)
session = aiohttp.ClientSession(connector=connector)

gate = asyncio.Semaphore(20)

aiohttp defaults to 100 connections in total and no per-host cap, so without limit_per_host a list of URLs from one site can open 100 connections to it at once. That is how an async scraper earns a wall of 429 errors, and on a shared pool it gets IPs flagged for everyone. Start low, five or so per site, and raise it only while the error rate stays flat.

Does an asyncio proxy rotate IPs on every request?

Only if every request opens a new connection. For an HTTPS URL, your client sends CONNECT to our residential gateway, and the gateway builds a tunnel from one exit IP to the site. Everything that travels through that tunnel leaves from the same IP, because it is one TCP connection end to end. A new connection gets a new tunnel, and a new exit.

Both clients keep connections alive by default and reuse them for the next request to the same host. So a keep-alive client scraping one site can send hundreds of requests down a handful of tunnels, from a handful of IPs. That is why the httpx loop above probably printed one address three times.

When you want rotation, turn reuse off. In httpx that is max_keepalive_connections=0; in aiohttp, force_close=True on the connector. Both appear in the snippet above. The cost is a fresh TCP and TLS handshake per request: more latency, and a few kilobytes of handshake per request that go through the proxy and are metered like everything else.

When you want the same IP, do not lean on keep-alive to give it to you. A connection can drop at any moment and the next one leaves from somewhere new. For a login or a cart, use a sticky session (a setting in the dashboard) or a static ISP address. The rotating vs sticky guide has the patterns.

Retries with exponential backoff and Retry-After

Retry what can change on a second try: timeouts, connection errors, 429, 502, 503, 504. Do not retry 403 or 404; the answer will be the same and you pay for it again. When a site sends Retry-After with a 429, wait that long, up to a cap (two minutes below) so one odd header cannot park a coroutine for an hour. Otherwise back off exponentially with jitter, so a hundred coroutines that failed together do not all retry in the same millisecond.

import random
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

RETRY_ON = {429, 502, 503, 504}


def retry_after(headers):
    value = headers.get("Retry-After", "").strip()
    if value.isdigit():
        return float(value)
    try:
        when = parsedate_to_datetime(value)
    except (TypeError, ValueError):
        return None
    if when.tzinfo is None:
        when = when.replace(tzinfo=timezone.utc)
    return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())


def backoff(attempt, cap=60.0):
    return random.uniform(0, min(cap, 2 ** attempt))

Retry-After comes as either a number of seconds or an HTTP date, so the helper handles both. The backoff is “full jitter”: a random wait between zero and a ceiling that doubles each attempt.

Timeouts that suit a residential proxy

Residential exits are home connections, and some are slow. Set a short connect timeout, so a dead exit fails fast and the retry gets a new one, and a longer read timeout for the page itself.

timeout = httpx.Timeout(30.0, connect=10.0)
timeout = aiohttp.ClientTimeout(total=45, sock_connect=10, sock_read=30)

Also note that httpx does not follow redirects unless you pass follow_redirects=True. aiohttp does follow them. Each redirect is another request through the proxy.

Count bytes the way a residential proxy meters them

We bill request bytes plus response bytes, headers included, on the wire. So the counter needs the request line and headers you sent, the status line and headers you got, and the body as it arrived, compressed if the site compressed it. httpx’s num_bytes_downloaded is exactly that: the raw body bytes before decoding.

def head_bytes(start_line, raw_headers):
    return len(start_line) + 2 + sum(len(k) + len(v) + 4 for k, v in raw_headers) + 2


def wire_bytes(response):
    req = response.request
    sent = head_bytes(f"{req.method} {req.url.raw_path.decode()} HTTP/1.1", req.headers.raw)
    got = head_bytes(f"HTTP/1.1 {response.status_code} {response.reason_phrase}", response.headers.raw)
    return sent + len(req.content) + got + response.num_bytes_downloaded

In aiohttp, use resp.request_info.headers for what you sent, resp.raw_headers for what came back, and resp.content.total_raw_bytes for the body before decompression. That last one is in the current aiohttp docs; if your installed version lacks it, upgrade rather than counting the decoded body, which overcounts.

Expect our figure to be a few percent above yours. TLS handshakes and tunnel set-up are real bytes on the proxy connection that a client-side counter does not see. The same arithmetic, with curl, is on the curl setup page, and the bill audit guide shows how to diff a count like this against your usage export.

All together: an async httpx scraper

import asyncio
import random
import sys
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

import httpx

PROXY = "http://USER:[email protected]:8000"
CONCURRENCY = 10
MAX_ATTEMPTS = 4
MAX_WAIT = 120
RETRY_ON = {429, 502, 503, 504}
HEADERS = {
    "User-Agent": ("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                   "(KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36"),
    "Accept-Language": "en-GB,en;q=0.9",
}
moved = {"bytes": 0, "requests": 0}


def head_bytes(start_line, raw_headers):
    return len(start_line) + 2 + sum(len(k) + len(v) + 4 for k, v in raw_headers) + 2


def count(response):
    req = response.request
    sent = head_bytes(f"{req.method} {req.url.raw_path.decode()} HTTP/1.1", req.headers.raw)
    got = head_bytes(f"HTTP/1.1 {response.status_code} {response.reason_phrase}", response.headers.raw)
    moved["bytes"] += sent + len(req.content) + got + response.num_bytes_downloaded
    moved["requests"] += 1


def retry_after(headers):
    value = headers.get("Retry-After", "").strip()
    if value.isdigit():
        return float(value)
    try:
        when = parsedate_to_datetime(value)
    except (TypeError, ValueError):
        return None
    if when.tzinfo is None:
        when = when.replace(tzinfo=timezone.utc)
    return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())


def backoff(attempt, cap=60.0):
    return random.uniform(0, min(cap, 2 ** attempt))


async def fetch(client, gate, url):
    for attempt in range(MAX_ATTEMPTS):
        try:
            async with gate:
                response = await client.get(url)
        except httpx.TransportError as exc:
            wait = backoff(attempt)
            print(f"retry in {wait:.1f}s ({type(exc).__name__})  {url}")
            await asyncio.sleep(wait)
            continue
        count(response)
        if response.status_code in RETRY_ON:
            wait = min(retry_after(response.headers) or backoff(attempt), MAX_WAIT)
            print(f"retry in {wait:.1f}s ({response.status_code})  {url}")
            await asyncio.sleep(wait)
            continue
        print(f"{response.status_code}  {url}")
        return response
    return None


async def main(path):
    urls = [line.strip() for line in open(path) if line.strip()]
    gate = asyncio.Semaphore(CONCURRENCY)
    limits = httpx.Limits(max_connections=CONCURRENCY, max_keepalive_connections=0)
    timeout = httpx.Timeout(30.0, connect=10.0)

    async with httpx.AsyncClient(proxy=PROXY, headers=HEADERS, limits=limits,
                                 timeout=timeout) as client:
        results = await asyncio.gather(*(fetch(client, gate, u) for u in urls))

    ok = sum(1 for r in results if r is not None and r.status_code == 200)
    print(f"{ok}/{len(urls)} ok, {moved['requests']} requests, "
          f"{moved['bytes'] / 1e6:.2f} MB through the proxy")


if __name__ == "__main__":
    asyncio.run(main(sys.argv[1]))

The sleep sits outside the semaphore on purpose: a coroutine waiting to retry should not hold a slot another URL could use, and MAX_WAIT caps any single wait, including one a Retry-After header asked for. For tens of thousands of URLs, swap the single gather for a queue and a fixed set of workers, so you are not holding a coroutine per URL in memory. Every retry is counted, because every retry is billed.

When to prefer ISP or datacenter IPs for async scraping

Which proxy line suits which async scraping job
JobLineWhy
Many independent public pagesResidentialA new IP per connection spreads per-IP limits; you pay per byte
Large pages, lenient siteDatacenterA dedicated IP is a flat price, so page weight stops mattering
A session or loginISPOne consumer-network address that stays put
High volume on one APISeveral dedicated static IPsRound-robin with a per-IP gap, no per-GB meter

Async clients make it easy to move a lot of data, which is exactly when per-gigabyte billing starts to bite. Residential starts at $5.50/GB for a single gigabyte. A dedicated datacenter address is $3.20 for 30 days with no traffic allowance to run out of. If the site answers a datacenter IP with real pages, try that line first; the residential vs datacenter test takes five minutes, and the datacenter pricing page has the full ladder.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk