Tutorial · 8 min read

Rotating proxies in Scrapy: middleware, retries and bans, done properly

A Scrapy setup for rotating and sticky proxies: the built-in HttpProxyMiddleware, a small custom middleware, retries, ban detection and a byte count.

Scrapy has had proxy support built in for years, and a one-line version works on the first try. The trouble starts at scale: every request leaves from the same few addresses, the site starts answering with 429s, and the retry middleware cheerfully pays to fetch the same block page three times.

This guide builds the setup properly, checked against Scrapy 2.19: the built-in middleware, what decides which IP each request uses, a small middleware that picks residential or static per domain, ban detection that slows down instead of hammering, and a byte count you can hold against your bill.

How do I use a proxy in Scrapy?

Put a proxy URL, credentials included, in request.meta["proxy"]. Scrapy’s HttpProxyMiddleware is on by default, reads that URL, strips the username and password out of it and sends them as a Proxy-Authorization header. Nothing to install.

import scrapy


class PricesSpider(scrapy.Spider):
    name = "prices"

    async def start(self):
        yield scrapy.Request(
            "https://httpbin.org/ip",
            meta={"proxy": "http://USER:[email protected]:8000"},
        )

    def parse(self, response):
        self.logger.info(response.text)

Three things trip people up here.

  • start_requests() is gone. Scrapy 2.13 added async def start(), and 2.16 stopped calling start_requests() at all. A spider that only defines the old method now opens, sends nothing and closes, without an error.
  • Special characters need percent-encoding. The middleware unquotes the username and password, so a password with @ or : in it goes into the URL as %40 or %3A. Get it wrong and you get a 407.
  • Keep credentials in the URL. Since 2.6.2 (July 2022), the middleware removes the Proxy-Authorization header whenever a request’s proxy URL changes, so one proxy’s password is never sent to another. A header you set by hand survives only while the proxy URL stays exactly the same. Credentials in the URL always follow the proxy.

Setting the proxy per request also means every request your callbacks yield needs it too, because response.follow does not copy the parent’s meta. The Scrapy setup page has the five-line middleware that fixes that; the one further down this page is its bigger sibling.

Does Scrapy rotate the proxy IP on every request?

Not by itself. Our residential gateway gives each new connection a new exit IP, and Scrapy tries hard not to open new connections. Its HTTP/1.1 handler keeps a persistent pool, up to CONCURRENT_REQUESTS_PER_DOMAIN connections per host. That is 8 by default and 1 in a project made with scrapy startproject. For an HTTPS site behind a proxy, each pooled connection is a CONNECT tunnel, keyed by the target host, the proxy and the credentials, and every request down a tunnel leaves from the same IP.

We checked. With AutoThrottle spacing the requests out, six requests to one site went down a single tunnel. With the project template’s limit of one connection per host, a whole crawl of one site can ride one tunnel, and to the site that is one visitor from one address.

To get a fresh IP per request, send Connection: close. The site closes the connection after answering, the tunnel goes with it, and the next request opens a new one. The same six requests then opened six tunnels. The price is a new TCP and TLS handshake per request: slower, and a few kilobytes more per page through the proxy.

You can watch it happen. This spider asks for its own address five times through the residential gateway:

import scrapy


class WhoAmISpider(scrapy.Spider):
    name = "whoami"

    async def start(self):
        for _ in range(5):
            yield scrapy.Request(
                "https://httpbin.org/ip",
                headers={"Connection": "close"},
                meta={"proxy": "http://USER:[email protected]:8000"},
                dont_filter=True,
            )

    def parse(self, response):
        self.logger.info("exit IP: %s", response.json()["origin"])

With the header we saw five tunnels, one per request. Without it, the same five requests shared three tunnels at the default limit of 8 per host, and one tunnel at the template’s limit of 1. dont_filter=True matters too: without it, Scrapy’s duplicate filter drops four of the five identical requests before they are sent.

The opposite case, keeping one IP, should not lean on keep-alive at all. A pooled connection can drop at any moment and the next one exits somewhere new. For a login or a cart, turn on a sticky session in the dashboard or use a static ISP address. The rotating vs sticky guide covers which one when.

A Scrapy proxy middleware that picks residential or static per domain

Some sites answer a datacenter IP with real pages. Those are cheaper on a static address, which is priced per IP, not per gigabyte. The rest go through residential. One middleware decides per request:

# monkeyshop/middlewares.py
from scrapy.utils.httpobj import urlparse_cached


class ProxyPickerMiddleware:
    def __init__(self, settings):
        self.residential = settings["PROXY_RESIDENTIAL"]
        self.static = settings["PROXY_STATIC"]
        self.static_domains = set(settings.getlist("PROXY_STATIC_DOMAINS"))
        self.rotate = settings.getbool("PROXY_ROTATE_EVERY_REQUEST")

    @classmethod
    def from_crawler(cls, crawler):
        return cls(crawler.settings)

    def uses_static(self, host):
        return any(host == d or host.endswith("." + d) for d in self.static_domains)

    def process_request(self, request, spider=None):
        if "proxy" in request.meta:
            return None
        host = urlparse_cached(request).hostname or ""
        if self.uses_static(host):
            request.meta["proxy"] = self.static
        else:
            request.meta["proxy"] = self.residential
            if self.rotate:
                request.headers.setdefault("Connection", "close")
        return None

It runs at 350, before the built-in proxy middleware at 750, and it leaves alone any request that already has a proxy, so one spider or one request can still override it. The optional spider=None suits both sides of a change: Scrapy 2.14 and later warn about middleware methods that require a spider argument, and older versions still pass one.

Retries and ban detection in Scrapy

RetryMiddleware is on by default: two retries, on 500, 502, 503, 504, 522, 524, 408 and 429. That last one is the problem. A 429 means “slow down”, and an immediate retry does the opposite. A 403 is dropped without a second try, and a 200 that is really a block page goes straight to your parser.

So take 429 out of RETRY_HTTP_CODES and handle bans in one place that also backs off the whole domain:

# monkeyshop/middlewares.py, continued
import logging

from scrapy.downloadermiddlewares.retry import get_retry_request
from scrapy.exceptions import IgnoreRequest

logger = logging.getLogger(__name__)


class BanBackoffMiddleware:
    def __init__(self, crawler):
        self.crawler = crawler
        s = crawler.settings
        self.codes = set(s.getlist("BAN_HTTP_CODES", [403, 429]))
        self.markers = [m.lower().encode() for m in s.getlist("BAN_MARKERS")]
        self.min_delay = s.getfloat("BAN_MIN_DELAY", 5.0)
        self.max_delay = s.getfloat("AUTOTHROTTLE_MAX_DELAY", 60.0)

    @classmethod
    def from_crawler(cls, crawler):
        return cls(crawler)

    def looks_banned(self, response):
        if response.status in self.codes:
            return True
        head = response.body[:50_000].lower()
        return any(marker in head for marker in self.markers)

    def slow_down(self, request, response):
        key = request.meta.get("download_slot")
        slot = self.crawler.engine.downloader.slots.get(key)
        if slot is None:
            return
        wait = response.headers.get(b"Retry-After", b"").decode()
        asked = float(wait) if wait.isdigit() else 0.0
        slot.delay = min(max(slot.delay * 2, self.min_delay, asked), self.max_delay)
        logger.info("backing off %s: %.1fs between requests", key, slot.delay)

    def process_response(self, request, response, spider=None):
        if not self.looks_banned(response):
            return response
        self.crawler.stats.inc_value("ban/detected")
        self.slow_down(request, response)
        retry = get_retry_request(request, spider=self.crawler.spider, reason="ban")
        if retry is None:
            raise IgnoreRequest(f"still blocked after retries: {request.url}")
        return retry

At priority 560 it sees each response just before RetryMiddleware (550) does. get_retry_request shares the same retry counter, so a URL gets RETRY_TIMES attempts in total, not twice that. A ban doubles the delay for that domain’s download slot, honours a numeric Retry-After, and caps both at AUTOTHROTTLE_MAX_DELAY. A URL still blocked after its retries is dropped, not parsed.

Fill BAN_MARKERS with a phrase from your target’s block page. If that page is a CAPTCHA, the site is asking for a human; the middleware’s job is to slow down and, eventually, give up. Solving it is not on the menu.

Scrapy settings for a proxy crawl, AutoThrottle included

# monkeyshop/settings.py
PROXY_RESIDENTIAL = "http://USER:[email protected]:8000"
PROXY_STATIC = "http://USER:PASS@IP:PORT"
PROXY_STATIC_DOMAINS = ["example.com"]
PROXY_ROTATE_EVERY_REQUEST = True

DOWNLOADER_MIDDLEWARES = {
    "monkeyshop.middlewares.ProxyPickerMiddleware": 350,
    "monkeyshop.middlewares.BanBackoffMiddleware": 560,
}
EXTENSIONS = {"monkeyshop.extensions.ProxyTraffic": 500}

RETRY_TIMES = 2
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408]
BAN_HTTP_CODES = [403, 429]
BAN_MARKERS = ["a phrase from the block page"]

AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 5.0
AUTOTHROTTLE_MAX_DELAY = 60.0
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
CONCURRENT_REQUESTS_PER_DOMAIN = 4
DOWNLOAD_TIMEOUT = 30
ROBOTSTXT_OBEY = True

For PROXY_STATIC, use the IP, port and credentials your dashboard lists for the order. AutoThrottle sets each domain’s delay from how fast it answers, and it never lowers the delay on a non-200 response, so the backoff above sticks until the site starts answering normally again. DOWNLOAD_TIMEOUT defaults to 180 seconds, far too long to wait on a dead residential exit; at 30 the retry gets a new one sooner. ROBOTSTXT_OBEY is on in the project template but off in Scrapy’s defaults, so say it out loud.

How many bytes did my Scrapy crawl use?

Scrapy already counts them. DownloaderStats adds up downloader/request_bytes (method, URL, headers, body) and downloader/response_bytes (status line, headers, body). It sits at 850, on the network side of the compression middleware, so the body it counts is the compressed one that crossed the wire. That matches how we meter: request bytes plus response bytes, headers included. A small extension turns the two into one line at the end of each run:

# monkeyshop/extensions.py
from scrapy import signals


class ProxyTraffic:
    def __init__(self, stats):
        self.stats = stats

    @classmethod
    def from_crawler(cls, crawler):
        ext = cls(crawler.stats)
        crawler.signals.connect(ext.report, signal=signals.spider_closed)
        return ext

    def report(self, spider):
        sent = self.stats.get_value("downloader/request_bytes", 0)
        got = self.stats.get_value("downloader/response_bytes", 0)
        responses = self.stats.get_value("downloader/response_count", 0) or 1
        total = sent + got
        spider.logger.info(
            "proxy traffic: %.2f MB, %.1f KB per response",
            total / 1e6, total / 1e3 / responses,
        )

Retries and ban pages are in there too, because they crossed the proxy and were billed. Our figure will still come out a little higher: tunnel set-up, TLS handshakes and TLS framing are real bytes on the proxy connection that Scrapy never sees, and Connection: close adds a handshake per request. The bill audit guide shows how big that gap should be. The count also includes static-IP traffic, which is not metered per gigabyte, so if you want the residential figure alone, run static domains as their own spider.

Then do the arithmetic before the big run. If a crawl averaged 80 KB per response (an example; your log line has the real number), 100,000 pages would come to about 8 GB, which is $36.80 on our current residential ladder.

When the spider needs a real browser

If the data only appears after JavaScript runs, keep the spider and hand those pages to a browser. The Scrapy-Playwright setup covers the proxy side. Budget for it: a browser fetches scripts, styles and images that a plain request never touches, and every one of them goes through the meter.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk