Python · setup guide

Scrapy proxy setup

Scrapy already ships HttpProxyMiddleware, which reads request.meta["proxy"] and turns the credentials in that URL into a Proxy-Authorization header. All you add is a way to put the proxy on every request, including the ones your callbacks create.
You will need

Residentialresi.proxymonkey.io:8000

ISP or datacenterIP:PORT

CredentialsUSER:PASS

Copy your own from the dashboard, which also lists the host and port for every order. Where it differs from this page, the dashboard is right.

Install

Inside a virtualenv, then start a project if you do not have one.

terminal
pip install scrapy && scrapy startproject monkeyshop

Rotating residential

A tiny middleware sets meta["proxy"] unless the request already has one. It runs at 350, before the built-in proxy middleware at 750, which is what makes the credentials in the URL work.

monkeyshop/middlewares.py + settings.py
# middlewares.py
RESIDENTIAL = "http://USER:[email protected]:8000"


class ProxyMiddleware:
    def process_request(self, request, spider):
        request.meta.setdefault("proxy", getattr(spider, "proxy", RESIDENTIAL))


# settings.py
DOWNLOADER_MIDDLEWARES = {
    "monkeyshop.middlewares.ProxyMiddleware": 350,
}

A static datacenter or ISP IP

The middleware above reads a proxy attribute off the spider, so a spider that should run on your static datacenter IP just declares one. For an ISP order, use the host and port from your dashboard.

monkeyshop/spiders/prices.py
import scrapy


class PricesSpider(scrapy.Spider):
    name = "prices"
    proxy = "http://USER:PASS@IP:PORT"
    start_urls = ["https://example.com/catalogue"]

    def parse(self, response):
        for href in response.css("a.product::attr(href)").getall():
            yield response.follow(href, self.parse_product)

    def parse_product(self, response):
        yield {"url": response.url, "price": response.css(".price::text").get()}

Keeping one identity

Scrapy keeps one cookie jar per spider by default, while rotating residential changes the address under it. A site that ties cookies to an IP will notice the mismatch. Either set COOKIES_ENABLED = False for stateless crawls, or keep stateful flows on a static IP or a residential sticky session; the session setting for your account is in the dashboard.

Specific to Scrapy

Things worth knowing

Why the middleware and not meta in start_requests

Requests you yield from a callback, including response.follow, do not inherit the parent request's meta. Setting the proxy only in start_requests means page two goes out from your own IP. The middleware catches every request.

Cache while you write selectors

You will re-run a spider twenty times while fixing a CSS selector. HTTPCACHE_ENABLED = True stores responses on disk so reruns cost nothing. Turn it off for the real crawl.

Settings that decide your bill

DOWNLOAD_TIMEOUT defaults to 180 seconds, which is a long time to wait on a dead residential exit; 30 is plenty. RETRY_TIMES and RETRY_HTTP_CODES decide how many times a blocked page is fetched again, and every retry is metered. AUTOTHROTTLE_ENABLED = True backs off when the target slows down.

settings.py
DOWNLOAD_TIMEOUT = 30
RETRY_TIMES = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 4
AUTOTHROTTLE_ENABLED = True
HTTPCACHE_ENABLED = True
When it breaks

Common errors and fixes

  • TunnelError: Could not open CONNECT tunnel with proxy ... [{'status': 407, ...}]

    Why
    The gateway rejected the credentials on an HTTPS target.
    Fix
    Check the URL in meta["proxy"], quote special characters in the password, and confirm your middleware runs before 750.
  • Ignoring response <407 http://...>: HTTP status code is not handled

    Why
    The same auth failure on a plain-HTTP target, where the 407 comes back as a response and gets filtered.
    Fix
    Same fix as above. Seeing it on http:// URLs only usually means the proxy was set by hand on some requests and not others.
  • TimeoutError: User timeout caused connection failure

    Why
    The download took longer than DOWNLOAD_TIMEOUT.
    Fix
    Lower the timeout so dead exits fail fast, and let the retry middleware send it again through a new address.
  • TLS handshake failures on some sites

    Why
    Scrapy does not verify certificates by default, so these are handshake problems with the target, often an old TLS version or a strict cipher list.
    Fix
    Look at DOWNLOADER_CLIENT_TLS_METHOD and DOWNLOADER_CLIENT_TLS_CIPHERS. The proxy is not involved in the handshake.

A failed connection that moved no data is not billed. Retries you send are billed like any other request, and each shows as its own line in your usage log.

The community layer

Scrapy still misbehaving?

Paste the error and the few lines that set up the proxy into Discord, with the password taken out. Someone there has seen it before.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk