Python · setup guide

scrapy-playwright proxy setup

scrapy-playwright hands marked requests to a real browser, and the browser does not read meta["proxy"]. The proxy goes where Playwright expects it: in PLAYWRIGHT_LAUNCH_OPTIONS for the whole browser, or in a browser context's arguments for one identity, with the credentials in their own username and password fields.
You will need

Residentialresi.proxymonkey.io:8000

ISP or datacenterIP:PORT

CredentialsUSER:PASS

Copy your own from the dashboard, which also lists the host and port for every order. Where it differs from this page, the dashboard is right.

Install

Python 3.10 or newer and Scrapy 2.7 or later. Playwright comes as a dependency; the browser is a separate download.

terminal
pip install scrapy-playwright && playwright install chromium

Rotating residential

Route downloads through the plugin, set the residential gateway at launch, and mark requests with playwright. Chromium keeps its tunnels open inside a browser context, so pages in one context tend to share an exit. Here each request gets its own context, closed after use, so each starts on fresh connections and a fresh address. The browser wraps JSON in a <pre> tag, which is why the callback reads that.

settings.py + spiders/ip.py
# settings.py
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
PLAYWRIGHT_LAUNCH_OPTIONS = {
    "proxy": {
        "server": "http://resi.proxymonkey.io:8000",
        "username": "USER",
        "password": "PASS",
    },
}


# spiders/ip.py
import json

import scrapy


class IpSpider(scrapy.Spider):
    name = "ip"

    async def start(self):
        for i in range(3):
            yield scrapy.Request(
                "https://httpbin.org/ip",
                meta={
                    "playwright": True,
                    "playwright_context": f"ip-{i}",
                    "playwright_include_page": True,
                },
                dont_filter=True,
            )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        await page.close()
        await page.context.close()
        yield json.loads(response.css("pre::text").get())

A static datacenter or ISP IP

PLAYWRIGHT_CONTEXTS creates named contexts at startup, and a proxy there overrides the launch proxy for that context. Send a request to one with playwright_context. Here the static context runs on your datacenter IP; for an ISP order, use the host and port from your dashboard. On Scrapy older than 2.13, rename start to start_requests and drop the async.

settings.py + spiders/account.py
# settings.py
PLAYWRIGHT_CONTEXTS = {
    "static": {
        "proxy": {
            "server": "http://IP:PORT",
            "username": "USER",
            "password": "PASS",
        },
    },
}


# spiders/account.py
import scrapy


class AccountSpider(scrapy.Spider):
    name = "account"

    async def start(self):
        yield scrapy.Request(
            "https://httpbin.org/ip",
            meta={"playwright": True, "playwright_context": "static"},
        )

    def parse(self, response):
        yield {"body": response.css("pre::text").get()}

Keeping one identity

In scrapy-playwright a browser context is the identity: its own cookies and storage and, with a proxy in its arguments, its own exit. Send every request of a login flow to the same named context and it stays one visitor. On the rotating gateway a context can still change address once a connection closes, so for a flow that must keep one IP use a context on a static IP, or a residential sticky session; the session setting for your account is in the dashboard.

Specific to scrapy-playwright

Things worth knowing

Why Scrapy's HttpProxyMiddleware does not apply

HttpProxyMiddleware works inside Scrapy's own downloader, which reads meta["proxy"]. A request marked playwright is fetched by the browser instead, and the plugin lists per-request proxy meta as unsupported. Requests without the mark still use Scrapy's downloader, so a middleware like the one in the Scrapy guide keeps covering them; have it skip browser requests so the two setups never disagree about which IP a request used.

middlewares.py
RESIDENTIAL = "http://USER:[email protected]:8000"


class ProxyMiddleware:
    def process_request(self, request, spider=None):
        if not request.meta.get("playwright"):
            request.meta.setdefault("proxy", RESIDENTIAL)

A new context per request, created while crawling

A context named in playwright_context that does not exist yet is created on the spot from playwright_context_kwargs, proxy included. The arguments only count the first time: if a context with that name already exists, they are ignored. Close contexts you are done with, page first, or PLAYWRIGHT_MAX_CONTEXTS can stall the crawl.

Count what the browser moved

Every image, script and API call the page loads is metered, request plus response bytes, headers included, and Scrapy sees none of them: the plugin does not fire bytes_received. Playwright's request.sizes() reports all four parts for each finished request; add them up per page. Aborted requests never finish, so PLAYWRIGHT_ABORT_REQUEST is the way to cut the total. TLS overhead makes our figure slightly higher.

spiders/prices.py
import scrapy


class PricesSpider(scrapy.Spider):
    name = "prices"
    wire_bytes = 0

    async def start(self):
        yield scrapy.Request(
            "https://example.com/",
            meta={
                "playwright": True,
                "playwright_page_event_handlers": {"requestfinished": "count_bytes"},
            },
        )

    async def count_bytes(self, request):
        self.wire_bytes += sum((await request.sizes()).values())

    def parse(self, response):
        self.logger.info("bytes so far: %d", self.wire_bytes)
When it breaks

Common errors and fixes

  • Pages come back as 407 Proxy Authentication Required, or page.goto fails with net::ERR_TUNNEL_CONNECTION_FAILED

    Why
    The gateway rejected the credentials, or they were written into the server string, where they do not belong.
    Fix
    Keep server to scheme, host and port, and put the pair in username and password. Test the same pair with curl.
  • Browser requests leave from your own IP although meta["proxy"] is set

    Why
    The browser never reads that key. Without a proxy in PLAYWRIGHT_LAUNCH_OPTIONS or the context, it connects directly.
    Fix
    Set the proxy at launch or on the context as shown above, and check with a request to https://httpbin.org/ip.
  • Page.goto: Timeout 30000ms exceeded in the crawl log

    Why
    Playwright's default 30-second navigation timeout, waiting for the load event on a heavy page through a slow exit.
    Fix
    Block heavy resources with PLAYWRIGHT_ABORT_REQUEST, and set PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT in milliseconds if you need longer.
  • Certificate errors such as net::ERR_CERT_DATE_INVALID on browser requests only

    Why
    Scrapy's own downloader does not verify certificates by default and the browser does, so the same URL can pass without the playwright mark and fail with it. TLS runs through the tunnel; the proxy is not part of it.
    Fix
    "ignore_https_errors": True in the context's arguments is for testing only.

A failed connection that moved no data is not billed. Retries you send are billed like any other request, and each shows as its own line in your usage log.

The community layer

scrapy-playwright still misbehaving?

Paste the error and the few lines that set up the proxy into Discord, with the password taken out. Someone there has seen it before.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk