Install
Python 3.10 or newer and Scrapy 2.7 or later. Playwright comes as a dependency; the browser is a separate download.
pip install scrapy-playwright && playwright install chromiumRotating residential
Route downloads through the plugin, set the residential gateway at launch, and mark requests with playwright. Chromium keeps its tunnels open inside a browser context, so pages in one context tend to share an exit. Here each request gets its own context, closed after use, so each starts on fresh connections and a fresh address. The browser wraps JSON in a <pre> tag, which is why the callback reads that.
# settings.py
DOWNLOAD_HANDLERS = {
"http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
"https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
PLAYWRIGHT_LAUNCH_OPTIONS = {
"proxy": {
"server": "http://resi.proxymonkey.io:8000",
"username": "USER",
"password": "PASS",
},
}
# spiders/ip.py
import json
import scrapy
class IpSpider(scrapy.Spider):
name = "ip"
async def start(self):
for i in range(3):
yield scrapy.Request(
"https://httpbin.org/ip",
meta={
"playwright": True,
"playwright_context": f"ip-{i}",
"playwright_include_page": True,
},
dont_filter=True,
)
async def parse(self, response):
page = response.meta["playwright_page"]
await page.close()
await page.context.close()
yield json.loads(response.css("pre::text").get())A static datacenter or ISP IP
PLAYWRIGHT_CONTEXTS creates named contexts at startup, and a proxy there overrides the launch proxy for that context. Send a request to one with playwright_context. Here the static context runs on your datacenter IP; for an ISP order, use the host and port from your dashboard. On Scrapy older than 2.13, rename start to start_requests and drop the async.
# settings.py
PLAYWRIGHT_CONTEXTS = {
"static": {
"proxy": {
"server": "http://IP:PORT",
"username": "USER",
"password": "PASS",
},
},
}
# spiders/account.py
import scrapy
class AccountSpider(scrapy.Spider):
name = "account"
async def start(self):
yield scrapy.Request(
"https://httpbin.org/ip",
meta={"playwright": True, "playwright_context": "static"},
)
def parse(self, response):
yield {"body": response.css("pre::text").get()}Keeping one identity
In scrapy-playwright a browser context is the identity: its own cookies and storage and, with a proxy in its arguments, its own exit. Send every request of a login flow to the same named context and it stays one visitor. On the rotating gateway a context can still change address once a connection closes, so for a flow that must keep one IP use a context on a static IP, or a residential sticky session; the session setting for your account is in the dashboard.