Install
Inside a virtualenv, then start a project if you do not have one.
pip install scrapy && scrapy startproject monkeyshopRotating residential
A tiny middleware sets meta["proxy"] unless the request already has one. It runs at 350, before the built-in proxy middleware at 750, which is what makes the credentials in the URL work.
# middlewares.py
RESIDENTIAL = "http://USER:[email protected]:8000"
class ProxyMiddleware:
def process_request(self, request, spider):
request.meta.setdefault("proxy", getattr(spider, "proxy", RESIDENTIAL))
# settings.py
DOWNLOADER_MIDDLEWARES = {
"monkeyshop.middlewares.ProxyMiddleware": 350,
}A static datacenter or ISP IP
The middleware above reads a proxy attribute off the spider, so a spider that should run on your static datacenter IP just declares one. For an ISP order, use the host and port from your dashboard.
import scrapy
class PricesSpider(scrapy.Spider):
name = "prices"
proxy = "http://USER:PASS@IP:PORT"
start_urls = ["https://example.com/catalogue"]
def parse(self, response):
for href in response.css("a.product::attr(href)").getall():
yield response.follow(href, self.parse_product)
def parse_product(self, response):
yield {"url": response.url, "price": response.css(".price::text").get()}Keeping one identity
Scrapy keeps one cookie jar per spider by default, while rotating residential changes the address under it. A site that ties cookies to an IP will notice the mismatch. Either set COOKIES_ENABLED = False for stateless crawls, or keep stateful flows on a static IP or a residential sticky session; the session setting for your account is in the dashboard.