Install
Crawlee 3, the current release, on Node.js. The second line is only for PlaywrightCrawler. Crawlee for Python has its own ProxyConfiguration, which takes proxy_urls and pairs sessions with proxies the same way.
npm i crawlee
npm i playwright && npx playwright install chromiumRotating residential
List the residential gateway as the only entry in proxyUrls. Crawlee drops a URL it has already seen, so each request gets its own uniqueKey. The gateway picks the exit for each connection, and proxyInfo in the handler says which proxy URL the request went through.
import { HttpCrawler, ProxyConfiguration } from "crawlee";
const proxyConfiguration = new ProxyConfiguration({
proxyUrls: ["http://USER:[email protected]:8000"],
});
const crawler = new HttpCrawler({
proxyConfiguration,
navigationTimeoutSecs: 30,
async requestHandler({ json, proxyInfo }) {
console.log(json.origin, "via", proxyInfo.hostname);
},
});
await crawler.run(
[0, 1, 2].map((i) => ({ url: "https://httpbin.org/ip", uniqueKey: `ip-${i}` })),
);A static datacenter or ISP IP
Each crawler takes its own proxyConfiguration, so one script can run an HTTP crawler on residential and a browser crawler on your static datacenter IP. With several static IPs, list them all: Crawlee hands them out round-robin and pins each session to one. For an ISP order, use the host and port your dashboard lists for it.
import { PlaywrightCrawler, ProxyConfiguration } from "crawlee";
const crawler = new PlaywrightCrawler({
proxyConfiguration: new ProxyConfiguration({
proxyUrls: ["http://USER:PASS@IP:PORT"],
}),
maxConcurrency: 4,
async requestHandler({ page, request, pushData }) {
await pushData({ url: request.url, title: await page.title() });
},
});
await crawler.run(["https://example.com/"]);Keeping one identity
The session pool is on by default. Each session keeps its own cookies, which crawlers save and send back (persistCookiesPerSession), and is pinned to one entry of proxyUrls. A session serves up to 50 requests and is retired early after errors or a 401, 403 or 429. With one rotating gateway URL the pin holds the URL, not the address: the gateway can pick a new exit for each connection, so a session's cookies may travel across IPs. For a flow that must keep one address, give that crawler a static IP, or use a residential sticky session; the session setting for your account is in the dashboard.