JavaScript and TypeScript · setup guide

Crawlee proxy setup

Crawlee takes proxies through one object. Build a ProxyConfiguration, pass it as proxyConfiguration to any crawler, HttpCrawler, CheerioCrawler or PlaywrightCrawler, and every request goes through it. Crawlee rotates through the proxyUrls you list; with ProxyMonkey that list can be a single gateway URL, because the residential gateway rotates behind it.
You will need

Residentialresi.proxymonkey.io:8000

ISP or datacenterIP:PORT

CredentialsUSER:PASS

Copy your own from the dashboard, which also lists the host and port for every order. Where it differs from this page, the dashboard is right.

Install

Crawlee 3, the current release, on Node.js. The second line is only for PlaywrightCrawler. Crawlee for Python has its own ProxyConfiguration, which takes proxy_urls and pairs sessions with proxies the same way.

terminal
npm i crawlee
npm i playwright && npx playwright install chromium

Rotating residential

List the residential gateway as the only entry in proxyUrls. Crawlee drops a URL it has already seen, so each request gets its own uniqueKey. The gateway picks the exit for each connection, and proxyInfo in the handler says which proxy URL the request went through.

rotate.mjs
import { HttpCrawler, ProxyConfiguration } from "crawlee";

const proxyConfiguration = new ProxyConfiguration({
  proxyUrls: ["http://USER:[email protected]:8000"],
});

const crawler = new HttpCrawler({
  proxyConfiguration,
  navigationTimeoutSecs: 30,
  async requestHandler({ json, proxyInfo }) {
    console.log(json.origin, "via", proxyInfo.hostname);
  },
});

await crawler.run(
  [0, 1, 2].map((i) => ({ url: "https://httpbin.org/ip", uniqueKey: `ip-${i}` })),
);

A static datacenter or ISP IP

Each crawler takes its own proxyConfiguration, so one script can run an HTTP crawler on residential and a browser crawler on your static datacenter IP. With several static IPs, list them all: Crawlee hands them out round-robin and pins each session to one. For an ISP order, use the host and port your dashboard lists for it.

static.mjs
import { PlaywrightCrawler, ProxyConfiguration } from "crawlee";

const crawler = new PlaywrightCrawler({
  proxyConfiguration: new ProxyConfiguration({
    proxyUrls: ["http://USER:PASS@IP:PORT"],
  }),
  maxConcurrency: 4,
  async requestHandler({ page, request, pushData }) {
    await pushData({ url: request.url, title: await page.title() });
  },
});

await crawler.run(["https://example.com/"]);

Keeping one identity

The session pool is on by default. Each session keeps its own cookies, which crawlers save and send back (persistCookiesPerSession), and is pinned to one entry of proxyUrls. A session serves up to 50 requests and is retired early after errors or a 401, 403 or 429. With one rotating gateway URL the pin holds the URL, not the address: the gateway can pick a new exit for each connection, so a session's cookies may travel across IPs. For a flow that must keep one address, give that crawler a static IP, or use a residential sticky session; the session setting for your account is in the dashboard.

Specific to Crawlee

Things worth knowing

Pick the proxy per request with newUrlFunction

In place of proxyUrls, a newUrlFunction returns the proxy URL to use, or null for none. In Crawlee 3 it receives the session id and, where there is one, the request, so a label can route logins to a static IP and everything else to residential. Crawlee 4 calls it with no arguments, once per session, and routes requests through named sessions instead.

route.mjs
const proxyConfiguration = new ProxyConfiguration({
  newUrlFunction: (sessionId, { request } = {}) =>
    request?.label === "LOGIN" ? "http://USER:PASS@IP:PORT" : "http://USER:[email protected]:8000",
});

Block what the browser does not need

We meter request plus response bytes for every image, font and stylesheet a browser crawler loads. blockRequests in a pre-navigation hook drops them before they are sent; by default it blocks .css, common image formats, .woff, .pdf and .zip, and extraUrlPatterns adds your own. It works in Chromium only.

block.mjs
const crawler = new PlaywrightCrawler({
  proxyConfiguration,
  preNavigationHooks: [
    async ({ blockRequests }) => {
      await blockRequests({ extraUrlPatterns: [".webp", ".mp4"] });
    },
  ],
  async requestHandler({ page }) {
    console.log(await page.title());
  },
});

Estimate a crawl's bytes before you run it

Crawlee does not total bytes for you, and the body in an HTTP crawler's handler has already been decompressed, so its length overstates what a compressed page moved. Measure one typical page with curl as in the curl guide and multiply by the size of your queue. We meter request bytes plus response bytes, headers included; TLS overhead makes our figure slightly higher than curl's.

When it breaks

Common errors and fixes

  • Every request fails with a 407 or Proxy Authentication Required in the error

    Why
    The gateway rejected the credentials, or a password with @, : or / broke the proxy URL. Crawlee retries requests that fail on proxy errors on a fresh session, up to maxSessionRotations (10 by default), and other failures up to maxRequestRetries (3); neither helps when the credentials are wrong.
    Fix
    Build the URL with encodeURIComponent around the username and password. proxyInfo.username shows what Crawlee decoded from it; compare it with the dashboard.
  • request timed out after 30 seconds. from an HTTP crawler

    Why
    The download ran past navigationTimeoutSecs, 30 by default, on a slow residential exit or a heavy response. Depending on which timer fires, the message can quote the handler timeout, 60 seconds, instead; both mean the download stalled. Browser crawlers allow 60 seconds for navigation.
    Fix
    Keep the timeout short so dead exits fail fast and let the retry go out on a new connection, which on the rotating gateway leaves from a new address. Every retry is metered.
  • Request blocked - received 403 status code.

    Why
    The session pool treats 401, 403 and 429 as blocks: it retires the session and retries the request on another one.
    Fix
    Lower maxConcurrency or set maxRequestsPerMinute before adding retries. On the rotating gateway the retry leaves from a new address; on one static IP it cannot, so slow down instead.
  • net::ERR_CERT_COMMON_NAME_INVALID in a PlaywrightCrawler for a URL a CheerioCrawler fetched without complaint

    Why
    Crawlee's HTTP crawlers default to ignoreSslErrors: true, so they never report a bad certificate; the browser does. Neither involves the proxy: TLS runs through the tunnel to the target.
    Fix
    Treat it as a problem with the target. Set ignoreSslErrors: false on HTTP crawlers if you want them to fail the same way.

A failed connection that moved no data is not billed. Retries you send are billed like any other request, and each shows as its own line in your usage log.

The community layer

Crawlee still misbehaving?

Paste the error and the few lines that set up the proxy into Discord, with the password taken out. Someone there has seen it before.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk