Moving from requests to a real browser fixes a lot: the TLS fingerprint is a browser’s, JavaScript runs, and challenge pages that only check for a working browser go away. It also multiplies the bytes behind every page, and on a per-gigabyte proxy that is money.
This guide sets up Playwright with an authenticated proxy in Python and Node, then covers the parts that decide whether the site treats you as a considerate visitor: pace, consistency and what you download.
What “without getting flagged” means here
It means looking like one consistent, unhurried visitor, because that is what you are being. It does not mean solving CAPTCHAs, getting past a site that has told you to stop, or logging in to accounts that are not yours. Our acceptable use policy is short: public data at a considerate rate is fine, credential stuffing and overloading a site are not, and scraping behind a login you hold is a grey area to ask us about first. If a site puts a CAPTCHA in front of you, it is asking for a human. This guide stops there.
Python: launch Chromium through the proxy
pip install playwright
playwright install chromiumfrom playwright.sync_api import sync_playwright
PROXY = {
"server": "http://resi.proxymonkey.io:8000",
"username": "USER",
"password": "PASS",
}
with sync_playwright() as p:
browser = p.chromium.launch(proxy=PROXY)
page = browser.new_page()
page.goto("https://httpbin.org/ip")
print(page.inner_text("body"))
browser.close()The credentials go in their own fields. Chromium ignores a username and password written into the server URL, and the symptom is a 407 that looks like a wrong password. The error field guide covers the rest of the 407 family.
Node: the same thing
import { chromium } from "playwright";
const browser = await chromium.launch({
proxy: {
server: "http://resi.proxymonkey.io:8000",
username: "USER",
password: "PASS",
},
});
const page = await browser.newPage();
await page.goto("https://httpbin.org/ip");
console.log(await page.textContent("body"));
await browser.close();Rotating or static, in a browser
One page load is dozens of requests: the HTML, then scripts, styles, images and API calls. On the rotating residential endpoint those can leave from different addresses, and a page whose API calls arrive from somewhere other than its HTML is odd to a site that checks. For scraping independent public pages that is often fine. For anything with a session, use a static ISP or datacenter address and launch one browser per address; the rotating vs sticky guide has the reasoning and the code for pinning.
Count the bytes the browser moves
We meter request bytes plus response bytes, including headers, and a browser sends and receives far more of both than you would guess. Playwright can report the sizes of every finished request:
moved = {"bytes": 0, "requests": 0}
def count(request):
sizes = request.sizes()
moved["requests"] += 1
moved["bytes"] += sum(
max(0, sizes[key])
for key in ("requestHeadersSize", "requestBodySize",
"responseHeadersSize", "responseBodySize")
)
page.on("requestfinished", count)Print moved after a handful of pages and you have the real bytes per page for your target. For scale: if your pages came out at 2 MB each (an example figure; measure yours), 1,000 of them would need 2 GB, which is $10.40 on our current residential ladder. The bandwidth guide has the full arithmetic.
Block what you do not need
Images, video and fonts are usually most of a page’s weight and none of the data you want. A request aborted inside the browser never reaches the proxy, so it costs nothing:
BLOCK = {"image", "media", "font"}
def block_heavy(route):
if route.request.resource_type in BLOCK:
route.abort()
else:
route.continue_()
page.route("**/*", block_heavy)Leave stylesheets and scripts alone unless you have checked the page still renders the data without them. And watch the result: a few sites notice a browser that never fetches an image. If challenge pages start appearing after you add the block, take images out of the set.
Stay consistent, and slow
- One context, one identity. Keep the viewport, locale and user agent fixed for the life of a context. Changing them per page makes you less like a person.
- Match the address. An ISP or datacenter address is bought for a country you pick at checkout, so set
localeandtimezone_idon the context to match it. A German address with a browser insisting it is in California is a mismatch a site can see. - Check what headless announces. Run
page.evaluate("navigator.userAgent")in your own headless session and read it. If it names a headless browser, every request you send says so too. - Pause like a reader. A random few seconds between navigations, one page at a time per site. Twenty tabs on one hostname is the pattern that gets IPs banned for everyone who shares the pool.
Ask robots.txt first
robots.txt is the site’s stated preference for automated visitors, including how fast. Python can read it with the standard library:
from urllib import robotparser
from urllib.parse import urlsplit
def robots_for(url):
parts = urlsplit(url)
rp = robotparser.RobotFileParser(f"{parts.scheme}://{parts.netloc}/robots.txt")
rp.read()
return rpAll together
import random
import sys
import time
from urllib import robotparser
from urllib.parse import urlsplit
from playwright.sync_api import sync_playwright
PROXY = {"server": "http://resi.proxymonkey.io:8000", "username": "USER", "password": "PASS"}
BLOCK = {"image", "media", "font"}
moved = {"bytes": 0, "requests": 0}
def count(request):
sizes = request.sizes()
moved["requests"] += 1
moved["bytes"] += sum(max(0, v) for v in sizes.values())
def block_heavy(route):
if route.request.resource_type in BLOCK:
route.abort()
else:
route.continue_()
def main(urls):
host = urlsplit(urls[0])
robots = robotparser.RobotFileParser(f"{host.scheme}://{host.netloc}/robots.txt")
robots.read()
delay = robots.crawl_delay("*") or 3
with sync_playwright() as p:
browser = p.chromium.launch(proxy=PROXY)
context = browser.new_context(locale="en-GB", viewport={"width": 1366, "height": 850})
page = context.new_page()
page.route("**/*", block_heavy)
page.on("requestfinished", count)
for url in urls:
if not robots.can_fetch("*", url):
print(f"skip (robots.txt) {url}")
continue
response = page.goto(url, wait_until="domcontentloaded", timeout=45_000)
print(f"{response.status if response else '---'} {page.title()[:60]} {url}")
time.sleep(delay + random.uniform(0, 2))
browser.close()
print(f"{moved['requests']} requests, {moved['bytes'] / 1e6:.2f} MB through the proxy")
if __name__ == "__main__":
main([line.strip() for line in open(sys.argv[1]) if line.strip()])It reads URLs from a file (one site per run, since it reads one robots.txt), skips what robots.txt disallows, waits at least the crawl delay the site asks for, and tells you what the run cost in bytes. That last line is the one to compare against your usage log.
Stuck on a site that keeps challenging you after all of this? Ask in Discord before buying anything. The answer is sometimes a different proxy type and sometimes a different approach to the data.
Top-ups start at $5.
One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.