Tutorial · 8 min read

How much bandwidth will my scrape eat? Measure it before you pay

Measure what one page costs on the wire from your own connection, multiply it out, and know the gigabytes before you buy any.

Residential proxies are billed by the byte, and most people buy their first gigabytes with no idea how many their job needs. Then they either run out halfway or sit on a balance they bought on a guess.

You can find out before spending anything. A page weighs about the same whoever downloads it, so you can measure it from your own connection, multiply, and walk into the pricing page knowing the number.

What gets counted

Our meter counts request bytes plus response bytes, including headers and the protocol overhead on the tunnel. That is the rule in our terms. So measure both directions: what you send (the request line and headers, plus any body) and what comes back (the status line, headers and body as they crossed the wire, which is compressed if the site compressed it).

Two things add bytes on top of a single clean fetch. Redirects are separate requests. Retries you start are billed again, and they show up as their own rows in your usage log. Both belong in the estimate.

Step 1: measure a sample of pages

Put twenty to fifty URLs that look like your real job in sample.txt. Mix the page types: a listing page and a detail page can differ by an order of magnitude. Save this as pagesize.py and run it from your own machine, with no proxy:

import sys
import requests

HEADERS = {
    "User-Agent": (
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
        "(KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36"
    ),
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-GB,en;q=0.9",
    "Accept-Encoding": "gzip, deflate",
}


def header_bytes(start_line, headers):
    lines = [start_line] + [f"{k}: {v}" for k, v in headers.items()]
    return sum(len(line) + 2 for line in lines) + 2


def wire_size(url):
    with requests.get(url, headers=HEADERS, stream=True, timeout=30) as r:
        body = sum(len(chunk) for chunk in r.raw.stream(65536, decode_content=False))
        sent = header_bytes(f"GET {r.request.path_url} HTTP/1.1", r.request.headers)
        received = header_bytes(f"HTTP/1.1 {r.status_code} {r.reason}", r.raw.headers)
        return r.status_code, len(r.history), sent, received + body


urls = [line.strip() for line in open(sys.argv[1]) if line.strip()]
total = 0
for url in urls:
    status, hops, sent, received = wire_size(url)
    total += sent + received
    note = f"  ({hops} redirect)" if hops else ""
    print(f"{status}  {sent:>6} out  {received:>9} in  {url}{note}")

print(f"\n{len(urls)} pages, {total:,} bytes, {total / len(urls):,.0f} bytes per page")
python pagesize.py sample.txt

A few notes on what it does and does not see:

  • decode_content=False counts the body as it arrived. If the site gzipped it, you get the gzipped size, which is what a meter sees.
  • Send the same Accept-Encoding your scraper will send. A scraper that asks for no compression gets bigger responses, and pays for them.
  • Only the final response is counted. If the script flags redirects, each hop is another request and response on the real bill.
  • TLS handshakes are not counted. Expect a real meter, ours included, to read a few percent above this script.

Step 2: if you scrape with a browser, measure the browser

A page loaded in Playwright is the HTML plus every script, stylesheet, image, font and API call it pulls in, and all of it goes through the proxy. The HTML-only number from step 1 will be far too low. The Playwright guide has a byte counter that hooks every request the browser makes, plus asset blocking that brings the number down.

Step 3: multiply it out

Every figure below is yours to replace. Nothing here is a typical value; the point is the arithmetic.

bytes_per_page = 85_000   # from step 1
pages_per_run = 2_000     # how many pages one run fetches
runs_per_month = 30       # how often it runs
retry_share = 0.10        # share of requests you expect to retry

monthly_bytes = bytes_per_page * pages_per_run * runs_per_month * (1 + retry_share)
print(f"{monthly_bytes / 1e9:.2f} GB a month")

This uses 1 GB = 1,000,000,000 bytes to keep the numbers readable. If you want to know which gigabyte your bill uses, the bill audit guide shows how to work it out from your own usage export.

What that costs on residential

Here is the arithmetic for 10,000 pages at a few page sizes. The sizes are made up for illustration; the prices are our residential ladder when this page was built. You buy whole gigabytes, and the rate is set by the size of the order.

Residential cost of 10,000 pages at example page sizes
If one page costs youTrafficSmallest orderRateOrder total
50 KB0.50 GB1 GB$5.50/GB$5.50
250 KB2.50 GB3 GB$5.20/GB$15.60
1 MB10 GB10 GB$3.95/GB$39.50
3 MB30 GB30 GB$3.40/GB$102.00
Example page sizes only. Run step 1 on your own target and use your number.

If the target lets datacenter IPs through, compare that with a flat price. A dedicated datacenter address for 30 days is $3.20 and has no traffic allowance to run out of. A shared one is $2.10 with 1 GB included; more traffic is $0.35/GB when you add it with the order and $0.40/GB if you add it later. Your measured GB figure tells you which of the three is cheapest for your job.

Cutting the number before you buy

  • Ask for compression. Run step 1 with and without Accept-Encoding: gzip. On HTML the difference is often large, and it costs you nothing.
  • Look for the JSON. Open the browser’s network tab on the target. Many pages fill themselves from a JSON endpoint that is a fraction of the HTML’s size. If it serves the public data you want, fetch that.
  • Do not download what you do not parse. In a browser, block images, media and fonts. Requests you abort never reach the proxy.
  • Skip unchanged pages. Conditional requests get a bodiless 304 when nothing changed. The price tracking guide shows the headers.
  • Stop retrying 403s. Every retry is a billed request, and a 403 rarely changes on the second ask. The error field guide covers which codes are worth retrying.

Once you have a monthly figure, the residential pricing page shows where it lands on the ladder. Buy that, and keep running the byte counter so you notice when the job grows.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk