On a residential proxy you pay for bytes, not pages. So the cheapest way to reduce proxy bandwidth is not a better rate; it is fetching the same pages with fewer bytes. Most scrapers download a lot they never parse: uncompressed HTML, hero images, fonts, analytics scripts, the same page twice.
Here are nine techniques that cut proxy bandwidth usage, with code for each. None of them needs a new tool. Measure before and after, because which ones help depends entirely on your target.
How to reduce proxy bandwidth: measure first
Count the bytes per page now, change one thing, count again. We bill request bytes plus response bytes, headers included, so the counter has to include both directions and the body as it crossed the wire. The bandwidth measuring guide has a full script; the core of it is this:
import requests
def header_bytes(start_line, headers):
lines = [start_line] + [f"{k}: {v}" for k, v in headers.items()]
return sum(len(line) + 2 for line in lines) + 2
def wire_size(session, url, **kwargs):
with session.get(url, stream=True, timeout=30, **kwargs) as r:
body = sum(len(c) for c in r.raw.stream(65536, decode_content=False))
sent = header_bytes(f"GET {r.request.path_url} HTTP/1.1", r.request.headers)
got = header_bytes(f"HTTP/1.1 {r.status_code} {r.reason}", r.raw.headers)
return sent + got + bodyRun it over twenty or so representative URLs from your own connection, no proxy needed: a page weighs about the same whoever fetches it. For a browser-based scraper, use the requestfinished counter from the Playwright guide instead.
Techniques that shrink every response
1. Ask for compression, and check you got it
HTML and JSON compress well, and the meter sees the compressed size. requests and httpx ask for gzip by default. The trap is custom headers: a scraper that copies a header set from somewhere and sends Accept-Encoding: identity, or leaves the header out, gets the full-size body. Check what arrived:
curl -s -o /dev/null -w '%{size_download}\n' https://example.com/
curl -s -o /dev/null -w '%{size_download}\n' --compressed https://example.com/If the two numbers are the same, the site does not compress that page and there is nothing to win. If they differ, look at the Content-Encoding header on your scraper’s own responses. One more trap: if you send br in Accept-Encoding, install the brotli package, or your client gets a body it cannot decode.
2. Block images, media, fonts and stylesheets in Playwright
In a browser, the HTML is often the small part. Abort what you do not parse, and it never reaches the proxy:
from urllib.parse import urlsplit
BLOCK_TYPES = {"image", "media", "font", "stylesheet"}
FIRST_PARTY = "example.com"
def lean(route):
request = route.request
host = urlsplit(request.url).hostname or ""
third_party = not (host == FIRST_PARTY or host.endswith("." + FIRST_PARTY))
if request.resource_type in BLOCK_TYPES or third_party:
route.abort()
else:
route.continue_()
page.route("**/*", lean)Stylesheets are in the list here, and third-party hosts are blocked too (that is technique 7, below). Test with each: some pages only fill in their data after a stylesheet or a third-party script loads. If a site behaves differently once you block its resources, respect that: slow down or fetch those resources, and consider whether you should be scraping it at all.
3. Find the JSON endpoint instead of rendering the page
Many pages are an empty shell that fills itself from an API. Open the page in your browser, then DevTools, Network, filter by Fetch/XHR, and reload. If one of those responses has the public data you want, fetch that directly:
import requests
PROXY = "http://USER:[email protected]:8000"
r = requests.get(
"https://example.com/api/products",
params={"category": "shoes", "page": 2},
headers={"Accept": "application/json"},
proxies={"http": PROXY, "https": PROXY},
timeout=25,
)
items = r.json()That swaps a browser, its scripts and a rendered page for one compressed JSON response. Keep to the endpoints the public page itself calls, at the pace a visitor would.
Techniques that skip downloads entirely
4. Conditional requests: ETag and If-Modified-Since
On a recrawl, most pages have not changed. Send back the validators the server gave you last time, and an unchanged page comes back as a bodiless 304 Not Modified:
import json
import pathlib
import requests
STORE = pathlib.Path("validators.json")
seen = json.loads(STORE.read_text()) if STORE.exists() else {}
def fetch_if_changed(session, url):
headers = {}
if url in seen:
if seen[url].get("etag"):
headers["If-None-Match"] = seen[url]["etag"]
if seen[url].get("modified"):
headers["If-Modified-Since"] = seen[url]["modified"]
r = session.get(url, headers=headers, timeout=25)
if r.status_code == 304:
return None
seen[url] = {"etag": r.headers.get("ETag"), "modified": r.headers.get("Last-Modified")}
STORE.write_text(json.dumps(seen))
return rA 304 still costs a request and response header, so it is small, not free. Not every site sends validators; the ones that do are usually the ones serving big pages. The price tracking guide uses this on a daily recrawl.
5. HEAD and Range requests
When you only need to know whether a file changed or how big it is, a HEAD request returns the headers without the body. When the data you want sits at the start of a large response, ask for just that part:
head = session.head(url, allow_redirects=True, timeout=25)
size = int(head.headers.get("Content-Length", 0))
with session.get(url, headers={"Range": "bytes=0-32767"}, stream=True, timeout=25) as r:
start = b""
for chunk in r.iter_content(8192):
start += chunk
if len(start) >= 32768:
breakA server that supports ranges answers 206 Partial Content. One that ignores the header sends 200 and the full body, which is why the loop stops reading after 32 KB and closes the connection. Bytes already in flight when you close are still bytes the proxy carried, so the saving is real but not exact.
6. Cache and dedupe URLs
The same page often appears under several URLs: tracking parameters, fragments, different parameter order. Normalise before you queue:
from urllib.parse import parse_qsl, urlencode, urlsplit, urlunsplit
DROP = {"utm_source", "utm_medium", "utm_campaign", "utm_term", "utm_content",
"gclid", "fbclid", "ref"}
def normalise(url):
parts = urlsplit(url)
query = sorted((k, v) for k, v in parse_qsl(parts.query) if k not in DROP)
return urlunsplit((parts.scheme, parts.netloc.lower(), parts.path or "/",
urlencode(query), ""))
queue = list(dict.fromkeys(normalise(u) for u in urls))Then cache responses for the length of a run. For requests, requests-cache does it without changing your fetch code. Only the parameters you are sure change nothing belong in the drop set; on some sites ref is a real product ID.
Techniques that stop waste
7. Do not follow redirects you could have skipped
Every redirect is another request and response through the proxy. If your URL list says http:// and the site sends everyone to https://www., you pay for a hop on every page. Check once, then fix the list:
r = requests.get(url, allow_redirects=False, timeout=25)
if r.is_redirect:
print(r.status_code, url, "->", r.headers["Location"])In a browser, the third-party block in technique 2 stops the ad, analytics and tracking requests that a page fans out to, and the redirect chains some of them start.
8. Retry smartly
A retry is a full, billed request. Retry what can change (timeouts, 429, 502, 503, 504), back off between attempts, and leave 403 and 404 alone:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util import Retry
retry = Retry(
total=3,
backoff_factor=1.0,
status_forcelist=[429, 502, 503, 504],
allowed_methods=["GET", "HEAD"],
respect_retry_after_header=True,
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))Which codes are worth a second try, and why a 403 rarely is, is in the error field guide.
Save residential proxy data by splitting traffic
9. Send only the pages that need residential through it
Plenty of sites serve their pages to a datacenter IP. Those pages do not need to be on the per-gigabyte meter at all. Try the static line first per host, and move a host to residential only when it refuses:
from urllib.parse import urlsplit
import requests
DATACENTER = "http://USER:PASS@IP:PORT"
RESIDENTIAL = "http://USER:[email protected]:8000"
needs_residential = set()
def fetch(session, url):
host = urlsplit(url).hostname
if host not in needs_residential:
r = session.get(url, proxies={"http": DATACENTER, "https": DATACENTER}, timeout=25)
if r.status_code != 403:
return r
needs_residential.add(host)
return session.get(url, proxies={"http": RESIDENTIAL, "https": RESIDENTIAL}, timeout=25)Per host is coarse; some sites block datacenter IPs on search pages and not on product pages, so split by path if yours does. The five-minute test tells you which side of the line a site is on. A dedicated datacenter address is $3.20 for 30 days with no traffic allowance to run out of, and every page it serves is a page you did not pay $5.50/GB (or $1.75/GB at volume) to move.
Which technique to try first
| Technique | Effort | Typical impact | Applies to |
|---|---|---|---|
| 1. Ask for compression | Minutes | Large on HTML and JSON | Every client |
| 2. Block images, media, fonts | Minutes | Often the largest | Browsers |
| 3. Use the JSON endpoint | An hour of digging | Large when one exists | Browsers and HTTP |
| 4. Conditional requests | Small | Large on repeat runs | Recrawls |
| 5. HEAD and Range | Small | Situational | Big files, change checks |
| 6. Cache and dedupe URLs | Small | Depends on your URL list | Every client |
| 7. Skip redirects and third parties | Small | Moderate | Browsers mostly |
| 8. Retry smartly | Small | Moderate to large on bad days | Every client |
| 9. Split traffic by line | Medium | Takes pages off the meter | Mixed targets |
Check the saving before you believe it
Run the byte counter over the same URL list with and without each change, and compare bytes per page. Change one thing at a time; two at once and you will not know which helped. Then watch your usage log for a day of real traffic, because the counter does not see TLS handshakes and our meter does. The bandwidth calculator turns your new bytes-per-page figure into a monthly estimate.
Top-ups start at $5.
One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.