When a request through a proxy fails, the error came from one of two places: the proxy, or the site behind it. The fixes have nothing in common, and most of the time people lose is spent fixing the wrong one. So before anything else, find out who answered.
Step 0: who said no?
For an https:// URL, your client first asks the proxy to open a tunnel with a CONNECT request. If the proxy refuses, the site never saw you. curl can print both answers separately: %{http_connect} is the proxy’s reply to the tunnel request and %{http_code} is the site’s reply.
curl -s -o /dev/null \
-w 'proxy said: %{http_connect} site said: %{http_code}\n' \
-x http://USER:[email protected]:8000 \
https://httpbin.org/status/429proxy said: 200 site said: 429means the tunnel opened fine and the site is rate-limiting you.proxy said: 407 site said: 000means the proxy rejected your credentials and the request never left. curl also exits with code 56, “CONNECT tunnel failed”.
Add -v when you want the full conversation. Lines starting with < before the TLS handshake are the proxy talking; after it, the site.
407 Proxy Authentication Required
Always the proxy, never the site. Three causes cover nearly every case:
- Wrong username or password. Copy them again. A trailing space from a copy-paste counts.
- A special character broke the URL. An
@,:,/,#or%in a password changes how the proxy URL is parsed. Percent-encode both parts:
from urllib.parse import quote
USER = "your-username"
PASSWORD = "p@ss:word/with#junk"
PROXY = f"http://{quote(USER, safe='')}:{quote(PASSWORD, safe='')}@resi.proxymonkey.io:8000"- Your tool ignores credentials in the URL. Browsers are the usual culprit. Chromium does not accept a username and password inside the proxy server string, so Playwright and Puppeteer take them separately. The Playwright guide shows the shape.
Two related traps. In Python, if you set only proxies={"http": ...} and request an https:// URL, requests finds no "https" key and sends the request without your proxy: straight from your own IP, or through whatever proxy your environment variables name. Set both keys. And the proxy URL starts with http:// even when the site is https://. Writing https://USER:[email protected]:8000 makes your client try to speak TLS to the proxy itself, and you get an SSL error that looks like it has nothing to do with proxies.
403 Forbidden
If the proxy said 200 and the site said 403, the site refused you on purpose. The usual reasons, cheapest first: your headers announce a script, your TLS fingerprint does not match the browser you claim to be, or the site filters datacenter ranges. The blocking checklist goes through them in order, and the residential vs datacenter test tells you whether the address type is the problem.
Do not retry a 403 in a loop. It rarely changes on the second ask, and every retry you start is billed as its own request: our terms list them as separate entries in your usage log. Read the body of the first one instead. A challenge page and a plain “access denied” are different problems.
If it was the proxy that returned the 4xx to your CONNECT, the site is not involved and the question is for us. Post the hostname and the curl output in Discord.
429 Too Many Requests
The site’s rate limit. It often tells you how long to wait in a Retry-After header, as a number of seconds or as a date. Honour it, and fall back to exponential backoff with jitter when it is missing:
import random
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
def retry_after_seconds(response, attempt):
value = response.headers.get("Retry-After", "").strip()
if value.isdigit():
return int(value)
if value:
try:
when = parsedate_to_datetime(value)
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
except (TypeError, ValueError):
pass
return 2 ** attempt + random.random()On the rotating residential endpoint your next request leaves from a different address anyway, so a 429 there usually means the limit is not per IP. It may be per session cookie, per account, or across the whole site. More IPs will not fix that. Slowing down will.
502, 503, 504 and timeouts
Check the tunnel code again. From the proxy, these mean the exit could not reach the site or gave up waiting; retry once after a pause. From the site, the site is struggling, and hammering it harder is how a scraper ends up on the wrong side of our acceptable use policy, which rules out volume high enough to degrade a service for its real users.
On billing: connections that fail before transferring any data are not billed. Retries you start are, each as its own entry. Both rules are in the terms, next to the main one: we meter request bytes plus response bytes.
A diagnosis function you can paste
This turns the common failures into a one-line verdict. Drop it into a scraper’s error path and log what it returns.
import requests
PROXY = "http://USER:[email protected]:8000"
PROXIES = {"http": PROXY, "https": PROXY}
def diagnose(url):
try:
r = requests.get(url, proxies=PROXIES, timeout=25)
except requests.exceptions.ProxyError as exc:
if "407" in str(exc):
return "proxy 407: check credentials and percent-encoding"
return f"proxy refused or unreachable: {exc}"
except requests.exceptions.SSLError as exc:
return f"TLS error, is the proxy URL http:// ? {exc}"
except requests.exceptions.ConnectTimeout:
return "timed out connecting: retry once, then ask in Discord"
except requests.exceptions.ReadTimeout:
return "site was slow: raise the timeout before assuming a block"
if r.status_code == 407:
return "proxy 407 on a plain http:// URL: check credentials"
if r.status_code == 403:
return "site 403: fix headers, then test datacenter vs residential"
if r.status_code == 429:
return f"site 429: wait {r.headers.get('Retry-After', 'and back off')}"
if r.status_code >= 500:
return f"site {r.status_code}: retry later with backoff"
return f"ok {r.status_code}"
print(diagnose("https://httpbin.org/status/429"))For https:// URLs a proxy 407 arrives as a ProxyError whose message contains “Tunnel connection failed: 407”, which is why the function checks the text. For plain http:// URLs there is no tunnel, and the 407 comes back as an ordinary response.
If none of this matches what you are seeing, paste the curl -v output in Discord with your password removed. With that output, whoever answers can see who said no.
Top-ups start at $5.
One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.