Tutorial · 7 min read

Free proxy lists: what they actually cost you

A script that tests a free proxy list for you: how many answer, how many leak your real IP or rewrite pages, and when free is enough and when it is not.

Search for a free proxy list and you will find thousands of ip:port lines, refreshed hourly, sorted by country and “anonymity level”. They cost nothing to download. This post is about what they cost to use, and it comes with a script so you can measure that on the list in front of you instead of taking our word for it.

We sell proxies, so you are right to read this with one eyebrow up. That is why the script is the point: it prints numbers from your own list, and the conclusions are yours.

What is on a free proxy list, really?

Mostly machines that answer proxy requests from anyone: open proxies left running by mistake, misconfigured routers and servers, and a few run on purpose by people whose motives you cannot see. Nobody is paid to keep them up, so they come and go by the hour, and whoever runs one can see and change the traffic that passes through it.

Many of the owners never meant to offer a proxy at all. So the rule is plain: only use a proxy whose operator deliberately offers it to the public. Routing your traffic through someone’s misconfigured machine without their permission can be unauthorised access in many places, whatever the list that published it calls itself. A list rarely says which entries are which, and that alone is a reason to be wary of it. Our own acceptable use policy draws the same line for anyone using our network: nothing you are not authorised to use or test.

Are free proxies safe?

For anything involving a login, a payment or personal data, no. A proxy sits in the middle of your connection, and a free one is run by a stranger with no reason to protect you. For throwaway requests to public pages, the risk is smaller, but it is not zero.

  • Plain HTTP is fully visible. On an HTTP proxy, every unencrypted request and response passes through in the clear. The operator can read it, log it and rewrite it.
  • Injected content. Rewriting pages is easy on plain HTTP: an extra script tag, swapped ad code, a changed download link. You will not notice unless you compare.
  • HTTPS protects the content, not the fact of your visit. Through a tunnel the proxy cannot read the page, but it still sees which hosts you connect to and when. If you ever click past a certificate warning, the thing the warning was about may be exactly this: someone in the middle presenting their own certificate.
  • Logging. Assume everything that passes through is recorded, including proxy credentials if you reuse them elsewhere.
  • They may not hide you. Some forward your real IP to the site in a header, which defeats the reason you used them.
  • They die. A large share of any public list will not answer at all, and the ones that do are shared with everyone else who downloaded the same list, so many are already blocked where you want to go.

How to test a free proxy list yourself

The script below takes a text file of ip:port lines and checks each proxy a few at a time. It asks only the questions that matter most for a free list: does the proxy answer at all, did your exit IP actually change, did the proxy pass your real IP along anyway, and does a known page come back altered. Run it only on proxies whose operators publish them for public use; testing is still using.

It is deliberately short. Latency, the network each exit IP belongs to, and a full header-leak test are what our proxy checker tutorial covers, and it works on any list, free or paid.

It deliberately uses plain HTTP for the checks. Over HTTPS the proxy only relays an encrypted tunnel and cannot add headers or rewrite anything, so plain HTTP is where misbehaviour shows. That also means: do not send anything through these proxies while testing that you would mind a stranger reading.

Install the one dependency and save the script:

pip install aiohttp
import asyncio
import hashlib
import sys

import aiohttp

IP_URL = "http://httpbin.org/ip"
PAGE_URL = "http://example.com/"
TIMEOUT = aiohttp.ClientTimeout(total=10)
CONCURRENCY = 10


def hops(origin):
    return [h.strip() for h in origin.split(",") if h.strip()]


async def direct_baseline(session):
    async with session.get(IP_URL) as r:
        my_ip = hops((await r.json())["origin"])[-1]
    async with session.get(PAGE_URL) as r:
        page_hash = hashlib.sha256(await r.read()).hexdigest()
    return my_ip, page_hash


async def check(session, proxy, my_ip, page_hash, limit):
    url = f"http://{proxy}"
    async with limit:
        try:
            async with session.get(IP_URL, proxy=url) as r:
                chain = hops((await r.json(content_type=None))["origin"])
            async with session.get(PAGE_URL, proxy=url) as r:
                body = await r.read()
        except Exception:
            return {"proxy": proxy, "alive": False}

    return {
        "proxy": proxy,
        "alive": True,
        "new_ip": chain[-1] != my_ip,
        "leaks_you": my_ip in chain[:-1],
        "altered": hashlib.sha256(body).hexdigest() != page_hash,
    }


async def main(path):
    proxies = [line.strip() for line in open(path) if line.strip()]
    if not proxies:
        sys.exit(f"no proxies found in {path}")
    limit = asyncio.Semaphore(CONCURRENCY)
    async with aiohttp.ClientSession(timeout=TIMEOUT) as session:
        my_ip, page_hash = await direct_baseline(session)
        results = await asyncio.gather(
            *(check(session, p, my_ip, page_hash, limit) for p in proxies)
        )

    alive = [r for r in results if r["alive"]]
    print(f"tested {len(results)}, alive {len(alive)} ({len(alive) / len(results):.0%})")
    if not alive:
        return
    for key, label in [
        ("new_ip", "changed your exit IP"),
        ("leaks_you", "passed your real IP along"),
        ("altered", "returned a different example.com"),
    ]:
        hits = [r["proxy"] for r in alive if r[key]]
        print(f"{label}: {len(hits)} of {len(alive)}")
    clean = [r for r in alive if r["new_ip"] and not (r["leaks_you"] or r["altered"])]
    print(f"clean on every check: {len(clean)}")
    for r in clean:
        print(f"  {r['proxy']}")


if __name__ == "__main__":
    asyncio.run(main(sys.argv[1]))

Put the list in a file, one proxy per line, and run it:

python check_free.py proxies.txt

A few details that are easy to get wrong if you adapt it:

  • /ip reports the whole forwarding chain, comma separated. The last entry is the address that actually connected to httpbin, so that is the exit IP. Anything before it came from a header the proxy sent, and a transparent proxy puts your address there. Taking the first entry, or checking whether your IP appears anywhere, would report that proxy as not changing your IP at all, when it did change it and then told the site who you are.
  • The example.com check compares a hash of the whole page with the copy you fetched directly. The page is tiny and rarely changes, so a mismatch is worth a closer look, not a false alarm to shrug off.
  • CONCURRENCY is ten on purpose. httpbin and example.com are free public services, and hammering them with a thousand-line list at once is not a considerate rate.
  • aiohttp speaks to HTTP proxies. If your list is SOCKS, add the aiohttp-socks package and use its connector; the SOCKS5 glossary entry covers how the two protocols differ.
  • Ten seconds is generous. Anything slower than that is not a proxy you would want to scrape through, and it shows up as dead here, the same way it would as a connection timeout in a real job.

Reading the results

  • Alive percentage. How much of the list is usable right now. Run it again in an hour and compare; the churn is part of the cost.
  • Changed your exit IP. The bare minimum. If it did not, the proxy is doing nothing for you.
  • Passed your real IP along. The site you are visiting can see who you are. Drop these.
  • Returned a different example.com. The page is short and stable, so any difference is the proxy changing what you receive. Never send anything that matters through one of these.

Whatever is left in “clean on every check” is the honest size of the list. Then point that remainder at the site you actually care about; it may well block them, since everyone else found the same list. A proxy can also pass all three checks and still announce itself with a header like Via, or crawl; the full checker measures both.

Free vs paid proxies: what the money buys

Free proxy lists compared with the cheapest paid options
What you getFree listDatacenter, sharedDatacenter, dedicated
CostNothing$2.10 per IP for 30 days$3.20 per IP for 30 days
Who else uses itAnyone who downloaded the listUp to 3 customersOnly you
Who runs itUnknownA company you can emailA company you can email
Still there tomorrowMaybeFor the whole termFor the whole term
TrafficWhatever it survives1 GB, then $0.35/GBNo allowance to run out of
Our prices when this page was built; the datacenter pricing page has the current ones. Paid proxies can still be blocked by a site that dislikes datacenter ranges. What they stop is the stranger in the middle.

When free is enough, and when it isn’t

Everything here assumes a proxy whose operator deliberately offers it to the public. That rules out most of a scraped list before any of the cases below apply.

  • Learning. Seeing how a proxy changes your exit IP, or why a request fails, works just as well on a deliberately public proxy. Keep it to public test endpoints like httpbin.
  • A one-off look. Checking how a public page looks from somewhere else, once, through a proxy its operator offers openly, with nothing logged in.
  • Not enough: anything with a login, a payment or personal data, anything you need running tomorrow, and any machine you cannot tell was offered on purpose.

The rule underneath all of it: nothing you send through a free proxy should be anything you would mind a stranger reading or changing.

The cheapest paid step

When free stops being worth the time, the next step up is usually datacenter, not residential. A shared datacenter IP is $2.10 for 30 days, and the minimum top-up is $5, so that is your real first outlay. It is fast, it stays put, and a company answers for it. Residential starts at $5.50/GB and is worth it only when a site blocks datacenter IPs; test which one your target needs before you pay for the pricier one.

If you are not sure you need a proxy at all, four questions settle which type, if any. And if you want to vet paid proxies with more rigour than this script, the proxy checker tutorial adds latency, header leaks and the network each exit IP belongs to, which is how you check a provider’s labels, and it is built on aiohttp.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk