Tor is free and gives you a different address every few minutes, so it looks like a scraping proxy that costs nothing. It is a bad one. Every Tor exit address is published in a list that websites load straight into their firewalls, a request takes many times longer than it would direct, and the network is run by volunteers whose bandwidth is meant for people who need privacy, not for your crawl. Tor is the right tool for keeping a person private. For collecting pages, use something built for it.
The behaviour below comes from running Tor 0.4.9 in a container and sending curl through it. We are a proxy company, so weigh our conclusion accordingly; the commands are here so you can check it.
Why Tor is a poor choice for web scraping
Every exit is on a public list
Traffic leaves the Tor network through exit relays, and the Tor Project publishes their addresses, one per line, for anyone to download:
curl -s https://check.torproject.org/torbulkexitlist | head -3That list exists so services can recognise Tor visitors, and plenty use it to block them, put them behind a challenge, or lock an account that suddenly signs in from one. Across our test requests we saw eight different exit addresses, and every one of them was on the list. Getting a new circuit does not help, because the next exit is on the same list.
It is slow, by design
A Tor circuit bounces your traffic through three relays run by different people, often in different countries. In our test, a small page took between half a second and 1.3 seconds through Tor and about 35 milliseconds direct, and that was with circuits already built. Multiply by a crawl and the difference is hours.
You cannot hold an address
Tor reuses a circuit for new connections for MaxCircuitDirtiness, 600 seconds by default, then moves on. That is good for privacy and bad for anything that logs in, keeps a cart, or needs to look like the same visitor tomorrow. There is no such thing as a static Tor address you can come back to.
It is somebody else’s donated bandwidth
Relays are run by volunteers, and the Tor Project’s own relay pages say the network is small compared with the number of people who need it. A crawl takes capacity from journalists and people on hostile networks who have no alternative, and the abuse complaints for what leaves an exit go to the volunteer who runs it. That is rude, even when it is legal.
Tor vs a proxy, side by side
| Tor | A paid proxy | Your own VPS proxy | |
|---|---|---|---|
| Costs you | Nothing | Per GB or per IP | Your server bill |
| Who runs the exit | A volunteer | The provider you pay | You |
| Exit addresses | Every one in a public list | The pool you buy | One, your server’s |
| Keep one address | No; circuits change every 10 minutes or so | Yes, on static IPs | Yes |
| Speed in our test | Half a second to 1.3 s for a small page | Depends on the line | Close to direct |
| Built for | A person’s privacy | Your traffic, billed to you | Your traffic |
If “free” was the attraction, the honest free options are your own connection with polite pacing, or a VPS you already have; the DIY proxy server guide sets one up, and free proxy lists explains why the other free route costs more than it looks.
When Tor is the right tool
- Reading and research where you do not want to be identified, by the site or by the network you are on. That is what Tor Browser is for, and it does the job far better than any proxy.
- Reaching onion services, which only exist inside the Tor network.
- Checking how your own site treats Tor visitors. If you block or challenge Tor, look at what a Tor user actually gets, from Tor.
- A one-off look at a page from somewhere other than home, at human speed.
What these have in common: one person, a few pages, privacy as the point. A scraper is none of those.
How to run Tor’s SOCKS5 proxy locally
For the jobs above, or to test your own site, run the Tor client and talk to its SOCKS port. On Debian or Ubuntu the package is tor, and the service listens on 127.0.0.1:9050. Tor Browser runs its own copy on port 9150 instead. Check it:
sudo apt install tor
curl -s -x socks5h://127.0.0.1:9050 https://check.torproject.org/api/ipThe answer is a line of JSON with "IsTor":true and the exit address. Keep the SOCKS port on 127.0.0.1; bound to a public address, it is a free door into Tor for anyone who finds it. These are the settings we used in /etc/tor/torrc:
SocksPort 127.0.0.1:9050
HTTPTunnelPort 127.0.0.1:9080Three behaviours worth knowing
- The SOCKS port is not an HTTP proxy. Point an HTTP proxy setting at 9050 and Tor answers
501with a page titled This is a SOCKS Proxy, Not An HTTP Proxy, and logsSocks version 71 not recognized. (This port is not an HTTP proxy; did you want to use HTTPTunnelPort?). - HTTPTunnelPort only tunnels. With the line above,
curl -x http://127.0.0.1:9080worked for anhttps://address, which uses CONNECT. A plainhttp://request through the same port got an empty reply. For tools that only speak HTTP proxy, Privoxy’sforward-socks5t / 127.0.0.1:9050 .line, from its own sample config, fills the gap; we tested that too. - Use socks5h, not socks5. With
socks5://, curl looked the name up itself and handed Tor only an IP, and Tor loggedYour application (using socks5 to port 443) is giving Tor only an IP address. Applications that do DNS resolves themselves may leak information.The socks5h scheme sends the name through Tor instead.
Tor also keeps different SOCKS usernames on different circuits, which is how privacy tools keep one identity’s traffic away from another’s; three usernames got us three different exits. And its control port has a NEWNYM signal that switches new connections to clean circuits, which the control protocol notes Tor may rate-limit. Neither changes the verdict above: every circuit still ends at an address on the public list.
What to scrape with instead
Start with no proxy at all, a sensible delay and a real User-Agent; not getting blocked is the checklist. When a site starts limiting your own address, move to something that is for this job. With us, that means datacenter IPs for sites that accept hosting ranges, and residential ones, billed by the gigabyte, for sites that do not. The proxy type flowchart picks between them in four questions. Leave Tor to the people who need it; Mo insists, and he is right about most things that are not bananas.
Top-ups start at $5.
One shared datacenter IP for 30 days is $1.25. A single gigabyte of residential is $5.50. The balance never expires.