Proxies for AI training data
What changes with one
Datacenter for the bulk
Most sites let datacenter IPs through. A handful of rented IPs, each with its own concurrency limit, gets through a large crawl for a fixed price per term.
Residential only for the stubborn ones
Send just the domains that block datacenter through residential, and pay per GB for that slice alone.
See which domains cost you
Every request is in your usage log with the bytes billed. Sort by host and drop the domains that cost more than they are worth to the dataset.
When it will not help
A proxy says nothing about whether you may use the text. Copyright and licences apply to the content, and robots.txt and a site's terms tell you what its owner wants. Content behind a login is out of reach for an honest crawler either way.
What to run it on
Datacenter
The bulk of a broad crawl hits sites that never check where you come from. Rented IPs cost the same however many pages they fetch.
from $1.50/IP/mo
Residential
For the minority of domains that block hosting ranges. Route only those through it so the per-GB bill stays small.
from $1.75/GB
What it would cost
Worked out from today’s price list with the assumptions shown. Your pages will weigh something else, so measure fifty of them and redo the sum.
The bulk of a broad text crawl, spread across a few rented IPs with a per-domain rate limit
IPs, dedicated plan5
Term30 days
Rate at that size$2.94/IP/mo
Cost for the term$14.70
Priced from the per-IP ladder for a 30-day term, before optional add-ons for speed and threads, which checkout lists. Longer terms take up to 20% off. The shared plan costs $2.10/IP/mo, is shared by up to 3 customers, and includes 1 GB of traffic, then $0.35/GB. The country of each datacenter IP is chosen at checkout.
The slice of that crawl that blocks datacenter: twenty thousand pages
Requests20,000
Weight each~80 KB
Traffic≈ 1.6 GB
Top-up that covers it2 GB
Rate at that size$5.20/GB
Cost$10.40
Assumes text-heavy HTML, compressed on the wire, no images, headers counted. We meter request bytes plus response bytes. The rate is set by the size of one top-up and applies to all of it, and the balance does not expire.
What is off limits
- Crawling a small site at a volume that slows it down for its readers
- Collecting information about particular people to track or contact them. Harassment and spam are both on the not-fine list.
- High-volume crawling of a small site is on the grey list. Throttle per domain, and ask if unsure.
From our acceptable use policy. We close accounts that break it, and refund the unused balance when we do.
AI training data questions
Should I respect robots.txt?
We would. It is the cheapest way to learn what a site's owner is fine with, and it keeps your crawl at the considerate rate our acceptable use policy asks for. Crawlers that ignore it are how ranges end up on blocklists for everyone.
How do I keep costs down on a big crawl?
Fetch HTML only, deduplicate URLs before fetching, and cache everything while you develop. Send only the domains that block datacenter through residential. Check the usage log after the first thousand pages and extrapolate before you commit to the rest.
Is there anything cheaper than crawling it myself?
Often, yes. If the pages you need are in a public crawl archive such as Common Crawl, downloading that costs less than any proxy. Crawl yourself for the gaps and for fresh pages.
Is this allowed for a university project?
Yes. Academic research is on the fine list in our acceptable use policy. Your university may have its own ethics rules for collected data, and those are yours to follow.
Other use cases
Ad verification
See the ads, landing pages and redirects a real visitor in another country gets.
Read →Geo testing
Check your own site's redirects, currencies, cookie banners and CDN from another country.
Read →Web scraping
Your scraper works on your laptop and dies at page 300. What fixes that, and what it costs.
Read →
Stuck on AI training data?
Post the target and the error in Discord. Somebody has usually hit the same wall and will tell you whether a proxy is even the fix.
Join the Discord4,200+monkeys in the Discord
Help from humans
Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.
A status bot that tells on us
Pool health, incidents and maintenance posted automatically. Including the bad days.
Deals and free traffic
Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.