Use case

Proxies for AI training data

You are fine-tuning a small model, building a RAG demo, or collecting a dataset for a thesis. That means a lot of pages from a lot of sites, most of which do not care who is asking. The trick is to spend proxy money only on the ones that do.
Sound familiar?

What goes wrong

A corpus crawl touches thousands of domains. Most serve anyone. Some block hosting ranges, and a few big ones block nearly everything automated.

  • A long crawl stalls halfway because a few domains start returning 403s
  • Your cloud VM's range is on a blocklist for some sites
  • Rate limits on one IP turn a week-long crawl into a month-long one
  • Retries against blocked pages quietly eat the budget
How a proxy helps

What changes with one

Datacenter for the bulk

Most sites let datacenter IPs through. A handful of rented IPs, each with its own concurrency limit, gets through a large crawl for a fixed price per term.

Residential only for the stubborn ones

Send just the domains that block datacenter through residential, and pay per GB for that slice alone.

See which domains cost you

Every request is in your usage log with the bytes billed. Sort by host and drop the domains that cost more than they are worth to the dataset.

When it will not help

A proxy says nothing about whether you may use the text. Copyright and licences apply to the content, and robots.txt and a site's terms tell you what its owner wants. Content behind a login is out of reach for an honest crawler either way.

Residential vs datacenter: the five-minute test →

Which line

What to run it on

Start here

Datacenter

The bulk of a broad crawl hits sites that never check where you come from. Rented IPs cost the same however many pages they fetch.

from $1.50/IP/mo

Or

Residential

For the minority of domains that block hosting ranges. Route only those through it so the per-GB bill stays small.

from $1.75/GB

Back of the envelope

What it would cost

Worked out from today’s price list with the assumptions shown. Your pages will weigh something else, so measure fifty of them and redo the sum.

Datacenter, by the IP

The bulk of a broad text crawl, spread across a few rented IPs with a per-domain rate limit

IPs, dedicated plan5

Term30 days

Rate at that size$2.94/IP/mo

Cost for the term$14.70

Priced from the per-IP ladder for a 30-day term, before optional add-ons for speed and threads, which checkout lists. Longer terms take up to 20% off. The shared plan costs $2.10/IP/mo, is shared by up to 3 customers, and includes 1 GB of traffic, then $0.35/GB. The country of each datacenter IP is chosen at checkout.

Residential, by the GB

The slice of that crawl that blocks datacenter: twenty thousand pages

Requests20,000

Weight each~80 KB

Traffic≈ 1.6 GB

Top-up that covers it2 GB

Rate at that size$5.20/GB

Cost$10.40

Assumes text-heavy HTML, compressed on the wire, no images, headers counted. We meter request bytes plus response bytes. The rate is set by the size of one top-up and applies to all of it, and the balance does not expire.

In code

The short version

Swap in your own credentials from the dashboard. It also lists the host and port for each order; where they differ from this example, the dashboard is right.

Full Go (Golang) setup →

main.go
proxy, _ := url.Parse("http://USER:PASS@IP:PORT")
client := &http.Client{
	Timeout:   30 * time.Second,
	Transport: &http.Transport{Proxy: http.ProxyURL(proxy)},
}

resp, err := client.Get("https://example.com/article/1")
if err != nil {
	log.Fatal(err)
}
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
fmt.Println(resp.StatusCode, len(body))

What is off limits

  • Crawling a small site at a volume that slows it down for its readers
  • Collecting information about particular people to track or contact them. Harassment and spam are both on the not-fine list.
  • High-volume crawling of a small site is on the grey list. Throttle per domain, and ask if unsure.

From our acceptable use policy. We close accounts that break it, and refund the unused balance when we do.

AI training data questions

Should I respect robots.txt?

We would. It is the cheapest way to learn what a site's owner is fine with, and it keeps your crawl at the considerate rate our acceptable use policy asks for. Crawlers that ignore it are how ranges end up on blocklists for everyone.

How do I keep costs down on a big crawl?

Fetch HTML only, deduplicate URLs before fetching, and cache everything while you develop. Send only the domains that block datacenter through residential. Check the usage log after the first thousand pages and extrapolate before you commit to the rest.

Is there anything cheaper than crawling it myself?

Often, yes. If the pages you need are in a public crawl archive such as Common Crawl, downloading that costs less than any proxy. Crawl yourself for the gaps and for fresh pages.

Is this allowed for a university project?

Yes. Academic research is on the fine list in our acceptable use policy. Your university may have its own ethics rules for collected data, and those are yours to follow.

The community layer

Stuck on AI training data?

Post the target and the error in Discord. Somebody has usually hit the same wall and will tell you whether a proxy is even the fix.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk