An AI agent that browses the web sends its traffic from wherever it runs, usually a cloud server, and plenty of sites treat cloud addresses as bots before the first page loads. A proxy for AI agents moves that traffic onto an address the site will talk to.
The catch is that “the agent” is usually three things making requests: a browser it drives, a crawler library, and plain HTTP tools it calls. Each takes its proxy in a different place. This guide covers all three, then the part that matters on a per-gigabyte bill: stopping a looping agent from downloading the internet on your balance.
Where does an LLM agent’s web traffic come from?
From three layers, and each needs the proxy set separately. A browser agent (browser-use, Playwright MCP, your own Playwright loop) takes it at browser launch. A crawler like Crawl4AI takes it in its run config. Plain HTTP tools read it from HTTPS_PROXY or a client argument. Miss one and that layer goes out from your server’s own IP.
Playwright-driven agents: set the proxy at launch
If you wrote the agent loop yourself on top of Playwright, the proxy is a launch option with the credentials in their own fields:
import asyncio
from playwright.async_api import async_playwright
PROXY = {
"server": "http://resi.proxymonkey.io:8000",
"username": "USER",
"password": "PASS",
}
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(proxy=PROXY)
page = await browser.new_page()
await page.goto("https://httpbin.org/ip")
print(await page.inner_text("body"))
await browser.close()
asyncio.run(main())Chromium ignores a username and password written into the server URL, so keep them separate or you get a 407. Everything else about running Playwright through a proxy, including pace and fingerprint consistency, is in the Playwright guide.
browser-use
browser-use takes a ProxySettings on its Browser, and the agent takes the browser:
import asyncio
from browser_use import Agent, Browser, ChatBrowserUse
from browser_use.browser.profile import ProxySettings
browser = Browser(
proxy=ProxySettings(
server="http://resi.proxymonkey.io:8000",
username="USER",
password="PASS",
bypass="localhost,127.0.0.1",
),
args=["--blink-settings=imagesEnabled=false"],
)
agent = Agent(
task="Find the opening hours on the public contact page of example.com",
browser=browser,
llm=ChatBrowserUse(),
)
async def main():
await agent.run()
asyncio.run(main())That matches browser-use’s documentation at the time of writing; swap ChatBrowserUse for whichever model you use. The args line is a Chromium switch that stops images loading, which comes up again under costs.
Playwright MCP
The Playwright MCP server has a --proxy-server flag, but no flag for proxy credentials. Its config file, though, passes launchOptions straight to Playwright, so the same proxy object works there. Save this as playwright-mcp.json:
{
"browser": {
"launchOptions": {
"proxy": {
"server": "http://resi.proxymonkey.io:8000",
"username": "USER",
"password": "PASS"
}
}
}
}And point your MCP client’s server entry at it:
{
"mcpServers": {
"playwright": {
"command": "npx",
"args": ["@playwright/mcp@latest", "--config", "/path/to/playwright-mcp.json"]
}
}
}Its --blocked-origins flag takes a semicolon-separated list of origins the browser must not request, which is handy for ad and analytics hosts. The project’s README notes that it is not a security boundary and does not affect redirects.
Crawl4AI proxy configuration
Crawl4AI’s docs now recommend setting the proxy per run with CrawlerRunConfig(proxy_config=...); the old proxy argument on BrowserConfig is marked deprecated.
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig
proxy = ProxyConfig(
server="http://resi.proxymonkey.io:8000",
username="USER",
password="PASS",
)
browser_conf = BrowserConfig(headless=True, text_mode=True)
run_conf = CrawlerRunConfig(
proxy_config=proxy,
cache_mode=CacheMode.ENABLED,
check_robots_txt=True,
)
async def main():
async with AsyncWebCrawler(config=browser_conf) as crawler:
result = await crawler.arun(url="https://example.com/", config=run_conf)
if result.success:
print(result.markdown[:500])
asyncio.run(main())text_mode tells the browser to skip images and other heavy content. check_robots_txt makes Crawl4AI read and obey the site’s robots.txt. CacheMode.ENABLED reuses pages it has already fetched instead of fetching them again through the proxy. Those three settings do more for your bill than anything else on this page.
Plain HTTP tools: HTTPS_PROXY, and keep the model API off it
Tools the agent calls directly (a fetch-URL tool, a search wrapper) usually use requests or httpx, and both read the proxy from the environment:
export HTTPS_PROXY=http://USER:[email protected]:8000
export HTTP_PROXY=http://USER:[email protected]:8000
export NO_PROXY=localhost,127.0.0.1,api.anthropic.com,api.openai.comThe NO_PROXY line matters. The environment applies to every client in the process, including the SDK that talks to your model provider. Without it, every prompt and completion goes through the residential proxy and is metered like page traffic. List whichever API hosts your agent calls. aiohttp is the exception: it ignores these variables unless the session is created with trust_env=True.
Check what the site sees before the agent starts
Before handing the agent a real task, give it a boring one: open https://httpbin.org/ip and report the address. Then run the same check from each layer separately (the browser, the crawler, a plain HTTP tool) and confirm all three come back with a proxy address rather than your server’s. A layer that prints your own IP is a layer you forgot to configure.
Two more things are worth knowing about a browser agent on the rotating gateway. First, the gateway rotates per connection, not per request, and a single page load opens several connections, so the HTML and its API calls may arrive at the site from different places. For reading public pages that is usually fine. Second, a multi-step flow on one site, like a search followed by a click-through and a form, keeps its address only while the browser reuses a kept-alive connection, and can see it change mid-flow whenever a new one opens. If that breaks the task, use a sticky session or a static address for that agent. To check a list of proxies in one go, including which network each address belongs to, the proxy checker tutorial has a script.
A byte budget for a runaway agent
Agents loop. One that keeps re-opening the same heavy page, or wanders into an image gallery, can move gigabytes before anyone looks. At $5.50/GB on a single residential gigabyte, a hard cap is worth ten minutes of setup.
The simplest cap that works for every layer is a tiny local relay that counts bytes and shuts off at a limit. Point the agent at it instead of at us; it forwards every connection to the gateway untouched, so the agent still sends its own credentials and every connection still gets a fresh IP.
import asyncio
import sys
UPSTREAM_HOST, UPSTREAM_PORT = "resi.proxymonkey.io:8000".split(":")
LISTEN = ("127.0.0.1", 8899)
BUDGET = int(float(sys.argv[1]) * 1_000_000) if len(sys.argv) > 1 else 200_000_000
used = 0
async def pipe(reader, writer):
global used
try:
while chunk := await reader.read(65536):
used += len(chunk)
if used > BUDGET:
break
writer.write(chunk)
await writer.drain()
except OSError:
pass
finally:
writer.close()
async def handle(client_reader, client_writer):
if used > BUDGET:
client_writer.close()
return
try:
up_reader, up_writer = await asyncio.open_connection(UPSTREAM_HOST, int(UPSTREAM_PORT))
except OSError:
client_writer.close()
return
await asyncio.gather(pipe(client_reader, up_writer), pipe(up_reader, client_writer))
async def main():
server = await asyncio.start_server(handle, *LISTEN)
print(f"relay on {LISTEN[0]}:{LISTEN[1]}, budget {BUDGET / 1e6:.1f} MB")
async with server:
while used <= BUDGET:
await asyncio.sleep(1)
print(f"budget spent: {used / 1e6:.1f} MB, relay closed")
asyncio.run(main())python relay.py 200Then use http://127.0.0.1:8899 as the proxy server in every snippet above, with the same username and password. Because it counts raw bytes on the connection, TLS included, it sees roughly what our meter sees. When the budget runs out it stops accepting new connections and cuts each open one the next time data flows on it; an idle connection stays open until then, but carries nothing more. The agent’s next page load fails loudly instead of quietly spending. Run it on the same machine as the agent; it listens on localhost only and has no authentication of its own.
Block heavy content, cache, and respect robots.txt
An agent reads text. Images, video and fonts are most of a page’s weight and nothing the model uses, unless it works from screenshots. In a Playwright agent loop, abort them before they reach the proxy:
BLOCK = {"image", "media", "font"}
async def block_heavy(route):
if route.request.resource_type in BLOCK:
await route.abort()
else:
await route.continue_()
async def lighten(page):
await page.route("**/*", block_heavy)In browser-use the Chromium switch above does the images; in Crawl4AI, text_mode. For HTTP tools, cache: an agent that fetches the same documentation page six times in one task should pay for it once. requests-cache does this for requests in one line, and Crawl4AI has its own cache mode. There are more techniques in nine ways to cut proxy bandwidth.
An agent is still a client, and the site’s rules still apply. Check robots.txt (Crawl4AI will, if you ask), keep the pace of a person, and do not point an agent at accounts that are not yours. Our acceptable use policy is fine with public data at a considerate rate, treats scraping behind a login you legitimately hold as a grey area to ask us about first, and closes accounts that overload a site. If a site puts a CAPTCHA in front of your agent, it is asking for a human; let one answer.
Which proxy for AI agents: residential or static ISP
| Agent task | Line | Why |
|---|---|---|
| Broad research across many sites | Residential | Consumer IPs, a fresh one per connection, billed by the byte |
| Logged-in session on your own account | Static ISP | One consumer-network address for the whole session |
| Heavy pages on lenient sites | Datacenter | A dedicated IP is priced per IP, so page weight does not drive the bill |
| Bulk dataset collection | Depends on the site | See the AI training data use case |
For logged-in agent sessions, a static ISP address keeps the account seeing one visitor instead of a new city on every page. For browsing many unrelated sites, residential rotation is the better fit, with the byte cap in front of it. If you are not sure which your agent does, the proxy chooser asks four questions.
Top-ups start at $5.
One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.