Tutorial · 7 min read

Giving an AI agent a proxy: browser agents, Crawl4AI and plain HTTP

Route an AI agent’s browsing through a proxy: browser agents, Crawl4AI and plain HTTP tools, plus a byte budget so a runaway loop cannot drain your balance.

An AI agent that browses the web sends its traffic from wherever it runs, usually a cloud server, and plenty of sites treat cloud addresses as bots before the first page loads. A proxy for AI agents moves that traffic onto an address the site will talk to.

The catch is that “the agent” is usually three things making requests: a browser it drives, a crawler library, and plain HTTP tools it calls. Each takes its proxy in a different place. This guide covers all three, then the part that matters on a per-gigabyte bill: stopping a looping agent from downloading the internet on your balance.

Where does an LLM agent’s web traffic come from?

From three layers, and each needs the proxy set separately. A browser agent (browser-use, Playwright MCP, your own Playwright loop) takes it at browser launch. A crawler like Crawl4AI takes it in its run config. Plain HTTP tools read it from HTTPS_PROXY or a client argument. Miss one and that layer goes out from your server’s own IP.

Playwright-driven agents: set the proxy at launch

If you wrote the agent loop yourself on top of Playwright, the proxy is a launch option with the credentials in their own fields:

import asyncio
from playwright.async_api import async_playwright

PROXY = {
    "server": "http://resi.proxymonkey.io:8000",
    "username": "USER",
    "password": "PASS",
}


async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(proxy=PROXY)
        page = await browser.new_page()
        await page.goto("https://httpbin.org/ip")
        print(await page.inner_text("body"))
        await browser.close()


asyncio.run(main())

Chromium ignores a username and password written into the server URL, so keep them separate or you get a 407. Everything else about running Playwright through a proxy, including pace and fingerprint consistency, is in the Playwright guide.

browser-use

browser-use takes a ProxySettings on its Browser, and the agent takes the browser:

import asyncio
from browser_use import Agent, Browser, ChatBrowserUse
from browser_use.browser.profile import ProxySettings

browser = Browser(
    proxy=ProxySettings(
        server="http://resi.proxymonkey.io:8000",
        username="USER",
        password="PASS",
        bypass="localhost,127.0.0.1",
    ),
    args=["--blink-settings=imagesEnabled=false"],
)

agent = Agent(
    task="Find the opening hours on the public contact page of example.com",
    browser=browser,
    llm=ChatBrowserUse(),
)


async def main():
    await agent.run()


asyncio.run(main())

That matches browser-use’s documentation at the time of writing; swap ChatBrowserUse for whichever model you use. The args line is a Chromium switch that stops images loading, which comes up again under costs.

Playwright MCP

The Playwright MCP server has a --proxy-server flag, but no flag for proxy credentials. Its config file, though, passes launchOptions straight to Playwright, so the same proxy object works there. Save this as playwright-mcp.json:

{
  "browser": {
    "launchOptions": {
      "proxy": {
        "server": "http://resi.proxymonkey.io:8000",
        "username": "USER",
        "password": "PASS"
      }
    }
  }
}

And point your MCP client’s server entry at it:

{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest", "--config", "/path/to/playwright-mcp.json"]
    }
  }
}

Its --blocked-origins flag takes a semicolon-separated list of origins the browser must not request, which is handy for ad and analytics hosts. The project’s README notes that it is not a security boundary and does not affect redirects.

Crawl4AI proxy configuration

Crawl4AI’s docs now recommend setting the proxy per run with CrawlerRunConfig(proxy_config=...); the old proxy argument on BrowserConfig is marked deprecated.

import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig

proxy = ProxyConfig(
    server="http://resi.proxymonkey.io:8000",
    username="USER",
    password="PASS",
)

browser_conf = BrowserConfig(headless=True, text_mode=True)
run_conf = CrawlerRunConfig(
    proxy_config=proxy,
    cache_mode=CacheMode.ENABLED,
    check_robots_txt=True,
)


async def main():
    async with AsyncWebCrawler(config=browser_conf) as crawler:
        result = await crawler.arun(url="https://example.com/", config=run_conf)
        if result.success:
            print(result.markdown[:500])


asyncio.run(main())

text_mode tells the browser to skip images and other heavy content. check_robots_txt makes Crawl4AI read and obey the site’s robots.txt. CacheMode.ENABLED reuses pages it has already fetched instead of fetching them again through the proxy. Those three settings do more for your bill than anything else on this page.

Plain HTTP tools: HTTPS_PROXY, and keep the model API off it

Tools the agent calls directly (a fetch-URL tool, a search wrapper) usually use requests or httpx, and both read the proxy from the environment:

export HTTPS_PROXY=http://USER:[email protected]:8000
export HTTP_PROXY=http://USER:[email protected]:8000
export NO_PROXY=localhost,127.0.0.1,api.anthropic.com,api.openai.com

The NO_PROXY line matters. The environment applies to every client in the process, including the SDK that talks to your model provider. Without it, every prompt and completion goes through the residential proxy and is metered like page traffic. List whichever API hosts your agent calls. aiohttp is the exception: it ignores these variables unless the session is created with trust_env=True.

Check what the site sees before the agent starts

Before handing the agent a real task, give it a boring one: open https://httpbin.org/ip and report the address. Then run the same check from each layer separately (the browser, the crawler, a plain HTTP tool) and confirm all three come back with a proxy address rather than your server’s. A layer that prints your own IP is a layer you forgot to configure.

Two more things are worth knowing about a browser agent on the rotating gateway. First, the gateway rotates per connection, not per request, and a single page load opens several connections, so the HTML and its API calls may arrive at the site from different places. For reading public pages that is usually fine. Second, a multi-step flow on one site, like a search followed by a click-through and a form, keeps its address only while the browser reuses a kept-alive connection, and can see it change mid-flow whenever a new one opens. If that breaks the task, use a sticky session or a static address for that agent. To check a list of proxies in one go, including which network each address belongs to, the proxy checker tutorial has a script.

A byte budget for a runaway agent

Agents loop. One that keeps re-opening the same heavy page, or wanders into an image gallery, can move gigabytes before anyone looks. At $5.50/GB on a single residential gigabyte, a hard cap is worth ten minutes of setup.

The simplest cap that works for every layer is a tiny local relay that counts bytes and shuts off at a limit. Point the agent at it instead of at us; it forwards every connection to the gateway untouched, so the agent still sends its own credentials and every connection still gets a fresh IP.

import asyncio
import sys

UPSTREAM_HOST, UPSTREAM_PORT = "resi.proxymonkey.io:8000".split(":")
LISTEN = ("127.0.0.1", 8899)
BUDGET = int(float(sys.argv[1]) * 1_000_000) if len(sys.argv) > 1 else 200_000_000
used = 0


async def pipe(reader, writer):
    global used
    try:
        while chunk := await reader.read(65536):
            used += len(chunk)
            if used > BUDGET:
                break
            writer.write(chunk)
            await writer.drain()
    except OSError:
        pass
    finally:
        writer.close()


async def handle(client_reader, client_writer):
    if used > BUDGET:
        client_writer.close()
        return
    try:
        up_reader, up_writer = await asyncio.open_connection(UPSTREAM_HOST, int(UPSTREAM_PORT))
    except OSError:
        client_writer.close()
        return
    await asyncio.gather(pipe(client_reader, up_writer), pipe(up_reader, client_writer))


async def main():
    server = await asyncio.start_server(handle, *LISTEN)
    print(f"relay on {LISTEN[0]}:{LISTEN[1]}, budget {BUDGET / 1e6:.1f} MB")
    async with server:
        while used <= BUDGET:
            await asyncio.sleep(1)
    print(f"budget spent: {used / 1e6:.1f} MB, relay closed")


asyncio.run(main())
python relay.py 200

Then use http://127.0.0.1:8899 as the proxy server in every snippet above, with the same username and password. Because it counts raw bytes on the connection, TLS included, it sees roughly what our meter sees. When the budget runs out it stops accepting new connections and cuts each open one the next time data flows on it; an idle connection stays open until then, but carries nothing more. The agent’s next page load fails loudly instead of quietly spending. Run it on the same machine as the agent; it listens on localhost only and has no authentication of its own.

Block heavy content, cache, and respect robots.txt

An agent reads text. Images, video and fonts are most of a page’s weight and nothing the model uses, unless it works from screenshots. In a Playwright agent loop, abort them before they reach the proxy:

BLOCK = {"image", "media", "font"}


async def block_heavy(route):
    if route.request.resource_type in BLOCK:
        await route.abort()
    else:
        await route.continue_()


async def lighten(page):
    await page.route("**/*", block_heavy)

In browser-use the Chromium switch above does the images; in Crawl4AI, text_mode. For HTTP tools, cache: an agent that fetches the same documentation page six times in one task should pay for it once. requests-cache does this for requests in one line, and Crawl4AI has its own cache mode. There are more techniques in nine ways to cut proxy bandwidth.

An agent is still a client, and the site’s rules still apply. Check robots.txt (Crawl4AI will, if you ask), keep the pace of a person, and do not point an agent at accounts that are not yours. Our acceptable use policy is fine with public data at a considerate rate, treats scraping behind a login you legitimately hold as a grey area to ask us about first, and closes accounts that overload a site. If a site puts a CAPTCHA in front of your agent, it is asking for a human; let one answer.

Which proxy for AI agents: residential or static ISP

Which proxy line suits which kind of agent task
Agent taskLineWhy
Broad research across many sitesResidentialConsumer IPs, a fresh one per connection, billed by the byte
Logged-in session on your own accountStatic ISPOne consumer-network address for the whole session
Heavy pages on lenient sitesDatacenterA dedicated IP is priced per IP, so page weight does not drive the bill
Bulk dataset collectionDepends on the siteSee the AI training data use case

For logged-in agent sessions, a static ISP address keeps the account seeing one visitor instead of a new city on every page. For browsing many unrelated sites, residential rotation is the better fit, with the byte cap in front of it. If you are not sure which your agent does, the proxy chooser asks four questions.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk