Tutorial · 10 min read

Scraping Amazon prices with proxies, without hammering Amazon

Track Amazon prices on a budget: the official API route first, then a polite Python scraper with proxies, parsing, pacing and the cost per 1,000 pages.

You want to know when the thing in your basket gets cheaper, or you sell on Amazon and want to see what the listing next to yours is doing. Both come down to reading a price off a product page on a schedule. This guide covers the official route first, then a small, deliberately slow Python scraper for a handful of products, and what it costs to run through a residential proxy.

Is it legal to scrape Amazon prices?

Short answer: Amazon says no, the law says “it depends”, and this is not legal advice. Amazon’s Conditions of Use (last updated August 14, 2026 when we read them) grant a licence for personal, non-commercial use, and then carve out, word for word, “any collection and use of any product listings, descriptions, or prices” and “any use of data mining, robots, or similar data gathering and extraction tools.” A price scraper is both of those things.

Whether breaking those terms makes scraping public prices unlawful is contested, and the answer changes from country to country and with what you do with the data afterwards. It is not risk-free. If this is for a business, ask a lawyer where you are.

On our side, price monitoring is on the “fine by us” list in our acceptable use policy, at a considerate rate. Volume that degrades a site for its real users is not. Amazon’s terms are between you and Amazon.

Is there an official Amazon price API?

Yes, with a catch: you have to be selling things for Amazon first. The old Product Advertising API 5.0 (PA-API) is gone. When we checked on 29 September 2026, its documentation address redirected to a deprecation notice saying PA-API 5 “has been deprecated and is being replaced by the Creators API”, and that calls to it now get a 403 with the message “Product Advertising API is deprecated.” The notice gives no date.

The Creators API has the familiar operations (GetItems, SearchItems, GetVariations, GetBrowseNodes). To use it you must be enrolled in Amazon Associates for the marketplace you want and have “at least 10 qualifying sales within the past 30 days”. Prices come through its OffersV2 resource; the old Offers resource is listed as not available in the migration guide. We are not quoting request limits here; read them in the docs when you sign up.

If you are a seller, the Selling Partner API covers your own listings and has pricing endpoints; check there before scraping anything. If you qualify for either API, use it. It is sanctioned, it does not break when the page layout changes, and it costs no proxy traffic. Everything below is for people who do not qualify and want to watch a few products for themselves.

What does Amazon’s robots.txt allow?

We read amazon.com/robots.txt on 29 September 2026. For the generic User-agent: * group:

  • Product pages at /dp/<ASIN> and /gp/product/<ASIN> are not disallowed. What is disallowed are helper paths underneath them, such as /dp/shipping/, /dp/product-availability/ and /gp/product/product-availability.
  • /gp/offer-listing/, the all-sellers view, is disallowed, apart from two narrow Allow lines for paths starting B000 and 9000. The cart, sign-in and wishlist paths are disallowed too.
  • There is no Crawl-delay, so the file does not tell you a safe pace. You have to pick one.
  • A long list of named bots is disallowed from the whole site, including AI crawlers and Scrapy. If you build on Scrapy, its default identity is shut out entirely, and changing the User-Agent to get around that is not something we would recommend.

So robots.txt lets a generic agent read product pages, and the Conditions of Use say no robots at all. robots.txt is crawler etiquette; the Conditions are the contract. Our scraper reads robots.txt at the start of each run, stops if what comes back is a robot check rather than a robots file, and skips any product the generic (*) rules disallow. It also sends a browser’s User-Agent, so it presents itself to Amazon as a browser rather than a bot. That is itself a way past Amazon’s bot filtering, and it is part of the risk you accept by running it. Passing the robots.txt check does not mean Amazon has agreed to anything.

Which proxies work for Amazon?

Amazon tends to answer datacenter ranges with its robot check page rather than a product. That is a tendency, not a rule, and it moves, so run the five-minute test against twenty of your own product URLs before paying for anything. One Amazon-specific tweak: the robot check can arrive with a 200 status, so also search each response for validateCaptcha rather than trusting the status code.

Residential is where most people end up for this target. Ours rotates across a pool and cannot be pointed at a country, which matters more on Amazon than on most shops. Each marketplace is its own domain (amazon.com, amazon.de, amazon.co.uk and so on), and the page shows a “Deliver to” location that it guesses from your IP address unless you set one. When we fetched two amazon.com product pages from a connection in Europe while writing this, the header read “Deliver to Lithuania”, Amazon set a cookie switching the display currency to euros, and one of the two pages said “This item cannot be shipped to your selected delivery location” with no price at all.

Through a rotating pool, each request can land in a different country, so the same product can come back with a different currency, a different delivery estimate, or no price. The scraper below records the delivery location and currency next to every price so you can compare like with like and ignore the rest. If you need one country’s prices every time, a static ISP address, whose country you pick at checkout, is the way to hold that steady. Put it through the same test first.

An Amazon price scraper in Python

This is a tracker for a short list of products you care about, not a crawler. It fetches one product page at a time, waits 20 to 60 seconds between them, and retries a 429 or 503 at most twice, waiting as long as Amazon’s Retry-After header asks (up to five minutes) or a growing pause if there is none. It stops the whole run if the retries run out, and the moment Amazon shows a robot check. It never tries to solve one. A CAPTCHA is Amazon asking for a human, and the polite reply is to go away for a few hours.

pip install requests beautifulsoup4

Save this as amazon_prices.py next to an asins.txt with one ASIN per line (the ten-character code in the product URL after /dp/):

import json
import random
import re
import sqlite3
import sys
import time
import zlib
from datetime import datetime, timezone
from urllib import robotparser

import requests
from bs4 import BeautifulSoup

MARKET = "https://www.amazon.com"
PROXY = "http://USER:[email protected]:8000"
PROXIES = {"http": PROXY, "https": PROXY}
HEADERS = {
    "User-Agent": (
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
        "(KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36"
    ),
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept-Encoding": "gzip, deflate",
}
GAP_SECONDS = (20, 60)
MAX_TRIES = 3
MAX_WAIT_SECONDS = 300
ROBOT_CHECK = ("validateCaptcha", "[email protected]")
PRICE_SELECTORS = (
    "#corePriceDisplay_desktop_feature_div .a-price .a-offscreen",
    "#corePrice_desktop .a-price .a-offscreen",
    "#corePrice_feature_div .a-price .a-offscreen",
    "#apex_desktop .a-price .a-offscreen",
    "#priceblock_ourprice",
    "#priceblock_dealprice",
)


def header_bytes(start_line, headers):
    lines = [start_line] + [f"{k}: {v}" for k, v in headers.items()]
    return sum(len(line) + 2 for line in lines) + 2


def fetch(url):
    with requests.get(url, headers=HEADERS, proxies=PROXIES, stream=True, timeout=30) as r:
        raw = b"".join(r.raw.stream(65536, decode_content=False))
        wire = header_bytes(f"GET {r.request.path_url} HTTP/1.1", r.request.headers)
        wire += header_bytes(f"HTTP/1.1 {r.status_code} {r.reason}", r.raw.headers) + len(raw)
        if r.headers.get("Content-Encoding", "") in ("gzip", "deflate"):
            raw = zlib.decompress(raw, 47)
        html = raw.decode(r.encoding or "utf-8", errors="replace")
        return r.status_code, html, wire, r.headers.get("Retry-After")


def text_of(soup, selector):
    node = soup.select_one(selector)
    return node.get_text(" ", strip=True) if node else None


def price_from_ld_json(soup):
    for tag in soup.select('script[type="application/ld+json"]'):
        try:
            data = json.loads(tag.string or "")
        except ValueError:
            continue
        for node in data if isinstance(data, list) else [data]:
            offers = node.get("offers") if isinstance(node, dict) else None
            if isinstance(offers, list):
                offers = offers[0] if offers else None
            if isinstance(offers, dict) and "price" in offers:
                return f"{offers['price']} {offers.get('priceCurrency', '')}".strip()
    return None


def to_number(text):
    match = re.search(r"\d[\d.,]*", text)
    if not match:
        return None
    digits = match.group().rstrip(".,")
    decimal = re.search(r"[.,](\d{1,2})$", digits)
    if decimal:
        whole = re.sub(r"[.,]", "", digits[: decimal.start()])
        return float(f"{whole}.{decimal.group(1)}")
    return float(re.sub(r"[.,]", "", digits))


def parse(html):
    soup = BeautifulSoup(html, "html.parser")
    found = {"title": text_of(soup, "#productTitle"), "deliver_to": text_of(soup, "#glow-ingress-line2")}
    for selector in PRICE_SELECTORS:
        shown = text_of(soup, selector)
        if shown and to_number(shown):
            return {**found, "shown": shown, "source": selector}
    shown = price_from_ld_json(soup)
    return {**found, "shown": shown, "source": "ld+json" if shown else None}


def pause(retry_after, attempt):
    try:
        return min(max(float(retry_after), 0), MAX_WAIT_SECONDS)
    except (TypeError, ValueError):
        return 30 * 2 ** attempt + random.uniform(0, 10)


def check(asin):
    url = f"{MARKET}/dp/{asin}"
    spent = 0
    for attempt in range(MAX_TRIES):
        status, html, wire, retry_after = fetch(url)
        spent += wire
        if any(marker in html for marker in ROBOT_CHECK):
            return {"status": "robot check"}, spent
        if status in (429, 503):
            if attempt == MAX_TRIES - 1:
                return {"status": "rate limited"}, spent
            time.sleep(pause(retry_after, attempt))
            continue
        if status != 200:
            return {"status": f"http {status}"}, spent
        result = parse(html)
        result["status"] = "ok" if result["shown"] else "no price"
        return result, spent


def open_db(path="amazon_prices.db"):
    db = sqlite3.connect(path)
    db.execute(
        "CREATE TABLE IF NOT EXISTS prices (asin TEXT, checked_at TEXT, status TEXT, "
        "title TEXT, shown TEXT, amount REAL, currency TEXT, deliver_to TEXT, "
        "source TEXT, wire_bytes INTEGER)"
    )
    return db


def load_robots():
    r = requests.get(f"{MARKET}/robots.txt", headers=HEADERS, proxies=PROXIES, timeout=30)
    r.raise_for_status()
    if "validateCaptcha" in r.text or "user-agent:" not in r.text.lower():
        sys.exit("robots.txt came back as something other than a robots file. Stopping; try again in a few hours.")
    robots = robotparser.RobotFileParser()
    robots.parse(r.text.splitlines())
    return robots


def main(asins):
    db = open_db()
    robots = load_robots()
    total = 0
    for i, asin in enumerate(asins):
        if not robots.can_fetch("*", f"{MARKET}/dp/{asin}"):
            print(f"skip (robots.txt)  {asin}")
            continue
        if i:
            time.sleep(random.uniform(*GAP_SECONDS))
        result, spent = check(asin)
        total += spent
        shown = result.get("shown") or ""
        db.execute(
            "INSERT INTO prices VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
            (
                asin,
                datetime.now(timezone.utc).isoformat(timespec="seconds"),
                result["status"],
                result.get("title"),
                shown,
                to_number(shown) if shown else None,
                re.sub(r"[\d.,\s]", "", shown) or None,
                result.get("deliver_to"),
                result.get("source"),
                spent,
            ),
        )
        db.commit()
        print(f"{result['status']:<12} {shown:<14} {result.get('deliver_to') or '-':<16} {asin}")
        if result["status"] == "robot check":
            print("Amazon asked for a human. Stopping this run; try again in a few hours.")
            break
        if result["status"] == "rate limited":
            print("Amazon is still refusing after retries. Stopping this run; try again later.")
            break
    print(f"{total:,} bytes through the proxy this run")


if __name__ == "__main__":
    main([line.strip() for line in open(sys.argv[1]) if line.strip()])
python amazon_prices.py asins.txt
sqlite3 -header -csv amazon_prices.db "SELECT * FROM prices" > prices.csv

Why Amazon price selectors break, and how this one copes

Amazon’s product page is huge, assembled from many widgets, and it changes. The price has lived under several IDs over the years, which is why the script tries a chain of selectors in order and records which one matched in the source column. When prices go missing, that column tells you which selector stopped working.

  • The selectors are a snapshot. We confirmed #productTitle, the corePrice and apex_desktop containers and the “Deliver to” line in pages we fetched on 29 September 2026. We could not confirm the price element inside them, because both of our pages showed no price to a European visitor. Expect to open a saved page and fix the chain from time to time.
  • No catch-all. A bare .a-price .a-offscreen would also match prices in the “customers also bought” carousels, which is worse than no price. The chain only looks inside the main price block.
  • Structured data is a fallback, not a plan. Many shops publish schema.org Offer data, and the generic price tracker relies on it. The two Amazon pages we checked had no ld+json block at all. The fallback costs nothing to keep.
  • The price is text. to_number handles both 1,299.00 and 1.299,00, and the currency is whatever symbol is left over. Keep the raw shown column so you can re-parse history if the rules change.

How much does it cost to scrape 1,000 Amazon pages?

Residential is billed by the gigabyte, and we count request bytes plus response bytes, headers included. The script’s byte counter adds up the same HTTP bytes, but our meter also counts the TLS and tunnel overhead around them, so expect our figure to come out a few percent higher; checking a provider’s meter covers where the gap comes from. The two product pages we fetched by hand came to roughly 250 KB each over the wire, gzipped. Two pages are not a sample, and page weight varies by product and by what Amazon decides to show you, so treat 250 KB as an assumption and measure your own before buying.

Residential cost of 1,000 Amazon product pages at assumed page weights
If one page weighsTraffic per 1,000 pagesAt the entry ratePages in the smallest order
100 KB0.10 GB$0.5510,000
250 KB (our assumption)0.25 GB$1.384,000
500 KB0.50 GB$2.752,000
Assumed page weights, 1 GB counted as 10^9 bytes, at $5.50/GB. You buy whole gigabytes: the smallest residential order is 1 GB for $5.50. Retries and robot-check pages are billed too.

For scale, 20 products checked 2 times a day for 30 days at 250 KB a page is 0.30 GB, which fits in a 1 GB order at $5.50. Amazon pages do not shrink much: there is no browser here to block images in, so the HTML is the bill. The bandwidth-cutting guide covers what is left: keep compression on, do not retry what can wait for the next run, and check less often. A price that is twelve hours old is fine for most watch lists.

Running it as an Amazon price tracker

Put it on cron twice a day and read the table when you like. To get alerts, compare each new amount with the last ok row for the same ASIN, the same deliver_to and the same currency; the side-project tracker has a Discord webhook you can lift straight across.

# crontab -e: 07:17 and 19:17, off the round hour
17 7,19 * * * cd /home/you/amazon && /usr/bin/python3 amazon_prices.py asins.txt >> run.log 2>&1

Keep the list short and the pace slow. If robot checks start showing up in run.log, that is the signal to check less often, not to push harder. And if your list keeps growing into hundreds of products, that is the point to get the sales that unlock the Creators API, or to look at what a proper feed costs. Our price monitoring page covers which proxy line suits which kind of feed.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk