Tutorial · 8 min read

Proxies for price tracking on a side project

A price tracker that runs on cron, reads prices from structured data, skips unchanged pages, and runs on the cheapest proxy plan we sell.

Watching a few hundred product pages for price drops is a classic first scraping project, and one of the cheapest to run, if you build it so that it barely downloads anything. This guide builds a tracker that runs on cron, reads prices from the structured data shops already publish, skips pages that have not changed, and pings you when something gets cheaper.

Price monitoring is on the “fine by us” list in our acceptable use policy. The same policy rules out volume that degrades a site for its real users, which is one more reason to check each page a few times a day and no more.

Which proxy

Product pages on smaller shops are usually fine from a datacenter address. Check yours with the five-minute test before assuming otherwise. If it passes, the cheapest setup we sell is one shared datacenter IP: $2.10 for 30 days, with 1 GB of traffic included. A dedicated address is $3.20 and has no traffic allowance to run out of. If the shop blocks hosting ranges, residential is billed by the gigabyte instead, and the budget below becomes the thing to watch.

How far the included traffic goes

How many checks fit depends on how much one product page weighs on the wire, which you can measure for free with the script in the bandwidth guide. The sizes below are examples; use your own.

Page checks that fit in the shared plan's included traffic
If one page costs youChecks in the included trafficPages you can check 4 times a day for 30 days
50 KB20,000166
150 KB6,66655
500 KB2,00016
Example page sizes, 1 GB counted as 10^9 bytes. Traffic past the allowance is $0.35/GB added with the order, $0.40/GB added later. A 304 from the conditional requests below costs only headers, so the real number of checks is higher.

Where the price lives

Most shops embed schema.org Product data in a <script type="application/ld+json"> block, because search engines read it for rich results. The price sits in an Offer as price, or in an AggregateOffer as lowPrice. Reading that is far sturdier than a CSS selector, which breaks the next time the shop redesigns.

The tracker

Standard library plus requests. Save it as tracker.py next to a urls.txt with one product URL per line:

import json
import random
import re
import sqlite3
import sys
import time

import requests

PROXY = "http://USER:PASS@IP:PORT"  # your datacenter address
PROXIES = {"http": PROXY, "https": PROXY}
HEADERS = {
    "User-Agent": (
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
        "(KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36"
    ),
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-GB,en;q=0.9",
    "Accept-Encoding": "gzip, deflate",
}
LD_JSON = re.compile(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", re.S | re.I)
WEBHOOK = ""  # optional: a Discord channel webhook URL

db = sqlite3.connect("prices.db")
db.execute(
    "CREATE TABLE IF NOT EXISTS pages (url TEXT PRIMARY KEY, etag TEXT, "
    "modified TEXT, price REAL, currency TEXT, checked TEXT)"
)


def find_price(node):
    if isinstance(node, list):
        for item in node:
            found = find_price(item)
            if found:
                return found
    elif isinstance(node, dict):
        kinds = node.get("@type")
        kinds = kinds if isinstance(kinds, list) else [kinds]
        if "Offer" in kinds and "price" in node:
            return float(node["price"]), node.get("priceCurrency", "")
        if "AggregateOffer" in kinds and "lowPrice" in node:
            return float(node["lowPrice"]), node.get("priceCurrency", "")
        for value in node.values():
            found = find_price(value)
            if found:
                return found
    return None


def check(url):
    row = db.execute("SELECT etag, modified, price FROM pages WHERE url = ?", (url,)).fetchone()
    headers = dict(HEADERS)
    if row and row[0]:
        headers["If-None-Match"] = row[0]
    if row and row[1]:
        headers["If-Modified-Since"] = row[1]

    r = requests.get(url, headers=headers, proxies=PROXIES, timeout=25)
    if r.status_code == 304:
        return "unchanged"
    if r.status_code != 200:
        return f"status {r.status_code}"

    price = None
    for block in LD_JSON.findall(r.text):
        try:
            price = find_price(json.loads(block))
        except ValueError:
            continue
        if price:
            break
    if not price:
        return "no structured price on this page"

    db.execute(
        "INSERT INTO pages VALUES (?, ?, ?, ?, ?, datetime('now')) "
        "ON CONFLICT(url) DO UPDATE SET etag = excluded.etag, modified = excluded.modified, "
        "price = excluded.price, currency = excluded.currency, checked = excluded.checked",
        (url, r.headers.get("ETag"), r.headers.get("Last-Modified"), price[0], price[1]),
    )
    db.commit()

    old = row[2] if row else None
    if old is not None and price[0] < old:
        message = f"Price drop: {old} -> {price[0]} {price[1]}"
        if WEBHOOK:
            requests.post(WEBHOOK, json={"content": f"{message}  {url}"}, timeout=10)
        return message
    return f"{price[0]} {price[1]}"


if __name__ == "__main__":
    urls = [line.strip() for line in open(sys.argv[1]) if line.strip()]
    random.shuffle(urls)
    for url in urls:
        try:
            print(f"{check(url)}  {url}")
        except requests.RequestException as exc:
            print(f"error {exc}  {url}")
        time.sleep(random.uniform(2, 5))

What makes it cheap

  • Conditional requests. The tracker stores each page’s ETag and Last-Modified and sends them back. When nothing changed, the shop answers 304 Not Modified with no body. Not every shop sends those headers for product pages; where they are missing, you download the full page every time.
  • Compression. Accept-Encoding: gzip keeps the pages that do come back small on the wire.
  • No browser. The JSON-LD is in the HTML the server sends, so there is no need to render the page. If a shop only builds its price with JavaScript, that page will say “no structured price” and you can decide whether it is worth a browser.
  • No retries. A failed check waits for the next run. Retries you start are billed as separate requests, and a price that is six hours stale is fine for a side project.

Run it on a schedule

# crontab -e: every six hours, on the hour
0 */6 * * * cd /home/you/tracker && /usr/bin/python3 tracker.py urls.txt >> tracker.log 2>&1

The webhook call goes straight to Discord without touching the proxy, so it costs nothing on the meter. Paste a channel webhook URL into WEBHOOK and drops land in your phone’s notifications.

When it grows

A few hundred pages from one address, spread over a run with a pause between each, is modest traffic for most shops. If you add shops or check more often and start seeing 429s, spread the load over a few static addresses with the round-robin in rotating vs sticky, or check less often. And when a shop starts answering 403, the error field guide will tell you whether it is the shop or us.

Try it while you read

Top-ups start at $5.

One shared datacenter IP for 30 days is $2.10. A single gigabyte of residential is $5.50. The balance never expires.

Published

Filed under

Found a mistake? Tell us in Discord and we will fix the post.

The community layer

Stuck halfway through?

Paste the error in Discord. Someone has hit it before and the answer is usually one message long.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk