Proxy glossary

What is robots.txt?

robots.txt is a plain-text file at the root of a website that tells crawlers which paths the owner would like them to leave alone.

robots.txt explained

It lives at /robots.txt on each host. Rules are grouped by user agent, so Disallow: /cart/ under User-agent: * applies to every crawler. Some sites add a Crawl-delay asking for a pause between requests; Google ignores it, plenty of other crawlers honour it.

It is a request from the site owner, and nothing technically stops a client from fetching a disallowed path. Scrapy obeys it when ROBOTSTXT_OBEY is on, which new Scrapy projects set by default.

For example

Before crawling a shop, open its robots.txt. A Sitemap: line often points at an XML file listing every product URL, so your spider can skip the category pages entirely.

The community layer

Not sure how it applies to you?

Post what you are trying to do in Discord. Someone who has hit the same thing will explain it against your setup.

Join the Discord

4,200+monkeys in the Discord

  • Help from humans

    Post your error, get an answer. Usually in minutes, usually from someone who has hit the same wall.

  • A status bot that tells on us

    Pool health, incidents and maintenance posted automatically. Including the bad days.

  • Deals and free traffic

    Bonus GB drops, early access to new pools, and the occasional giveaway for a good bug report.

Join the Discord4,200+ monkeys, free to lurk