The previous three lessons were about what you are permitted to do. This one is about how you do it, which is separate and in day-to-day terms more consequential — because conduct is what determines whether you stay unblocked, and whether a complaint ever reaches anybody's desk.
It is also the part most within your control. You usually cannot change whether a site has a login. You can always change how often you ask.
What robots.txt actually is
A text file at the root of a domain, in which the site operator states which paths they would prefer automated clients not to request. It is a convention from the early web, it is honoured voluntarily, and it is not a technical control — nothing stops you ignoring it.
Whether it carries legal weight varies and is genuinely unsettled. The practical case for respecting it does not depend on that. It is the only channel a site has for saying what it wants, it costs almost nothing to honour, and disregarding it is the clearest possible evidence of bad faith if anybody ever looks at your behaviour.
Read it with some judgement, though. A blanket disallow aimed at search engine indexing is a different statement from a specific disallow on a product path. If a file forbids everything and you believe your use is legitimate, the right move is to ask the operator — not to decide unilaterally that the file did not mean you.
Choosing a rate
Rules of thumb, not standards. When in doubt, be slower — nobody has ever been blocked for being too polite.
| Site type | Reasonable concurrency | Reasonable gap | Why |
|---|---|---|---|
| Large retailer or marketplace | 2 to 4 | 1 to 2 seconds | They serve far more than this per second already |
| Mid-sized e-commerce | 1 to 2 | 2 to 5 seconds | Your traffic is visible in their analytics |
| Small or independent site | 1 | 5 to 10 seconds | You could be a noticeable share of their load |
| Anything that has slowed down | Back off immediately | Exponential | A 429 or a rising latency is a request to stop |
Frequency is a separate question from rate, and it is the bigger one
Rate is how fast you go through a run. Frequency is how often you do the run, and it is where most unnecessary load comes from.
The honest test: how often does the underlying value actually change? A price that moves twice a week does not need hourly collection, and most schedules are set by what felt responsive at the time rather than by any measurement. Halving your frequency halves your load on the source, halves your cost, and in most cases changes nothing anybody downstream would notice.
Better still, make the frequency match the volatility. Fast-moving categories daily, stable ones weekly. This is both more considerate and cheaper, which is a rare combination.
Identify yourself
The default instinct is to blend in — copy a browser's user agent and hope nobody looks. It is understandable and it is usually the wrong call.
A user agent naming your organisation with a contact URL means that when somebody does look at their logs, they find a named, contactable, well-behaved client. The realistic outcomes are an allowlist, an email asking you to slow down, or nothing at all. The outcomes available to an anonymous client failing a fingerprint check are a ban and, if it ever escalates, a much worse story.
There is a real tension here, because some protection layers are more hostile to a declared crawler than to a convincing browser impersonation. Our position is that the long game favours being identifiable: impersonation is an arms race you lose eventually, and a relationship is durable in a way that a working header set is not.
Conduct that keeps you welcome
- Back off on 429 and 503 rather than retrying immediately — those are explicit requests to slow down
- Run heavy jobs outside the source's peak hours
- Cache, and never request a page twice when once would do
- Request only what you need — no images, no assets, no pages outside the scope
- Publish a contact address and answer it
- Stop promptly if asked, and talk to them before resuming
- Review the schedule periodically. Jobs set up two years ago are usually running far more often than anyone currently needs.
Worked example: the same 38,600 pages at four different rates
One competitor, 38,600 product URLs, run once a day. The rate you pick decides both how long the run takes and how visible you are in their logs. Their own traffic is the number that matters: a mid-size retailer serving roughly 40,000 page views a day is handling about 0.5 requests per second on average.
| Rate | Run duration | Share of their average traffic | How it reads in their logs |
|---|---|---|---|
| 10 req/s | 64 minutes | About 20x their own average rate | A spike. Someone will look, and they should. |
| 2 req/s | 5 hours 22 minutes | About 4x | Still the loudest client they have that hour. |
| 0.5 req/s | 21 hours 26 minutes | About 1x | Indistinguishable from a busy customer, but it no longer fits in a day. |
| 1 req/s, split across 2 nightly windows | 10 hours 43 minutes, off-peak | About 2x, during their quietest hours | Invisible in practice. This is the one to pick. |
What usually goes wrong
Most of the damage comes from the schedule rather than the rate.
- Tuning the rate carefully and then running six scrapers in parallel against the same host, so the real figure is six times the one on the dashboard.
- Running hourly because the scheduler made it easy, when prices on that site change twice a week. Frequency multiplies everything.
- Retrying failures immediately, so the moment a site is struggling is exactly the moment your traffic triples.
- Running at the top of the hour like everyone else's cron, which stacks your load onto theirs.
- Using a generic browser user agent with no contact address, which turns a solvable conversation into a block. Being identifiable is what lets someone email you instead of banning you.
- Scraping the same page five times to get five fields, because the extraction was written field by field rather than page by page.
The underlying principle
Be the kind of client a site operator would not bother blocking, because blocking you would be more effort than tolerating you.
That is not only an ethical stance, it is the most reliable operational strategy available. The scrapers that get blocked are the ones that are expensive, anonymous and relentless. The ones that run for years are polite, identifiable and ask for exactly what they need.
Finally: turning all of this into the one page that gets you an answer instead of a reflexive no.
Lesson 5: what to put in front of your legal team