Skip to content

About our crawler

If you found this address in a server log, this page is for you.

What we fetch, and why

Laddoo publishes a plain-language reading of Indian food-labelling regulations, and every rule we publish cites the instrument it came from. To make that citation checkable rather than merely stated, we keep our own copy of each cited document and record a SHA-256 of exactly the bytes the publisher served.

We fetch only documents we already cite by name - published regulations, amendments and guidance. We do not crawl a site looking for things, we do not follow links, and we do not index anything. Each request is for one document an operator has already identified.

We re-fetch on a schedule for one reason: so that when a document changes, a person is told and reviews the amendment. The crawler never changes a rule. It raises a task for a human.

How to recognise us

User-Agent
LaddooArtefactVerifier/0.1 (+https://laddoo.in/about/provenance)
What it requests
Published regulation documents - PDFs we already cite by name
How often
Weekly, by default, per document
Rate
At most one request every 3 seconds to any one host
robots.txt
Read and obeyed, including Crawl-delay
Backs off on
429 and 5xx, honouring Retry-After
Maximum download
100 MB per document

Asking us to stop, or to slow down

A Disallow for our User-Agent in your robots.txt stops us, and a Crawl-delay slows us. We read both before every fetch and we obey them.

You can also write to crawler@laddoo.in and a person will answer. If our traffic is causing you a problem, tell us and we will stop first and discuss afterwards.

If a document we hold should not have been published, or contains something it should not, tell us at the same address. We keep our copies to make a citation checkable, not to republish anything you have withdrawn.

The rules we publish, and the documents behind each of them, are at /rules. The machine-readable version is at /rules.json.

About our crawler - Laddoo