Skip to content

What we hold, and for how long

Generated from the code that enforces it, so it cannot drift from what the system actually does.
DataPurposeLawful basisRetention
Hashed IP address of an anonymous scanRate limiting and abuse handling on a free, unauthenticated endpointLegitimate interest in keeping the service availableCleared 30 days after the scan
Salted hash of a signed-in user id, sent to Amplitude (S.97)Recognising the same person across sessions and devices in product analytics, which is what makes funnels and retention readableConsent, and only consent: Amplitude is the one provider behind the cookie gate, so this is processed for visitors who granted it and for nobody elseHeld by Amplitude under their retention; we store nothing and can send no id at all by declining
Email address supplied to unlock a scan breakdownDelivering the requested breakdown and, separately, alerts the person subscribed toPerformance of a contract, and consent for anything beyond itUntil the person asks us to delete it
Directory and listing measurementsThe published indexLegitimate interest; the data is factual metadata about public pagesKept indefinitely, this history is the product
Workflow run logsOperating and debugging the crawlerLegitimate interestPruned after 90 days
Aggregate page analytics: page path, referrer, country, device typeKnowing which pages are used, and reading the M1 gateLegitimate interest; no cookie is set and no visitor is identifiedHeld in aggregate by our analytics providers, not in our database. Private report and share URLs are excluded before they are sent

The short version

A scan you run without giving us an email is recorded so we can rate limit the endpoint. The IP is hashed when it is written, never stored raw, and the hash is deleted after 30 days. The row stays, because the count of scans is how we know whether the thing works.

If you give us an email, we use it for what you asked for. Anything beyond that needs your consent separately, and you can have it deleted by asking.

Everything else we store is measurements of public pages: which directory links to which product, what the link's rel attribute said, whether the page is in Google's index. That is factual metadata about pages anyone can visit, and it is kept indefinitely because the history is the product.

Our crawler

It identifies itself, respects robots.txt, and fetches politely. If a site blocks us we record the block and publish it rather than working around it. We store the links we parsed, not copies of anyone's pages. See the methodology for how each number is produced.

Cookies

Plausible and Vercel Web Analytics measure traffic without cookies and run on every visit. Amplitude, when you allow it, sets an analytics cookie so we can see how visits fit together. Visitors in the EU, EEA, UK and Switzerland are asked before it is set; you can change your answer here at any time, and declining also deletes what it stored.

Asking us to delete something

Email the address on the account page and we will remove your personal data. Measurements of public directory pages are not personal data and stay.

This page describes what the system does. It has not yet been reviewed by a lawyer, and it says so rather than implying it has.