Where Rank Data Comes From
Four routes, with different costs, coverage, legal positions and failure modes. Most operations end up with two of them and should understand both.
Rank data reaches an analyst by one of four routes, and the choice determines cost, reliability and what happens when something breaks.
Official APIs
What they are. Marketplace-provided programmatic access, generally requiring a seller or developer account and approval.
Advantages. Reliable, documented, permitted, structured, and stable across site redesigns.
Limitations. Coverage is usually restricted to your own products or to a subset of fields. Rate limits are real. Approval processes can be slow and terms can change.
Where they fit. First-party data for your own catalogue. This should always be the primary source for your own products, because it is the only one that is unambiguously permitted and reliable.
Commercial data providers
What they are. Firms that collect at scale and sell access, by subscription or per query.
Advantages. Broad coverage, historical depth you cannot build retrospectively, no collection infrastructure to maintain, and someone else absorbs the breakage when a site changes.
Limitations. Cost. Opacity about method and coverage. Their history is their asset and export is frequently restricted. Coverage gaps in smaller categories and marketplaces are rarely disclosed.
The questions to ask: how is it collected, at what cadence, what is the coverage by category and marketplace, what is the gap rate, how far back does history go, and can I export.
Where they fit. Competitive and market data, particularly where you need history you did not collect.
Your own collection
What it is. Fetching and parsing pages yourself.
Advantages. Exactly the coverage you want, at your cadence, with the context fields you choose. Full control and full export.
Limitations. It breaks. Site structures change, blocks appear, rate limits bite. It requires ongoing maintenance rather than a one-time build. And it raises the terms-of-service question covered separately.
Where it fits. Narrow, well-defined watch lists where commercial coverage is poor or too expensive, and where you have someone to maintain it.
Manual observation
What it is. A person looking and recording.
Advantages. No infrastructure, no permission question, and it catches context a scraper misses — badges, layout changes, promotional treatments.
Limitations. Does not scale, cadence is poor, and consistency depends on discipline.
Where it fits. Small watch lists, spot checks against automated data, and investigating anomalies. Underrated as a verification method.
The combination that works
Official API for your own products. Authoritative and permitted.
A commercial provider for market and competitive history. Buys the history you cannot build.
Targeted manual checks for verification and context.
Own collection only where there is a specific gap that the other three do not fill, and only with someone assigned to maintain it.
Verification between sources
If you use more than one, compare them.
Take twenty products and check the same day's value across sources. They will differ, because of timing, of rank type, and of collection method.
Understand why before trusting either. A systematic offset is usually a timing or rank-type difference and is manageable. Random large disagreement means one source is unreliable.
Do this at the start and periodically, because providers change their methods without announcing it.
The history problem
The single most important property of this data is that history cannot be acquired retrospectively.
Marketplaces do not publish archives. If you did not collect it, and no provider did, last year's rank series does not exist.
This argues for starting collection before you need it, even at a low cadence and a modest watch list. Daily rank for fifty products costs almost nothing to collect and store, and in two years it is an asset that cannot be bought.
It also argues for valuing a provider's history correctly. The current data is a commodity; the archive is the product.
Comparing two sources against each other
If you use more than one source, reconciling them once at the start prevents a year of confusion.
Take twenty products and compare the same day's values.
Expect differences. Timing, rank type and method all differ between collectors.
A consistent offset is manageable. Note it and adjust, or pick one source as canonical.
Random large disagreement means one source is unreliable and you need to establish which by manual observation.
Check across rank bands. Sources frequently agree at the top and diverge in the tail, where refresh is slower.
Repeat annually, because providers change methods silently.
Never mix sources within a single series without flagging it. A series that switches provider mid-way has a discontinuity that will be read as a market event.