Market Sizing From Rank Data, and Its Limits
The most requested application and the least reliable. What a bounded estimate looks like, and why the full-category version should usually be refused.
Someone will ask how big a category is. Rank data can contribute to an answer and cannot produce one on its own, and the difference between those two statements is where the credibility of an analytics function is spent.
Why the naive method fails
Take every product in a category, estimate units from rank, multiply by price, sum.
The errors compound multiplicatively. Each product's unit estimate carries a factor-of-two uncertainty; summing thousands of them does not average the error away, because the errors are correlated — they all come from the same conversion curve.
The long tail dominates the count and is where estimates are worst. Most products in any category are in the deep tail where rank barely maps to units.
Coverage is incomplete. You cannot enumerate every product, and the ones you miss are systematically the obscure ones.
Price is not revenue. Discounts, variants, bundles and multipacks all distort it.
The output can be wrong by an order of magnitude and it will be presented as a specific number, because that is what spreadsheets produce.
The bounded estimate that works
Restrict to the top of the category, where estimates are least bad and coverage is complete.
"The top 100 products in this category represent approximately X to Y units per month" is defensible if the conversion basis is stated.
State explicitly that it excludes the tail and give a sense of what proportion of the category that is by product count.
Do not extrapolate to the full category using an assumed tail distribution, which is where the honesty is usually lost.
Better approaches to the same question
Industry data. Trade associations, market research firms, regulatory filings. Frequently exists and is frequently better than anything derivable from rank.
Public company disclosures for categories dominated by listed firms.
Marketplace-published aggregates where they exist.
Your own share, inverted. If you know your units and can estimate your share from rank position relative to the category, that inverts to a category estimate with error you can characterise. This is usually the best available method for a seller.
Triangulation. Two independent methods agreeing is worth more than either alone. Disagreeing is informative too and should be reported rather than resolved by picking the convenient one.
When rank data is genuinely the best available
Emerging categories with no industry coverage.
Marketplace-specific sizing, where the question is about that channel rather than the whole market.
Relative sizing — is category A bigger than category B — which is far more robust than absolute sizing because the conversion errors partly cancel.
Trend — is the category growing — using a fixed basket over time, which avoids the coverage problem entirely.
That last one is the strongest application in this article. Tracking a fixed set of products over quarters tells you the direction of a category with much better reliability than any absolute estimate.
Presenting it
Ranges, always.
The basis, named and dated.
The scope, stated: which products, which marketplace, which period.
The exclusions, stated: the tail, unavailable products, variants.
A sensitivity note: what the figure becomes if the conversion basis is off by a factor of two, which it might be.
And a recommendation about what would improve it, because the honest position is usually that rank data alone is insufficient and something else would help.
When to refuse
Investment cases and valuations. The error is too wide and the consequences are too large.
Anything that will be quoted onward without the range. If you cannot control how the number travels, give a band or a ratio instead.
Full-category absolute sizing in a category with a long tail, which is most of them.
Refusing is a professional act, not an evasion. A function that produces a number for every request loses the ability to be believed about the ones it can answer well.
The triangulation write-up
Where two methods are available, presenting both is stronger than reconciling them.
Method one, with its figure and range.
Method two, with its figure and range.
Whether they overlap.
If they overlap, state the intersection as the working estimate and note that two independent approaches agree, which is genuinely reassuring.
If they do not overlap, say so and investigate rather than averaging. Non-overlapping estimates mean at least one method has a problem, and finding it is more valuable than producing a number.
Never quietly pick the convenient one. This is the single most common integrity failure in market sizing and it is invisible in the output.
Where only one method is available, say that too. A single-method estimate with a stated range is honest; the same estimate presented as though triangulated is not.