Estimating Units From Rank, With Error Bars
It can be done defensibly. What that requires, what the realistic error is, and how to present a number that will not embarrass you six months later.
Unit estimation from rank is legitimate when the method and its uncertainty are stated. It becomes indefensible the moment the range is dropped.
What a defensible estimate needs
A conversion basis with a stated source and date. A published table, a vendor's model, or your own calibration.
Category match. The basis must have been built for the category you are applying it to.
Marketplace match. A US curve does not apply to a smaller national marketplace.
A stated rank band. Estimates are least bad in the band where the basis had data, which is usually mid-catalogue.
A range, not a point. A factor of two either way is a realistic minimum for most bands. Wider at the extremes.
The observation window, because a daily rate estimated from one reading is a different claim from one estimated over a month.
Building your own calibration
The strongest approach if you sell on the marketplace.
Record your own daily units and your own daily rank, at a fixed time, for a quarter.
Plot units against rank on a log-log scale. The relationship is approximately a power law over a range, which is why log-log is the right view.
Fit within the band you observed. Do not extrapolate. A calibration built between rank 10,000 and 60,000 says nothing about rank 200.
Note the scatter. The spread around the fit is your error, and it is what you report.
Recalibrate every few months. The relationship drifts as the catalogue grows.
With several products across several bands you get a wider curve, which is the practical reason a multi-product seller has better estimates than an outside analyst.
Realistic error
Within your own calibrated band, with your own data: perhaps ±30 to 50 percent. Good enough for planning.
Using a published table in a matched category: a factor of two either way is optimistic in many bands.
Top of the catalogue: poor, because published data there is sparse.
Deep long tail: poor, because rank is dominated by decay from single sales and the mapping is nearly meaningless.
Cross-category application: unquantifiable. Do not do it.
Presenting the number
"Roughly 40 to 120 units per day, based on [table], [date], applied to overall rank in [category], averaged over 28 days."
That sentence contains the estimate, the range, the basis, the date, the rank type, the category and the window. It is long and it is defensible.
Compare with "approximately 73 units per day", which is the same analysis with the honesty removed and is what most reports contain.
In tables, show the range in the cell, not in a footnote. Footnoted uncertainty is uncertainty that will be dropped when the figure is quoted onward.
Where the error compounds
Revenue = units × price. Two uncertainties multiply.
Annual revenue = daily estimate × 365, which assumes today's velocity holds all year and it will not.
Market size = sum over many products, each with its own error, and dominated by the long tail where estimates are worst.
Share = one uncertain estimate divided by another.
By the fourth step the figure can be wrong by an order of magnitude, and the spreadsheet will present it to the cent.
The discipline is to carry the range through every operation and present the final range. If the final range is too wide to support the decision, the honest conclusion is that rank data cannot answer the question.
When to refuse
When the audience will strip the range. If a figure will be quoted onward as a fact, and you cannot prevent that, consider presenting a ratio or a band instead of a number.
When the decision needs precision the method cannot give. Investment cases, valuations and legal claims need better evidence than a fitted curve on volunteered data.
When you cannot state the basis. An estimate whose provenance you cannot describe should not leave your machine.
Recording your calibration
A calibration is only useful if it can be audited later, which means recording more than the curve.
The raw pairs: date, product, rank observed, units sold that day. Keep these, not just the fit.
The rank band covered, stated explicitly as the range of validity.
The category and marketplace.
The period, because the relationship drifts.
The scatter, as the error you will quote.
The fitting method, even if it is a log-log regression, so that someone can reproduce it.
Re-fit quarterly and keep the old versions. Comparing successive calibrations shows you how fast the relationship is moving in your category, which is itself a useful finding and is unavailable to anyone who overwrote the previous fit.