Survivorship Bias in Bestseller Analysis
Every study of what top products have in common is drawn from a sample that excludes the failures. The conclusions are unsupported and they are published constantly.
The most popular analysis in this field takes the top hundred products in a category, looks for common attributes, and concludes those attributes cause success.
The sample was selected on the outcome. Nothing about causation can be concluded from it.
The structure of the error
Products are selected because they are at the top.
Attributes are observed in that sample.
The failures with the same attributes are invisible, because they are not in a bestseller list.
Without the failures, the base rate is unknown. If 90 percent of top products have a particular listing style, and 90 percent of all products in the category have it, the attribute explains nothing.
This is not a subtlety. It invalidates the conclusion entirely.
What it produces
"Bestsellers have an average of 47 reviews." So do most products at that price point, possibly.
"Top products use video in their listings." As do many failed products; video is a marker of a seller with a budget.
"Successful launches use this pricing strategy." The failed launches that used it are not in the sample.
"Winning products have these keyword patterns." Everyone copies the visible winners, so the pattern is downstream of imitation.
The correction
Include a comparison sample. Products from the same category, same period, not selected on rank. A random sample of the middle and the tail.
Compare rates, not counts. The question is whether the attribute is more common in the top than in the general population.
Report the base rate. Without it, no attribute finding means anything.
Acknowledge you still have not established causation, only association with a comparison — which is more than the original analysis had and less than a causal claim.
Why the corrected version is rare
The comparison sample is hard to draw. Marketplaces expose the top of a category readily and the tail poorly. Enumerating a random sample of mid-catalogue products is real work.
It weakens the finding. Most attributes lose their apparent effect once a base rate is included, and a report saying "we found nothing distinguishing" is harder to publish than one with seven success factors.
Nobody asks. Audiences do not routinely request the base rate, so its absence goes unnoticed.
The related error: bestseller lists as a research population
Bestseller lists are also self-selecting over time. Products that were briefly at the top and then failed drop out, so a list examined today contains the ones that sustained.
Historical analysis of "products that were bestsellers in year X" is better, if you have the history — which is one of the practical arguments for long-run collection.
A list of everything that entered the top 100 over three years, including the ones that fell out, is a far more honest population than a snapshot of today's top 100.
Where this bites commercially
Copying competitors. A seller who copies the visible attributes of top products is copying attributes that may be irrelevant, on the basis of an analysis with no comparison group.
Vendor "success factor" reports, which are almost universally constructed this way.
Internal post-mortems on winners with no equivalent study of losers.
Case studies, which are survivorship bias in narrative form.
The question to ask of any such analysis
"What did the products that failed look like?"
If the answer is that they were not examined, the analysis describes the top of a distribution and explains nothing about how anything got there.
That question takes five seconds and it correctly dismisses a large proportion of published marketplace research.
Drawing a comparison sample
The hard part of correcting survivorship bias, and it is more tractable than it appears.
You need products from the same category that are not selected on rank.
Browse pages beyond the first, which most collectors never do. Category listings extend well past the top and sampling from deeper pages gets you the middle.
Sample from search results for category terms, taking results from positions well down the list.
Use the new-releases listing, which is selected on date rather than performance and therefore includes future failures.
Sample from a competitor's full catalogue, which contains their winners and their duds.
None of these is a true random sample and all of them are enormously better than a bestseller list.
State the sampling method in the write-up. A reader can then judge how much the base rate is worth, which is more than they can do with a bestseller-only study.