Skip to content
Rank Tracer

Notes  ยท  Pitfalls

The Correlation Traps in Rank Analysis

Five specific ways rank data produces confident wrong conclusions, all of them common enough to appear in published analysis regularly.

Rank data is unusually good at producing correlations that look causal. Five patterns account for most of them.

Reverse causation

The claim: reviews drive sales.

The reality: sales produce reviews. The correlation is strong and the arrow points the other way.

Same structure elsewhere: advertising spend and rank, where sellers increase spend on products already selling. Price and rank, where sellers reprice in response to performance. Listing quality and rank, where successful sellers invest in listings.

The test: did the supposed cause precede the effect, or follow it? Timestamps answer this and are rarely checked.

The common cause

The claim: products with feature X sell better.

The reality: established brands both include feature X and have distribution, marketing and reputation. The feature is a marker of the kind of company that also does everything else.

Very common in category analysis, where any attribute correlated with being a large seller appears to drive sales.

The test: is there a plausible third factor that produces both? Usually there is, and it is usually the seller's size.

Market-wide movement

The claim: our campaign improved rank 30 percent.

The reality: the category rose after a seasonal trough.

Covered at length elsewhere and it belongs on this list because it is the most frequent single error.

The test: the control basket. Nothing else resolves it.

Survivorship

The claim: bestsellers share these characteristics, therefore these characteristics produce bestsellers.

The reality: the products with those characteristics that failed are not in the sample, because the sample was drawn from the top of a ranking.

Every "what do top products have in common" analysis has this problem and most do not mention it.

The test: can you see the failures? If your sample is a bestseller list, you cannot, and the conclusion is unsupported regardless of how many bestsellers you examined.

Regression to the mean

The claim: our intervention on underperforming products worked.

The reality: products selected because they were at an extreme move back toward normal without any intervention, because part of the extreme was noise.

This is why interventions on the worst performers always appear to work and interventions on the best performers always appear to fail.

The test: a control group of similarly-selected products that received no intervention. Without one, the effect is unmeasured.

The structure they share

All five come from the same source: rank is an observational measure of an uncontrolled system, and observational data supports association rather than causation.

The remedies are the same too:

A control set, which handles market movement and regression.

Timestamps, which handle reverse causation.

Asking what else could produce both, which handles common causes.

Asking what is missing from the sample, which handles survivorship.

Where causation is actually available

Your own interventions, with a control. A price test on some products and not others, with the untested products as control, supports a causal claim.

Natural experiments. A competitor's stock-out is a randomised removal you did not have to run. These are the most underused source of causal evidence in this field.

Staggered rollouts. Changing a listing on half the catalogue first.

Anything where you controlled the variable and had a comparison.

Everything else is association, and describing it as such is not weakness. It is the difference between an analysis that survives someone checking it and one that does not.

Designing a test that would settle it

When a correlation is interesting, the useful next step is describing what would establish causation, even if you cannot run it.

What is the intervention? A price change, a listing change, an advertising pause.

What is the control? Comparable products receiving no intervention, selected before the fact.

What is the outcome measure, and over what window?

What result would falsify the hypothesis?

What confounders would remain, and can they be observed?

Writing this down has two benefits. It frequently reveals that the test is cheap and could simply be run. And where it cannot be run, it makes explicit what the observational finding is missing, which is a far more useful thing to report than a correlation presented as a mechanism.