Quant Memo
Core

Sample Size Per Name and Signal Reliability

Why a trading signal tested on only a handful of stocks, or with only a few years of history per stock, produces much noisier performance estimates than the same signal tested across a broad universe over a long history.

A signal's backtested performance is itself just an estimate, and like any estimate its reliability depends on how much independent data went into it. Testing a factor on 20 stocks over 3 years gives far fewer effectively independent observations than testing it on 2,000 stocks over 15 years — even though both are "a backtest," the first is built on a small, noisy sample and the second on a much larger one, so the first result's confidence interval around any measured Sharpe ratio or hit rate is correspondingly much wider.

The standard error of an average return shrinks roughly with the square root of the number of independent observations, so quadrupling the sample size only halves the noise — meaning a small universe needs a dramatically stronger true effect before it's distinguishable from luck. A signal that looks good on 20 names might simply be a few lucky stock picks; the same edge measured across thousands of names, with results averaging out idiosyncratic noise, is far more likely to reflect something real.

Cross-sectional correlation compounds the problem: stocks in the same sector or under the same macro regime move together, so 2,000 "names" are not 2,000 independent draws — the effective sample size is smaller than the name count suggests, and researchers typically account for this before trusting a signal built on a narrow or highly correlated universe.

A signal tested on a small or highly correlated set of names carries a much wider margin of error than one tested broadly, so the apparent strength of a backtest result must be judged relative to how many effectively independent observations actually built it — not just how good the headline number looks.

Related concepts

Practice in interviews

Further reading

  • Bailey & Lopez de Prado, The Deflated Sharpe Ratio
ShareTwitterLinkedIn