Separating Luck from Skill
A three-year track record with a Sharpe of 1.8 sounds like skill. The math of decomposing observed variance into a skill component and a luck component often says most of that number is noise you'd expect from a coin-flipper too.
Prerequisites: Sharpe Ratio, The Standard Error of the Sharpe Ratio
An allocator is choosing between two portfolio managers. Manager A ran a Sharpe of 1.8 for three years. Manager B ran a Sharpe of 1.1 for twelve years. Most instinct picks A — the number is bigger. The decomposition below asks a different question: given how much data each track record actually contains, how confident can you be that the true long-run Sharpe of each manager is above zero at all? For A, at three years, the answer is "not very." For B, at twelve, it's "quite."
The decomposition
An observed Sharpe ratio is an estimate of a true, unobservable Sharpe, and every estimate from a finite sample has a standard error. For daily returns over years (annualized, ignoring skew and autocorrelation for a first pass):
In words: your observed Sharpe is the true, unknowable Sharpe plus estimation noise, and that noise shrinks only with the square root of the number of years of track record — not months, not trades. A manager's t-statistic on skill is simply , restated more directly as to first order.
Worked example: Manager A, three years
, . First-order standard error, ignoring the Sharpe-squared correction for a quick estimate: . A rough 95% confidence interval on the true Sharpe: , i.e. roughly 0.67 to 2.93. The interval is wide enough that "true Sharpe of 0.7" and "true Sharpe of 2.9" are both statistically plausible from the same three-year track record — a nine-fold range of skill, all consistent with the same observed number.
Worked example: Manager B, twelve years
, . . 95% interval: , roughly 0.53 to 1.67. Lower headline number, but a far tighter band, and critically the entire interval sits comfortably above zero. Manager B's number is smaller and more trustworthy at the same time — those are different axes, and the decomposition is what lets you see both.
Putting them side by side
| Manager | Observed SR | Years | SE | 95% CI | t-stat on skill (SR·√T) |
|---|---|---|---|---|---|
| A | 1.8 | 3 | 0.577 | 0.67 – 2.93 | 3.12 |
| B | 1.1 | 12 | 0.289 | 0.53 – 1.67 | 3.81 |
Manager B actually has the higher t-statistic — more confidently distinguishable from zero true skill — despite the lower headline Sharpe, purely because there are four times as many years of evidence behind it. This is the whole point of the decomposition: it separates "how good does the number look" from "how much do I actually know," and a shorter, flashier track record routinely loses on the second axis even while winning on the first.
Switch the sample-size slider above between small and large: the sampling distribution of a mean tightens exactly the way a Sharpe ratio's confidence interval tightens with more years of track record, and that tightening — not the headline value — is what the luck/skill decomposition is really measuring.
An observed Sharpe ratio always comes with a standard error that shrinks with in years. Compare confidence intervals, not point estimates, when ranking track records — a higher Sharpe over a shorter window can be statistically weaker evidence of skill than a lower Sharpe over a longer one.
The classic confusion: treating a short, hot track record as strong evidence because the point estimate is large. Point estimates from small samples are exactly as noisy as the math says — a manager with only a few years of history has, by definition, not yet given you enough data to rule out that the entire result is luck.
What this means in practice
Report every track record with its confidence interval, not just its point estimate, and prefer the manager whose interval clears zero by the widest margin — a function of both Sharpe and sample length, not Sharpe alone. See Why Maximum Drawdown Is a Noisy Statistic for the same problem in an even noisier statistic.
Related concepts
Practice in interviews
Further reading
- Bailey & López de Prado, The Sharpe Ratio Efficient Frontier
- Grinold & Kahn, Active Portfolio Management