Quant Memo
Core

Harvey, Liu and Zhu: The Multiple-Testing Hurdle

Campbell Harvey, Yan Liu and Heqing Zhu counted hundreds of published factors and asked the obvious uncomfortable question: if you test that many ideas, how many would look significant by pure chance? Their answer was that the standard significance bar finance uses is far too low, and most "discovered" factors should not have cleared it.

Prerequisites: p-values and Multiple Testing

By 2015, academic finance had published somewhere around 300 factors claiming to predict which stocks beat the market. Campbell Harvey, Yan Liu and Heqing Zhu looked at that pile and asked a question referees had mostly been skipping: across 300 tests, how many "significant" results would you expect from a coin flip? Their 2016 paper, bluntly titled after its punchline, argued that most published factors fail a properly adjusted significance test — the factor zoo is mostly noise dressed up as discovery.

The problem: one test vs. three hundred tests

A single hypothesis test with the standard 5% significance threshold accepts a 1-in-20 chance of a false positive by pure luck. That is the deal researchers implicitly sign up for when they say "significant at the 5% level." The trouble is that finance did not run one test — it ran hundreds, across different research groups, different data windows, and different definitions of essentially the same idea, and only the ones that cleared the bar got published. If you test 300 genuinely useless signals against a 5% threshold, you should expect roughly 300×0.05=15300 \times 0.05 = 15 of them to look "significant" by chance alone, with no economic content whatsoever. Publication bias then guarantees that the 15 lucky draws are the ones that make it into a journal, not the 285 that failed.

Worked example. Suppose 200 candidate factors are truly worthless — no genuine predictive power — and each is tested independently at a 5% significance threshold. The expected count of false "discoveries" is 200×0.05=10200 \times 0.05 = 10 factors flagged significant purely by chance. If the finance literature has in fact published around 300 factors and a large share show only marginal significance under the original 5% bar, Harvey, Liu and Zhu's point is that a meaningful fraction of that literature is statistically indistinguishable from these 10 chance false positives — you cannot tell, from a single study's t-statistic alone, which bucket a given factor is really in.

old bar: t ≈ 2.0 HLZ bar: t ≈ 3.0 many published factors cluster here — significant under the old bar, not under HLZ's
Harvey, Liu and Zhu argue the profession's habitual t-statistic cutoff near 2.0 is far too permissive once you account for how many factors were tried. Their recommended bar, adjusted for multiple testing, sits closer to 3.0 — and a large cluster of "significant" published factors falls in the gap between the two.

"Statistically significant at 5%" only means what you think it means if it is the only test you ran. Run 300 tests and report the winners, and the 5% threshold stops controlling your error rate at all — the multiple-testing correction is not optional statistical pedantry, it is the difference between a discovery and a lottery ticket.

What the fix looks like

Harvey, Liu and Zhu don't just diagnose the problem — they propose raising the significance bar to account for how many tests have effectively been run across the literature's history, arriving at a recommended minimum t-statistic near 3.0 rather than the conventional 2.0. Applying that stricter bar, they estimate a large share of the previously published factors fail to clear it, meaning a substantial fraction of the "factor zoo" cannot be statistically distinguished from noise once multiple testing is properly accounted for.

Why this mattered beyond one paper

The paper landed at a moment when the number of published factors was still climbing and quant shops were building products around ever more granular signals. It reframed the burden of proof: a new factor now needs to explain why it should survive a stricter bar, not just clear the traditional one, and it pushed the field toward requiring out-of-sample or cross-market replication rather than a single significant coefficient in the original dataset. Harvey's later Presidential Address extended the argument into a broader critique of how financial economics does empirical work, arguing for pre-registration and higher evidentiary standards generally.

A common misreading is "Harvey, Liu and Zhu proved most factors are fake." They didn't prove any specific factor is fake — they showed that, in aggregate, the published literature's significance claims are not trustworthy at face value given how the tests were selected. Distinguishing a true discovery from a lucky one still requires case-by-case evidence: out-of-sample tests, different markets, different time periods.

In interviews

Be ready to reproduce the arithmetic: NN tests at a 5% threshold produce roughly 0.05N0.05N expected false positives with zero true signal, and publication bias selects for exactly those false positives. Then state the fix (raise the bar, roughly to a t-statistic near 3.0) and the honest caveat (this flags aggregate over-claiming, it does not individually falsify any one factor). This is one of the cleanest places to demonstrate you understand multiple-testing correction in a finance-specific context rather than an abstract statistics one.

Related concepts

Practice in interviews

Further reading

  • Harvey, Liu & Zhu (2016), ...and the Cross-Section of Expected Returns
  • Harvey (2017), Presidential Address: The Scientific Outlook in Financial Economics
ShareTwitterLinkedIn