The Factor Zoo and the Replication Crisis
Hundreds of "factors" claiming to predict stock returns have been published, more than any economic story can plausibly justify. This is the portfolio-construction side of that problem — what it means for anyone actually building a factor portfolio, not just for the academics arguing about it.
Prerequisites: Factor Investing
Open a factor-investing textbook from 2025 and you'll find a bibliography with several hundred entries, each claiming to have found something that predicts which stocks beat the market. Value, momentum, quality, low volatility, accruals, share issuance, dozens more, each backed by a published paper with a statistically significant result. John Cochrane's name for this pile is the "factor zoo," and the name stuck because it captures the problem exactly: a zoo has too many species to remember, and no obvious organizing principle for which ones are real.
Why a portfolio builder should care, not just an academic
The Replication Crisis in Factor Research covers the research side, how independent teams re-tested these factors and found many failed to hold up. This page is about what that means once you're not writing a paper but actually allocating money to factor strategies.
The practical risk isn't abstract. If you build a multi-factor portfolio out of ten "factors" and three of them are statistical artifacts of the original study rather than real, persistent sources of return, your portfolio is quietly carrying extra noise and extra trading costs for no expected payoff. Worse, you can't tell which three from the outside, the published paper looked exactly as convincing as the seven that hold up.
Why so many factors got published in the first place
The mechanism is simple multiple testing. If a researcher tries 300 candidate signals against the same stock-return data and reports only the ones that clear a conventional significance bar, a good number will clear that bar by chance alone, even if none of them contain real information. The published literature is, by construction, the surviving tail of a much larger search that never gets reported — nobody publishes "I tried a signal and it did nothing." That selective reporting is what produces the zoo: hundreds of "discoveries," each individually plausible, collectively far more than an efficient market should contain.
| Signal that looked genuine turned out to be |
|---|
| A trading-cost artifact — real in raw returns, gone after realistic costs |
| Sample-specific — worked in the original data window, vanished out of sample |
| A repackaging of an existing, already-known factor under a new name |
| Genuinely real, but shrunk substantially once retested carefully |
A concrete illustration of the multiple-testing problem
Suppose 300 researchers each independently test a signal that, in truth, predicts nothing at all — pure noise. Using a conventional 5% significance threshold, each test has a 5% chance of falsely appearing significant purely by luck. Across 300 independent noise signals, the expected number of false "discoveries" is . If only the positive results get written up and submitted, a reader sees fifteen papers, each reporting a seemingly real, statistically significant factor, with zero indication that 285 other attempts failed and were never mentioned. This is exactly why Harvey, Liu and Zhu argued that a new factor should need a t-statistic near 3.0 rather than the traditional 2.0 threshold — the traditional bar was calibrated for a world where researchers tested one hypothesis, not hundreds.
What happens after publication
Two well-documented effects hit a factor once it's published. Its future live returns tend to shrink, because sophisticated capital crowds into the trade the paper just revealed, competing away part of the edge — this is Factor Crowding. And its apparent past returns, when re-examined, often shrink too, because the original result benefited from data mining that a clean re-test strips out. Both effects push the same direction: published factor returns are, on average, weaker than the paper that introduced them suggested.
A published factor's t-statistic tells you it looked good in one dataset, examined by one team, possibly after trying many variants. It does not tell you whether it will keep working, or whether it was one of several hundred tries that happened to clear the bar.
What this means for building a factor portfolio
- Prefer factors with an economic story that predicts the sign in advance, not one fitted after seeing the data. A factor motivated by risk-bearing or a known behavioral bias (value, momentum) has survived more scrutiny than one discovered by scanning thousands of accounting ratios.
- Demand out-of-sample and out-of-country evidence. A factor that works in US equities from 1963–1990 and also in international markets or a later, untouched sample is far more credible than one tested once.
- Diversify across a small number of well-established factors rather than many similar-sounding ones. Factor Risk Models often reveals that a portfolio's twelve "factors" reduce to three or four genuinely independent sources of return once correlations are accounted for.
- Expect shrinkage, and size positions accordingly. Treat a paper's reported Sharpe ratio as an upper bound on what to expect live, not a forecast.
A fast filter: if a factor's name references a specific accounting line item most investors have never heard of, and its supporting sample is a single narrow time window, treat it as an unreplicated zoo entry until proven otherwise.
The honest state of the field is a shorter list than the zoo suggests: a handful of factors — value, momentum, quality, low-volatility, size with caveats — have survived decades of adversarial re-testing across markets. Everything else deserves the same skepticism you'd apply to any single study, no matter how sharp its published t-statistic looked.
Related concepts
Practice in interviews
Further reading
- Cochrane, Presidential Address: Discount Rates (JF 2011)
- Harvey, Liu & Zhu, …and the Cross-Section of Expected Returns (RFS 2016)
- McLean & Pontiff, Does Academic Research Destroy Stock Return Predictability? (JF 2016)