The Replication Crisis in Factor Research
Hundreds of published stock-return predictors have been re-tested by independent teams. Most of them shrink badly and many vanish entirely. This page explains why, and how to read a factor paper without inheriting its optimism.
Prerequisites: Data-Snooping Bias, p-values and Multiple Testing
By the mid-2010s the academic finance literature had published several hundred variables that supposedly predicted the cross-section of stock returns. John Cochrane called it a "zoo of factors". Each one arrived with a t-statistic above 2, a plausible economic story and a chart of cumulative returns pointing up and to the right.
Then several teams did the obvious thing: they re-ran the zoo. Same data sources, same definitions, published code where it existed. The results were bad enough that the field now talks about them as a crisis rather than a disagreement.
What the re-runs found
| Study | What it re-ran | Headline result |
|---|---|---|
| Harvey, Liu & Zhu (2016) | 316 published factors | Once you account for how many things were tried, a new factor needs a t-statistic near 3.0, not 2.0. Most published t-stats sit between 2 and 3. |
| McLean & Pontiff (2016) | 97 predictors | Returns fall about 26% in the years after the sample ends but before publication, and about 58% after publication. |
| Hou, Xue & Zhang (2020) | 452 anomalies | About 65% fail to clear t = 1.96 once microcaps are handled properly — value weighting and NYSE size breakpoints. |
| Chen & Zimmermann (2020s) | 200+ predictors, code published | Most do replicate in their original samples, but the average live return is roughly half the published number. |
Read together they say two separate things. Harvey–Liu–Zhu and Hou–Xue–Zhang say a large share of published effects were never really there. McLean–Pontiff and Chen–Zimmermann say that of the ones that were there, a lot got traded away or were always smaller than the headline. Both problems apply to the same paper at the same time.
Why it happens
Search. Nobody publishes the 400 signals that did not work. A researcher who tries 400 candidates and reports the best one will find a t of 2.5 by luck alone; the reported statistic describes one draw from a search, not one draw from nature. See Data-Snooping Bias and Publication Bias and the File Drawer.
Weighting and universe. Roughly half of US listed names are microcaps. They are about 3% of total market value and they are where nearly every anomaly is strongest — because they are illiquid, hard to short and expensive to trade. Equal-weighting the cross-section hands them the same influence as Apple. This one choice explains a large fraction of the Hou–Xue–Zhang failures.
Decay after publication. Once a signal is public, capital arrives. Some of the drop is arbitrage; some is that the published number was inflated to begin with. You cannot tell them apart from the outside, and for a trading decision you do not need to.
Worked example: should you trade the paper on your desk?
A 2014 paper reports a long-short quintile spread of 0.55% per month, t = 2.4, over 1972–2012, equal-weighted across all CRSP names.
Step one, fix the construction. Re-run value-weighted, dropping everything below the NYSE 20th-percentile market cap. Suppose the spread falls to 0.28% per month, t = 1.3. This is typical, not pessimistic.
Step two, apply the out-of-sample discount. It is now more than a decade past publication, so the McLean–Pontiff post-publication haircut of ~58% applies: about 0.12% per month.
Step three, pay for it. A quintile spread rebalanced monthly runs high turnover. At 25 bps round-trip on the liquid names that survived step one, costs are on the order of 0.08% per month.
Verdict: about 0.04% a month, or half a percent a year, before borrow fees. That is not a strategy. It is not zero either — as one weakly-correlated input inside a blend of twenty, it may still earn its keep. What it can never justify is a dedicated book, and the paper's 0.55% never had anything to do with what you could have collected.
Treat a published t-statistic as an upper bound on a number you have not yet computed, not as evidence. The honest sequence is: rebuild the signal yourself, run it value-weighted with microcaps excluded, halve whatever survives for publication decay, then subtract costs. Whatever is left is the thing you are actually deciding about.
What survived
The wreckage is not total. A short list replicates almost everywhere people have looked — across decades, across countries and across asset classes: market exposure, value, momentum, profitability, investment/asset growth, and low risk. These share a property the failures do not: they show up in markets that were never part of anyone's original search, which is the closest thing factor research has to an out-of-sample test. See Cross-Market Replication Tests.
The opposite over-correction is just as expensive. "Everything is overfit, so nothing works" is not scepticism, it is a different kind of laziness — it also happens to be contradicted by the survivors above. And the crisis is not something that happened to academics. Your own research pipeline is a factor zoo with one contributor, no referees and no file drawer you are willing to look inside.
Three cheap filters before you spend a month on someone's paper. Does it hold value-weighted with microcaps excluded? Does it appear in a market that was not in the original sample? Was the sign of the effect predicted by the story before the data was seen, or attached afterwards? A signal that fails all three has never been tested against anything.
Related concepts
Practice in interviews
Further reading
- Harvey, Liu & Zhu, …and the Cross-Section of Expected Returns (RFS 2016)
- McLean & Pontiff, Does Academic Research Destroy Stock Return Predictability? (JF 2016)
- Hou, Xue & Zhang, Replicating Anomalies (RFS 2020)
- Chen & Zimmermann, Open Source Cross-Sectional Asset Pricing