Hou, Xue and Zhang's Anomaly Replication
Kewei Hou, Chen Xue and Lu Zhang took hundreds of published stock-return anomalies and re-ran every one with the exact data and methods the original papers used. Roughly half failed to replicate at standard significance, and the paper became the single most cited stress test of the factor literature.
Prerequisites: p-values and Multiple Testing
If Harvey, Liu and Zhu asked "how many published factors would you expect to be noise, statistically?", Kewei Hou, Chen Xue and Lu Zhang asked a more direct question: "if I literally rebuild each published anomaly, using the same variable definitions and the same test as the original paper, does it still show up?" Their 2020 paper, Replicating Anomalies, attempted to reconstruct 447 stock-return anomalies from the published literature and re-test each one. Barely half survived.
What "replication" meant here
This was not a philosophical replication crisis argument — it was a spreadsheet-level exercise. For each of 447 anomalies, Hou, Xue and Zhang built the trading signal exactly as the original paper described, formed the long-short portfolio the same way, and tested whether the return was still statistically significant using the same significance threshold the original authors used (a t-statistic of about 1.96, corresponding to 5%) and then again using a stricter bar closer to the 3.0 cutoff that Harvey, Liu and Zhu had separately argued for.
Worked example. Suppose the original paper on some anomaly reported a monthly long-short return of 0.6% with a t-statistic of 2.3 — comfortably above the conventional 1.96 threshold, so "significant" as published. Hou, Xue and Zhang rebuild the same portfolio using a larger, more current dataset and a more careful treatment of microcaps (excluding or value-weighting the smallest, least liquid stocks that dominate equal-weighted portfolios but are expensive to actually trade). The replicated monthly return comes out to 0.3% with a t-statistic of 1.4 — below even the loose 1.96 bar. Nothing about the underlying economic story changed; what changed was that the original result depended heavily on tiny, illiquid stocks that an equal-weighted test overweights and that a real portfolio could barely trade.
Failing to replicate does not mean the original researcher lied or made an error. Small, defensible choices — how you weight portfolios, which stocks you exclude, how you handle missing data — compound across hundreds of anomalies, and the ones that survive publication are disproportionately the choices that happened to work.
Why so many results were fragile
Two mechanical culprits explain most of the gap between original and replicated results:
- Microcap weighting. Many anomaly papers form portfolios by equal-weighting all stocks in a decile, which gives a stock with a $20m market cap the same influence as one worth $20bn. Microcaps are volatile, have wide bid-ask spreads, and are often hard to borrow for the short leg — an equal-weighted test can show a strong anomaly that is nearly impossible to trade at scale. Value-weighting, or simply excluding the smallest stocks, erased a large share of the anomalies Hou, Xue and Zhang tested.
- Look-back and construction choices. Details like how many months of data are required before a stock enters the sample, or exactly how a ratio is calculated, were rarely specified precisely enough in original papers to be unambiguous, and different reasonable choices produced different significance results on the same underlying idea.
What survived
Not everything failed. Momentum, several profitability measures, and a handful of value-related anomalies replicated robustly across specifications, which is part of why they remain core factors in commercial risk models today. The paper's contribution was separating the anomaly literature into a durable core and a much larger, fragile periphery — useful information for anyone deciding which published signal is worth building a strategy around versus which is worth skepticism.
"It replicated in Hou, Xue and Zhang" is not the same as "it will make money after costs." Their test is about statistical robustness to reasonable construction choices, not about transaction costs, capacity, or crowding. A factor can clear their bar and still be unprofitable to trade at any meaningful size.
In interviews
Use this paper as your concrete example of a replication study, distinct from Harvey, Liu and Zhu's more theoretical multiple-testing argument. Be specific about the mechanism: equal-weighting overweights illiquid microcaps, and reasonable-looking construction choices compound into different significance conclusions. If asked which anomalies held up, name momentum and profitability as examples that survived — it signals you read past the headline "half of anomalies are fake" summary.
Related concepts
Practice in interviews
Further reading
- Hou, Xue & Zhang (2020), Replicating Anomalies
- McLean & Pontiff (2016), Does Academic Research Destroy Stock Return Predictability?