Why Replications Fail
Most failed replications aren't failures of the original result — they're the replication using a different universe, a different point-in-time data source, or a different cost model than the paper did, and mistaking that mismatch for disproof.
Prerequisites: Reading a Paper Adversarially
A researcher tries to reproduce a well-cited anomaly and gets nothing — a decile spread near zero where the paper reported 5% a year. The instinctive conclusion is "the anomaly isn't real." That conclusion is right less often than it feels like it should be. In the largest systematic effort to reproduce published anomalies, a majority of published results replicated at least directionally once the replicators matched the original methodology closely — the failures that get reported informally on trading desks are disproportionately failures to match the original setup, not failures of the underlying effect.
The five places a replication quietly diverges from the original
A different universe. The paper tests "all CRSP common stocks"; the replication runs on the desk's usual top-1,500 liquid names. If the effect is concentrated in microcaps — common for anomalies built on illiquid, under-covered names — the replication will show nothing, correctly, on a universe the paper never claimed to describe.
A different definition of the same variable. "Book-to-price" sounds unambiguous. It isn't: does book value include or exclude intangibles, is it the most recent quarter or trailing twelve months, is price measured at the same date as the accounting data or with a lag? Two reasonable definitions can produce meaningfully different portfolios from the same underlying data.
Non-point-in-time data. The paper's authors, working with academic databases, sometimes use restated fundamentals without realizing it — the effect they measured partly reflects information investors didn't have at the time. A replication using genuinely point-in-time data can legitimately get a smaller number, and that's not the replication's failure — it's a more honest one. See Point-in-Time Data.
A different sample period. Running the exact original window first, before extending to more recent data, separates two different questions: "did I reproduce their result" from "does the effect still exist today." Skipping straight to a fresh sample conflates the two, and a null result stops being informative about either one.
A different cost model, or none at all. A paper reporting gross returns and a replication that immediately subtracts realistic costs will disagree even with identical portfolios. That's not a failed replication — the two are answering different questions, and only one of them is the one a trading desk cares about.
The discipline that fixes this
Run the replication in two stages, never one. First, match the paper's own universe, definitions, period and (if stated) cost assumptions as closely as possible, and confirm the number comes close to theirs. Only once that's confirmed, change one thing at a time — swap in the desk's universe, then point-in-time data, then realistic costs — and track which change moves the number and by how much. A replication that jumps straight to "our universe, our data, our costs" and gets zero has thrown away the information needed to tell whether the original effect is real but untradeable, or never existed at all.
A null replication answers "does this survive under my conditions," not "was the original wrong." Match the paper's own setup first and confirm you can reproduce their number before changing anything — otherwise a mismatched universe or a missing cost model gets mistaken for evidence the effect isn't real.
A concrete case
A desk tries to replicate a published post-earnings-announcement drift result and gets nothing on their standard universe and cost model. Before writing it off, a researcher reruns it matching the paper exactly: same universe (all CRSP stocks, no size filter), same definition of the surprise variable, same sample period, no cost adjustment. That version reproduces the paper's number closely. Changing one variable at a time from there shows the effect survives switching to point-in-time data, survives extending to the last five years, but collapses once the desk's realistic cost model is applied — the raw effect concentrates in small, high-cost names. The conclusion isn't "the paper is wrong." It's "the effect is real, and untradeable at the desk's cost structure" — a genuinely different, more useful finding than either "it replicates" or "it doesn't."
Announcing a failed replication without first matching the original setup is a common way for a real, if impractical, effect to get incorrectly labeled "doesn't hold up" on a desk — and then nobody checks it again for years.
Related concepts
Practice in interviews
Further reading
- Hou, Xue & Zhang (2020), Replicating Anomalies
- Harvey, Liu & Zhu (2016), ...and the Cross-Section of Expected Returns