Quant Memo
Core

Sample Period Red Flags in Published Anomalies

Why the exact start date, end date, and any gaps in a paper's sample period deserve scrutiny — a convenient window can manufacture a result that a longer or later one erases.

Prerequisites: Reading a Paper Adversarially

Two questions to ask about any published anomaly's sample period, before reading a single result: why does the sample start when it does, and why does it end when it does? Both dates are choices, and both can be chosen — consciously or not — because they happen to produce the strongest result.

A start date is suspicious when it skips an inconvenient period without explanation. A momentum strategy that starts its sample in 1965 rather than 1926, when good U.S. equity data actually exists back to 1926, might simply be avoiding decades that would have weakened the finding. An end date is suspicious when it stops right before a period that would plausibly test the strategy under different conditions — a volatility strategy whose sample ends in 2019, conveniently missing the 2020 crash, invites the question of whether the author looked and didn't like what they saw. Gaps inside the sample are a further flag: a paper that excludes 2008–2009 "due to unusual market conditions" is removing exactly the period most likely to break a fragile strategy.

None of these, on their own, prove anything is wrong — sometimes data genuinely doesn't exist before a certain date, or a crisis period genuinely behaves differently enough to warrant separate treatment. The point is that an undisclosed reason for a convenient window is a red flag worth chasing down, and the way to chase it is simple: pull the same signal on the full available history, including the periods the paper left out, and see if the result survives.

It's also worth checking whether the sample period lines up suspiciously well with when the underlying data became conveniently available to the author, rather than with any economic reasoning. A paper using a proprietary dataset that starts in 2010 because that's when the vendor's coverage begins isn't necessarily doing anything wrong, but the strategy's true performance before 2010 is simply unknown, and readers should treat the reported sample as the full extent of the evidence rather than assuming it would have looked the same further back.

Any sample period that starts, ends, or has gaps at a moment that happens to help the result — without a stated, non-circular reason — should be treated as unverified until re-tested on the fuller history.

"We excluded the financial crisis because it was unusual" is one of the most common ways a fragile result gets dressed up as robust. Crises are exactly when strategies are supposed to be tested, not excused from testing.

Related concepts

Further reading

  • McLean & Pontiff, 'Does Academic Research Destroy Stock Return Predictability?'
ShareTwitterLinkedIn