Requiring an Economic Rationale
A backtest that passes every statistical test but can't explain why the market would pay for the signal is still probably noise. Requiring a rationale before the data mining starts filters most of it out for free.
Feed a machine enough columns of historical data and it will find something that "worked." A well-known illustration: monthly US stock returns correlated, for a multi-year stretch, with butter production in Bangladesh at a level that would clear most statistical significance tests used in factor research. Nobody trades that signal, not because the correlation wasn't real in-sample, but because there is no mechanism connecting Bangladeshi dairy output to the price of US equities. The correlation was real and worthless at the same time — a distinction pure statistics can't make on its own.
Why "it's significant" isn't enough
Statistical significance answers "could this pattern have arisen by chance, given how much I searched?" It does not answer "why would this pattern persist once other people notice it?" Only an economic rationale answers the second question, and the second question is the one that determines whether the pattern is tradeable next year or gone. A rationale doesn't have to be exotic — "index rebalancing forces mechanical buying that dealers front-run," "earnings announcements are underreacted to because most investors are attention-constrained," "small caps are illiquid and undercovered, so mispricings persist longer" — but it has to identify a reason a profit opportunity would survive being found, whether that's a structural friction, a behavioural bias, or a risk premium someone is paid to bear.
Worked example
A researcher screens 400 candidate signals built from combinations of price, volume, and calendar features against five years of daily US equity data. Roughly 20 clear the standard bar of a backtested Sharpe above 1.0 with a t-statistic above 2 — exactly the number multiple-testing theory predicts will pass by chance alone at that sample size and search breadth, before a single one is inspected for a mechanism.
Requiring a written rationale before backtesting cuts the candidate list from 400 to 35 — the researcher only builds and tests signals they can already explain in one sentence tied to a market mechanism. Of those 35, 6 clear the same statistical bar. Six real candidates from 35 rationale-first tries is a far healthier hit rate than 20 from 400 blind ones, and — this is the part that matters — the 6 that survive are far more likely to still work out of sample, because they weren't selected purely for fitting noise well.
| Screening approach | Candidates tested | Statistically significant | Plausible economic rationale |
|---|---|---|---|
| Unrestricted data mining | 400 | 20 | Unknown until checked after the fact |
| Rationale required before testing | 35 | 6 | 6 (by construction) |
A rationale is a filter applied before the statistics, not a story invented after a good backtest to make it sound respectable. If you can't state the mechanism in one sentence before running the test, the eventual significance is much more likely to be noise.
The trap is writing the rationale after seeing the result — "it works because momentum investors chase recent winners" bolted onto a signal that was actually discovered by unrestricted search. A post-hoc rationale can always be constructed for any pattern; that flexibility is exactly why it provides no filtering power. The rationale only does its job if it existed, in writing, before the backtest ran.
A useful discipline: before writing a single line of backtest code, write down who is on the other side of the trade and why they'd keep losing to it. If you can't name a counterparty and a reason, the signal doesn't get built yet.
An economic rationale doesn't replace statistical validation — a plausible story with no significant backtest result is just a story. It's a precondition that should run first, cheaply, and reject most candidate ideas before they ever reach the expensive, careful validation protocol described in Designing a Validation Protocol Up Front. See The False Strategy Theorem for the statistical machinery behind why unrestricted search produces exactly this many false positives, and The Replication Crisis in Factor Research for what happens at scale when the field skips this filter.
Related concepts
Practice in interviews
Further reading
- Harvey, Liu & Zhu, ...and the Cross-Section of Expected Returns
- Israel & Moskowitz, The Role of Shorting, Firm Size, and Time on Market Anomalies