Publication Bias and the File Drawer Problem
The distortion that appears when only statistically significant results get published or acted on, while null results quietly disappear into a file drawer — inflating how strong an effect looks in the visible record.
Prerequisites: p-values and Multiple Testing
Imagine 20 independent researchers each test whether a slightly different accounting ratio predicts stock returns, and by pure chance one gets a statistically significant result at the 5% level even if none of the ratios actually work. If only the one significant finding gets published — and the 19 null results get filed away and never mentioned — anyone reading the published literature sees a 100% "hit rate" instead of the true 1-in-20 chance result. This is the file drawer problem: the visible record is a biased sample of all the tests actually run, skewed toward the ones that happened to look good.
In quant finance the same mechanism appears whenever an anomaly gets published in an academic journal: journals favour significant, novel results, so a factor that "worked" in-sample after researchers quietly tried and discarded several variations looks far more robust in print than it deserves to. This is one reason many published anomalies show weaker or vanished returns once traded live, out of sample, after publication — the original result partly reflected which version of the test survived to publication, not a true and stable effect.
A concrete illustration: suppose 100 candidate signals are each genuinely useless (true effect exactly zero), but tested at a 5% significance threshold — by chance, about 5 will appear "significant." If those 5 get written up and the other 95 get discarded, a reader sees only apparently strong findings and no hint that 95 similar tests failed.
Publication bias means the record of published results overrepresents lucky, significant findings because null results are rarely reported, so a published effect's true strength is systematically weaker than it appears in print.
Related concepts
Practice in interviews
Further reading
- Rosenthal, The File Drawer Problem and Tolerance for Null Results, Psychological Bulletin (1979)