Publication Bias and the File Drawer
Academic finance only publishes what worked: the strategies that failed to beat a coin flip stay in a researcher's file drawer, which quietly inflates how good published factors look.
Prerequisites: The Factor Zoo and the Replication Crisis
Picture a hundred PhD students each testing a different accounting ratio against future stock returns. Five of them, purely by chance, find a ratio that beats the market with a statistically significant t-stat. Those five write papers and get published. The other ninety-five find nothing, and their results never see daylight — no journal wants a paper whose conclusion is "this didn't work." That silent pile of negative results is the file drawer, and it is why the published factor literature looks far more impressive than the underlying reality.
Every published factor was drawn from a much larger pool of tested-and-discarded ideas. Because only the winners get published, the published t-statistics systematically overstate how strong the true effect is.
Why this isn't just bad luck
This is not accusing anyone of fraud. A researcher who finds nothing publishable simply moves to a different project — there's no incentive to spend months writing up a null result, and few journals would accept it anyway. But from the outside, all a reader ever sees is the finished, significant papers. If a hundred independent tests are run at a 5% significance threshold and the true effect is zero, roughly five will look "significant" by pure chance. Those five get written up. The file drawer problem is this selection process operating silently across the entire academic industry, compounding with the multiple-testing problem of any one researcher trying many variables.
Worked example
Suppose 200 accounting ratios are tested independently against stock returns, and none of them actually predicts anything — the true population is pure noise. At a 5% significance cutoff, chance alone produces about ratios that clear the bar with an impressive-looking t-stat above roughly 2.0. If ten different research teams each publish one of those ten "discoveries" as a standalone paper, a reader scanning the journals sees ten confirmed factors and no failures — because the 190 non-findings were never written up. The reader has no way to know the true hit rate was 10 out of 200, not 10 out of 10.
What this means in practice
A practitioner reading a new factor paper should treat the published t-stat as an upper bound on plausibility, not a fair estimate. Firms with large internal research operations see the file drawer directly: for every strategy that reaches production, dozens were tested and discarded, and that ratio is a far more honest guide to base rates than anything in a journal. This is one reason serious factor investors demand out-of-sample and out-of-universe evidence before trusting a new signal, and why replication studies routinely find published effects shrink once the file drawer is accounted for.
Publication bias is easy to confuse with p-hacking, but they are different problems. P-hacking is one researcher trying many specifications on the same data until something works. The file drawer is the field-wide version: many different researchers each trying one idea, with only the successes reaching print. Correcting for one does not correct for the other, and most real factor claims suffer from both at once.
Related concepts
Practice in interviews
Further reading
- Harvey, Liu, Zhu, '...and the Cross-Section of Expected Returns' (Review of Financial Studies)
- McLean, Pontiff, 'Does Academic Research Destroy Stock Return Predictability?' (Journal of Finance)