Reporting What You Tried and Discarded
Disclosing the variants, parameters, and approaches that didn't make it into the final research note is what lets a reader judge how much of the reported result might be the product of search rather than signal.
A research note that presents only the winning strategy variant, with no mention of what else was tried, gives a reader no way to judge how much of the search was involved in getting there. A signal that was the first thing tried and worked is a much stronger piece of evidence than a signal that was the best of forty variants tried over three months — even if the final backtest numbers look identical, the second one has had far more chances for noise to look like a pattern. Without disclosure, a reader can't tell which situation they're looking at.
The discipline is to keep, and eventually share, a record of what was attempted and discarded: other lookback windows that were tested and underperformed, other universes the signal didn't work in, other formulations of the same idea that seemed reasonable but didn't pan out. This doesn't need to be exhaustive detail in the main write-up — a short section or appendix listing "also tested: 10-day and 60-day variants of this signal (both weaker), applied to European equities (no effect found), combined with a volume filter (no improvement)" gives a reader a real sense of how much searching produced the final result, without turning the note into a lab notebook.
This kind of disclosure is uncomfortable because it can make a result look less impressive — "we tried twelve things and this one worked" is a less clean story than "we hypothesized this and it worked." But the discomfort is the point: a reader who only sees the clean story has no way to apply the discount that multiple testing actually warrants, and will trust the reported Sharpe ratio more than the underlying process earned. Disclosing the search is what lets the reader, not just the author, decide how much to trust the result.
Reporting only the winning variant hides exactly the information a reader needs to judge overfitting risk — how many things were tried. A short account of what was tested and discarded lets a reader apply their own discount rather than trusting an undisclosed search.
"We didn't keep track of everything we tried" is a common excuse, but it's a process failure worth fixing going forward — without some record, neither the author nor the reader can ever separate a real finding from the best of an unknown number of attempts.
Further reading
- Bailey, Borwein, López de Prado & Zhu, 'The Probability of Backtest Overfitting'