Defining Success Criteria Before You Look
Deciding, before running a test, exactly what result would count as a success — because a bar set after seeing the number always gets set to fit whatever the number happened to be.
Prerequisites: Writing a Research Brief Before You Start
A researcher runs a backtest and gets a Sharpe ratio of 0.6. Is that good? Without a bar set in advance, the honest answer is "it depends on what you were hoping for" — and that's exactly the problem. Left unconstrained, most people's judgment of "good enough" quietly adjusts to match whatever number they got, especially after weeks of work on an idea they want to succeed. A Sharpe of 0.6 gets called "promising, worth refining" if that's the number in hand, even though the same researcher, asked in advance, might have said a Sharpe below 1.0 wasn't worth pursuing further.
Setting the bar before the test
Defining success criteria upfront means writing down, before the test runs, a specific threshold and the reasoning behind it: "this needs a Sharpe above 1.0 net of estimated transaction costs to be worth building out, because that's roughly what it would take to be additive to the existing portfolio rather than just diversifying noise." The number matters less than the discipline of committing to one before the result is known — a bar chosen after the fact isn't really a bar at all, it's a description of whatever happened, dressed up to look like a decision rule.
This matters most for judgment calls that don't have an obvious external answer: how big a sample counts as "enough," how much of a drawdown is "acceptable," whether a result needs to hold across sub-periods to count as robust. Each of these is a place where post-hoc reasoning quietly bends toward whatever conclusion the researcher already wanted, unless the bar was fixed before looking.
Worked example
Two researchers each test a signal and both get a marginal result: a small positive effect that's not quite statistically significant at conventional levels. The one who set a bar in advance — "needs to clear a specific significance threshold, tested out of sample, to move forward" — has a clean answer: this doesn't clear the bar, park it. The one without a pre-set bar spends an afternoon trying different sample windows and small variations until one of them happens to clear significance, then writes it up as a positive finding — a result manufactured by the search itself, not by the underlying idea.
What this means in practice
Committing to success criteria before looking at results is one of the cheapest defenses against fooling yourself in research, and it costs nothing but a few minutes of writing before the test runs. The discipline is easiest to keep when the bar is written somewhere durable — in the research brief, shared with a colleague — rather than held only in the researcher's head, where it's free to quietly move.
A success bar set after seeing the result isn't a bar — it's a rationalization. Writing down what would count as success before running a test, and sharing that bar with someone else, is a simple and effective defense against unconsciously moving the goalposts.
Further reading
- Grinold and Kahn, Active Portfolio Management (ch. 1)