Strategy Promotion Criteria
A strategy moving from research to paper trading to live capital needs a written bar to clear at each step, decided before anyone is emotionally invested in a specific number.
A strategy passes its backtest with a Sharpe of 1.8. There's no written rule for what happens next, so it gets a informal nod from a portfolio manager who liked the memo, skips paper trading because "the backtest already proved it," and goes live with $5 million on a Friday. Three weeks later it's down 4%, well within normal variance for a Sharpe-1.8 strategy over three weeks — but with no promotion criteria written down, nobody agreed in advance on how much early drawdown was tolerable, so the desk debates in real time whether to cut it, and cuts it at the worst possible moment, locking in a loss a strategy with genuine long-run edge would have recovered from.
What a promotion ladder does
Strategy promotion criteria define, in writing and before a specific strategy exists to argue about, the bar each stage must clear before capital or size increases: backtest passes validation at stage one, paper trading (or a small live sleeve) for a fixed minimum period at stage two, full allocation at stage three, and a size increase at stage four — each with numeric pass/fail conditions decided in advance. The value isn't the specific thresholds. It's that the thresholds exist before a strategy with a specific, emotionally persuasive backtest number is standing in front of the committee asking to skip a step.
Worked example
| Stage | Requirement | Minimum duration | Fail condition |
|---|---|---|---|
| Backtest validation | Passes written protocol, held-out Sharpe > 0.8 | — | Fails on first held-out run |
| Paper trading | Live signals, simulated fills | 3 months | Live-vs-backtest tracking error > 30% |
| Small live (10% of target size) | Real fills, real slippage | 3 months | Drawdown exceeds 1.5x backtest-implied volatility |
| Full allocation | Target size | Ongoing | Rolling 6-month Sharpe below 0.3 for two consecutive checks |
A strategy that skips paper trading — as in the opening example — skips the one stage designed to catch exactly the gap between backtested fills and real ones. Desks that enforce the ladder consistently see a large fraction of backtest-validated strategies fail at the paper-trading stage specifically because live fills, real latency, and actual liquidity degrade the edge in ways no backtest fully captures; that failure is the ladder working as intended, not a sign the process is too slow.
Promotion criteria only work if they're written and numeric before a specific strategy needs to clear them. A threshold decided in the room, after seeing the strategy's own early results, isn't a threshold — it's a rationalisation with a number attached.
The common failure is granting an exception for a strategy that "obviously" deserves to skip a stage because the backtest was unusually strong. A strong backtest is exactly the case the ladder exists to check, not the case that earns an exemption from it — an impressive backtest and a strategy that degrades badly on live fills are not mutually exclusive.
Write the fail condition, not just the pass condition, for every stage — "cut if rolling Sharpe falls below X for N periods" is enforceable in the moment; "keep going as long as it looks fine" is not, because it invites exactly the real-time debate the criteria were meant to prevent.
Setting the minimum duration honestly
The duration at each stage is as important as the numeric threshold, and it's the parameter most often shortened under pressure. Three months of paper trading is enough to see a strategy trade through a reasonable range of market conditions and enough live-fill data to compare against the backtest's cost assumptions; three weeks is not, because three weeks of any market-neutral strategy's returns is mostly noise, and a good or bad three-week stretch tells you almost nothing about the strategy's real Sharpe. The temptation to shorten the stage is strongest for exactly the strategies most likely to need the full duration — ones with a backtest so strong that everyone wants to see it running with real capital sooner. Treating the duration as fixed regardless of how the early results look is what keeps the ladder from becoming a formality that gets waived whenever a number looks good enough to make waiving it feel safe.
It also helps to separate the criteria that gate an increase in size from the criteria that gate a cut, even though they live on the same ladder. A strategy stuck at its paper-trading stage past the minimum duration without clearly failing isn't necessarily broken — it may simply need more data before a size decision either way is statistically defensible, and treating "no verdict yet" as a silent pass is how strategies with thin evidence end up fully funded by default.
Promotion criteria and a strategy kill switch answer different questions — promotion asks "should this get more capital," a kill switch asks "should this get cut right now" — but they share the same discipline of deciding the numeric bar before the emotionally loaded moment arrives. See Designing a Strategy Kill Switch and Live Versus Backtest Reconciliation for the monitoring that feeds the pass/fail checks at each stage.
Related concepts
Practice in interviews
Further reading
- López de Prado, Advances in Financial Machine Learning (ch. 8, 20)