Order Fill Modeling in a Backtest
The rule your simulator uses to decide whether an order filled, and at what price, usually matters more than the signal. Four defensible fill rules on the same strategy can produce four completely different businesses.
Prerequisites: Market vs. Limit Orders, Transaction Costs
A backtest reports the P&L of two things bolted together: a signal, and an assumption about how orders turn into trades. Researchers spend weeks on the first and thirty seconds on the second. That ratio is backwards. For anything faster than weekly rebalancing, the fill assumption is usually the larger term, and it is the one that will not survive contact with a real broker.
Historical data does not contain your fills. It contains what happened in a market you were not in. The simulator invents the counterfactual, and every way of inventing it embeds a liquidity claim you never had to justify.
The ladder of fill rules, from fantasy to defensible
- Fill at the signal bar's close. You decided using the close and traded at the close. Free money; a pure one-bar look-ahead. Never acceptable.
- Fill at the next bar's open. The minimum honest rule for daily work, and still optimistic: it ignores the spread, the opening auction imbalance, and the overnight gap where many mean-reversion signals get their apparent edge.
- Fill at the next bar's VWAP or mid, plus the half-spread. Reasonable for medium-frequency equity work, if your size is small relative to that bar's volume.
- Mid, plus half-spread, plus a size-dependent impact term. The standard is the square-root law, — cost grows with your order size relative to daily volume , but sublinearly, so doubling the order costs about 1.4 times as much rather than twice. See The Square-Root Impact Law.
- Queue-simulated limit orders. Required if you post passively. You sit behind whatever was already resting at your price, and fill only once enough volume clears the queue ahead of you.
Rank the same strategy under rules 1 through 5 before you believe anything. If the ordering of your variants changes as the fill rule gets stricter, you have been ranking execution assumptions, not signals.
Worked example: the bar that touched both
A trend strategy on 5-minute bars, long, with a 50 bps take-profit and a 50 bps stop. On one bar the prices are open 100.00, high 100.60, low 99.40, close 100.10. Your target sat at 100.50 and your stop at 99.50. Both were touched inside that bar. The four numbers do not say which came first.
If the simulator resolves the tie by checking the target first, the trade books +50 bps. If it checks the stop first, it books −50 bps. That is a one-basis-point coding decision worth a full 100 bps per occurrence.
Over a three-year backtest, 480 of 4,100 trades land on such double-touch bars — about 12%. Target-first: +240 bps a year from those bars alone, and a headline Sharpe of 1.9. Stop-first: −240 bps a year and a Sharpe of 0.4. Reality is worse than the midpoint, because double-touch bars are volatile bars, and that is exactly when a stop slips through instead of filling at its limit.
The explorer makes the same point with random walks. Drag resample and watch how differently the paths wander between similar endpoints; push volatility up and count how many would have tripped a stop on the way.
Worked example: the limit order that never filled
A passive mean-reversion strategy posts bids 2 ticks below the mid. The simulator's rule is the usual one: if the bar's low touched my price, I filled.
Backtest: 68% of orders fill, average +2.1 bps per fill, 40,000 fills a year, gross 840 bps.
Now simulate the queue. Your order joins behind an average of 38,000 shares resting at that price; you post 1,000. To fill, roughly 38,000 shares must trade at or through your level while it stays the best bid. Historically that happened on 22% of the touches, not 68%.
Worse, the fills are not a random 22%. You fill when a genuine seller keeps pushing — precisely when the price is about to keep falling. Conditional on actually being filled, the average outcome is −0.4 bps, not +2.1. Post-queue: 13,000 fills a year at −0.4 bps, gross −52 bps before fees. The strategy was never positive; the simulator was handing it the good half of a distribution it would never have seen. See Adverse Selection and Queue Position and Priority.
"The low touched my limit, so I filled" is the most common single lie in retail and academic backtests. A touch is one trade at your price, not clearance of the queue ahead of you — and the touches that do clear the queue are the ones where you did not want the fill.
Run every strategy once under a deliberately pessimistic fill rule: stop-first on ambiguous bars, next-bar VWAP plus the full spread, and passive fills only when the price trades through your level rather than to it. If the strategy survives that, the honest number is somewhere between. If it does not survive, you have saved yourself a live pilot.
In practice
Report P&L as a function of the fill rule, the way you would report sensitivity to any other parameter — a table of net Sharpe under rules 2 through 5 says far more than one Sharpe under an unstated assumption. Then check it against reality: once even a little live trading exists, plot realised fill prices against the arrival mid and against the simulator's prediction. That gap, in basis points per trade, tells you whether your research process is calibrated. See Implementation Shortfall.
Related concepts
Practice in interviews
Further reading
- López de Prado, Advances in Financial Machine Learning (Ch. 13)
- Bouchaud, Bonart, Donier & Gould, Trades, Quotes and Prices
- Almgren & Chriss, Optimal Execution of Portfolio Transactions