Gap-Fill Statistics and How to Condition Them
A stock that opens sharply away from its previous close sometimes drifts back to fill the gap and sometimes keeps running. The raw fill rate is close to a coin flip — the edge is in conditioning it on gap size, direction, volume, and what caused the gap in the first place.
Prerequisites: Overnight vs Intraday Returns
A stretched rubber band and a rope both look "displaced from rest" the moment you pull one end, but they behave completely differently once you let go: the rubber band snaps back, the rope just stays where you dragged it. A stock's overnight price gap looks the same way to a naive observer — a jump away from yesterday's close — but whether it snaps back ("fills the gap") or keeps drifting depends entirely on why it gapped, and treating every gap the same way is the single biggest mistake in gap-fill trading.
The unconditional rate is close to useless
Define the gap as the overnight change from yesterday's close to today's open, and "fill" as the price returning to yesterday's close at some point during the trading day:
In words: how far today's opening price sits above or below yesterday's closing price, as a percentage. Pooling every stock and every day, the unconditional probability that a gap fills by end of day sits close to 50%, which is roughly what you'd expect if opens were a random walk continuation of closes with no special structure — not a tradeable edge on its own. The edge, when there is one, only shows up after conditioning: splitting gaps by size, direction, the news behind them, and recent volatility, and measuring the fill rate separately within each bucket.
Worked example 1. Suppose historical data on small gaps (under 1%, no identifiable news catalyst) shows a fill rate of 54% — a modest edge over the coin-flip baseline, consistent with those gaps being mostly noise from overnight order imbalance that mean-reverts once the full order book reopens. Now condition on large gaps (over 4%) driven by an identifiable earnings surprise: the fill rate for these drops to roughly 30%. In words, a plain-language restatement: small, unexplained gaps behave like noise and tend to close; large gaps with a real information catalyst behave like a genuine repricing and tend to persist. A strategy that fades every gap indiscriminately is implicitly betting the 54% edge applies everywhere, when the data shows it inverts for exactly the largest, most tempting-looking gaps.
Compare paths here to a genuine trending path in your head — an unconditioned gap-fade strategy is betting every gap behaves like the mean-reverting path shown; conditioning on gap size and catalyst is how you find out, case by case, which gaps actually do.
Building the conditioning table
Worked example 2. A more complete conditioning scheme buckets gaps by size and by whether volume in the first 5 minutes is elevated relative to the stock's normal opening volume — high early volume is a proxy for how much new information is actually being absorbed. Suppose the data shows: small gap (under 1.5%) + normal volume → 58% fill rate; small gap + elevated volume (2x normal) → 47% fill rate; large gap (over 3%) + normal volume → 44% fill rate; large gap + elevated volume → 27% fill rate. Trading only the first bucket — small gaps with unremarkable volume, betting on a fill — with an average fill move of 0.8% and using a stop that caps the loss on a non-fill at 0.6%, the expected value per trade is — a real, positive edge per trade before costs, built entirely by conditioning down to the specific bucket where it exists, and specifically avoiding the buckets where the raw statistics say it doesn't.
The unconditional gap-fill rate is close to 50% and not tradeable. The edge lives entirely in the conditioning: gap size, direction, and a proxy for how much real information is behind the move (news catalyst, early volume) split a noisy 50/50 average into buckets that behave very differently — some fadeable, some not.
What this means in practice
Building a robust conditioning table takes years of clean intraday data across many names and requires care that the buckets used live were decided before looking at the results they produced, not carved out after the fact to make a backtest look good (see Overnight Gaps: Continuation or Fade for the closely related continuation side of this same trade). The strategy also needs enough liquidity in the first minutes of trading to actually execute at prices close to the open — gap trades are, almost by definition, entered into some of the most volatile and widest-spread minutes of the trading day.
The classic confusion: conditioning on so many variables that each bucket has too few historical observations to trust. Splitting gaps by size, direction, sector, day of week, and volume can leave a bucket with only a few dozen historical instances — enough to produce an eye-catching fill rate by chance, not enough to trust with real capital. Check the sample size in each bucket before trading its statistics, the same discipline as any other data-mined subgroup.
Related concepts
Practice in interviews
Further reading
- Fung & Lo (adapted), Intraday Pattern Studies
- Jegadeesh (1990), Evidence of Predictable Behavior of Security Returns