Quant Memo
Core

Learning Inventory Skew From Fill Data

How a market maker can learn, from historical fill data, how much to skew quotes away from fair value as inventory builds up — rather than relying on a single hand-set skew parameter.

Prerequisites: Adverse Selection, Market Impact

A market maker who has just bought a large position doesn't want to keep buying at the same eager price — every additional share adds to a position that's already uncomfortably large, and if the market moves against that inventory before it can be unwound, the loss compounds. The standard response is to skew quotes: shift both the bid and ask down when long, making it less attractive for others to sell to the maker and more attractive for them to buy, gently steering inventory back toward flat. How aggressively to skew, though, is a genuine quantity to get right — too little and inventory risk accumulates unchecked; too much and the maker gives away edge on every quote just to manage a position that might have mean-reverted on its own.

The classic formula and its limits

A well-known theoretical result, from the Avellaneda-Stoikov market-making framework, derives an optimal skew proportional to current inventory, risk aversion, and volatility — a clean formula that gives real intuition (skew more when volatility is high and risk tolerance is low) but rests on simplifying assumptions: that order flow arrives in a particular idealized way, and that the maker's own quotes have no effect on the pattern of arrivals. Real fill data lets a maker check whether those assumptions hold for a specific instrument and market regime, and if not, learn a skew that fits the data actually observed rather than the idealized model.

Learning skew empirically

The approach is to look at historical episodes across a range of inventory levels and measure the realized outcome — subsequent markout and time-to-flatten — as a function of the skew applied at each level. If empirically doubling the theoretical skew at large inventory levels shows meaningfully better subsequent markouts (because it more effectively attracts offsetting flow) without costing much in foregone trades, that's a signal the theoretical formula is under-skewing for this instrument. Features beyond inventory itself — recent realized volatility, current spread, and the maker's own recent fill rate — let the learned skew respond to conditions the closed-form formula doesn't capture, like a temporarily one-sided order-flow imbalance that calls for extra caution even at moderate inventory.

Worked example: comparing two skew rules at the same inventory

A market maker is long 5,000 shares against a normal trading unit of 1,000. The theoretical formula prescribes a skew of 2 ticks. Historical data shows that at this inventory level, episodes where the desk applied a 2-tick skew took an average of 14 minutes to return to flat inventory, with an average markout of 3-3bp during that period as the price occasionally drifted against the position while it was being worked off. Episodes where the desk (for unrelated reasons) applied a 3.5-tick skew at similar inventory returned to flat in an average of 8 minutes, with markout of 1-1bp — faster mean reversion in inventory with less exposure to adverse drift, at the modest cost of slightly worse average fill prices on the way down. This comparison suggests the theoretical 2-tick skew is too conservative for this instrument's actual order-flow sensitivity, and a learned rule would shift toward something closer to 3–3.5 ticks at this inventory level.

Time to flatten (min) 2-tick: 14 3.5-tick: 8 Markout (bp) 2-tick: -3 3.5-tick: -1
At the same 5,000-share inventory, the more aggressive empirical skew flattens position faster with a smaller adverse markout, at the cost of a slightly worse average fill price.

What this means in practice

Learning skew from data doesn't replace the theoretical framework — it calibrates it, replacing an assumed proportionality constant with one fitted to how order flow actually responds on that instrument, at that liquidity, in that regime. Because skew directly trades off giving away edge (wider effective spread on one side) against inventory risk, and because both markout and time-to-flatten are noisy, this calibration needs a meaningful volume of historical episodes across a range of inventory levels to be trustworthy, and should be re-checked periodically as market conditions shift.

The optimal quote skew as inventory builds up can be learned empirically from historical fill data — measuring time-to-flatten and subsequent markout at different skew levels and inventory sizes — as a calibration on top of theoretical formulas like Avellaneda-Stoikov, which rest on idealized assumptions about order-flow arrival that real markets don't always satisfy.

Related concepts

Practice in interviews

Further reading

  • Guéant, The Financial Mathematics of Market Making, ch. 2
  • Avellaneda, Stoikov, High-Frequency Trading in a Limit Order Book
ShareTwitterLinkedIn