Quant Memo
Core

Choosing an Execution Benchmark

The same trade can be scored as a 22 bps loss or a 40 bps win depending on what you compare it to. Picking the benchmark is not a reporting decision — it is an instruction to the algo about what to optimise.

Prerequisites: Implementation Shortfall, TWAP, VWAP & POV

A portfolio manager decides to buy 200,000 shares when the stock is at $40.00. The order reaches the trading desk a few minutes later at $40.02. The desk works it through the day, averaging $40.09 a share. The stock closes at $40.25, and the volume-weighted average price over the working interval was $40.11.

Now grade the trader.

BenchmarkLevelSlippage on the buy
Decision price40.00+22.5 bps (cost)
Arrival price40.02+17.5 bps (cost)
Interval VWAP40.11−5.0 bps (beat it)
Closing price40.25−39.8 bps (beat it)

One execution, four verdicts, spanning sixty basis points. Nothing about the trade changed. Only the question did.

trading window decision arrival VWAP avg fill close 09:30 16:00
The same average fill sits below two reference levels and above two others. Which line you draw the comparison against decides whether the day was a success.

What each benchmark is actually asking

  • Decision price — "what did the whole idea cost, end to end?" It includes the delay between the manager's decision and the order reaching the desk, which the trader usually does not control. Honest for the fund, unfair as a trader scorecard.
  • Arrival price — "from the moment you owned this order, how much did you give away?" This is Implementation Shortfall as traders are normally measured on it. It charges you for both market impact and the price drifting away while you were patient.
  • Interval VWAP — "did you beat the average participant?" It forgives drift entirely: if the stock ran 25 cents, VWAP ran with it.
  • Closing price — "did you get the mark?" The right question when your fund is valued at the close, or when the trade exists to track an index that rebalances there.
  • TWAP and PWP-n (the price of the first n×n \times your size traded after arrival) are variations on "average", used where volume data is thin or manipulable.

A benchmark is not a measurement, it is an instruction. Measure a desk on VWAP and it will trade with the volume curve and happily accept drift risk. Measure it on arrival price and it will front-load and pay impact. Both are rational responses to the scorecard you handed over.

Why VWAP flatters you: the self-reference problem

Interval VWAP includes your own trades. Here the interval printed 500,000 shares: your 200,000 at $40.09 and everyone else's 300,000 at $40.1233, giving

VWAP=200,000×40.09+300,000×40.1233500,000=40.11.\text{VWAP} = \frac{200{,}000 \times 40.09 + 300{,}000 \times 40.1233}{500{,}000} = 40.11.

Now suppose the desk had traded badly and averaged $40.20 instead — 11 cents worse, $22,000 of real money. The VWAP itself moves to (200,000×40.20+300,000×40.1233)/500,000=40.154(200{,}000 \times 40.20 + 300{,}000 \times 40.1233)/500{,}000 = 40.154, so the measured slippage goes from −5.0 bps to only +11.5 bps. A 27.5 bps deterioration in reality shows up as 16.5 bps on the report.

That damping factor is exactly 1participation rate=10.40=0.601 - \text{participation rate} = 1 - 0.40 = 0.60. At 40% of the volume you can only ever be graded on 60% of your own mistakes; at 80% participation, VWAP is almost a mirror.

Matching the benchmark to the reason for the trade

Why you are tradingUseBecause
Fast-decaying alpha signalArrival priceDelay is a real cost and must be charged
Index tracking, fund flowsClosing priceYour NAV is struck there
Patient rebalance, no timing viewInterval VWAP / TWAPBeing average is genuinely the goal
Broker risk transfer / guaranteed VWAPArrival, with the guarantee priced inRisk was transferred, so price it up front
Illiquid names, thin printsTWAP or PWPVolume-based benchmarks are too easy to distort

Report against two benchmarks, not one: arrival price and interval VWAP. The gap between them is a clean read on the day's drift — how much of the outcome was execution skill versus which way the stock happened to go while you worked.

Every benchmark is gameable and the gaming is usually legal. VWAP is diluted by your own participation and by choosing when the "interval" starts and ends. Arrival price is gamed by timestamping arrival late, after the price has already moved. Closing benchmarks are gamed by parking everything in the auction and calling the resulting impact "the mark". Worst of all, benchmarks that only score filled shares reward not trading — always carry the opportunity cost of the unfilled portion into the number, or a trader who cancels the hard half looks like a hero.

In interviews

The reliable question is "you beat VWAP by 5 bps — good day?" The expected answer is it depends what the trade was for, followed by the self-reference point: if you were 40% of the interval's volume, beating VWAP by 5 bps is close to beating yourself. Then contrast it with arrival price, where the same execution cost 17.5 bps, and say plainly which one you would put on a scorecard and why. Being able to name the incentive each benchmark creates — VWAP makes you patient, arrival makes you aggressive — is the part that shows you have thought about it as a control system, not a report. See TWAP, VWAP & POV for the algos and Implementation Shortfall for the full cost decomposition.

Related concepts

Practice in interviews

Further reading

  • Perold (1988), The Implementation Shortfall: Paper versus Reality
  • Kissell, The Science of Algorithmic Trading and Portfolio Management
  • Berkowitz, Logue & Noser (1988), The Total Cost of Transactions on the NYSE
ShareTwitterLinkedIn