Quant Memo
Advanced

Capacity-Constrained Backtesting

A normal backtest trades an unlimited amount at the printed price, so its Sharpe is the Sharpe at zero dollars of capital. Capacity-constrained backtesting re-runs the strategy at real book sizes and reports the curve instead of the number.

Prerequisites: Market Impact, Portfolio Capacity

An ordinary backtest has no size. It buys as much as it likes at whatever price the tape printed, as if the market would have absorbed any quantity without flinching. That is a fine approximation when you are trading a hundred thousand dollars of Apple. It is a fantasy when you are trading fifty million dollars of micro-caps, and the gap between the two is not a detail — it is usually the difference between a strategy and a hobby.

Think of a restaurant with a superb recipe and four tables. The recipe's quality per plate is real and measurable. It tells you nothing about how much money the restaurant can make, because at some point you are seating people on the pavement and the food gets cold. Capacity is the question of how many tables the market gives you.

Where the money leaks

Cost per trade does not stay put as you scale. The empirical workhorse is the square-root law: pushing QQ shares through a market that trades VV shares a day costs roughly

impact    κσQ/V,\text{impact} \;\approx\; \kappa\,\sigma\,\sqrt{Q/V},

which in words says the price moves against you in proportion to the square root of your participation rate, scaled by the name's volatility σ\sigma and a fitted constant κ\kappa. Trading four times as much costs about twice as much per share, so eight times as much in total. Your gross edge per trade, meanwhile, is a fixed number of basis points. Somewhere those two lines cross. See The Square-Root Impact Law.

Worked example: the small-cap reversal that dies at 200 million

A five-day reversal signal on US small and mid caps. Equal-weight, 300 names, the median name trades $20 million a day, and the book turns over 50 round trips a year. Gross edge measured in a zero-size backtest: 40 bps per round trip, which annualises to 20% gross on 8% volatility — a gross Sharpe of 2.5.

Take the impact model to be 100 bps times the square root of your participation rate pp (your daily order divided by that name's daily volume), and remember you pay it twice, once in and once out. Net edge per round trip is

net  =  402×100pbps.\text{net} \;=\; 40 - 2 \times 100\sqrt{p} \quad \text{bps}.

Now walk the book size up. A $60 million book spread over 300 names is $200,000 per name, which is 1% of a $20 million daily volume.

BookParticipationImpact both waysNet per tradeNet annualNet Sharpe
$5m0.08%5.7 bps34.3 bps17.2%2.1
$60m1.0%20 bps20 bps10.0%1.25
$120m2.0%28.3 bps11.7 bps5.9%0.73
$240m4.0%40 bps0 bps0%0
$400m6.7%51.6 bps−11.6 bps−5.8%−0.72

The headline Sharpe of 2.5 is only available to someone trading five million dollars. At $240 million the strategy is exactly break-even, and past that it pays the market to take its own idea.

Now the surprising part. Dollar profit is book size times net return, and that product peaks before break-even but long after the Sharpe has degraded. Maximising W×(40200p)W \times (40 - 200\sqrt{p}) over book size WW puts the peak near $107 million, earning 6.7% net — about $7.1 million a year, at a Sharpe of 0.83. The size that makes the most money and the size that looks best on a tear sheet are nowhere near each other, and neither is the size the backtest reported.

max dollars break-even net Sharpe annual profit book size
Sharpe falls monotonically with size; dollar profit rises, peaks, then collapses. A single backtest number is the value of the green line at the far left edge, where nobody actually trades.

The Sharpe in your backtest is the Sharpe at zero assets. The deliverable is a capacity curve — net Sharpe and net dollars against book size — plus the two sizes that matter: where profit is maximised and where the edge hits zero.

The trap is treating capacity as a final sanity check rather than a design input. Every choice made against the zero-size backtest — universe, holding period, rebalance frequency — was optimised for a market that does not push back. A slower version of this strategy with 12 round trips a year instead of 50 shows a lower paper Sharpe and a break-even capacity several times higher. Rank your variants at your intended book size, not at zero.

Ten-second estimate: cap yourself at 5% of each name's daily volume, so book 0.05×N×ADV\le 0.05 \times N \times \text{ADV}. Here that is $300 million — above the $240 million break-even, which is the lesson. The participation cap tells you what is executable; the impact model tells you what is profitable, and the second binds first.

Doing it properly

Re-run the whole simulation at several book sizes rather than scaling one P&L series after the fact, because the constraints bind unevenly. Liquidity filters knock names out at large size; the position you wanted is capped, so the portfolio drifts away from the target weights; a multi-day unwind means you hold stale positions the fast version exited on day one. Also let the strategy's own trading move prices — if your book is 5% of a name's volume, the price you get tomorrow reflects what you did today. And check where the capacity sits by regime: a strategy whose signal fires hardest in a liquidity crisis has its worst capacity precisely when it most wants to trade. See Portfolio Capacity and Alpha Decay.

Related concepts

Practice in interviews

Further reading

  • Almgren & Chriss, Optimal Execution of Portfolio Transactions
  • Frazzini, Israel & Moskowitz, Trading Costs of Asset Pricing Anomalies
  • Kyle & Obizhaeva, Market Microstructure Invariance
ShareTwitterLinkedIn