Transaction-Cost-Aware Portfolio Optimization
The best portfolio to hold and the best portfolio to trade into are different things. Putting trading costs inside the optimizer, rather than subtracting them afterwards, produces a no-trade region and a partial-adjustment rule instead of a chase.
Prerequisites: Pitfalls of Mean-Variance Optimization, Transaction Costs, Market Impact
Here is a failure mode that has killed more live strategies than bad forecasting. A researcher builds a signal, feeds expected returns and a covariance matrix into an optimizer, gets a set of target weights, and rebalances to them every day. The backtest looks excellent. Then they subtract an estimate of trading costs at the end and the whole edge evaporates. So they try again with a lower turnover cap, and again with a slower signal, and each time the answer feels arbitrary.
The problem is not the cost estimate. The problem is where the costs were applied. Subtracting costs after optimising asks "what does my ideal portfolio cost to reach?" The right question is "given that trading costs money, what should I actually own?" Those give different answers, and only the second one is a plan.
An analogy before any symbols
You live in a decent apartment. A better one comes on the market across town: nicer light, shorter commute, slightly cheaper. Do you move?
Not necessarily. Moving costs you a deposit, a van, two days off work, and a week of chaos. So you only move if the new place is better by more than the cost of moving. If it is only a little better, you stay put, even though you know a better apartment exists and you know exactly where it is.
Now notice what this implies. There is a whole range of apartments that are better than yours but not better enough to justify the move. Inside that range you do nothing. That range is the no-trade region, and it is the single most important consequence of taking costs seriously. Notice also that if you do decide to move, you might not move all the way across town, you might take a place halfway that captures most of the gain for a fraction of the disruption. That is partial adjustment, the second consequence.
Putting costs inside the objective
Standard mean-variance picks weights to maximise expected return minus a penalty for risk. Write for what you currently hold and for the trade. The cost-aware version adds a third term:
In words: pick the holding that gives the best combination of forecast return and portfolio risk, after paying for the trades needed to reach it. The only new ingredient is , and the shape you assume for it changes the answer completely.
The objective now depends on . That is the conceptual break with Markowitz: there is no longer a single optimal portfolio, only an optimal portfolio given where you are standing. Two managers with identical forecasts and identical risk models should hold different things if they started from different places.
Linear costs give you a no-trade region
Spreads, commissions and taxes are roughly proportional to how much you trade: , where is the round-trip cost of trading one unit of asset . In words: every dollar you push through the market costs a fixed number of basis points.
The absolute value matters enormously. has a kink at zero, it is V-shaped, not smooth. A kink in the cost means the marginal cost of the very first dollar you trade is already , not zero. So there is a threshold: unless the improvement from trading exceeds per dollar right away, the optimum is to trade nothing at all. Small mispositionings simply are not worth fixing.
Impact costs give you partial adjustment
Your own trading moves prices, and the bigger the order the worse the fill. A convenient stand-in is a quadratic cost, , in words: cost grows with the square of trade size, so doubling the order roughly quadruples the impact bill. (Empirically the exponent is closer to 1.5 than 2, see The Square-Root Impact Law, but quadratic keeps the problem solvable in closed form.)
Quadratic costs are smooth at zero, so there is no threshold, you always trade a little. But the marginal cost rises as you push harder, so you stop short of the target. In the one-asset case the optimum is a simple weighted average:
which in words says: land somewhere between where you are and where you want to be, closer to the target when risk-adjusted alpha is compelling ( large) and closer to home when trading is expensive ( large). Repeat this each day and your holding closes a constant fraction of the remaining gap per period, decaying toward the target geometrically rather than jumping.
Costs turn the optimizer's answer from a point into a region, and turn rebalancing from a jump into a glide. Wider spreads mean a wider band; faster-decaying signals mean a slower glide is not affordable and the position may not be worth taking at all.
Worked example: the no-trade band, by hand
One asset. Your model says its expected excess return is , its volatility is , and your risk-aversion parameter is .
Step 1, the frictionless target. Maximising means setting the derivative to zero: , so
A 10% weight. Note , hold onto that number, it is the "curvature" of the objective and it sets everything else.
Step 2, the band. Add a one-way cost of 20 basis points, . Trading is worthwhile only while the marginal benefit exceeds . Solving gives the upper edge, and the lower edge, so the band is
The no-trade region is 9% to 11%. Anywhere in there, sit still.
Step 3, use it. You currently hold 10.6%. That is inside the band, so you do nothing, even though the "optimal" weight is 10%.
Step 4, the signal strengthens. Alpha rises to 2.6%, so moves to and the band becomes 12% to 14%. You are at 10.6%, now outside. Do you trade to 13%?
No. Trade to 12%, the near edge. Check it numerically. Write for the gross objective.
- Stay at 10.6%: .
- Move to 13%: , a gross gain of . Cost is . Net gain 0.000010.
- Move to 12%: , a gross gain of . Cost is . Net gain 0.000020.
Stopping at the edge is twice as good as going all the way. The last 1% of the trade costs more than it earns, and the naive optimizer would have made it.
Worked example: splitting a trade under impact costs
Now a $100m book that needs to move a position by 10 percentage points of weight, and impact cost that grows with the square of the daily trade: with (so a one-shot 10-point trade costs , i.e. 50 basis points, or $500,000).
Suppose the optimizer's adjustment rate works out to 40% of the remaining gap per day. Then:
| Day | Gap at start | Trade | Cost (bp) |
|---|---|---|---|
| 1 | 10.00 pp | 4.00 pp | 8.0 |
| 2 | 6.00 pp | 2.40 pp | 2.9 |
| 3 | 3.60 pp | 1.44 pp | 1.0 |
| 4 | 2.16 pp | 0.86 pp | 0.4 |
| 5 | 1.30 pp | 0.52 pp | 0.1 |
Costs come from with in weight units: day 1 is bp, day 2 is bp, and so on. Total: about 12.4 bp versus 50 bp for the one-shot trade, a saving of roughly $376,000, while ending day 5 with 92% of the position on.
The remaining 0.78 percentage points is the price of the schedule: for five days you were carrying less of the position than you wanted, and any return the signal delivered in that window was partly missed. That missed return is the other half of the trade-off, and it is what the risk-and-alpha terms in the objective are weighing against the cost term. If your signal decays in two days, this schedule is far too slow and the honest conclusion is that the trade is not worth doing, see Alpha Decay.
For reference, here is the frictionless picture the optimizer would hand you if costs did not exist. Drag the weight slider along the curve and remember that every step along it is a trade you would have to pay for, a cost this chart cannot show:
What this means in practice
- Aim in front of the target. Gârleanu and Pedersen's result: when signals are persistent, you should trade toward a weighted average of today's target and the targets you expect tomorrow, weighting fast-decaying signals down because you will not be able to trade into them cheaply enough to profit.
- The cost model is part of the alpha model. A per-name cost estimate that ignores liquidity will let the optimizer pile into illiquid small caps. Costs should scale with the name's spread and with your order as a share of its daily volume.
- Convexity keeps it solvable. Both the absolute-value and quadratic cost forms are convex, so the whole problem stays a convex program and solves reliably, see Convex Optimization. Fixed per-trade costs and minimum lot sizes are not convex and turn it into a much harder integer problem.
- Capacity is a cost statement. Ask how large the strategy can get before impact eats the edge and you are asking where the cost term overwhelms the alpha term, see Portfolio Capacity.
Subtracting costs from a backtest is not the same as optimising with costs. The first tells you your ideal portfolio was unaffordable. The second gives you a different, affordable portfolio, usually with materially lower turnover and only slightly lower gross alpha. Strategies that look dead after a post-hoc cost haircut frequently come back to life when the costs go inside the objective.
Two related traps. First, do not tune the cost coefficient until turnover looks acceptable, that is fitting, not modelling; estimate from your own fills. Second, do not replace the linear cost with a quadratic one just because it is differentiable. The kink at zero is not a mathematical nuisance, it is the no-trade region. Smooth it away and you get a model that always trades a little, every day, in every name, which is precisely the behaviour you were trying to stop.
Practice
- Recompute the no-trade band with a 5 bp cost instead of 20 bp. How much does the band narrow, and what does that imply about trading liquid futures versus small-cap equities?
- Your signal's information decays with a half-life of three days. Using the five-day glide path above, roughly what fraction of the signal is left by the time you are fully positioned?
- Show that with proportional costs the band half-width widens when volatility falls. Explain in one sentence why a calmer asset should be traded less often, not more.
- Add a fixed $50 ticket charge per trade to the one-asset problem. Why can you no longer solve it by setting a derivative to zero?
Related concepts
Practice in interviews
Further reading
- Gârleanu & Pedersen (2013), Dynamic Trading with Predictable Returns and Transaction Costs
- Grinold & Kahn, Active Portfolio Management (Ch. 16)
- Boyd et al. (2017), Multi-Period Trading via Convex Optimization