Quant Memo
Advanced

Multi-Period Portfolio Choice

Optimizing one period at a time and just repeating the answer is not the same as optimizing across a lifetime. Once you can rebalance and your risk appetite or opportunities can change, the single-period allocation is often the wrong one to hold.

Prerequisites: Utility Functions and Risk Aversion, The Kelly Criterion

Solve the one-period problem, "how should I split my money between stocks and bonds today," and it is tempting to just re-solve it identically every year for the next thirty. But a thirty-year investor can rebalance, react to a market crash while it happens, and needs money at different future dates for different reasons. Treating thirty independent one-period decisions as if they were one long-run decision throws away information that changes the right answer.

The analogy before any symbols

A road trip planned only one exit at a time, always heading toward whichever direction currently looks fastest, can drive you in circles: locally optimal choices need not add up to the best route to the destination. A GPS instead plans the whole trip backward from the destination, so it is willing to take a slower road now because it sets up a faster one later. Multi-period portfolio choice is the GPS version of investing: the allocation today should already account for how you will react to tomorrow's news, not just today's expected return and risk.

The mechanics

In the single-period Markowitz problem, you pick weights ww once to maximize E[U(W1)]E[U(W_1)] given a fixed horizon of one step, see Utility Functions and Risk Aversion. The multi-period problem instead maximizes utility of terminal wealth WTW_T over a sequence of decisions w1,w2,,wTw_1, w_2, \ldots, w_T, where each wtw_t can depend on everything observed up to time tt:

maxw1,,wT  E[U(WT)],Wt+1=Wt(1+wtrt+1).\max_{w_1,\ldots,w_T} \; E\big[U(W_T)\big], \qquad W_{t+1} = W_t\big(1 + w_t^\top r_{t+1}\big).

In words: wealth compounds forward by whatever return the chosen weights earn each period, and you are choosing an entire policy, a rule for each period's weight given what has happened so far, not a single fixed weight. Solved by backward induction (dynamic programming): figure out the best last-period decision for every possible wealth level, then the best second-to-last decision given that you will behave optimally afterward, and so on to today.

Merton's classic continuous-time result shows that under constant investment opportunities (returns and volatility never change) and CRRA utility, the optimal weight is constant over time and equals the single-period myopic answer, w\*=μ/(γσ2)w^\* = \mu/(\gamma\sigma^2). Multi-period and single-period agree exactly in that special case. The moment opportunities can change, say volatility clusters or expected returns are predictable, an extra hedging demand term appears, adjusting the weight to hedge against future shifts in the investment menu, not just current risk and return.

Path explorer
13055time →
end (bold path) 100.38spread of ends 58.966 independent paths, same settings

Watch how a single simulated wealth path can wander far from its average trend; a multi-period investor who only ever looks at expected return and volatility at time zero never sees, or plans for, the mid-path detours this explorer makes visible.

today vol spikes → cut w vol calms → raise w a fixed single-period w cannot react to either branch
A multi-period policy is a rule, react to volatility as it changes, not one number chosen once and held blindly.

Worked example: rebalancing beats buy-and-hold under mean reversion

Two assets, each starting at $100, mean-reverting around that level. Over one period asset A rises to $120 and asset B falls to $80. A buy-and-hold 50/50 investor now holds $120 + $80 = $200, still 60% in A. A rebalancing investor sells enough A to restore 50/50: $100 in each, realizing a small profit locked in from A's rise. If the assets mean-revert, A falls back toward $100 and B rises back toward $100 next period. Buy-and-hold, overweight A, gives back A's gain and misses B's rebound; the rebalanced investor benefits from the reversal on both legs. The rebalancing premium is a genuinely multi-period effect, invisible to any single-period analysis.

Worked example: the myopic weight is not always the target

With μ=8%\mu = 8\%, σ=25%\sigma = 25\%, γ=3\gamma = 3: the myopic weight is w=0.08/(3×0.0625)=42.7%w = 0.08/(3\times0.0625) = 42.7\%. Now suppose returns are known to be predictable, today's high dividend yield forecasts higher future expected returns after a bad year. A long-horizon investor optimally tilts above 42.7% today, accepting more short-term risk, because a bad year now improves next period's opportunity set and partially offsets today's loss. That extra tilt is the hedging demand Merton's model adds on top of the myopic term, and it has no counterpart in a single-period optimizer.

What this means in practice

Target-date funds, glide paths, and any "rebalance to target weights" policy are implicitly multi-period strategies; a purely single-period optimizer re-run each year would ignore both rebalancing effects and predictability in opportunities. Backtests that treat every period independently, refitting Markowitz weights from scratch with no regard for what dynamic policy generated them, systematically understate the value (and the risk) of strategies that depend on path, not just endpoint.

A multi-period optimum is a policy, a plan for how to react at every future date, not a single weight computed once from today's inputs. It equals the single-period myopic weight only when investment opportunities never change.

The common error is assuming multi-period optimization is just "single-period optimization done more often." Rebalancing costs, mean reversion, and time-varying volatility all break that equivalence, and predictable returns add a hedging-demand term the single-period formula has no way to represent. A strategy that looks suboptimal one period at a time can still be optimal over the full horizon, and vice versa: a sequence of individually optimal decisions is not automatically the best path to the destination.

Related concepts

Practice in interviews

Further reading

  • Merton (1969), Lifetime Portfolio Selection Under Uncertainty
  • Campbell & Viceira, Strategic Asset Allocation
ShareTwitterLinkedIn