Multi-Period Portfolio Choice
Optimizing one period at a time and just repeating the answer is not the same as optimizing across a lifetime. Once you can rebalance and your risk appetite or opportunities can change, the single-period allocation is often the wrong one to hold.
Prerequisites: Utility Functions and Risk Aversion, The Kelly Criterion
Solve the one-period problem, "how should I split my money between stocks and bonds today," and it is tempting to just re-solve it identically every year for the next thirty. But a thirty-year investor can rebalance, react to a market crash while it happens, and needs money at different future dates for different reasons. Treating thirty independent one-period decisions as if they were one long-run decision throws away information that changes the right answer.
The analogy before any symbols
A road trip planned only one exit at a time, always heading toward whichever direction currently looks fastest, can drive you in circles: locally optimal choices need not add up to the best route to the destination. A GPS instead plans the whole trip backward from the destination, so it is willing to take a slower road now because it sets up a faster one later. Multi-period portfolio choice is the GPS version of investing: the allocation today should already account for how you will react to tomorrow's news, not just today's expected return and risk.
The mechanics
In the single-period Markowitz problem, you pick weights once to maximize given a fixed horizon of one step, see Utility Functions and Risk Aversion. The multi-period problem instead maximizes utility of terminal wealth over a sequence of decisions , where each can depend on everything observed up to time :
In words: wealth compounds forward by whatever return the chosen weights earn each period, and you are choosing an entire policy, a rule for each period's weight given what has happened so far, not a single fixed weight. Solved by backward induction (dynamic programming): figure out the best last-period decision for every possible wealth level, then the best second-to-last decision given that you will behave optimally afterward, and so on to today.
Merton's classic continuous-time result shows that under constant investment opportunities (returns and volatility never change) and CRRA utility, the optimal weight is constant over time and equals the single-period myopic answer, . Multi-period and single-period agree exactly in that special case. The moment opportunities can change, say volatility clusters or expected returns are predictable, an extra hedging demand term appears, adjusting the weight to hedge against future shifts in the investment menu, not just current risk and return.
Watch how a single simulated wealth path can wander far from its average trend; a multi-period investor who only ever looks at expected return and volatility at time zero never sees, or plans for, the mid-path detours this explorer makes visible.
Worked example: rebalancing beats buy-and-hold under mean reversion
Two assets, each starting at $100, mean-reverting around that level. Over one period asset A rises to $120 and asset B falls to $80. A buy-and-hold 50/50 investor now holds $120 + $80 = $200, still 60% in A. A rebalancing investor sells enough A to restore 50/50: $100 in each, realizing a small profit locked in from A's rise. If the assets mean-revert, A falls back toward $100 and B rises back toward $100 next period. Buy-and-hold, overweight A, gives back A's gain and misses B's rebound; the rebalanced investor benefits from the reversal on both legs. The rebalancing premium is a genuinely multi-period effect, invisible to any single-period analysis.
Worked example: the myopic weight is not always the target
With , , : the myopic weight is . Now suppose returns are known to be predictable, today's high dividend yield forecasts higher future expected returns after a bad year. A long-horizon investor optimally tilts above 42.7% today, accepting more short-term risk, because a bad year now improves next period's opportunity set and partially offsets today's loss. That extra tilt is the hedging demand Merton's model adds on top of the myopic term, and it has no counterpart in a single-period optimizer.
What this means in practice
Target-date funds, glide paths, and any "rebalance to target weights" policy are implicitly multi-period strategies; a purely single-period optimizer re-run each year would ignore both rebalancing effects and predictability in opportunities. Backtests that treat every period independently, refitting Markowitz weights from scratch with no regard for what dynamic policy generated them, systematically understate the value (and the risk) of strategies that depend on path, not just endpoint.
A multi-period optimum is a policy, a plan for how to react at every future date, not a single weight computed once from today's inputs. It equals the single-period myopic weight only when investment opportunities never change.
The common error is assuming multi-period optimization is just "single-period optimization done more often." Rebalancing costs, mean reversion, and time-varying volatility all break that equivalence, and predictable returns add a hedging-demand term the single-period formula has no way to represent. A strategy that looks suboptimal one period at a time can still be optimal over the full horizon, and vice versa: a sequence of individually optimal decisions is not automatically the best path to the destination.
Related concepts
Practice in interviews
Further reading
- Merton (1969), Lifetime Portfolio Selection Under Uncertainty
- Campbell & Viceira, Strategic Asset Allocation