Quant Memo
Advanced

The Merton Portfolio Problem

If you can rebalance continuously and your appetite for risk does not change with how rich you are, the optimal fraction of your wealth to hold in a risky asset is a single constant, edge divided by risk aversion times variance, no matter your horizon or your bank balance.

Prerequisites: Geometric Brownian Motion, The Kelly Criterion, The Risk-Return Tradeoff

Markowitz answers a one-shot question: given a pot of money and one period ahead, how should you split it? But nobody invests once. You invest, the market moves, your wealth changes, and you get to decide again, tomorrow and every day after that for forty years. The question a real investor faces is not "what is the best split today" but "what is the best rule for the split, at every moment, for the rest of my life?"

That is a much harder question, and the surprising thing about Robert Merton's 1969 answer is how simple it turns out to be.

An analogy before any symbols

Think of a thermostat. It does not care what the weather did last week, how long until spring, or how big the house is. It has one setting, and it does whatever is needed at each moment to hold that setting. When the room drifts warm it cools; when it drifts cold it heats.

Merton's result is that, under a specific set of assumptions, the best investment policy is a thermostat. You pick one number, the fraction of your wealth that sits in the risky asset, and you hold it there forever. Not a fraction that ramps down as you age, not one that depends on how well you have done so far. One setting, held constant, with continual small adjustments to stay on it.

That is genuinely counterintuitive. Most people's instinct is that a young investor should take more risk, or that after a windfall you should get more aggressive. Merton's model says no, and understanding why it says no is more useful than the formula itself.

The setup

There is one risky asset whose price SS follows a Geometric Brownian Motion:

dStSt=μdt+σdWt.\frac{dS_t}{S_t} = \mu\,dt + \sigma\,dW_t.

In words: over any short instant, the asset drifts up at rate μ\mu and jitters randomly with size σ\sigma, where WW is the random shove and dtdt is the tiny slice of time. There is also cash paying a certain rate rr.

You choose πt\pi_t, the fraction of your wealth in the risky asset at time tt. Not a dollar amount, a fraction. If π=0.6\pi = 0.6 you are 60% invested and 40% in cash; if π=1.5\pi = 1.5 you have borrowed 50% of your wealth to invest more. Your wealth XX then evolves as

dXtXt=[r+πt(μr)]dt+πtσdWt.\frac{dX_t}{X_t} = \big[r + \pi_t(\mu - r)\big]dt + \pi_t \sigma\, dW_t.

In words: you always earn the cash rate, plus your share π\pi of the extra return the risky asset offers over cash, and you carry your share π\pi of its jitter.

Finally you need a way to say what you want. Merton uses constant relative risk aversion (CRRA) utility,

U(X)=X1γ1γ,U(X) = \frac{X^{1-\gamma}}{1-\gamma},

which is a mathematical way of writing "more money is better, but each extra dollar is worth less than the last, and how much less is governed by one number γ\gamma." Small γ\gamma means you barely flinch at risk; large γ\gamma means you hate it. The word "relative" is the crucial part: your feelings depend on percentage changes in wealth, not dollar changes. Losing half is equally painful whether you started with ten thousand or ten million.

The answer

Maximising expected utility of terminal wealth gives the Merton fraction:

π=μrγσ2.\pi^\star = \frac{\mu - r}{\gamma\,\sigma^2}.

In words: your edge over cash, divided by how much you hate risk times how violently the asset moves. Three readings worth holding onto:

  • Double the edge, double the position.
  • Double the volatility and you quarter the position, because σ\sigma is squared. Risk punishes you far harder than reward rewards you.
  • Nothing on the right-hand side is tt or XX. No time, no wealth. That is the constant-thermostat result.

Why does horizon drop out? Because under this model, returns in disjoint time intervals are independent and identically distributed, and CRRA preferences care only about percentage outcomes. A ten-year investor is just a one-year investor repeated ten times, facing the same problem each time, so they make the same choice each time. Why does wealth drop out? Because doubling your wealth doubles every possible outcome proportionally, and "relative" risk aversion is indifferent to that scaling.

With γ=1\gamma = 1 (log utility) the formula becomes (μr)/σ2(\mu - r)/\sigma^2, which is exactly the continuous-time The Kelly Criterion. Kelly is Merton with one specific risk appetite baked in.

π=(μr)/(γσ2)\pi^\star = (\mu - r)/(\gamma \sigma^2) is a fraction of wealth held constant through time, not a dollar amount set once. Holding it constant is an active policy: it forces you to sell into rallies and buy into declines, automatically.

Growth versus the size of the bet

Plug π\pi back into the wealth equation and the long-run growth rate of your money is

g(π)=r+π(μr)12π2σ2.g(\pi) = r + \pi(\mu - r) - \tfrac12\pi^2\sigma^2.

In words: you earn cash, plus your share of the edge, minus a variance drag that grows with the square of your bet. That subtraction is why leverage eventually turns on you: the reward term is linear in π\pi, the penalty is quadratic, so past some point the penalty wins.

fastest growth all cash twice the bet, no better than cash π* 2π* fraction invested π long-run growth rate
Growth rises, peaks, then falls. Doubling the log-optimal bet does not double your growth, it hands all of it back: at 2π* you earn the risk-free rate while carrying enormous volatility. Past that point you lose money on average.

To see the randomness this policy is steering through, run a few sampled paths. Set the drift and volatility, then generate again a few times, notice how wide the spread of outcomes is even when the average drift is comfortably positive:

Path explorer
13055time →
end (bold path) 100.38spread of ends 58.966 independent paths, same settings

Worked example: sizing a real position

An equity index with expected return μ=9%\mu = 9\%, cash at r=3%r = 3\%, volatility σ=18%\sigma = 18\%. So the edge is 6% and σ2=0.182=0.0324\sigma^2 = 0.18^2 = 0.0324.

Moderately risk-averse investor, γ=3\gamma = 3:

π=0.063×0.0324=0.060.0972=0.617.\pi^\star = \frac{0.06}{3 \times 0.0324} = \frac{0.06}{0.0972} = 0.617.

So 61.7% in the index, 38.3% in cash. On a $200,000 portfolio that is $123,400 in equities and $76,600 in cash. Its growth rate is 0.03+0.617(0.06)12(0.617)2(0.0324)=0.030+0.0370.006=6.1%0.03 + 0.617(0.06) - \tfrac12 (0.617)^2(0.0324) = 0.030 + 0.037 - 0.006 = 6.1\%.

Log-utility (Kelly) investor, γ=1\gamma = 1: π=0.06/0.0324=1.85\pi^\star = 0.06/0.0324 = 1.85, meaning 185% invested, borrowing 85% of wealth. Growth rate =0.03+0.1110.056=8.6%= 0.03 + 0.111 - 0.056 = 8.6\%, the highest achievable. But the portfolio's volatility is 1.85×18%=33%1.85 \times 18\% = 33\% a year, which very few people can hold through a bad decade.

Cautious investor, γ=5\gamma = 5: π=0.06/(5×0.0324)=0.370\pi^\star = 0.06/(5 \times 0.0324) = 0.370, so 37% invested. Growth =0.03+0.0220.002=5.0%= 0.03 + 0.022 - 0.002 = 5.0\%. Less growth, far smoother ride.

And the cliff. At π=2×1.85=3.70\pi = 2 \times 1.85 = 3.70 the growth rate is 0.03+0.2220.222=3.0%0.03 + 0.222 - 0.222 = 3.0\%, exactly the cash rate, with 67% annual volatility. You have taken on ruinous risk to earn what a savings account pays.

Worked example: what "constant" actually costs you

Start with $100,000 and a target of π=60%\pi = 60\%: $60,000 in the index, $40,000 in cash.

Day 1, the index rises 10%. Your equity sleeve becomes $66,000, cash stays $40,000, total $106,000. Your actual fraction is now 66/106=62.3%66/106 = 62.3\%, above target. Sixty percent of $106,000 is $63,600, so you sell $2,400 of equities.

60% 40% 62.3% 37.7% 60% 40% start index +10% rebalanced sell 2,400 upper block = risky asset, lower block = cash
Nothing was decided by a human here. The rally pushed the risky weight above target, and holding the fraction constant forced a sale. The same mechanism forces a purchase after a fall.

Day 2, the index falls 10%. Your $63,600 becomes $57,240, cash is $42,400, total $99,640. Your fraction is 57.24/99.64=57.4%57.24/99.64 = 57.4\%, below target. Sixty percent of $99,640 is $59,784, so you buy $2,544 of equities.

Two things to notice. First, the index went up 10% then down 10% and ended below where it started (1.10×0.90=0.991.10 \times 0.90 = 0.99), and so did you, at $99,640. Second, the policy sold high and bought low without being told to. A constant fraction is mechanically contrarian. That is a virtue in choppy markets and a cost in trending ones, which is exactly the trade-off buy-and-hold makes in reverse.

Also notice the friction: two trades in two days. Merton's model assumes trading is free and continuous. It is neither, which is why the practical version uses tolerance bands, see Turnover and Rebalancing and Transaction-Cost-Aware Portfolio Optimization.

What this means in practice

  • Volatility targeting is Merton in disguise. If μr\mu - r and γ\gamma are treated as roughly fixed, then π1/σ2\pi^\star \propto 1/\sigma^2, so when volatility doubles you cut exposure to a quarter. That is precisely what a Vol Targeting overlay does. See also Position Sizing.
  • Multi-asset version. With many risky assets, π=1γΣ1(μr1)\pi^\star = \frac{1}{\gamma}\Sigma^{-1}(\mu - r\mathbf{1}): the same shape, with the covariance matrix doing the work. The direction of that vector is the The Tangency Portfolio and the Capital Market Line; γ\gamma only sets its length. Merton and Markowitz agree on the recipe and differ only on the dose.
  • Adding consumption. Merton's 1971 extension lets you also spend. The answer keeps its shape: invest a constant fraction of wealth, and spend a constant fraction of wealth per year. This is where the "4% rule" of retirement planning comes from.
  • When the constant breaks. If μ\mu or σ\sigma move predictably over time, the constant-fraction result fails and an extra hedging demand appears, an adjustment for the fact that today's portfolio also insures you against tomorrow's worse investment opportunities. That is the interesting part of modern dynamic asset allocation.

The formula is only as good as μ\mu, and μ\mu is the one thing you cannot measure. The position size is directly proportional to the estimated edge, and estimating an equity risk premium to within a percentage point takes many decades of data. Volatility, by contrast, can be estimated well from a few months. So the numerator is a guess and the denominator is nearly a fact, yet the guess drives the answer.

The practical consequence: quants routinely run at a fraction of the Merton or Kelly number, often a half or a quarter. Given the growth curve above, that costs surprisingly little. Half-Kelly gives you about three-quarters of the maximum growth at half the volatility. Overshooting, on the other hand, is punished quadratically. The curve is nearly flat on the left of its peak and falls off a cliff on the right, so when in doubt, be small.

The other frequent misreading is "horizon does not matter, so stocks are no safer over the long run." The model says the optimal fraction does not depend on horizon. It does not say a long horizon reduces risk, that would require returns to mean-revert, which this model explicitly assumes they do not.

Practice

  1. With μ=9%\mu = 9\%, r=3%r = 3\%, σ=18%\sigma = 18\%, what γ\gamma makes π\pi^\star exactly 100%? What does that tell you about a fully-invested, unlevered equity investor's implied risk aversion?
  2. Volatility spikes from 18% to 30% with no change in expected return. By what factor should a Merton investor cut exposure? (Answer: about a third of the original position.)
  3. Compute g(π)g(\pi) for π=0.5,1.0,1.5,2.0\pi = 0.5, 1.0, 1.5, 2.0 using the numbers above and confirm the peak sits at 1.85.
  4. Two assets with the same Sharpe ratio but different volatilities. Show that the dollar risk Merton allocates to each is the same, even though the weights differ.

Related concepts

Practice in interviews

Further reading

  • Merton (1969), Lifetime Portfolio Selection under Uncertainty
  • Merton (1971), Optimum Consumption and Portfolio Rules in a Continuous-Time Model
  • Björk, Arbitrage Theory in Continuous Time (Ch. 20)
ShareTwitterLinkedIn