Local Stochastic Volatility Models
Local volatility gets every vanilla price right and the dynamics wrong. Stochastic volatility gets the dynamics right and the prices slightly wrong. LSV bolts a correction factor onto a stochastic volatility engine so it does both, and the correction factor has a formula.
Prerequisites: Local Volatility and Dupire's Formula, Stochastic Volatility and the Heston Model, The Dupire Equation
An exotics desk is stuck between two models and neither one is usable alone. Local volatility reprices every listed call and put to the tick, which is non-negotiable — you cannot sell a barrier option off a model that misprices the vanillas you hedge it with. But its volatility is welded to the index level, so it predicts the smile will move in ways it demonstrably does not, and barrier options and notes are priced almost entirely off how the smile moves. Stochastic volatility has believable movement, because volatility has its own random driver, but with five parameters it will miss a forty-strike surface by a volatility point here and there. A point is a lot when your hedges are quoted to a tenth of one.
A dimmer switch on a flickering bulb
Imagine a bulb that flickers on its own, in a realistic, lifelike way. You like the flicker. What you do not like is that the average brightness comes out wrong in each room — too dim in the hallway, too bright in the kitchen. You do not want to redesign the bulb, because then you lose the flicker. You want a dimmer in each room, set once, so that the average brightness matches your target exactly while the flicker is untouched.
That is the whole idea. The stochastic volatility model is the flickering bulb: it produces realistic volatility movement. The dimmer is a function of two things — where the index is and what date it is — that multiplies the model's volatility up or down at each point so that the vanilla prices come out exactly right. Its official name is the leverage function, and the remarkable part is that you do not have to search for its settings. There is a formula.
The model
Two equations. The index:
with drift left out for clarity. In plain English: the index moves by a random shove whose size is the stochastic volatility multiplied by the dimmer setting at the current level and date. Here is the index, is the instantaneous variance, is the index's random driver, and is the leverage function we have to find.
And the variance, in its usual Heston form:
In plain English: variance drifts back toward a long-run level at speed , gets shoved around by its own randomness with intensity (the vol-of-vol), and that randomness is correlated with the index by , which is negative for equities because volatility rises when markets fall.
Set and you have plain Heston. Set so variance never moves, and collapses to a deterministic function of level and date — plain local volatility. LSV sits in between and contains both.
The calibration identity
Gyöngy's theorem says something you would not guess: any complicated process can be mimicked, in the sense of matching every terminal distribution and hence every vanilla price, by a simple one-factor diffusion whose volatility is the conditional average of the complicated one. Applied here it gives the whole calibration in one line:
In plain English: at each price level and date , the dimmer is set to the ratio of two variances. The numerator is the local variance the market demands, read off the surface by Dupire. The denominator is the average variance the stochastic engine actually delivers among the paths that happen to be sitting at level on date — not the average over all paths, only the ones that got there. Divide and you have the correction factor.
Two things are worth pausing on. First, if the stochastic engine already produces the right variance at that node, the ratio is 1 and the dimmer does nothing. Second, the denominator is a conditional expectation inside the model you are calibrating, so this is circular: you need to simulate the paths, and you need the paths to compute . It is solved forward in time, one small time step at a time, which is what the particle method does.
LSV = stochastic volatility for the dynamics, times a leverage function for the prices. The leverage function is not fitted; it is local variance divided by the model's conditional expected variance, node by node.
Worked example 1: the dimmer at two nodes
One-year horizon, index at 100. Suppose Dupire's formula gives local volatility 22 percent at the 100 node and 30 percent at the 80 node — a steep downside, as equity surfaces have. Suppose the Heston engine, simulated forward, delivers conditional expected variances of 0.0441 at the 100 node and 0.0625 at the 80 node.
At the money. Local variance is . Conditional model variance is 0.0441, i.e. a volatility of 21 percent. So
Check it: percent. The dimmer scales the engine up by 4.8 percent and the node now prices correctly.
Twenty percent down. Local variance is . The engine's conditional variance there is 0.0625, i.e. 25 percent — already higher than its 21 percent at the money, because is negative and paths that fell are paths where volatility rose. So
Now compare the two. The market wants the downside node 36 percent more volatile than the at-the-money node (). The stochastic engine supplied 19 percent of that on its own (), and the dimmer only had to supply the remaining 15 percent (). Under pure local volatility the leverage function would have had to carry all of it, because there is no engine underneath doing any work. That split is the entire point of LSV: the more of the skew the stochastic part explains, the more the model's forward dynamics look like the real thing.
Worked example 2: computing the denominator by hand
The formula's denominator, , sounds abstract. In practice it is a bucket average over simulated paths, and the arithmetic is primary-school.
Simulate a batch of paths to . Collect the ones that landed near 100, and record each one's instantaneous variance at that moment. Say four paths landed in the bucket, carrying variances
Their sum is , so the average is , which is a conditional volatility of exactly 21 percent — the number used above. Divide the target 0.0484 by it and , as before.
Do the same for the bucket that landed near 80. Five paths, with variances
Sum: ; ; ; . Average , a conditional volatility of 25 percent. Divide 0.0900 by it and .
Notice what the buckets show. The paths that fell to 80 are carrying variances of 22 to 27 percent while the paths that stayed at 100 are carrying 19 to 23 percent. Nobody told the model to do that. It falls out of the negative correlation , and it is the part of the skew that LSV gets for free and local volatility has to be forced into.
Now the practical worry, which is also visible in these numbers. The 80 bucket had five paths. Real buckets in the deep wings can have a handful out of a hundred thousand, and a five-sample average of a spread-out quantity is noisy. Suppose the true conditional variance were 0.0625 but your bucket estimated 0.0550. Then , , and your model's effective volatility at that node is percent instead of 30 — a 2-point error, in the corner of the surface that barrier and autocall pricing depends on most. Bucket noise is the central engineering problem of LSV calibration, not the mathematics.
The paths below make the bucketing concrete. Set volatility high and hit resample: notice how few of six paths end up in any given narrow band. That thinness, multiplied across every node of a two-dimensional grid, is what the calibration has to average over.
What this means in practice
LSV is the production model for equity and FX exotics: barriers, autocallables, cliquets, anything whose value depends on how the smile behaves in the future rather than only today. The workflow is fixed. Fit an arbitrage-free surface, extract local volatility with Dupire, choose the stochastic volatility parameters to match the dynamics you care about — typically forward-starting or long-dated skew, or the volatility of volatility implied by options on volatility — and then let the leverage function absorb the residual so vanillas reprice exactly.
The dial that matters most is the mixing weight. Scale the vol-of-vol down toward zero and the leverage function has to do more, and the model slides back toward local volatility. Scale it up and the stochastic part dominates. Barrier and autocall prices move monotonically between the two extremes, often by several percent of the note's value, and no amount of vanilla data will pin it down — the vanillas are matched exactly for every setting. The choice has to come from products that are actually sensitive to forward smile, or from a view.
Every cell above is a constraint the calibrated model hits exactly, at every mixing weight. That is precisely why the vanilla surface cannot tell you which mixing weight is right.
Matching the vanilla surface is a constraint, not a validation. Every LSV calibration on the desk reprices every listed option perfectly, including the ones with wildly different mixing weights that disagree by 5 percent on the barrier you are about to sell. Do not read "the model calibrates" as "the model is right". The second trap is treating the leverage function as economically meaningful — a large in the wings usually means the surface extrapolation is bad or the bucket was empty, not that the market expects extreme volatility down there.
Practice
- At a node, Dupire's local volatility is 25 percent and the engine's conditional expected variance is 0.0400. What is the leverage factor?
- If the engine already produced conditional variance exactly equal to local variance everywhere, what would be, and what model would you have?
- A bucket at a deep downside node contains three paths with variances 0.10, 0.14 and 0.09. Local variance there is 0.16. Compute , then recompute it if a fourth path with variance 0.27 had landed in the bucket. What does the difference tell you?
Answers. (1) , so . (2) everywhere, and you would have pure stochastic volatility that happens to fit the surface — the situation LSV exists because we do not have. (3) Mean , , . With the fourth path, mean , , . One extra path moved the node's leverage by 14 percent — wing calibration is dominated by sampling noise.
Key terms
- Leverage function — the multiplicative correction that forces vanilla prices to match.
- Gyöngy's theorem — a complicated process is mimicked, distribution by distribution, by a one-factor diffusion using conditional average volatility.
- Conditional expected variance — the average of over only those paths sitting at a given level at a given date.
- Mixing weight — how much of the volatility dynamics comes from the stochastic engine versus the leverage function.
- Particle method — the forward, step-by-step simulation that resolves the circularity in the calibration identity.
Related concepts
Practice in interviews
Further reading
- Gyöngy (1986), Mimicking the One-Dimensional Marginal Distributions of Processes
- Guyon & Henry-Labordère (2012), Being Particular About Calibration
- Lipton (2002), The Vol Smile Problem