Quant Memo
Advanced

Bergomi Forward Variance Models

Instead of modeling instantaneous volatility and letting it imply a term structure, Bergomi models start from the whole forward variance curve directly and let each point on it evolve — matching what variance swaps actually trade.

Prerequisites: Forward Variance Curve, The Itô Integral

A weather app doesn't just tell you today's temperature and a drift rule — for planning a trip, you want the whole forecast curve: expected temperature for every day ahead. Classic stochastic volatility models like Heston work like the "temperature plus drift" approach: they model instantaneous variance and let a term structure emerge as a byproduct. Bergomi's forward variance models flip this around and start directly from the forecast curve itself.

The core object and how it moves

Define ξt(u)\xi_t(u) as the forward variance seen from today tt, for an instant uu in the future — the market's current price for "how much variance will be realized around date uu," analogous to a forward interest rate. This whole curve is directly observable today from variance swap quotes (see forward variance curve). The model specifies how each point on this curve moves as new information arrives:

dξt(u)=ωξt(u)dZt(u)d\xi_t(u) = \omega\, \xi_t(u)\, dZ_t(u)

In words: the forward variance for horizon uu, ξt(u)\xi_t(u), moves proportionally to its own current level, scaled by ω\omega (a volatility-of-volatility parameter — how jumpy the forecast itself is) and driven by a random shock dZt(u)dZ_t(u) that can be correlated across different horizons uu and correlated with the underlying's own returns (this correlation is what generates the volatility skew). It's a lognormal-style dynamic for variance forecasts themselves, not for the stock price directly — the stock price dynamics come along separately, using the instantaneous variance ξt(t)\xi_t(t) (the near-end of the curve) as the actual variance driving the underlying at each moment.

Worked example 1 — a shock at one horizon, felt everywhere correlated

Suppose today's forward variance curve gives ξ0(1m)=0.04\xi_0(1\text{m}) = 0.04 (20% vol, squared) and ξ0(6m)=0.0625\xi_0(6\text{m}) = 0.0625 (25% vol). A single-factor Bergomi model with ω=1.5\omega = 1.5 and near-perfect correlation across horizons means a shock that moves the 1-month point up 10% tends to move the 6-month point similarly: new ξ(1m)0.044\xi(1\text{m}) \approx 0.044, new ξ(6m)0.06875\xi(6\text{m}) \approx 0.06875. A two-factor version (a second, less correlated driver) lets the short and long ends move by different amounts — the front of the vol term structure is empirically more jumpy than the back, which a single flat-shock model can't reproduce.

shock size: near end vs far end 1-factor: equal move 2-factor: near > far
A single-factor shock moves the whole curve in parallel; a second factor lets the near end (front of the term structure) move more than the far end, matching what's observed empirically.

Worked example 2 — where the skew comes from

The correlation ρ\rho between the underlying's return shocks and the forward variance shocks dZt(u)dZ_t(u) generates the implied volatility skew, the same mechanism as in Heston applied point by point along the curve. With ρ=0.7\rho = -0.7 (typical for an equity index: stock down, vol up), a 5% drop in the underlying tends to come with an increase across the forward variance curve, pushing ξ0(1m)\xi_0(1\text{m}) from 0.04 to roughly 0.046. Because the model is calibrated so this response matches quoted option skews at every maturity simultaneously, a Bergomi model fit to today's whole surface reprices the skew correctly after a spot move in a way a single-instantaneous-vol model often cannot.

horizon u today's ξ(u) after spot drop
A spot drop shifts the whole forward variance curve up (negative correlation), but the near end reacts more than the far end — the pattern a multi-factor Bergomi model is built to reproduce.

What this means in practice

Bergomi-style models are the standard tool on exotic desks for pricing options whose payoff depends on the path of volatility (forward-starting options, cliquets, VIX options), because they're built to match the forward variance curve exactly by construction — a direct, observable market input, not something implied indirectly.

It's easy to read ξt(u)\xi_t(u) as a forecast of what realized variance will actually turn out to be on date uu — it isn't. It's a risk-neutral price embedding a variance risk premium (the market pays up for vol protection), so forward variance tends to sit above what variance realizes on average. Confusing "the market's current price for future variance" with "an unbiased prediction" is the most common misreading of this framework.

Bergomi models start from the observable forward variance curve and specify how each point on it moves, rather than modeling instantaneous variance and hoping a realistic term structure falls out — this makes them naturally consistent with the market's actual variance swap and vol swap prices.

Practice

  1. If a two-factor Bergomi model gives the short end of the forward variance curve a higher ω\omega (vol-of-vol) than the long end, what pattern in the term structure's jumpiness does that reproduce?
  2. Why is it a mistake to treat today's 6-month forward variance quote as a forecast that realized variance over the next six months will equal that number?

Related concepts

Practice in interviews

Further reading

  • Bergomi, Smile Dynamics II (Risk, 2005) / Stochastic Volatility Modeling (Ch. 5-7)
  • Gatheral, The Volatility Surface (Ch. 6)
ShareTwitterLinkedIn