Chapter 09
Medium

Mean-Reverting SDEs: The Ornstein–Uhlenbeck Process

00 · Symbol Glossary

$X_t$the OU process

Unlike GBM's StS_t, XtX_t can be negative — it models a deviation from a level (a spread, a rate, a log-volatility), not a price with limited liability.

$\theta$mean-reversion speed

Governs how fast XtX_t is pulled back toward its long-run level. Larger θ\theta means faster reversion and a shorter half-life (Section 03).

$\kappa$long-run mean (level)

The value XtX_t reverts toward. Some texts write μ\mu for this; κ\kappa is used here to avoid clashing with GBM's drift μ\mu from Chapter 08, which plays a completely different role.

$\sigma$volatility (additive, not multiplicative)

Constant, not scaled by XtX_t — additive noise in the classification of Chapter 07, which is exactly why the integrating factor solves this SDE directly with no transform needed first.


01 · The Model

Definition — Ornstein–Uhlenbeck (OU) Process

XtX_t follows the Ornstein–Uhlenbeck process if it solves

dXt=θ(κ−Xt) dt+σ dWt,θ>0, σ>0dX_t = \theta(\kappa - X_t)\,dt + \sigma\,dW_t, \qquad \theta>0,\ \sigma>0

Plain language: the drift θ(κ−Xt)\theta(\kappa-X_t) is positive when Xt<κX_t<\kappa and negative when Xt>κX_t>\kappa — a restoring force, proportional in strength to how far XtX_t currently is from κ\kappa, exactly like a spring pulling toward equilibrium. The noise term is additive and constant, unlike GBM's multiplicative noise, which is why XtX_t can go negative: there is nothing in the equation preventing it, and nothing in the intended application (a spread, a rate, a deviation) requiring it to stay positive.


02 · Solving via Integrating Factor

This is the linear SDE that Chapter 07 flagged as the integrating-factor case, worked out here in full.

Step-by-step — Solving the OU equation
1
Rewrite the drift: dXt=(θκ−θXt) dt+σ dWtdX_t = (\theta\kappa - \theta X_t)\,dt + \sigma\,dW_t, isolating the −θXt-\theta X_t term that the integrating factor must cancel.
2
Multiply by eθte^{\theta t}: compute d(eθtXt)=θeθtXt dt+eθt dXtd(e^{\theta t}X_t) = \theta e^{\theta t}X_t\,dt + e^{\theta t}\,dX_t (ordinary product rule — eθte^{\theta t} is deterministic, so it contributes no second-order Itô correction). Substituting dXtdX_t from Step 1, the θeθtXt\theta e^{\theta t}X_t terms cancel exactly, leaving d(eθtXt)=θκeθt dt+σeθt dWtd(e^{\theta t}X_t) = \theta\kappa e^{\theta t}\,dt + \sigma e^{\theta t}\,dW_t.
3
Integrate from 00 to tt: eθtXt−X0=κ(eθt−1)+σ∫0teθs dWse^{\theta t}X_t - X_0 = \kappa(e^{\theta t}-1) + \sigma\int_0^t e^{\theta s}\,dW_s.
4
Solve for XtX_t: divide by eθte^{\theta t}.
Xt=κ+(X0−κ)e−θt+σ∫0te−θ(t−s) dWsX_t = \kappa + (X_0-\kappa)e^{-\theta t} + \sigma\int_0^t e^{-\theta(t-s)}\,dW_s

The first two terms are the deterministic path a noiseless version of the equation would take — decaying exponentially from X0X_0 toward κ\kappa. The third term is a weighted sum of past noise, with weight e−θ(t−s)e^{-\theta(t-s)} decaying as the noise ages: recent shocks matter more than old ones, and old shocks fade geometrically rather than persisting.

Example — The distribution of $X_t$ for fixed $X_0$

Since ∫0te−θ(t−s) dWs\int_0^t e^{-\theta(t-s)}\,dW_s is an Itô integral of a deterministic integrand, it is Gaussian with mean 00 and, by the Itô isometry (Chapter 05), variance σ2∫0te−2θ(t−s) ds=σ22θ(1−e−2θt)\sigma^2\int_0^t e^{-2\theta(t-s)}\,ds = \dfrac{\sigma^2}{2\theta}\left(1-e^{-2\theta t}\right). So conditional on X0X_0,

Xt  ∣  X0  ∼  N ⁣(κ+(X0−κ)e−θt, σ22θ(1−e−2θt))X_t \;\Big|\; X_0 \;\sim\; N\!\left(\kappa+(X_0-\kappa)e^{-\theta t},\ \frac{\sigma^2}{2\theta}\left(1-e^{-2\theta t}\right)\right)

03 · Stationary Distribution and Half-Life

Let t→∞t\to\infty in the variance formula above: e−2θt→0e^{-2\theta t}\to 0, so the variance converges to a finite limit rather than growing without bound — a sharp contrast with Brownian motion's variance tt, which diverges.

Definition — Stationary Distribution and Half-Life

As t→∞t\to\infty (for fixed X0X_0),

Xt  →d  N ⁣(κ, σ22θ)(stationary distribution)X_t \;\xrightarrow{d}\; N\!\left(\kappa,\ \frac{\sigma^2}{2\theta}\right) \qquad \textbf{(stationary distribution)}

The half-life — the time for the deterministic component of a deviation from κ\kappa to decay by half — solves e−θt1/2=12e^{-\theta t_{1/2}}=\tfrac12:

t1/2=log⁡2θt_{1/2} = \frac{\log 2}{\theta}

Plain language: no matter where X0X_0 starts, the process eventually forgets it and settles into one fixed distribution centered at κ\kappa with a variance set by the balance of pull (θ\theta) against noise (σ\sigma) — stronger pull tightens the distribution, more noise widens it. The half-life converts the abstract rate θ\theta into a number with direct intuition: a spread with half-life of two weeks has largely closed within a month, one with a half-life of two years has not.

Example — Reading $\theta$ off a half-life target

If a trading desk wants a mean-reverting spread model with a two-week half-life (t1/2=14t_{1/2}=14 days), θ=log⁡2/14≈0.0495\theta = \log 2 / 14 \approx 0.0495 per day. This is the standard way calibration proceeds in practice: pick a half-life from data or intuition, back out θ\theta, then fit κ,σ\kappa,\sigma from the residual behavior around the reverting level.


04 · OU vs. the Discrete AR(1)

Analogy to Time Series, not a requirement

The stationary AR(1) process Yt=ϕYt−1+εtY_t=\phi Y_{t-1}+\varepsilon_t (Time Series Chapter 02) is the discrete-time cousin of OU — both have a single parameter controlling how fast deviations decay (ϕ\phi here, e−θΔte^{-\theta\Delta t} if OU is sampled at spacing Δt\Delta t), both converge to a stationary distribution, and both have a well-defined half-life. Sampling an OU process at a fixed frequency produces exactly an AR(1) with ϕ=e−θΔt\phi=e^{-\theta\Delta t} and Gaussian innovations — this is not a coincidence but an exact correspondence, and it is one clean bridge between the two subjects' otherwise separate toolkits. Nothing in this chapter requires that correspondence; it is offered as intuition, not as a solving technique.


05 · Quant Uses

OU is the standard building block wherever "reverts to a level, doesn't run away to infinity" is the right qualitative story: modeling interest rate deviations from a target (the short-rate models of Vasicek and CIR extend OU directly), a statistical-arbitrage spread between two cointegrated assets (Time Series Chapter 08 builds the discrete-time version of this same idea), or log-volatility itself, which tends to cluster around a long-run level rather than drift to ±∞\pm\infty.

❌ Using OU for a price level instead of a deviation

Modeling a stock price StS_t directly with dSt=θ(κ−St) dt+σ dWtdS_t=\theta(\kappa-S_t)\,dt+\sigma\,dW_t.

Why it breaks: OU's Gaussian increments allow St<0S_t<0 with positive probability at every finite tt (the stationary distribution in Section 03 is Gaussian, hence supported on all of R\mathbb{R}) — a structural contradiction for a price with limited liability, the same property GBM (Chapter 08) was built specifically to avoid.

Consequence: simulated paths eventually go negative, and any pricing formula relying on log⁡St\log S_t or St−1S_t^{-1} breaks down on those paths. OU is the right tool for a spread, rate deviation, or log-volatility, not for a raw asset price — that distinction is the main thing to get right before writing down the SDE.


06 · Exercises

EXERCISE 9.1

Differentiate eθtXte^{\theta t}X_t using the product rule and substitute the OU SDE for dXtdX_t; check which terms cancel.

d(eθtXt)=θeθtXt dt+eθt(θ(κ−Xt) dt+σ dWt)=θeθtXt dt+θκeθt dt−θeθtXt dt+σeθt dWt=θκeθt dt+σeθt dWtd(e^{\theta t}X_t) = \theta e^{\theta t}X_t\,dt + e^{\theta t}\big(\theta(\kappa-X_t)\,dt+\sigma\,dW_t\big) = \theta e^{\theta t}X_t\,dt + \theta\kappa e^{\theta t}\,dt - \theta e^{\theta t}X_t\,dt + \sigma e^{\theta t}\,dW_t = \theta\kappa e^{\theta t}\,dt + \sigma e^{\theta t}\,dW_t. The two θeθtXt dt\theta e^{\theta t}X_t\,dt terms cancel — exactly Step 2 of Section 02, shown explicitly.

Verify Step 2 of Section 02 directly: show that d(eθtXt)=θκeθt dt+σeθt dWtd(e^{\theta t}X_t) = \theta\kappa e^{\theta t}\,dt+\sigma e^{\theta t}\,dW_t, with the XtX_t-dependent terms cancelling.

EXERCISE 9.2

Use t1/2=log⁡2/θt_{1/2}=\log 2/\theta from Section 03, solved for θ\theta.

θ=log⁡2/t1/2=log⁡2/30≈0.0231\theta = \log 2/t_{1/2} = \log 2/30 \approx 0.0231 per day. Doubling the target half-life to 6060 days would halve θ\theta to about 0.01160.0116 — half-life and θ\theta are inversely proportional, so slower reversion always means a longer half-life, and the relationship is exact, not approximate.

A trader wants an OU spread model with a 3030-day half-life. Find θ\theta. If the target half-life is doubled to 6060 days, what happens to θ\theta?

EXERCISE 9.3

Compare the stationary variance formula from Section 03 to the sample variance of an AR(1) using the correspondence ϕ=e−θΔt\phi=e^{-\theta\Delta t} from Section 04.

The AR(1) stationary variance is γ(0)=σε2/(1−ϕ2)\gamma(0)=\sigma_\varepsilon^2/(1-\phi^2) (Time Series Chapter 02). Substituting ϕ=e−θΔt\phi=e^{-\theta\Delta t} and letting Δt→0\Delta t\to 0 with σε2\sigma_\varepsilon^2 scaled appropriately recovers σ2/(2θ)\sigma^2/(2\theta) — the continuous-time OU stationary variance from Section 03. The discrete formula is the exact statement; the continuous one is its Δt→0\Delta t\to 0 limit, matching the general relationship between the two subjects described in Chapter 01's closing note.

Using the AR(1)–OU correspondence of Section 04, show how the discrete stationary variance σε2/(1−ϕ2)\sigma_\varepsilon^2/(1-\phi^2) relates to the continuous-time stationary variance σ2/(2θ)\sigma^2/(2\theta) of Section 03.


07 · Chapter Summary

ConceptMeaning
OU SDEdXt=θ(κ−Xt) dt+σ dWtdX_t=\theta(\kappa-X_t)\,dt+\sigma\,dW_t
Solving methodintegrating factor eθte^{\theta t}
Closed formXt=κ+(X0−κ)e−θt+σ∫0te−θ(t−s) dWsX_t=\kappa+(X_0-\kappa)e^{-\theta t}+\sigma\int_0^t e^{-\theta(t-s)}\,dW_s
Stationary distributionN(κ, σ2/2θ)N(\kappa,\ \sigma^2/2\theta)
Half-lifet1/2=log⁡2/θt_{1/2}=\log 2/\theta
Discrete cousinsampled OU is exactly AR(1) with ϕ=e−θΔt\phi=e^{-\theta\Delta t}
Use forspreads, rate deviations, log-volatility — not raw prices

Next: Chapter 10 — Girsanov's Theorem and Change of Measure, which explains how a drift like OU's θ(κ−Xt)\theta(\kappa-X_t) can be removed entirely by changing the probability measure, the key step toward risk-neutral pricing.