Chapter 10
Rigorous

Girsanov's Theorem and Change of Measure

00 · Symbol Glossary

$\mathbb{P},\ \mathbb{Q}$P and Q — two probability measures

Two ways of assigning probabilities to the same set of paths. P\mathbb{P} is usually the real-world (physical) measure; Q\mathbb{Q} is a reweighted measure built for pricing. Both live on the same sample space and agree on which paths are possible — only the weight placed on each path differs.

$Z_T=\frac{d\mathbb{Q}}{d\mathbb{P}}$Radon–Nikodym derivative

A positive random variable that converts P\mathbb{P}-probabilities into Q\mathbb{Q}-probabilities: EQ[X]=EP[XZT]\mathbb{E}^{\mathbb{Q}}[X]=\mathbb{E}^{\mathbb{P}}[XZ_T] for every XX. Requires ZT>0Z_T>0 and EP[ZT]=1\mathbb{E}^{\mathbb{P}}[Z_T]=1, so Q\mathbb{Q} is a genuine probability measure.

$Z_t$density process

The P\mathbb{P}-martingale Zt=EP[ZT∣Ft]Z_t=\mathbb{E}^{\mathbb{P}}[Z_T\mid\mathcal{F}_t] (Chapter 04) that carries the reweighting through time. Z0=1Z_0=1, and ZtZ_t tells you the running likelihood ratio using only information available up to tt.

$\theta_t$Girsanov kernel

The adapted process controlling how much drift the change of measure absorbs. In the finance application below, θ=(μ−r)/σ\theta=(\mu-r)/\sigma — excess return per unit of volatility.

$\tilde W_t$W tilde — Brownian motion under Q

W~t=Wt+∫0tθs ds\tilde W_t=W_t+\int_0^t\theta_s\,ds. Under Q\mathbb{Q}, W~t\tilde W_t has exactly the defining properties Chapter 02 gave WtW_t under P\mathbb{P}: continuous paths, independent increments, W~t−W~s∼N(0,t−s)\tilde W_t-\tilde W_s\sim N(0,t-s).


01 · Same Noise, Different Drift

An SDE dXt=μt dt+σt dWtdX_t=\mu_t\,dt+\sigma_t\,dW_t (Chapter 07) has two ingredients that feel similar but behave completely differently under a change of probability measure. Quadratic variation (Chapter 03) is computed from a single observed path — sum the squared increments along the trajectory that actually happened, and the limit is ∫0tσs2 ds\int_0^t\sigma_s^2\,ds, no probabilities involved. Drift is the opposite: μt=lim⁡h→0E[Xt+h−Xt]/h\mu_t=\lim_{h\to0}\mathbb{E}[X_{t+h}-X_t]/h is an average over paths, and an average is exactly the kind of object that changes when you reweight which paths count more.

That asymmetry is the entire content of this chapter. Reweighting probabilities — assigning more mass to some paths and less to others, without altering the set of paths themselves — can turn one drift into a different drift on the same noise. It cannot touch volatility, because volatility is read off a single path and a change of measure never edits a path, only how likely it was.

A crowd, not a coin

Think of P\mathbb{P} and Q\mathbb{Q} as two different opinions about how likely each possible path is, held by two observers watching the same physical process unfold. Neither observer can change what actually happens. They can only disagree about which outcomes were probable in advance — which is precisely enough to disagree about means, but not enough to disagree about a quantity like quadratic variation that is determined by the realized path alone.

Practice — Which quantities can a measure change touch?

For dXt=μt dt+σt dWtdX_t=\mu_t\,dt+\sigma_t\,dW_t, list which of the following could differ between an equivalent P\mathbb{P} and Q\mathbb{Q}: the realized path itself; EP[Xt]\mathbb{E}^{\mathbb{P}}[X_t] versus EQ[Xt]\mathbb{E}^{\mathbb{Q}}[X_t]; the quadratic variation [X]t[X]_t; the set of paths assigned probability zero. Justify each answer in one sentence.


02 · The Radon–Nikodym Derivative, Informally

Definition — Equivalent Measures

Two measures P\mathbb{P} and Q\mathbb{Q} on the same space are equivalent (P∼Q\mathbb{P}\sim\mathbb{Q}) if they agree on which events are impossible: P(A)=0  ⟺  Q(A)=0\mathbb{P}(A)=0 \iff \mathbb{Q}(A)=0. Equivalence (not just similarity) is exactly the condition under which a Radon–Nikodym derivative ZT=dQ/dPZ_T=d\mathbb{Q}/d\mathbb{P} exists: a nonnegative random variable with EP[ZT]=1\mathbb{E}^{\mathbb{P}}[Z_T]=1 such that

EQ[X]=EP[X ZT]for every random variable X\mathbb{E}^{\mathbb{Q}}[X] = \mathbb{E}^{\mathbb{P}}[X\,Z_T] \quad \text{for every random variable } X

ZTZ_T is a likelihood ratio: on a path where ZTZ_T is large, Q\mathbb{Q} thinks that path was more likely than P\mathbb{P} did, and vice versa. Because expectations under Q\mathbb{Q} are just P\mathbb{P}-expectations with an extra factor of ZTZ_T inside, every Q\mathbb{Q}-probability statement can be computed by staying entirely inside P\mathbb{P} and multiplying by ZTZ_T.

Example — A two-path cartoon

Suppose only two paths are possible, ω1\omega_1 and ω2\omega_2, with P(ω1)=0.5\mathbb{P}(\omega_1)=0.5, P(ω2)=0.5\mathbb{P}(\omega_2)=0.5. Define Q(ω1)=0.8\mathbb{Q}(\omega_1)=0.8, Q(ω2)=0.2\mathbb{Q}(\omega_2)=0.2. Then Z(ω1)=0.8/0.5=1.6Z(\omega_1)=0.8/0.5=1.6 and Z(ω2)=0.2/0.5=0.4Z(\omega_2)=0.2/0.5=0.4, and EP[Z]=0.5(1.6)+0.5(0.4)=1\mathbb{E}^{\mathbb{P}}[Z]=0.5(1.6)+0.5(0.4)=1 as required. Nothing about ω1\omega_1 or ω2\omega_2 as paths changed — only how much probability mass sits on each one.

Why Z_t must be a martingale

Consistency across time forces Zt=EP[ZT∣Ft]Z_t=\mathbb{E}^{\mathbb{P}}[Z_T\mid\mathcal{F}_t] to be a P\mathbb{P}-martingale: today's best guess of the eventual likelihood ratio, updated as information arrives, is by construction unbiased under P\mathbb{P} (the tower property of conditional expectation, Chapter 04). This is what lets the measure change be applied gradually along a filtration rather than only at the terminal time TT.


03 · Girsanov's Theorem: Statement and Payoff

Definition — Girsanov's Theorem

Let θt\theta_t be adapted and satisfy Novikov's condition (EP[exp⁡(12∫0Tθs2 ds)]<∞\mathbb{E}^{\mathbb{P}}[\exp(\tfrac12\int_0^T\theta_s^2\,ds)]<\infty, enough to rule out pathological blow-ups). Define the density process

Zt=exp⁡(−∫0tθs dWs−12∫0tθs2 ds)Z_t = \exp\left(-\int_0^t\theta_s\,dW_s - \frac12\int_0^t\theta_s^2\,ds\right)

Then ZtZ_t is a P\mathbb{P}-martingale, and setting dQ=ZT dPd\mathbb{Q}=Z_T\,d\mathbb{P} defines an equivalent measure Q\mathbb{Q} under which

W~t=Wt+∫0tθs ds\tilde W_t = W_t + \int_0^t\theta_s\,ds

is a standard Brownian motion.

Rearranged, dWt=dW~t−θt dtdW_t = d\tilde W_t - \theta_t\,dt. Substitute this into any SDE written in terms of WtW_t and the diffusion term is untouched — σt dWt=σt dW~t−σtθt dt\sigma_t\,dW_t=\sigma_t\,d\tilde W_t-\sigma_t\theta_t\,dt moves a piece from noise into drift, but the coefficient multiplying the new Brownian motion is still exactly σt\sigma_t. This is the payoff: Girsanov lets you dial the drift of a diffusion to any adapted process you like by an appropriate choice of θt\theta_t, while the volatility σt\sigma_t is invariant — it appears unchanged on both sides of the substitution.

❌ Expecting Girsanov to change volatility

Hoping that some clever choice of θt\theta_t could turn dXt=μ dt+σ dWtdX_t=\mu\,dt+\sigma\,dW_t into a process with a different diffusion coefficient, say σ′≠σ\sigma'\neq\sigma, purely by a change of measure.

Why it breaks: equivalent measures share null sets by definition, and quadratic variation is a pathwise a.s. limit (Chapter 03). If P\mathbb{P} and Q\mathbb{Q} disagreed about [X]t=σ2t[X]_t=\sigma^2 t, that disagreement would itself be an event of probability 00 under one measure and positive probability under the other — contradicting equivalence.

Consequence: θt\theta_t can only ever move terms between the dtdt (drift) and dWtdW_t (noise) pieces of the same σt dWt\sigma_t\,dW_t; it can never rescale σt\sigma_t itself. A model that needs different volatility is a different model, not a different measure.


04 · Worked Example: Removing the Drift from GBM

Step-by-step — Shifting GBM's drift to r via Girsanov
1
Start from the GBM SDE (Chapter 08): dSt=μSt dt+σSt dWtdS_t=\mu S_t\,dt+\sigma S_t\,dW_t, with constants μ,σ\mu,\sigma and a constant risk-free rate rr.
2
Choose the constant kernel θ=μ−rσ\theta=\dfrac{\mu-r}{\sigma}, and define W~t=Wt+θt\tilde W_t=W_t+\theta t, a Q\mathbb{Q}-Brownian motion by Girsanov.
3
Substitute dWt=dW~t−θ dtdW_t=d\tilde W_t-\theta\,dt into the SDE: dSt=μSt dt+σSt(dW~t−θ dt)=(μ−σθ)St dt+σSt dW~tdS_t=\mu S_t\,dt+\sigma S_t(d\tilde W_t-\theta\,dt)=(\mu-\sigma\theta)S_t\,dt+\sigma S_t\,d\tilde W_t.
4
Simplify the drift: μ−σθ=μ−σ⋅μ−rσ=μ−(μ−r)=r\mu-\sigma\theta=\mu-\sigma\cdot\dfrac{\mu-r}{\sigma}=\mu-(\mu-r)=r.
5
Conclude: under Q\mathbb{Q}, dSt=rSt dt+σSt dW~tdS_t=rS_t\,dt+\sigma S_t\,d\tilde W_t — same GBM, same σ\sigma, drift replaced by rr.

Nothing about the asset's actual path changed; Q\mathbb{Q} simply weights paths differently so that, averaged under this new weighting, StS_t grows at the risk-free rate. A short calculation with Itô's lemma (Chapter 06) on e−rtSte^{-rt}S_t confirms this drift shift makes the discounted price a Q\mathbb{Q}-martingale — the property that later chapters build derivative pricing on.

This measure has a name

Q\mathbb{Q} here is the risk-neutral measure, and θ=(μ−r)/σ\theta=(\mu-r)/\sigma is the market price of risk. Nothing about risk preferences enters the mathematics of this section — the name describes how the resulting measure gets used (Chapters 12–13), not a step in the derivation.


05 · Exercises

EXERCISE 10.1

Apply Itô's lemma (Chapter 06) to ln⁡Zt\ln Z_t, where Zt=exp⁡(−∫0tθs dWs−12∫0tθs2 ds)Z_t=\exp(-\int_0^t\theta_s\,dW_s-\tfrac12\int_0^t\theta_s^2\,ds).

Let Yt=−∫0tθs dWs−12∫0tθs2 dsY_t=-\int_0^t\theta_s\,dW_s-\tfrac12\int_0^t\theta_s^2\,ds, so Zt=eYtZ_t=e^{Y_t} with dYt=−θt dWt−12θt2 dtdY_t=-\theta_t\,dW_t-\tfrac12\theta_t^2\,dt. Itô's lemma gives dZt=Zt dYt+12Zt (dYt)2=Zt(−θt dWt−12θt2 dt)+12Ztθt2 dt=−Ztθt dWtdZ_t=Z_t\,dY_t+\tfrac12Z_t\,(dY_t)^2=Z_t(-\theta_t\,dW_t-\tfrac12\theta_t^2\,dt)+\tfrac12Z_t\theta_t^2\,dt=-Z_t\theta_t\,dW_t. The dtdt terms exactly cancel, leaving pure diffusion with no drift — ZtZ_t is (at least) a local martingale, and Novikov's condition upgrades this to a true martingale.

Show that the density process ZtZ_t has zero drift, i.e. that dZt=−Ztθt dWtdZ_t=-Z_t\theta_t\,dW_t with no dtdt term.

EXERCISE 10.2

Quadratic variation is computed from squared increments of the realized path; check that substituting dWt=dW~t−θ dtdW_t=d\tilde W_t-\theta\,dt does not change the σ2St2 dt\sigma^2S_t^2\,dt piece.

[S]t=∫0tσ2Ss2 ds[S]_t=\int_0^t\sigma^2S_s^2\,ds under either the P\mathbb{P}-representation dSt=μSt dt+σSt dWtdS_t=\mu S_t\,dt+\sigma S_t\,dW_t or the Q\mathbb{Q}-representation dSt=rSt dt+σSt dW~tdS_t=rS_t\,dt+\sigma S_t\,d\tilde W_t, because quadratic variation only sees the coefficient multiplying the Brownian term, and that coefficient is σSt\sigma S_t in both. The dtdt terms (drift) never contribute to quadratic variation regardless of which measure produced them.

Confirm that StS_t's quadratic variation is the same whether computed from the P\mathbb{P}-SDE or the Q\mathbb{Q}-SDE of Section 04.

EXERCISE 10.3

Compute the quadratic variation of cWtcW_t for a constant c≠1c\neq1 and compare it to that of WtW_t.

[cW]t=c2t[cW]_t=c^2t, while [W]t=t[W]_t=t. If some measure change turned WtW_t (under P\mathbb{P}) into cWtcW_t (under an alleged Q\mathbb{Q}) with c≠1c\neq1, then P\mathbb{P} and Q\mathbb{Q} would disagree about the value of a pathwise-computable quantity (tt versus c2tc^2t) on a set of full probability — precisely the disagreement equivalence rules out. So no change of measure can rescale Brownian motion; Girsanov only ever shifts drift.

Explain, using quadratic variation, why no Girsanov-type change of measure can map WtW_t to cWtcW_t for a constant c≠1c\neq1.


06 · Chapter Summary

ConceptMeaning
Equivalent measuresP∼Q\mathbb{P}\sim\mathbb{Q}: same null sets; ZT=dQ/dPZ_T=d\mathbb{Q}/d\mathbb{P} exists
Density process ZtZ_tP\mathbb{P}-martingale carrying the reweighting through time
GirsanovdWt=dW~t−θt dtdW_t=d\tilde W_t-\theta_t\,dt; reshapes drift only
θ=(μ−r)/σ\theta=(\mu-r)/\sigmashifts GBM's drift from μ\mu to rr under Q\mathbb{Q}
Invariantσt\sigma_t and quadratic variation — identical under P\mathbb{P} and Q\mathbb{Q}
Changeddrift μt→μt−σtθt\mu_t\to\mu_t-\sigma_t\theta_t

Next: Chapter 11 — Feynman–Kac and the Black-Scholes PDE, where the risk-neutral drift built here turns a pricing expectation into a partial differential equation.