Chapter 03
Hard

Quadratic Variation and Why Calculus Breaks

00 · Symbol Glossary

$\Pi_n$partition of $[0,t]$

A finite set of points 0=t0<t1<⋯<tn=t0=t_0<t_1<\cdots<t_n=t dividing [0,t][0,t] into subintervals, used to build a Riemann-sum-style approximation. Mesh ∣Πn∣=max⁡i(ti−ti−1)|\Pi_n| = \max_i (t_i-t_{i-1}) is the width of the widest subinterval.

$V_t$total variation

Vt=sup⁡Π∑i∣f(ti)−f(ti−1)∣V_t = \sup_{\Pi} \sum_i |f(t_i)-f(t_{i-1})|, the supremum over all partitions of the sum of absolute increments — the total distance a path travels, not its net displacement.

$[W,W]_t$quadratic variation of $W$ on $[0,t]$

The limit in probability of ∑i(Wti−Wti−1)2\sum_i (W_{t_i}-W_{t_{i-1}})^2 over partitions with mesh →0\to 0. This chapter's central computation: [W,W]t=t[W,W]_t = t.

$(\Delta W_i)^2$squared increment

Shorthand for (Wti−Wti−1)2(W_{t_i}-W_{t_{i-1}})^2, the term summed to build quadratic variation. Each term has mean ti−ti−1t_i-t_{i-1} and is itself random — quadratic variation is a limit of a random sum, and its value tt is the fact that the randomness washes out in that particular limit.


01 · Total Variation vs. Quadratic Variation

A smooth deterministic path and a Brownian path can both be continuous, yet they measure "how much a path moves" in completely different ways once the increments are squared instead of just summed.

Definition — Total Variation

For a function ff on [0,t][0,t] and a partition Πn={0=t0<⋯<tn=t}\Pi_n=\{0=t_0<\cdots<t_n=t\}, the total variation is

Vt(f)=sup⁡Πn∑i=1n∣f(ti)−f(ti−1)∣V_t(f) = \sup_{\Pi_n} \sum_{i=1}^{n} \big| f(t_i)-f(t_{i-1}) \big|

ff has bounded variation if Vt(f)<∞V_t(f)<\infty. Every continuously differentiable function on [0,t][0,t] has bounded variation — Vt(f)=∫0t∣f′(s)∣ dsV_t(f)=\int_0^t |f'(s)|\,ds.

Plain language: total variation adds up the absolute size of every wiggle, so it grows large for a path that oscillates a lot, even if the oscillations cancel and the path ends up near where it started.

Definition — Quadratic Variation

The quadratic variation of ff on [0,t][0,t] is the limit, if it exists, of

[f,f]t=lim⁡∣Πn∣→0∑i=1n(f(ti)−f(ti−1))2[f,f]_t = \lim_{|\Pi_n|\to 0} \sum_{i=1}^{n} \big(f(t_i)-f(t_{i-1})\big)^2

as the mesh ∣Πn∣→0|\Pi_n|\to 0.

For a continuously differentiable ff, this limit is always 00: each squared increment is O((Δti)2)O((\Delta t_i)^2), and summing n∼t/∣Πn∣n\sim t/|\Pi_n| such terms gives a sum of order ∣Πn∣⋅t→0|\Pi_n|\cdot t \to 0. Smooth functions have zero quadratic variation and (generically) positive, finite total variation. Brownian motion reverses both of these facts.


02 · The Central Result: [W,W]t=t[W,W]_t = t

Theorem — Quadratic Variation of Brownian Motion

For a standard Wiener process WW and any sequence of partitions of [0,t][0,t] with mesh →0\to 0,

[W,W]t=lim⁡∣Πn∣→0∑i=1n(Wti−Wti−1)2=t(limit in probability, or a.s. along a suitable subsequence)[W,W]_t = \lim_{|\Pi_n|\to 0} \sum_{i=1}^{n} (W_{t_i}-W_{t_{i-1}})^2 = t \quad \text{(limit in probability, or a.s. along a suitable subsequence)}

The computation behind the theorem is a mean/variance argument, not a deep analytic one. Write Δi=Wti−Wti−1\Delta_i = W_{t_i}-W_{t_{i-1}} and Δti=ti−ti−1\Delta t_i = t_i - t_{i-1}.

Step-by-step — Why the random sum converges to the deterministic value $t$
1
Take the mean of the sum: E ⁣[∑iΔi2]=∑iE[Δi2]=∑iVar(Δi)=∑iΔti=t\mathbb{E}\!\left[\sum_i \Delta_i^2\right] = \sum_i \mathbb{E}[\Delta_i^2] = \sum_i \mathrm{Var}(\Delta_i) = \sum_i \Delta t_i = t, using Δi∼N(0,Δti)\Delta_i \sim N(0,\Delta t_i) from axiom (iii) of Chapter 02.
2
Take the variance of the sum: by independence of increments (axiom (ii)), Var ⁣(∑iΔi2)=∑iVar(Δi2)\mathrm{Var}\!\left(\sum_i \Delta_i^2\right) = \sum_i \mathrm{Var}(\Delta_i^2). For Z∼N(0,h)Z\sim N(0,h), Var(Z2)=2h2\mathrm{Var}(Z^2)=2h^2, so this equals ∑i2(Δti)2\sum_i 2(\Delta t_i)^2.
3
Bound the variance by the mesh: ∑i2(Δti)2≤2∣Πn∣∑iΔti=2∣Πn∣ t→0\sum_i 2(\Delta t_i)^2 \leq 2|\Pi_n|\sum_i \Delta t_i = 2|\Pi_n|\,t \to 0 as ∣Πn∣→0|\Pi_n|\to 0, since each Δti≤∣Πn∣\Delta t_i \le |\Pi_n|.
4
Conclude by Chebyshev: the sum has mean exactly tt (step 1) and variance shrinking to 00 (step 3), so the sum converges to the constant tt in probability — a random quantity with vanishing variance converges to its own mean.

Plain language: the sum of squared increments is random for any fixed partition, but its randomness averages away as the partition gets finer, leaving the deterministic answer tt. This is the opposite of total variation, which for Brownian motion is infinite almost surely — the path moves "a lot" in the total-variation sense while its squared movement converges to something perfectly deterministic.

Infinite total variation — stated, not proven

Vt(W)=∞V_t(W) = \infty almost surely for Brownian motion. This chapter does not reproduce that proof (it uses a similar partition argument combined with the law of the iterated logarithm); the deliberate omission does not affect anything downstream, since quadratic variation, not total variation, is the quantity stochastic calculus is built on.


03 · Why the Riemann–Stieltjes Integral Fails

Classical calculus defines ∫0tf dg\int_0^t f\,dg as a Riemann–Stieltjes integral: a limit of sums ∑if(ξi)(g(ti)−g(ti−1))\sum_i f(\xi_i)\big(g(t_i)-g(t_{i-1})\big) over shrinking partitions, for a sample point ξi\xi_i in each subinterval. That construction requires the integrator gg to have bounded variation — otherwise the sums do not converge to a value independent of how ξi\xi_i is chosen inside [ti−1,ti][t_{i-1},t_i].

❌ Defining $\int_0^t W_s\, dW_s$ pathwise as a Riemann–Stieltjes integral

Approximate ∫0tWs dWs\int_0^t W_s\,dW_s by ∑iWξi(Wti−Wti−1)\sum_i W_{\xi_i}(W_{t_i}-W_{t_{i-1}}) for two natural choices of sample point: left endpoint ξi=ti−1\xi_i = t_{i-1}, and midpoint ξi=(ti−1+ti)/2\xi_i = (t_{i-1}+t_i)/2.

Why it breaks: because Wt(ω)W_t(\omega) has infinite total variation for almost every path ω\omega (Section 02's NoteBlock), the Riemann–Stieltjes sums do not converge to a common value as ξi\xi_i ranges over [ti−1,ti][t_{i-1},t_i]. The left-endpoint sum converges (in probability) to 12Wt2−12t\tfrac12 W_t^2 - \tfrac12 t; the midpoint sum converges to 12Wt2\tfrac12 W_t^2. The two limits differ by exactly 12t\tfrac12 t, and no choice of ξi\xi_i is more "correct" than another under the classical definition.

Consequence: ∫0tWs dWs\int_0^t W_s\,dW_s has no meaning as an ordinary Riemann–Stieltjes integral — the answer depends on an arbitrary convention. Chapter 05 resolves this by defining the Itô integral with the left-endpoint convention fixed as part of the definition, which is what makes the Itô integral a specific object rather than an ill-posed one.


04 · Second-Order Terms Do Not Vanish

In ordinary calculus, a Taylor expansion of f(x+dx)f(x+dx) keeps the first-order term f′(x) dxf'(x)\,dx and discards (dx)2(dx)^2 as negligible — legitimate because for a smooth deterministic path, (dx)2=O(dt2)(dx)^2 = O(dt^2), which vanishes faster than dtdt itself when integrated.

Example — Why $(dx)^2$ is negligible for smooth paths

If xtx_t is differentiable with dx=x′(t) dtdx = x'(t)\,dt, then (dx)2=(x′(t))2 dt2(dx)^2 = (x'(t))^2\,dt^2. Summing over n=t/Δtn=t/\Delta t subintervals gives a total contribution of order n⋅Δt2=t⋅Δt→0n\cdot\Delta t^2 = t\cdot\Delta t \to 0. Discarding second-order terms in ordinary calculus is not an approximation — it is exact in the limit.

For Brownian motion, the same bookkeeping gives a different answer. (dW)2(dW)^2 is not O(dt2)O(dt^2); by Section 02, when summed over a partition, ∑i(ΔWi)2→t=∫0tds\sum_i (\Delta W_i)^2 \to t = \int_0^t ds. Heuristically, (dWt)2(dW_t)^2 behaves like dtdt itself, not like a negligible higher-order term.

dWt⋅dWt=dt(informal shorthand for [W,W]t=t)dW_t \cdot dW_t = dt \qquad \text{(informal shorthand for } [W,W]_t = t\text{)}

Plain language: the "square of a small random step" is the same order of magnitude as "a small step in time" — not smaller. Any Taylor expansion of f(Wt)f(W_t) that drops the (dWt)2(dW_t)^2 term the way ordinary calculus drops (dx)2(dx)^2 is discarding a term that does not actually vanish, which understates the drift of f(Wt)f(W_t) by exactly the term 12f′′(Wt) dt\tfrac12 f''(W_t)\,dt. Correcting for this missing term is precisely Itô's Lemma, covered in Chapter 06.


05 · Exercises

EXERCISE 3.1

Use E[Δi2]=Var(Δi)\mathbb{E}[\Delta_i^2]=\mathrm{Var}(\Delta_i) and sum over a partition with nn equal subintervals of length t/nt/n.

With nn equal subintervals, Δti=t/n\Delta t_i = t/n for each ii, so E[∑iΔi2]=∑i=1nΔti=n⋅(t/n)=t\mathbb{E}[\sum_i \Delta_i^2] = \sum_{i=1}^n \Delta t_i = n\cdot(t/n) = t regardless of nn. The expected value of the sum is exactly tt for every partition, not just in the limit — the limit theorem in Section 02 is about the sum's variance collapsing to 00, not about its mean changing.

For a partition of [0,t][0,t] into nn equal subintervals, compute E ⁣[∑i(ΔWi)2]\mathbb{E}\!\left[\sum_i (\Delta W_i)^2\right] exactly (not just in the limit) and explain why the answer does not depend on nn.

EXERCISE 3.2

Apply the variance bound from Step 3 of Section 02 with mesh ∣Πn∣=t/n|\Pi_n| = t/n.

Var ⁣(∑iΔi2)≤2∣Πn∣ t=2(t/n) t=2t2/n\mathrm{Var}\!\left(\sum_i \Delta_i^2\right) \leq 2|\Pi_n|\,t = 2(t/n)\,t = 2t^2/n. This tends to 00 as n→∞n\to\infty at rate 1/n1/n, confirming convergence in probability (and, with more work using a Borel–Cantelli argument along a subsequence like n=2kn=2^k, almost-sure convergence). The rate 1/n1/n also shows how fast a numerical simulation's estimate of quadratic variation improves as the partition is refined.

For nn equal subintervals of [0,t][0,t], bound Var ⁣(∑i(ΔWi)2)\mathrm{Var}\!\left(\sum_i (\Delta W_i)^2\right) in terms of nn and tt, and state the rate at which this variance shrinks as n→∞n\to\infty.

EXERCISE 3.3

Expand Wti(Wti−Wti−1)W_{t_i}(W_{t_i}-W_{t_{i-1}}) and Wti−1(Wti−Wti−1)W_{t_{i-1}}(W_{t_i}-W_{t_{i-1}}) using ab=12(a2+b2−(a−b)2)a b = \tfrac12(a^2+b^2-(a-b)^2) style identities, then sum by telescoping.

Left endpoint: ∑iWti−1Δi=∑iWti−1(Wti−Wti−1)\sum_i W_{t_{i-1}}\Delta_i = \sum_i W_{t_{i-1}}(W_{t_i}-W_{t_{i-1}}). Using ab=12[(a+b)2−a2−b2]ab=\tfrac12\big[(a+b)^2-a^2-b^2\big] with a=Wti−1a=W_{t_{i-1}}, b=Δib=\Delta_i: Wti−1Δi=12[Wti2−Wti−12−Δi2]W_{t_{i-1}}\Delta_i = \tfrac12\big[W_{t_i}^2 - W_{t_{i-1}}^2 - \Delta_i^2\big]. Summing telescopes the first two terms to 12Wt2\tfrac12 W_t^2, leaving 12∑iΔi2→12t\tfrac12\sum_i \Delta_i^2 \to \tfrac12 t. So the left-endpoint sum →12Wt2−12t\to \tfrac12 W_t^2 - \tfrac12 t. Midpoint sums instead symmetrize the increment and the quadratic-variation correction cancels, converging to 12Wt2\tfrac12 W_t^2. The two differ by 12t\tfrac12 t, matching Section 03's FailBlock.

Show algebraically that the left-endpoint Riemann–Stieltjes sum for ∫0tWs dWs\int_0^t W_s\,dW_s converges to 12Wt2−12t\tfrac12 W_t^2 - \tfrac12 t, using the telescoping identity for ∑iWti−1(Wti−Wti−1)\sum_i W_{t_{i-1}}(W_{t_i}-W_{t_{i-1}}).


06 · Chapter Summary

ConceptMeaning
Total variation VtV_tsum of absolute increments; infinite a.s. for WW
Quadratic variation [W,W]t[W,W]_tlimit of squared increments; equals tt exactly
Smooth functionszero quadratic variation, generically finite total variation
Brownian motioninfinite total variation, finite deterministic quadratic variation
Riemann–Stieltjes failsrequires bounded variation; ∫Ws dWs\int W_s\,dW_s depends on sample-point convention without one
(dWt)2=dt(dW_t)^2 = dtinformal shorthand — second-order terms do not vanish in Itô calculus

Next: Chapter 04 — Filtrations, Adapted Processes, and Martingales, which supplies the information structure (Ft\mathcal{F}_t, adaptedness) that the Itô integral of Chapter 05 needs before it can fix the sample-point convention this chapter left unresolved.