Chapter 05
Rigorous

The Itô Integral

00 · Symbol Glossary

$W_t$W sub t — standard Brownian motion

The driving noise from Chapter 02: W0=0W_0=0, independent increments, Wt−Ws∼N(0,t−s)W_t-W_s\sim N(0,t-s) for s<ts<t, continuous paths. Everything in this chapter integrates against WtW_t.

$H_t$integrand — an adapted process

The process being integrated. HtH_t must be adapted to the filtration Ft\mathcal{F}_t of Chapter 04 — it cannot peek at the future value of WW.

$\int_0^T H_t\,dW_t$Itô integral

A random variable built as a limit of sums ∑iHti(Wti+1−Wti)\sum_i H_{t_i}(W_{t_{i+1}}-W_{t_i}), evaluated at the left endpoint of each subinterval. The left-endpoint choice is not cosmetic — it is what makes the whole theory work.

$L^2$square-integrable processes

The space of processes with E ⁣∫0THt2 dt<∞\mathbb{E}\!\int_0^T H_t^2\,dt<\infty. The Itô integral is first built for simple processes, then extended to all of L2L^2 by a limiting argument.


01 · Why an Ordinary Integral Does Not Work

Chapter 03 showed that WtW_t has infinite total variation on every interval, no matter how short. A Riemann–Stieltjes integral ∫0THt dWt\int_0^T H_t\,dW_t is normally defined path by path, as a limit of sums that converges regardless of which point in each subinterval you sample — left endpoint, right endpoint, midpoint, all give the same answer when the integrator has finite variation. Brownian motion does not have finite variation, so that convergence fails: different sampling points give different limits, and none of them exist as an ordinary limit for almost every path.

The fix is not to abandon the sum. It is to fix the sampling point once and for all — always the left endpoint — and build convergence in a weaker sense than "for almost every path": convergence in L2L^2, i.e. in mean square across the randomness. That single choice is the entire content of Itô's construction.

Why left, specifically

The left endpoint makes HtiH_{t_i} known at the start of [ti,ti+1][t_i,t_{i+1}] — no information about the future increment Wti+1−WtiW_{t_{i+1}}-W_{t_i} leaks into the weight. That is exactly the "adapted, non-anticipating" requirement, and it is what produces the martingale and isometry properties in Section 03. Change the sampling point and both properties disappear.


02 · Step One: Simple (Piecewise-Constant) Processes

Build the integral first for the easiest integrands, where no limit is needed at all.

Definition — Simple Process

A process HtH_t is simple if there is a partition 0=t0<t1<⋯<tn=T0=t_0<t_1<\cdots<t_n=T and Fti\mathcal{F}_{t_i}-measurable random variables HiH_i (each with finite variance) such that

Ht=Hifor t∈[ti,ti+1)H_t = H_i \quad \text{for } t\in[t_i,t_{i+1})

HtH_t is constant on each subinterval, and that constant is known already at the start of the subinterval.

For a simple process, define the integral directly as the sum that motivated it:

∫0THt dWt  :=  ∑i=0n−1Hi (Wti+1−Wti)\int_0^T H_t\,dW_t \;:=\; \sum_{i=0}^{n-1} H_i\,(W_{t_{i+1}}-W_{t_i})

There is no limit here — it is a finite sum of random variables, each term a known weight times a Brownian increment. This object already has the two properties that will survive the general construction, and they are worth checking directly. HiH_i is measurable with respect to Fti\mathcal{F}_{t_i}, and Wti+1−WtiW_{t_{i+1}}-W_{t_i} is independent of Fti\mathcal{F}_{t_i} with mean zero, so E[Hi(Wti+1−Wti)∣Fti]=Hi⋅E[Wti+1−Wti∣Fti]=0\mathbb{E}[H_i(W_{t_{i+1}}-W_{t_i})\mid\mathcal{F}_{t_i}]=H_i\cdot\mathbb{E}[W_{t_{i+1}}-W_{t_i}\mid\mathcal{F}_{t_i}]=0. Sum that over ii and the whole integral has conditional mean zero at every stage — the martingale property, checked from the definition rather than assumed.

Example — A two-step simple integral

Partition [0,2][0,2] at t0=0,t1=1,t2=2t_0=0,t_1=1,t_2=2, with Ht=W0=0H_t=W_0=0 on [0,1)[0,1) and Ht=W1H_t=W_1 on [1,2)[1,2). Then

∫02Ht dWt=0⋅(W1−W0)+W1⋅(W2−W1)=W1(W2−W1)\int_0^2 H_t\,dW_t = 0\cdot(W_1-W_0) + W_1\cdot(W_2-W_1) = W_1(W_2-W_1)

E[W1(W2−W1)]=E[W1] E[W2−W1]=0\mathbb{E}[W_1(W_2-W_1)]=\mathbb{E}[W_1]\,\mathbb{E}[W_2-W_1]=0 by independence of the increment from W1W_1 — consistent with the martingale property above.


03 · Extension to L2L^2

Most processes worth integrating — Ht=WtH_t=W_t, Ht=f(Wt,t)H_t=f(W_t,t) — are not piecewise constant. The extension is a standard density argument: approximate, take a limit, check the limit does not depend on the approximating sequence.

Step-by-step — Building the general Itô integral
1
Restrict to L2L^2: require E∫0THt2 dt<∞\mathbb{E}\int_0^T H_t^2\,dt<\infty and HtH_t adapted. This is the class the construction covers; wider classes exist but need more machinery.
2
Approximate: for any such HH, there is a sequence of simple processes H(n)H^{(n)} with E∫0T(Ht(n)−Ht)2 dt→0\mathbb{E}\int_0^T (H_t^{(n)}-H_t)^2\,dt\to 0 — simple processes are dense in L2L^2 of adapted processes.
3
Define via the limit: ∫0THt(n) dWt\int_0^T H_t^{(n)}\,dW_t is already defined (Section 02) and forms a Cauchy sequence in L2(Ω)L^2(\Omega) because of the isometry below. Define ∫0THt dWt\int_0^T H_t\,dW_t as its L2L^2 limit.
4
Check well-posedness: any two approximating sequences give the same limit, so the integral is a well-defined random variable, not an artifact of the approximation chosen.

The isometry that Step 3 leans on is the single most useful computational fact in the whole subject.

Definition — Itô Isometry, Martingale Property, Zero Mean

For H∈L2H\in L^2 adapted,

E ⁣[(∫0THt dWt)2]=E ⁣∫0THt2 dt(Itoˆ isometry)\mathbb{E}\!\left[\left(\int_0^T H_t\,dW_t\right)^2\right] = \mathbb{E}\!\int_0^T H_t^2\,dt \qquad \textbf{(Itô isometry)}
E ⁣[∫0THt dWt]=0(zero mean)\mathbb{E}\!\left[\int_0^T H_t\,dW_t\right] = 0 \qquad \textbf{(zero mean)}

and Mt=∫0tHs dWsM_t=\int_0^t H_s\,dW_s is a martingale: E[Mt∣Fs]=Ms\mathbb{E}[M_t\mid\mathcal{F}_s]=M_s for s≤ts\leq t.

Plain language: the isometry converts a variance computation on a stochastic integral into an ordinary (deterministic-looking) integral of E[Ht2]\mathbb{E}[H_t^2] — no need to expand a double sum of Brownian increments by hand. The martingale property says the running integral MtM_t has no drift: knowing everything up to time ss, the best forecast of MtM_t is just MsM_s. Both facts trace back to the left-endpoint, no-look-ahead sampling of Section 02 — they are inherited, not new assumptions.

Example — Variance of $\int_0^T W_t\,dW_t$

By the isometry, Var ⁣(∫0TWt dWt)=E∫0TWt2 dt=∫0TE[Wt2] dt=∫0Tt dt=T2/2\mathrm{Var}\!\left(\int_0^T W_t\,dW_t\right)=\mathbb{E}\int_0^T W_t^2\,dt=\int_0^T \mathbb{E}[W_t^2]\,dt=\int_0^T t\,dt=T^2/2. Chapter 06 will show ∫0TWt dWt=12WT2−12T\int_0^T W_t\,dW_t=\tfrac12 W_T^2-\tfrac12 T exactly, and it is a good check to confirm Var(12WT2−12T)=12T2\mathrm{Var}(\tfrac12 W_T^2-\tfrac12 T)=\tfrac12 T^2 matches — a preview of Itô's Lemma.

❌ Treating $\int_0^T W_t\,dW_t$ as ordinary calculus

Naive substitution suggests ∫0TWt dWt=12WT2\int_0^T W_t\,dW_t=\tfrac12 W_T^2, by analogy with ∫0Tx dx=12x2\int_0^T x\,dx=\tfrac12 x^2.

Why it breaks: ordinary calculus assumes the integrator has finite variation and that the sampling point in the defining sum does not matter (Section 01). Both fail for WtW_t. The correction term comes from the nonzero quadratic variation [W]T=T[W]_T=T that Chapter 03 established.

Consequence: the correct identity is ∫0TWt dWt=12WT2−12T\int_0^T W_t\,dW_t=\tfrac12 W_T^2-\tfrac12 T. Dropping the −12T-\tfrac12 T misstates the mean (the naive guess has E[12WT2]=12T≠0\mathbb{E}[\tfrac12 W_T^2]=\tfrac12 T\neq 0, contradicting the zero-mean property above) and every downstream Greek or hedge computed from it.


04 · Itô vs. Stratonovich

Two integrals, two conventions

The Stratonovich integral samples at the interval midpoint instead of the left endpoint, and it obeys ordinary chain-rule calculus — no correction terms. It is used in physics, where noise is often a limit of smooth processes and the midpoint arises naturally from that limit. Finance almost always uses Itô: the left-endpoint convention matches "you trade at today's price using only information available today," which is exactly the non-anticipating requirement of Section 01. The two integrals differ by a deterministic correction term and are related by a conversion formula, but mixing them in one derivation — using an Itô SDE with Stratonovich calculus rules, or vice versa — silently reintroduces the finite-variation assumption that Section 01 showed fails for WtW_t.


05 · Exercises

EXERCISE 5.1

Apply the isometry directly with Ht=tH_t=t.

Var ⁣(∫0Tt dWt)=E∫0Tt2 dt=∫0Tt2 dt=T3/3\mathrm{Var}\!\left(\int_0^T t\,dW_t\right)=\mathbb{E}\int_0^T t^2\,dt=\int_0^T t^2\,dt=T^3/3. The mean is 00 by the zero-mean property, since Ht=tH_t=t is deterministic (hence trivially adapted) and square-integrable.

Compute Var ⁣(∫0Tt dWt)\mathrm{Var}\!\left(\int_0^T t\,dW_t\right) using the Itô isometry.

EXERCISE 5.2

Check the conditional-mean-zero argument of Section 02 for a two-step simple process, but change which endpoint of the subinterval is sampled.

Right-endpoint sampling would use Hi′=Hti+1H_i'=H_{t_{i+1}}, which is Fti+1\mathcal{F}_{t_{i+1}}-measurable, not Fti\mathcal{F}_{t_i}-measurable. Then E[Hti+1(Wti+1−Wti)∣Fti]\mathbb{E}[H_{t_{i+1}}(W_{t_{i+1}}-W_{t_i})\mid\mathcal{F}_{t_i}] no longer factors into a product of a known constant and a zero-mean increment, because Hti+1H_{t_{i+1}} and the increment can be correlated (e.g. Ht=WtH_t=W_t gives E[Wti+1(Wti+1−Wti)]=ti+1−ti≠0\mathbb{E}[W_{t_{i+1}}(W_{t_{i+1}}-W_{t_i})]=t_{i+1}-t_i\neq 0). The martingale property breaks.

Explain why replacing the left endpoint with the right endpoint in the defining sum of Section 02 destroys the martingale property, using the conditional-expectation argument from that section.

EXERCISE 5.3

Compare the two conventions' treatment of ∫0TWt dWt\int_0^T W_t\,dW_t; Stratonovich gives the ordinary-calculus answer.

Itô: ∫0TWt dWt=12WT2−12T\int_0^T W_t\,dW_t=\tfrac12 W_T^2-\tfrac12 T (Section 03). Stratonovich: ∫0TWt∘dWt=12WT2\int_0^T W_t\circ dW_t=\tfrac12 W_T^2, matching ordinary calculus, since the midpoint sampling absorbs exactly the quadratic-variation correction that left-endpoint sampling leaves exposed. The difference is 12T\tfrac12 T — a deterministic quantity, not noise — which is the general pattern relating the two integrals.

State the Stratonovich value of ∫0TWt∘dWt\int_0^T W_t\circ dW_t and compare it to the Itô value from Section 03. What kind of quantity is the difference?


06 · Chapter Summary

ConceptMeaning
Simple processpiecewise-constant, known at the start of each subinterval
Left-endpoint samplingencodes "no look-ahead"; source of the martingale property
L2L^2 extensionlimit of simple-process integrals in mean square
Itô isometryE[(∫H dW)2]=E∫H2 dt\mathbb{E}[(\int H\,dW)^2]=\mathbb{E}\int H^2\,dt
Zero mean / martingale∫0tHs dWs\int_0^t H_s\,dW_s has no drift
Itô vs. Stratonovichleft vs. midpoint sampling; finance uses Itô

Next: Chapter 06 — Itô's Lemma, the chain rule that accounts for the −12T-\tfrac12 T correction term seen in ∫0TWt dWt\int_0^T W_t\,dW_t and generalizes it to any smooth function of WtW_t.