Chapter 04
Hard

Filtrations, Adapted Processes, and Martingales

00 · Symbol Glossary

$\mathcal{F}_t$F sub t — information available at time $t$

An informal stand-in for "everything that can be determined by observing the process up to and including time tt." Formally a sigma-algebra; this chapter uses it only as a bookkeeping device for what is known when.

$(\mathcal{F}_t)_{t\geq 0}$filtration

The increasing family Fs⊆Ft\mathcal{F}_s \subseteq \mathcal{F}_t for s≤ts\leq t — information never gets forgotten as time passes. Read as "the flow of information over time."

$X_t$ is $\mathcal{F}_t$-adaptedadapted process

A process whose value at each time tt is known once Ft\mathcal{F}_t is known — it does not peek into the future. Every process actually observable in real time is adapted by construction.

$\mathbb{E}[X_t \mid \mathcal{F}_s]$conditional expectation given information up to $s$

The best prediction of XtX_t using only what is known at time s≤ts\leq t — the direct continuous-time analog of a conditional mean, and the object every martingale condition is written in terms of.

$M_t$M sub t — martingale

An adapted process whose best prediction of any future value, given the present information, is exactly its current value: E[Mt∣Fs]=Ms\mathbb{E}[M_t\mid\mathcal{F}_s]=M_s for s≤ts\leq t. The formal statement of "no forecastable drift."


01 · Filtrations: A Clock for Information

Every process considered since Chapter 01 has an implicit rule about what is known when — the walk's value at step nn depends on ξ1,…,ξn\xi_1,\ldots,\xi_n and nothing later. Filtrations make that rule explicit so later chapters can state precisely which quantities are allowed to depend on the past and which are not.

Definition — Filtration (Informal)

A filtration (Ft)t≥0(\mathcal{F}_t)_{t\geq 0} is an increasing family of "information sets": Fs⊆Ft\mathcal{F}_s\subseteq\mathcal{F}_t whenever s≤ts\leq t, where Ft\mathcal{F}_t represents everything determinable from observations up to time tt. The natural filtration of a process XX is FtX=σ(Xs:s≤t)\mathcal{F}_t^X = \sigma(X_s : s\leq t) — the smallest filtration that makes XX observable at every time, and no larger.

Plain language: Ft\mathcal{F}_t is a growing archive. At time tt you can answer any yes/no question about the path up to tt; you cannot answer anything that requires knowing what happens after tt. A trading desk's filtration at 10:00am contains every tick printed by 10:00am and nothing from 10:01am, no matter how predictable that next tick might feel.

Sigma-algebras are being kept informal here on purpose

A rigorous filtration is a family of sigma-algebras on a probability space, and Ft\mathcal{F}_t-measurability is the formal notion behind "determinable from information up to tt." The Probability & Statistics subject on this site would normally carry that machinery, but those chapters are not live yet. Nothing below requires more than the plain-language reading of Ft\mathcal{F}_t as "the information available at time tt" — that reading is sufficient for every definition and computation used in this subject.


02 · Adapted and Predictable Processes

Not every process one might write down is a legitimate candidate for a trading strategy or a model input — some formulas secretly require future information to evaluate.

Definition — Adapted Process

A process {Xt}t≥0\{X_t\}_{t\geq 0} is adapted to (Ft)t≥0(\mathcal{F}_t)_{t\geq 0} if XtX_t is determined by Ft\mathcal{F}_t for every tt — equivalently, XtX_t is Ft\mathcal{F}_t-measurable. Informally: at time tt, you can compute XtX_t from what you have observed so far, with no peeking ahead.

Example — Adapted vs. not adapted

WtW_t itself is adapted to its own natural filtration by construction. A running maximum Mt=sup⁡s≤tWsM_t=\sup_{s\leq t}W_s is adapted — computable from the path observed up to tt. But Yt=Wt+1Y_t = W_{t+1} (the value one unit of time in the future) is not adapted to FtW\mathcal{F}_t^W: computing YtY_t at time tt requires knowing Wt+1W_{t+1}, which is not yet observed.

❌ A trading strategy that uses tomorrow's price

A backtest defines position size at time tt as θt=1{Wt+1>Wt}\theta_t = \mathbb{1}\{W_{t+1} > W_t\} — go long today exactly when the price rises tomorrow.

Why it breaks: θt\theta_t depends on Wt+1W_{t+1}, so θt\theta_t is not Ft\mathcal{F}_t-adapted. It is not a legitimate trading strategy; it is a strategy that requires clairvoyance.

Consequence: any backtest that produces implausibly good returns should be checked for this exact bug — a signal computed with information not yet available at the decision time. The fix is not a modeling refinement; it is discarding the signal.

Stochastic integrals (Chapter 05) require a slightly stronger condition than adaptedness for the integrand, called predictability — informally, the integrand's value at time tt must be determined by information from strictly before tt, not merely by time tt itself. This distinction matters for the technical construction of the Itô integral; it does not change anything discussed in this chapter, and is deferred there.


03 · Conditional Expectation, Informally

Every martingale statement below is written using conditional expectation, so it is worth pinning down what that expression means without a full measure-theoretic derivation.

Definition — Conditional Expectation (Informal)

E[Xt∣Fs]\mathbb{E}[X_t\mid\mathcal{F}_s] for s≤ts\le t is the best prediction of XtX_t using only the information available at time ss — the analogue of E[Xt+h∣Xt,Xt−1,…]\mathbb{E}[X_{t+h}\mid X_t,X_{t-1},\ldots] from Time Series' forecasting chapters (ts-05), but conditioned on a whole information set rather than a finite list of past values. Two properties are used repeatedly in this text: tower property — E[E[Xt∣Fu] ∣ Fs]=E[Xt∣Fs]\mathbb{E}\big[\mathbb{E}[X_t\mid\mathcal{F}_u]\ \big|\ \mathcal{F}_s\big] = \mathbb{E}[X_t\mid\mathcal{F}_s] for s≤u≤ts\le u\le t (conditioning in two stages gives the same answer as conditioning once on the coarser information); and taking out what is known — if YY is Fs\mathcal{F}_s-measurable, E[YXt∣Fs]=Y E[Xt∣Fs]\mathbb{E}[Y X_t\mid\mathcal{F}_s] = Y\,\mathbb{E}[X_t\mid\mathcal{F}_s].

Plain language: conditioning on Fs\mathcal{F}_s freezes everything already known at time ss and averages only over what is still uncertain. If YY is already determined by time ss, it comes out of the conditional expectation like a constant — it is one.


04 · Martingales: Definition and Examples

Definition — Martingale

An adapted process {Mt}t≥0\{M_t\}_{t\geq 0} (with E∣Mt∣<∞\mathbb{E}|M_t|<\infty for every tt) is a martingale with respect to (Ft)t≥0(\mathcal{F}_t)_{t\geq0} if

E[Mt∣Fs]=Msfor all s≤t\mathbb{E}[M_t \mid \mathcal{F}_s] = M_s \qquad \text{for all } s\leq t

If E[Mt∣Fs]≥Ms\mathbb{E}[M_t\mid\mathcal{F}_s]\geq M_s, MM is a submartingale (tends to drift up); if ≤Ms\leq M_s, a supermartingale (tends to drift down).

Plain language: given everything known now, the best forecast of any future value is exactly the current value — no built-in tendency to rise or fall. This is the continuous-time version of the martingale-difference idea from ts-01: a return series with E[rt∣Ft−1]=0\mathbb{E}[r_t\mid\mathcal{F}_{t-1}]=0 has partial sums that form a martingale.

Example — Three martingales built from $W_t$

(1) WtW_t itself: E[Wt∣Fs]=E[Ws+(Wt−Ws)∣Fs]=Ws+0=Ws\mathbb{E}[W_t\mid\mathcal{F}_s] = \mathbb{E}[W_s + (W_t-W_s)\mid\mathcal{F}_s] = W_s + 0 = W_s, using that Wt−WsW_t-W_s is independent of Fs\mathcal{F}_s (Chapter 02) with mean 00. (2) Wt2−tW_t^2 - t: this compensates the martingale-breaking drift in Wt2W_t^2 exactly — E[Wt2∣Fs]=Ws2+(t−s)\mathbb{E}[W_t^2\mid\mathcal{F}_s] = W_s^2 + (t-s) (expand Wt2=(Ws+(Wt−Ws))2W_t^2=(W_s+(W_t-W_s))^2 and take expectations), so subtracting tt restores the martingale property. (3) exp⁡(σWt−12σ2t)\exp(\sigma W_t - \tfrac12\sigma^2 t) for any constant σ\sigma: the exponential martingale, which reappears as the building block of Girsanov's theorem (Chapter 10) and the driftless log-price under a risk-neutral measure (Chapter 08).

❌ Assuming $W_t^2$ is a martingale

A modeler checks that Wt2W_t^2 is adapted, nonnegative, and built from a martingale, and concludes it must itself be a martingale.

Why it breaks: E[Wt2∣Fs]=Ws2+(t−s)>Ws2\mathbb{E}[W_t^2\mid\mathcal{F}_s] = W_s^2 + (t-s) > W_s^2 for t>st>s — a strictly positive correction term survives. Wt2W_t^2 is a submartingale, not a martingale: it has a built-in upward drift of rate 11 per unit time, which is exactly the quadratic-variation rate from Chapter 03.

Consequence: a function of a martingale is not automatically a martingale — convexity (here, x↦x2x\mapsto x^2) introduces drift. This exact failure is what Itô's Lemma (Chapter 06) makes precise and general: any convex function of WtW_t picks up a positive drift term proportional to the second derivative and the quadratic variation rate.


05 · Why Martingales Are the Right Language for Pricing

A martingale formalizes "no forecastable drift." That property is not just a technical convenience — it is close to a restatement of the absence of arbitrage.

If a discounted asset price were a strict submartingale under the pricing measure (systematically expected to rise, after adjusting for the risk-free rate), a trader could borrow at the risk-free rate, buy the asset, and expect a positive profit with no offsetting risk premium required by the model — a mispricing the model itself would flag. Requiring the discounted price to be a martingale under an appropriately chosen probability measure is the mathematical content behind the (informal) statement "there is no free lunch": all expected excess drift has already been priced out.

What this chapter is not proving

This is a motivating sketch, not the Fundamental Theorem of Asset Pricing, which needs the Itô integral (Chapter 05) to build trading strategies, Girsanov's theorem (Chapter 10) to change the probability measure under which the martingale property holds, and a precise definition of arbitrage that this chapter has not given. Chapter 13 states the theorem properly once that machinery exists.


06 · Exercises

EXERCISE 4.1

Check whether Mt=sup⁡s≤tWsM_t = \sup_{s\le t} W_s can be computed from FtW\mathcal{F}_t^W alone.

Yes, Mt=sup⁡s≤tWsM_t = \sup_{s\leq t}W_s is adapted: at time tt, the entire path {Ws:s≤t}\{W_s : s\leq t\} is known (that is what FtW\mathcal{F}_t^W means), so its supremum is a deterministic function of that known path — no future information is used. Adaptedness never requires that a process be "nice" (here MtM_t is nondecreasing, not mean-reverting); it only requires the value to be computable from the observed past.

Is the running maximum Mt=sup⁡s≤tWsM_t=\sup_{s\leq t}W_s adapted to the natural filtration of WW? Justify using the definition, not intuition about what a maximum "should" do.

EXERCISE 4.2

Expand Wt2=(Ws+(Wt−Ws))2W_t^2 = (W_s + (W_t - W_s))^2 and take E[⋅∣Fs]\mathbb{E}[\cdot \mid \mathcal{F}_s] term by term, using independence of the increment from Fs\mathcal{F}_s.

Wt2=Ws2+2Ws(Wt−Ws)+(Wt−Ws)2W_t^2 = W_s^2 + 2W_s(W_t-W_s) + (W_t-W_s)^2. Conditioning on Fs\mathcal{F}_s: Ws2W_s^2 is already known, so it passes through unchanged; 2Ws E[Wt−Ws∣Fs]=2Ws⋅0=02W_s\,\mathbb{E}[W_t-W_s\mid\mathcal{F}_s] = 2W_s\cdot 0 = 0 since the increment is independent of Fs\mathcal{F}_s with mean 00; E[(Wt−Ws)2∣Fs]=Var(Wt−Ws)=t−s\mathbb{E}[(W_t-W_s)^2\mid\mathcal{F}_s] = \mathrm{Var}(W_t-W_s) = t-s. Summing: E[Wt2∣Fs]=Ws2+(t−s)\mathbb{E}[W_t^2\mid\mathcal{F}_s] = W_s^2 + (t-s). Subtracting t−st-s from both sides confirms Mt=Wt2−tM_t = W_t^2 - t satisfies E[Mt∣Fs]=Ws2−s=Ms\mathbb{E}[M_t\mid\mathcal{F}_s] = W_s^2 - s = M_s: a genuine martingale.

Derive E[Wt2∣Fs]\mathbb{E}[W_t^2\mid\mathcal{F}_s] for s≤ts\leq t step by step, and use the result to confirm that Mt=Wt2−tM_t = W_t^2 - t is a martingale.

EXERCISE 4.3

A submartingale drifts upward in expectation; think about what that implies for a trader who can borrow risk-free and buy the asset.

If the discounted price S~t\tilde S_t is a strict submartingale, E[S~t∣Fs]>S~s\mathbb{E}[\tilde S_t\mid\mathcal{F}_s] > \tilde S_s: the market expects the discounted price to rise, on average, above its current value, for free — no risk is being compensated by that expected gain because discounting has already removed the risk-free component. A trader could borrow at the risk-free rate, buy the asset, and expect a strictly positive payoff beyond the cost of borrowing, without bearing priced risk for it under the model's own probabilities. That expected free profit is exactly the informal definition of an arbitrage opportunity used in Section 05.

Explain, in one or two sentences and without invoking a formal arbitrage theorem, why a discounted asset price that is a strict submartingale (rather than a martingale) under the pricing measure would represent a problem for a no-arbitrage model.


07 · Chapter Summary

ConceptMeaning
Ft\mathcal{F}_tinformation available up to time tt
Filtrationincreasing family Fs⊆Ft\mathcal{F}_s\subseteq\mathcal{F}_t for s≤ts\le t
Adapted processXtX_t computable from Ft\mathcal{F}_t; no peeking at the future
Predictablestronger than adapted; needed for stochastic integrands (Chapter 05)
E[Xt∣Fs]\mathbb{E}[X_t\mid\mathcal{F}_s]best prediction of XtX_t using information up to ss
MartingaleE[Mt∣Fs]=Ms\mathbb{E}[M_t\mid\mathcal{F}_s]=M_s — no forecastable drift
Pricing linkdiscounted no-arbitrage prices should be martingales under the right measure

Next: Chapter 05 — The Itô Integral, which uses adapted (predictable) integrands and the left-endpoint convention motivated in Chapter 03 to define ∫0tXs dWs\int_0^t X_s\,dW_s rigorously, and shows that the result is itself a martingale.