Measure-Theoretic Probability
Introduction
Probability theory is measure theory with a normalisation: a probability space is a measure space of total mass one, a random variable is a measurable function, and expectation is the integral. The vocabulary is different from that of analysis, and it carries a century of intuition, but every statement about random variables has a statement about measurable functions behind it, and the theorems are proved by the theorems of the measure theory of Measure Theory and Integration. What the probabilistic formulation adds is a language for independence, for conditioning and for the convergence of distributions, and it is these notions — not the measure theory — that this part of the corpus develops.
This article sets up the frame and fixes the notation for the probability articles throughout. It defines the probability space, the random variable, the distribution, the expectation and the moments, records the standard inequalities, describes the modes of convergence of random variables and their relations, introduces the characteristic function as the Fourier transform of a distribution, and states the classical limit theorems in outline, leaving their proofs to the articles that own them. Three boundaries are held.
- The general measure theory — measurable spaces, measures, the integral, the monotone and dominated convergence theorems, the $L^p$ spaces, the Radon–Nikodym theorem, Fubini's theorem — is the subject of Measure Theory and Integration, and the comparison of the modes of convergence is Modes of Convergence; both are written, and this article uses them.
- Independence and conditional expectation are not covered here,which develops the product measures, the Borel–Cantelli lemmas, the Kolmogorov zero–one law, the conditional expectation as a projection, and the filtrations. This article defines independence only far enough to state the limit theorems; the proofs of the limit theorems are deferred.
- The Fourier transform on $\mathbb{R}^d$ — the Schwartz space, the inversion theorem, the Parseval identity — is Fourier Analysis on Euclidean Spaces, written; the characteristic function is the Fourier transform of a probability measure, and the general analytical theory is cited there. The specific distributions and their special functions are the subject of the Part V articles on the number systems , where the gamma function and its relatives are developed. No physics is invoked.
Throughout, $(\Omega, \mathcal{F}, \mathbb{P})$ is a probability space: $\mathcal{F}$ is a $\sigma$-algebra of subsets of $\Omega$, and $\mathbb{P}$ is a measure on $\mathcal{F}$ with $\mathbb{P}(\Omega) = 1$. A random variable is a measurable function $X : \Omega \to \mathbb{R}^d$ (or to a measurable space), and $\mathbb{E}[X] = \int_\Omega X \, d\mathbb{P}$ is its expectation; the notation $\int X\,d\mathbb{P}$, $\int X$, and $\mathbb{E}[X]$ are used interchangeably. The Borel $\sigma$-algebra of $\mathbb{R}^d$ is $\mathcal{B}(\mathbb{R}^d)$, and a statement holds almost surely (a.s.) when it fails on a $\mathbb{P}$-null set; this is the probabilistic name for the almost everywhere of Measure Theory and Integration, and the two phrases are used interchangeably across the corpus. The $\sigma$-algebra generated by a family $\{X_i\}$ is $\sigma(X_i)$.
Probability Spaces and Random Variables
The Definitions
Definition. A probability space is a measure space $(\Omega, \mathcal{F}, \mathbb{P})$ with $\mathbb{P}(\Omega) = 1$. Its elements are events — the measurable subsets of $\Omega$ — and $\mathbb{P}(A)$ is the probability of the event $A$.
The normalisation turns the measure-theoretic statements into statistical ones: $\mathbb{P}(A^c) = 1 - \mathbb{P}(A)$, $\mathbb{P}(\varnothing) = 0$, and the monotone and dominated convergence theorems of the measure theory become the theorems about the limits of expectations. A null set is a set of probability zero, and the completion of $\mathcal{F}$ with respect to $\mathbb{P}$ — the process of adjoining all subsets of null sets — is assumed without comment, so that almost-sure statements are insensitive to the choice of version of a random variable.
Definition. A random variable on $(\Omega,\mathcal{F},\mathbb{P})$ with values in a measurable space $(S,\mathcal{S})$ is a measurable map $X : \Omega \to S$; when $S = \mathbb{R}^d$ and $\mathcal{S} = \mathcal{B}(\mathbb{R}^d)$ one speaks of a real or vector-valued random variable. The $\sigma$-algebra generated by $X$ is $\sigma(X) = \{X^{-1}(B) : B \in \mathcal{S}\}$, the smallest $\sigma$-algebra making $X$ measurable. A family $\{X_i\}_{i\in I}$ is measurable when each $X_i$ is.
Example (the finite model). A fair coin is the space $\Omega = \{0,1\}$ with $\mathcal{F} = 2^\Omega$ and $\mathbb{P}(\{0\}) = \mathbb{P}(\{1\}) = 1/2$; the identity map is the random variable taking the values $0$ and $1$. A sequence of $n$ independent tosses is the product space $\Omega^n$ with the product measure, and the number of heads is the random variable $\sum_{i=1}^n X_i$. The product measure is constructed.
Example (the uniform model). The space $\Omega = [0,1]$ with the Borel $\sigma$-algebra and Lebesgue measure is a probability space; the identity is the uniform random variable, and every distribution on a Polish space is the image of the uniform distribution under a measurable map by the quantile construction below. This is the sense in which $[0,1]$ is the universal probability space for the standard distributions.
Generated $\sigma$-Algebras and Measurability
Proposition. A random variable $Y : \Omega \to \mathbb{R}$ is $\sigma(X)$-measurable if and only if there is a measurable function $f : \mathbb{R}^d \to \mathbb{R}$ with $Y = f(X)$.
Proof. If $Y = f(X)$ the measurability is clear. Conversely, if $Y$ is $\sigma(X)$-measurable, approximate $Y$ by simple $\sigma(X)$-measurable functions, write each as a finite linear combination of indicators of sets $X^{-1}(B)$, and pass to the limit; the resulting function of $X$ is the required $f$.
The proposition is the measure-theoretic form of the statement that the information carried by $X$ is exactly the information of its $\sigma$-algebra, and it is the first instance of the general principle that conditioning on $X$ is conditioning on $\sigma(X)$. The same statement with $Y$ vector-valued and $f$ measurable is proved by applying the real case coordinatewise.
Distributions and Their Properties
The Distribution of a Random Variable
Definition. The distribution (or law) of a random variable $X$ on $(\Omega,\mathcal{F},\mathbb{P})$ with values in $(S,\mathcal{S})$ is the pushforward measure
$$ \mu_X = \mathbb{P} \circ X^{-1}, \qquad \mu_X(B) = \mathbb{P}(X \in B) = \mathbb{P}\bigl(X^{-1}(B)\bigr), \quad B \in \mathcal{S}. $$
For $S = \mathbb{R}^d$ the distribution function is $F_X(x) = \mu_X\bigl((-\infty, x]\bigr) = \mathbb{P}(X \leq x)$, where the inequality is coordinatewise.
The distribution records everything probabilistic about $X$: two random variables with the same distribution have the same expectation of every measurable function, and one writes $X \stackrel{d}{=} Y$ for equality in distribution. The distribution function of a real random variable is nondecreasing, right-continuous, with limits $0$ at $-\infty$ and $1$ at $+\infty$, and every such function is the distribution function of a unique probability measure on $\mathbb{R}$; the correspondence between measures and distribution functions is a bijection.
Definition. A point $x \in \mathbb{R}$ is an atom of $\mu$ if $\mu(\{x\}) > 0$; the measure is discrete if it is a countable sum of atoms, absolutely continuous with density $f$ if $d\mu = f\,dx$ with $f \geq 0$ measurable and $\int f\,dx = 1$, and singular if it is supported on a Lebesgue-null set. By the Lebesgue decomposition theorem of Measure Theory and Integration every probability measure on $\mathbb{R}$ decomposes uniquely into an absolutely continuous part, a discrete part and a singular continuous part.
Example. The Bernoulli distribution with parameter $p$ is the measure $p\delta_1 + (1-p)\delta_0$ on $\{0,1\}$, discrete with atoms at $0$ and $1$. The binomial distribution $\mathrm{Bin}(n,p)$ is $\sum_{k=0}^{n}\binom{n}{k}p^k(1-p)^{n-k}\delta_k$. The uniform distribution on $[0,1]$ is absolutely continuous with density $\mathbf{1}_{[0,1]}$; the standard normal distribution $N(0,1)$ is absolutely continuous with density $(2\pi)^{-1/2}e^{-x^2/2}$; and the Cantor distribution is singular continuous, with distribution function the Cantor function.
The Quantile Function and the Universal Model
Definition. For a distribution function $F$ the quantile function is $F^{-1}(u) = \inf\{x \in \mathbb{R} : F(x) \geq u\}$ for $u \in (0,1)$, and $F^{-1}(0) = -\infty$, $F^{-1}(1) = +\infty$.
Theorem (Skorokhod representation). Let $F$ be a distribution function on $\mathbb{R}$ and let $U$ be uniform on $[0,1]$. Then $F^{-1}(U)$ has distribution function $F$. Consequently every probability distribution on $\mathbb{R}$ is the image of the uniform distribution under a measurable map.
Proof. The inequality $F^{-1}(u) \leq x$ is equivalent to $u \leq F(x)$ for $u \in (0,1)$; hence
$$ \mathbb{P}\bigl(F^{-1}(U) \leq x\bigr) = \mathbb{P}\bigl(U \leq F(x)\bigr) = F(x), $$
using the uniformity of $U$ and the fact that $F(x)$ lies in $[0,1]$.
The theorem is the concrete form of the statement that the uniform distribution on $[0,1]$ is universal, and it is the basis of the simulation of a distribution from uniform variates. It also shows that the arithmetic of a probability space can be reduced to the arithmetic of $[0,1]$, which is the classical construction of a probability space carrying a sequence of independent random variables of prescribed distributions.
Expectation and Moments
Expectation
Definition. Let $X$ be a real random variable. The expectation is the integral
$$ \mathbb{E}[X] = \int_\Omega X(\omega)\, d\mathbb{P}(\omega), $$
defined when $\int |X| \, d\mathbb{P} < \infty$, in which case $X$ is integrable; and $\mathbb{E}[X] = \int_{\mathbb{R}} x\, d\mu_X(x)$ by the change-of-variables theorem, so the expectation is computed from the distribution alone.
Proposition (properties). Expectation is linear, positive and monotone; it satisfies $\mathbb{E}[c] = c$ for a constant, and $|\mathbb{E}[X]| \leq \mathbb{E}[|X|]$. If $X \geq 0$ a.s. and $\mathbb{E}[X] = 0$ then $X = 0$ a.s. If $\{X_n\}$ is a sequence of nonnegative random variables then the monotone convergence theorem gives $\mathbb{E}[\lim_n X_n] = \lim_n \mathbb{E}[X_n]$; if $|X_n| \leq Y$ with $\mathbb{E}[Y] < \infty$ and $X_n \to X$ a.s. then the dominated convergence theorem gives $\mathbb{E}[X_n] \to \mathbb{E}[X]$. These are the convergence theorems of Measure Theory and Integration translated into probabilistic language.
Inequalities
The probabilistic inequalities are the quantitative content of the measure theory, and they are the tools of the limit theorems.
Theorem (Jensen). Let $\varphi : \mathbb{R} \to \mathbb{R}$ be convex and let $X$ be integrable with $\varphi(X)$ integrable. Then
$$ \varphi(\mathbb{E}[X]) \leq \mathbb{E}[\varphi(X)]; $$
for $\varphi(x) = |x|^p$, $p \geq 1$, this gives $\|X\|_1 \leq \|X\|_p$ and hence $L^p \subseteq L^1$.
Theorem (Markov). For a nonnegative random variable $X$ and $a > 0$,
$$ \mathbb{P}(X \geq a) \leq \frac{\mathbb{E}[X]}{a}. $$
Theorem (Chebyshev). For a random variable with mean $\mu$ and finite variance $\sigma^2$,
$$ \mathbb{P}\bigl(|X - \mu| \geq a\bigr) \leq \frac{\sigma^2}{a^2}, \qquad a > 0. $$
The three inequalities are the standard quantitative bounds, and they are proved by elementary manipulations: Jensen by the supporting line of a convex function, Markov by integrating the inequality $a\mathbf{1}_{\{X\geq a\}} \leq X$, and Chebyshev by applying Markov to $(X-\mu)^2$. The Chebyshev inequality is the engine of the weak law of large numbers, and the Markov inequality is the engine of the first-moment bounds.
Definition. For $X,Y \in L^2$, the variance is $\operatorname{Var}(X) = \mathbb{E}[(X-\mathbb{E}X)^2] = \mathbb{E}[X^2] - (\mathbb{E}X)^2$, the covariance is $\operatorname{Cov}(X,Y) = \mathbb{E}[(X-\mathbb{E}X)(Y-\mathbb{E}Y)]$, the correlation is $\rho(X,Y) = \operatorname{Cov}(X,Y)/(\sigma_X\sigma_Y)$, the moments are $\mathbb{E}[X^k]$, and the $L^p$ norm is $\|X\|_p = (\mathbb{E}[|X|^p])^{1/p}$.
Theorem (Hölder, Minkowski, Cauchy–Schwarz). For $p, q > 1$ with $1/p + 1/q = 1$,
$$ \mathbb{E}[|XY|] \leq \|X\|_p\|Y\|_q, \qquad \|X + Y\|_p \leq \|X\|_p + \|Y\|_p, \qquad \left|\mathbb{E}[XY]\right| \leq \|X\|_2\|Y\|_2, $$
the last being the case $p=q=2$ of the first. Equality in Hölder holds if and only if $|X|^p$ and $|Y|^q$ are proportional a.s.; equality in Cauchy–Schwarz holds if and only if $X$ and $Y$ are linearly dependent a.s. on the event $\{XY \neq 0\}$.
The Cauchy–Schwarz inequality makes $L^2$ an inner product space and is the reason the variance is the square of a norm and the covariance is an inner product; the geometric form of the theory, including the projection onto a subspace, is developed in Banach and Hilbert Spaces and is the structure behind conditional expectation.
Convergence of Random Variables
The Modes
Definition. Let $X, X_1, X_2, \dots$ be random variables on $(\Omega,\mathcal{F},\mathbb{P})$ with values in $\mathbb{R}^d$.
- $X_n \to X$ almost surely ($X_n \to X$ a.s.) if $\mathbb{P}(\lim_n X_n = X) = 1$.
- $X_n \to X$ in probability ($X_n \xrightarrow{\mathbb{P}} X$) if for every $\varepsilon > 0$, $\mathbb{P}(\|X_n - X\| > \varepsilon) \to 0$.
- $X_n \to X$ in $L^p$ for $p \geq 1$ if $\|X_n - X\|_p \to 0$.
- $X_n \to X$ in distribution ($X_n \xrightarrow{d} X$) if $\mathbb{E}[f(X_n)] \to \mathbb{E}[f(X)]$ for every bounded continuous $f$.
Theorem (the relations). Convergence almost sure and convergence in $L^p$ each imply convergence in probability, which implies convergence in distribution. None of the implications reverses. If $X_n \xrightarrow{d} X$ and $X$ is a.s. constant, then $X_n \xrightarrow{\mathbb{P}} X$; if $X_n \xrightarrow{\mathbb{P}} X$ then a subsequence converges a.s.; and $X_n \xrightarrow{d} X$ if and only if $X_n \xrightarrow{\mathbb{P}} X$ together with the uniform integrability of an appropriate family, by the theorem of Vitali.
Proof. Almost sure convergence implies convergence in probability by the continuity from above of the measure; convergence in probability implies convergence in distribution by the portmanteau theorem below and the boundedness of $f$; the strictly decreasing chain is witnessed by the standard examples. The remaining statements are the standard measure-theoretic arguments.
Example (the implications are strict). The typewriter sequence $X_n = \mathbf{1}_{[k/2^m, (k+1)/2^m]}$ on $[0,1]$, $n = 2^m + k$, converges in probability and in $L^1$ to $0$ but converges a.s. at no point. The sequence $X_n = 2^n\mathbf{1}_{[0,1/n]}$ converges to $0$ a.s. and in probability but not in $L^1$. And a sequence converging in distribution to a random limit need not converge in probability, as the example $X_n$ uniform on $\{0,1\}$ and $X = 1 - X_n$ shows.
The Portmanteau Theorem
Theorem (portmanteau). For probability measures $\mu, \mu_n$ on a metric space the following are equivalent: (i) $\mu_n \to \mu$ weakly, that is $\int f\,d\mu_n \to \int f\,d\mu$ for every bounded continuous $f$; (ii) $\liminf_n \mu_n(G) \geq \mu(G)$ for every open $G$; (iii) $\limsup_n \mu_n(F) \leq \mu(F)$ for every closed $F$; (iv) $\mu_n(A) \to \mu(A)$ for every Borel set $A$ with $\mu(\partial A) = 0$; (v) on $\mathbb{R}^d$, the distribution functions converge at every continuity point of $F_\mu$.
The theorem is the standard coherence statement for convergence in distribution, and its proof is a measure-theoretic approximation argument: the continuous functions are used to test against the closed and open sets, and the boundary of a set is the place where the two can disagree. In the probabilistic language it says that convergence in distribution is weak convergence of the laws, and the two phrases are used interchangeably.
Theorem (continuous mapping and Slutsky). If $X_n \xrightarrow{d} X$ and $g$ is continuous then $g(X_n) \xrightarrow{d} g(X)$. If $X_n \xrightarrow{d} X$ and $Y_n \xrightarrow{\mathbb{P}} c$ a constant, then $X_n + Y_n \xrightarrow{d} X + c$ and $X_nY_n \xrightarrow{d} cX$.
Slutsky's theorem is the working form of the continuous mapping principle: it permits the replacement of the deterministic constants in a limit theorem by consistent estimators, and it is used in every application of the central limit theorem in statistics. Its proof is a direct estimate against a bounded Lipschitz test function, using the tightness of $\{X_n\}$ and the convergence in probability of $Y_n$.
Characteristic Functions
The Definition and Basic Properties
Definition. The characteristic function of a random variable $X$ on $\mathbb{R}^d$ is
$$ \varphi_X(t) = \mathbb{E}\!\left[e^{i\langle t, X\rangle}\right] = \int_{\mathbb{R}^d} e^{i\langle t,x\rangle}\, d\mu_X(x), \qquad t \in \mathbb{R}^d, $$
the Fourier transform of the distribution $\mu_X$.
Theorem (basic properties). The characteristic function satisfies $\varphi_X(0) = 1$, $|\varphi_X(t)| \leq 1$, $\varphi_X(-t) = \overline{\varphi_X(t)}$, and $\varphi_X$ is uniformly continuous on $\mathbb{R}^d$. It is positive definite in the sense that $\sum_{j,k} c_j\overline{c_k}\varphi_X(t_j - t_k) \geq 0$ for all finite families $t_j \in \mathbb{R}^d$ and $c_j \in \mathbb{C}$; conversely, by Bochner's theorem, every continuous positive definite function with $\varphi(0) = 1$ is the characteristic function of a unique probability measure.
Proof. The first four properties are immediate from the definition and the dominated convergence theorem; uniform continuity follows from $\left|\varphi(t+h) - \varphi(t)\right| \leq \mathbb{E}\left|e^{i\langle h,X\rangle} - 1\right|$, whose right-hand side is independent of $t$ and tends to $0$ as $h \to 0$ by dominated convergence. Positive definiteness is the computation $\sum_{j,k}c_j\overline{c_k}\varphi(t_j-t_k) = \mathbb{E}\left|\sum_j c_j e^{i\langle t_j,X\rangle}\right|^2 \geq 0$. The converse is Bochner's theorem, whose proof is the inversion formula of Fourier Analysis on Euclidean Spaces.
Theorem (uniqueness and inversion). The characteristic function determines the distribution: if $\varphi_X = \varphi_Y$ then $\mu_X = \mu_Y$. More precisely, if $\varphi$ is integrable then $\mu_X$ has the continuous density
$$ f(x) = \frac{1}{(2\pi)^d}\int_{\mathbb{R}^d} e^{-i\langle t,x\rangle}\varphi(t)\, dt, $$
and in general the distribution is recovered from the characteristic function by the inversion formula for the Fourier transform of a finite measure.
The uniqueness theorem is the reason the characteristic function is the standard tool for limit theorems: it linearises the operation of passing to a limit of distributions, and it converts the convergence in distribution into the pointwise convergence of functions. The analytic theory behind the inversion formula — the Schwartz space, the Parseval identity and the extension of the Fourier transform to tempered distributions — is the subject of Fourier Analysis on Euclidean Spaces.
Example. The normal distribution $N(\mu,\sigma^2)$ has characteristic function $\varphi(t) = e^{i\mu t - \sigma^2t^2/2}$; the Poisson distribution with parameter $\lambda$ has $\varphi(t) = e^{\lambda(e^{it}-1)}$ for a variable on $\mathbb{Z}_{\geq0}$; the Cauchy distribution with density $1/(\pi(1+x^2))$ has $\varphi(t) = e^{-|t|}$; and the Bernoulli distribution has $\varphi(t) = 1 - p + pe^{it}$. The last three show that the characteristic function is generally complex-valued and need not be integrable, so the inversion formula must be applied in the distributional sense.
Moment Generating and Cumulants
Definition. When $X$ has finite moments of all orders the moment generating function is $M_X(t) = \mathbb{E}[e^{tX}]$, defined near $t=0$ when the expectation converges; the cumulants $\kappa_n$ are the coefficients in the expansion $\log \varphi_X(t) = \sum_{n\geq1}\kappa_n (it)^n/n!$. The moments are recovered from the derivatives of $\varphi_X$ at $0$: $\mathbb{E}[X^n] = i^{-n}\varphi_X^{(n)}(0)$ when the moment exists.
The cumulants are additive for sums of independent random variables, and this is the structural reason they appear in the central limit theorem: the $n$-th cumulant of a sum of independent variables is the sum of the $n$-th cumulants, so the asymptotic distribution of a sum is governed by its first two cumulants. The additivity is a formal consequence of the factorisation $\varphi_{X+Y} = \varphi_X\varphi_Y$ for independent $X,Y$, which is stated in the next section and proved.
The Limit Theorems in Outline
Independence, Briefly
Definition. The random variables $X_1, \dots, X_n$ are independent if for all Borel sets $B_1, \dots, B_n$,
$$ \mathbb{P}(X_1 \in B_1, \dots, X_n \in B_n) = \prod_{i=1}^n \mathbb{P}(X_i \in B_i), $$
and an infinite family is independent if every finite subfamily is. Equivalently, the joint distribution is the product of the marginal distributions, and the characteristic function factorises: $\varphi_{X_1+\cdots+X_n} = \prod_i \varphi_{X_i}$.
Independence is the probabilistic content of the product measure, and it is developed elsewhere, where the product measure is constructed, the Borel–Cantelli lemmas and the Kolmogorov zero–one law are proved, and the conditional expectation is defined as a projection. This article uses only the factorisation of the characteristic function and the additivity of the cumulants.
The Statements
Theorem (weak law of large numbers; Khinchin). Let $X_1, X_2, \dots$ be independent and identically distributed with $\mathbb{E}|X_1| < \infty$ and mean $\mu$. Then the sample means converge in probability:
$$ \frac{1}{n}\sum_{i=1}^n X_i \xrightarrow{\mathbb{P}} \mu. $$
Theorem (strong law of large numbers; Kolmogorov). Under the same hypotheses the sample means converge almost surely: $\frac{1}{n}\sum_{i=1}^n X_i \to \mu$ a.s.
Theorem (central limit theorem; Lindeberg–Lévy). Let $X_1, X_2, \dots$ be independent and identically distributed with mean $\mu$ and finite variance $\sigma^2 > 0$. Then
$$ \frac{1}{\sigma\sqrt{n}}\sum_{i=1}^n (X_i - \mu) \xrightarrow{d} N(0,1), $$
the standard normal distribution.
The three theorems are the classical limit theorems, and they are the subject, where they are proved, refined and generalised: the weak law follows from Chebyshev when the variance is finite and from the truncation argument of Khinchin in general; the strong law is proved by the maximal inequality of Kolmogorov; and the central limit theorem is proved by the convergence of characteristic functions, via the expansion $\varphi(t/\sqrt n)^n \to e^{-t^2/2}$ and the continuity theorem of Lévy. The rate of convergence in the central limit theorem is the Berry–Esseen theorem, and the local form is the local limit theorem; both are developed there.
The place of the limit theorems in the present article is structural. The weak law is the statement that the empirical mean is a consistent estimator of the mean; the strong law is the statement that the empirical measure converges to the distribution almost surely, and it is the probabilistic form of the pointwise ergodic theorem; and the central limit theorem is the statement that the fluctuations of the empirical mean around its limit are asymptotically Gaussian, and it is the reason the normal distribution is universal. The passage from the law of large numbers to ergodic theory is the passage from independent sequences to stationary ones, and it is developed.
Summary
A probability space is a measure space of total mass one; a random variable is a measurable function, its distribution is the pushforward $\mu_X = \mathbb{P}\circ X^{-1}$, the distribution function $F_X(x) = \mathbb{P}(X \le x)$ is nondecreasing and right-continuous and determines the measure, and the Lebesgue decomposition splits every distribution into an absolutely continuous, a discrete and a singular continuous part. Every distribution on $\mathbb{R}$ is the image of the uniform distribution under the quantile map $F^{-1}$, which makes $[0,1]$ the universal probability space. Expectation is the integral, computed from the distribution by change of variables, and it satisfies the monotone and dominated convergence theorems; the standard inequalities are Jensen, Markov and Chebyshev, together with Hölder, Minkowski and Cauchy–Schwarz, the last making $L^2$ an inner product space whose geometry underlies the covariance and the conditional expectation.
Convergence of random variables has four modes — almost surely, in probability, in $L^p$ and in distribution — with the implications a.s. and $L^p$ $\Rightarrow$ probability $\Rightarrow$ distribution, none reversible; the portmanteau theorem characterises convergence in distribution by the weak convergence of the laws, and Slutsky's theorem and the continuous mapping principle are the working tools. The characteristic function is the Fourier transform of the distribution: it is uniformly continuous, positive definite, bounded by $1$ and equal to $1$ at the orig, it determines the distribution by the inversion formula and Bochner's theorem, it factorises over independent sums, and its logarithm has additive cumulants. The weak and strong laws of large numbers and the central limit theorem are stated here in their classical form and proved; independence and conditional expectation are constructed, and the ergodic form of the law of large numbers is developed.
Summary of Notation
| Symbol | Meaning |
|---|---|
| $(\Omega, \mathcal{F}, \mathbb{P})$ | Probability space; $\mathbb{P}(\Omega)=1$ |
| a.s. | Almost surely; the probabilistic name for almost everywhere |
| $X$, $Y$ | Random variables (measurable maps) |
| $\sigma(X_i)$ | $\sigma$-algebra generated by the family $\{X_i\}$ |
| $\mu_X = \mathbb{P}\circ X^{-1}$ | Distribution (law) of $X$ |
| $F_X(x) = \mathbb{P}(X \le x)$ | Distribution function |
| $F^{-1}$ | Quantile function, $F^{-1}(u)=\inf\{x: F(x)\ge u\}$ |
| $\mathbb{E}[X]$ | Expectation, $\int_\Omega X\,d\mathbb{P}$ |
| $\operatorname{Var}$, $\operatorname{Cov}$, $\rho$ | Variance, covariance, correlation |
| $\|X\|_p$ | $L^p$ norm $(\mathbb{E}\|X\|^p)^{1/p}$ |
| $\mathbf{1}_A$, $\delta_x$ | Indicator of $A$, point mass at $x$ |
| $X_n \to X$ a.s., $\xrightarrow{\mathbb{P}}$, $\xrightarrow{L^p}$, $\xrightarrow{d}$ | The four modes of convergence |
| $\mu_n \to \mu$ weakly | $\int f\,d\mu_n \to \int f\,d\mu$ for bounded continuous $f$ |
| $\varphi_X(t) = \mathbb{E}[e^{i\langle t,X\rangle}]$ | Characteristic function |
| $M_X$, $\kappa_n$ | Moment generating function, cumulants |
| $N(\mu,\sigma^2)$, $\mathrm{Bin}(n,p)$, $\mathrm{Pois}(\lambda)$ | Normal, binomial, Poisson distributions |
| independent | joint law is the product of the marginals |
Further Reading
- Patrick Billingsley, Probability and Measure (Wiley, 3rd edition, 1995), for the measure-theoretic foundations of probability from first principles.
- Andrei N. Kolmogorov, Foundations of the Theory of Probability (Chelsea, 1950), for the original measure-theoretic axiomatisation.
- Kai Lai Chung, A Course in Probability Theory (Academic Press, 3rd edition, 2001), for the modes of convergence, the portmanteau theorem and the classical limit theorems.
- William Feller, An Introduction to Probability Theory and Its Applications, Vols. I and II (Wiley, 1968 and 1971), for the characteristic functions, the standard distributions and their limit behaviour.
- Michel Loève, Probability Theory (Springer, 4th edition, 1977), for the systematic development of distributions and their decompositions.
- David Williams, Probability with Martingales (Cambridge University Press, 1991), for a concise account that passes directly from the measure theory to the martingale theory.
- Elias M. Stein and Rami Shakarchi, Fourier Analysis: An Introduction (Princeton University Press, 2003), for the Fourier transform, the inversion formula and the positive-definite functions used in the characteristic-function theory.