Modular Forms

Introduction

A modular form is a holomorphic function on the upper half-plane which is almost invariant under the action of a discrete subgroup of the Möbius transformations, the discrepancy between the value at $z$ and the value at the image of $z$ being a prescribed power of the derivative of the transformation. The invariance is by a factor of automorphy rather than exact, and that factor is what makes the theory non-trivial: it produces finite-dimensional spaces of functions, a ring structure in favourable cases, a Fourier expansion with integral coefficients for arithmetic subgroups, and — through the Hecke operators — Dirichlet series with Euler products and functional equations. Modular forms are thus at once analytic objects, carrying growth conditions and an inner product; algebraic ones, forming graded rings with algebraic Fourier coefficients; and arithmetic ones, parametrising elliptic curves and encoding Galois representations.

This article develops the classical theory over $\mathbb{Q}$: the modular group and its congruence subgroups, the spaces $M_k(\Gamma)$ and $S_k(\Gamma)$ and their dimension, the structure theorem for the level-one ring with the discriminant and $j$-invariant, the Hecke operators with the theory of newforms and the Petersson inner product, the $L$-function of a modular form with its Euler product and functional equation, theta series and the half-integral weight case, modular curves with the Eichler–Shimura isomorphism, and the modularity theorem with its arithmetic consequences. The analytic input — the Fourier expansion, the growth of the coefficients, the convergence of the Eisenstein series, the theta transformation law — rests on the previous articles: the power series of Analytic Functions and Power Series, the vocabulary of Modes of Convergence, the integration on the upper half-plane of Measure Theory and Integration, and the theta functional equation of Zeta Functions. The arithmetic of the elliptic curves that the modularity theorem attaches to weight-two forms is that of Elliptic Curves; the class-field content of the $j$-invariant is that of Galois Theory and Algebraic Number Theory; and the arithmetic of the complex numbers used in the multiplier systems of half-integral weight is deferred to the synthetic study Galois Theory of $\mathbb{C}/\mathbb{R}$. The adelic re-reading of the whole theory, and the general linear group case, belong to Automorphic Forms, in the neighbouring category of this Part and written in parallel; the seam is fixed at the end of this article. The zeros of the $L$-functions of modular forms are treated in The Riemann Hypothesis, written above.

Throughout, $\mathbb{H} = \{z\in\mathbb{C} : \operatorname{Im}z>0\}$ is the upper half-plane, $\Gamma$ a subgroup of finite index of $SL_2(\mathbb{Z})$ inside $SL_2(\mathbb{R})$, and $q = e^{2\pi iz}$; the notation $\Gamma(1) = SL_2(\mathbb{Z})$ is used for the full group and $\bar\Gamma$ for its image in $PSL_2(\mathbb{Z})$. Weights $k$ are integers, positive unless stated. The identity matrix is $I$, the standard generators of $\Gamma(1)$ are $$ S = \begin{pmatrix}0&-1\\1&0\end{pmatrix}, \qquad T = \begin{pmatrix}1&1\\0&1\end{pmatrix}, $$ acting by $Sz = -1/z$ and $Tz = z+1$, and the modular group is $\bar\Gamma(1) = \Gamma(1)/\{\pm I\}$.

The Modular Group and its Fundamental Domain

The Action on the Upper Half-Plane

Proposition. $SL_2(\mathbb{R})$ acts on $\mathbb{H}$ by $$ \begin{pmatrix}a&b\\c&d\end{pmatrix}\cdot z = \frac{az+b}{cz+d}, $$ the action is transitive, and the stabiliser of $i$ is $SO(2)$. The kernel of the action is $\{\pm I\}$, so the group $\bar\Gamma(1)$ acts faithfully.

Proof. For $\operatorname{Im} z = y>0$ one computes $\operatorname{Im}(\gamma z) = y/\lvert cz+d\rvert^2>0$, so the action preserves $\mathbb{H}$; transitivity follows because $$ \begin{pmatrix}y^{1/2}&xy^{-1/2}\\0&y^{-1/2}\end{pmatrix}\cdot i = x+iy , $$ and the stabiliser computation is a direct matrix identity. The kernel is $\{\pm I\}$, since if $\gamma z = z$ for every $z$ then $\gamma$ is scalar, and the scalars in $SL_2$ are $\pm I$.

Theorem (the presentation). $\bar\Gamma(1)$ is generated by $S$ and $T$, with the presentation $\langle S,T \mid S^2 = 1,\ (ST)^3 = 1\rangle$, so that $\bar\Gamma(1)$ is the free product of $\mathbb{Z}/2\mathbb{Z}$ and $\mathbb{Z}/3\mathbb{Z}$. The element $ST$ is $z\mapsto-1/(z+1)$, of order $3$; the conjugate element $TS$ is $z\mapsto(z-1)/z$, of the same order.

Proof sketch. One reduces an arbitrary unimodular matrix by the Euclidean algorithm, right multiplication by the powers of $T$ (which reduce the size of the upper-right entry) and by $S$ (which swaps the entries), and concludes generation; the relations are checked by direct computation, and the absence of further relations by the action on a fundamental domain whose boundary is a union of translates by $S$ and $ST$.

The Fundamental Domain and the Elliptic Points

Definition. The standard fundamental domain is $$ \mathcal{F} = \Bigl\{z\in\mathbb{H} : \lvert\operatorname{Re}z\rvert\leq\tfrac12,\ \lvert z\rvert\geq1\Bigr\} \cup \Bigl\{z\in\mathbb{H}: \lvert\operatorname{Re}z\rvert = \tfrac12,\ \lvert z\rvert>1\Bigr\}, $$ with the convention that the boundary arcs are assigned to one of the two adjacent translates.

Theorem. $\mathcal{F}$ is a fundamental domain for $\bar\Gamma(1)$: every orbit meets $\mathcal{F}$, and no two interior points of $\mathcal{F}$ are equivalent. The quotient $\bar\Gamma(1)\backslash\mathbb{H}$ is homeomorphic to a sphere with one point removed, and it is compactified by adjoining the cusp $\infty$.

Proof sketch. For each $z$, the orbit contains a point of maximal imaginary part, among which one of minimal $\lvert\operatorname{Re}\rvert$ is in $\mathcal{F}$; two interior points of $\mathcal{F}$ cannot be equivalent, because a Möbius transformation with $c\neq0$ decreases the imaginary part of an interior point and one with $c=0$ shifts the real part by an integer. The quotient is topologically a disc with the point $\infty$ adjoined.

Definition. A point $z\in\mathbb{H}$ is elliptic for $\Gamma$ if some $\gamma\neq\pm I$ in $\Gamma$ fixes it; the order of the stabiliser is the order of the point. The cusps of $\Gamma$ are the orbits of $\Gamma$ on $\mathbb{P}^1(\mathbb{Q}) = \mathbb{Q}\cup\{\infty\}$, and $\mathbb{H}^* = \mathbb{H}\cup\mathbb{P}^1(\mathbb{Q})$.

Proposition. The elliptic points of $\bar\Gamma(1)$ are $i$, of order $2$, and $\rho = e^{2\pi i/3}$, of order $3$; the group has exactly one cusp, namely $\infty$. The stabiliser of $\infty$ in $\Gamma(1)$ is generated by $T$ and $-I$.

Proof. Elliptic points correspond to the fixed points of conjugates of the elements of finite order, which are the conjugates of $S$ (order $2$, fixed point $i$) and of $ST$ (order $3$, fixed point $\rho = e^{2\pi i/3}$); the transitivity of the action identifies the fixed points of the conjugates. The cusps are the rational boundary points, and the action of $\Gamma(1)$ on them is transitive.

Congruence Subgroups

Definition. For $N\geq1$ the principal congruence subgroup is $$ \Gamma(N) = \Bigl\{\gamma\in SL_2(\mathbb{Z}) : \gamma\equiv I \pmod N\Bigr\}, $$ and the groups $\Gamma_0(N)$ and $\Gamma_1(N)$ are the preimages in $SL_2(\mathbb{Z})$ of the upper triangular and of the unipotent upper triangular subgroups of $SL_2(\mathbb{Z}/N\mathbb{Z})$: $$ \Gamma_0(N) = \Bigl\{\begin{pmatrix}a&b\\c&d\end{pmatrix} : c\equiv0 \pmod N\Bigr\}, \qquad \Gamma_1(N) = \Bigl\{\begin{pmatrix}a&b\\c&d\end{pmatrix} : c\equiv0,\ a\equiv d\equiv1 \pmod N\Bigr\}. $$ A subgroup is a congruence subgroup if it contains $\Gamma(N)$ for some $N$.

Theorem (the indices). $\Gamma(N)$ is normal in $\Gamma(1)$ with quotient $SL_2(\mathbb{Z}/N\mathbb{Z})$, and $$ [\Gamma(1):\Gamma(N)] = N^3\prod_{p\mid N}\Bigl(1-\frac{1}{p^2}\Bigr), \qquad [\Gamma(1):\Gamma_0(N)] = N\prod_{p\mid N}\Bigl(1+\frac1p\Bigr), \qquad [\Gamma_0(N):\Gamma_1(N)] = \varphi(N), \qquad [\Gamma_1(N):\Gamma(N)] = N . $$ Proof. The reduction map $SL_2(\mathbb{Z})\to SL_2(\mathbb{Z}/N\mathbb{Z})$ is surjective — for $N$ a prime power by lifting a preimage of a column and completing it to a unimodular matrix, and for general $N$ by the Chinese remainder theorem — and its kernel is $\Gamma(N)$; the order of $SL_2(\mathbb{Z}/N\mathbb{Z})$ is the displayed product, computed from the prime-power case by counting the primitive vectors. The remaining indices are the orders of the corresponding quotients: $\Gamma_0(N)/\Gamma_1(N)$ is the diagonal subgroup and $\Gamma_1(N)/\Gamma(N)$ the unipotent one.

Remark (the congruence condition). Not every subgroup of finite index contains a principal congruence subgroup, so not every finite-index subgroup is a congruence subgroup. The Fourier-coefficient integrality, the Hecke theory and the congruence relations below are proved for congruence subgroups, and all the groups used here are of that kind: $\Gamma_0(N)$ and $\Gamma_1(N)$ are preimages in $SL_2(\mathbb{Z})$ and so contain $\Gamma(N)$.

The Compactified Quotient and its Genus

Definition. For a congruence subgroup $\Gamma$ the quotient $X_\Gamma = \bar\Gamma\backslash\mathbb{H}^*$ is the modular curve of $\Gamma$, a compact Riemann surface, and $Y_\Gamma = \bar\Gamma\backslash\mathbb{H}$ is its affine part. Its genus $g(\Gamma)$ is the genus of $X_\Gamma$.

Theorem (the genus; the case of prime level). Let $N$ be a prime. Then $\Gamma_0(N)$ has exactly two cusps, its count of elliptic points of order $2$ is $$ \nu_2 = \begin{cases} 1 & N = 2,\\ 2 & N\equiv1 \pmod 4,\\ 0 & N\equiv3 \pmod 4,\end{cases} \qquad\text{and its count of order }3\text{ is}\qquad \nu_3 = \begin{cases} 1 & N = 3,\\ 2 & N\equiv1 \pmod 3,\\ 0 & N\equiv2 \pmod 3.\end{cases} $$ Consequently $$ g(\Gamma_0(N)) = 1 + \frac{N+1}{12} - 1 - \frac{\nu_2}{4} - \frac{\nu_3}{3}, $$ so that $X_0(N)$ has genus $0$ for $N = 2,3,5,7,13$, genus $1$ for $N = 11,17,19,\dots$, and genus $2$ for $N=37$.

Proof sketch. The Riemann–Roch formula for the orbifold $\Gamma\backslash\mathbb{H}^*$ gives $$ g = 1 + \frac{[\Gamma(1):\Gamma]}{12} - \frac{\nu_\infty}{2} - \frac{\nu_2}{4} - \frac{\nu_3}{3}, $$ with $\nu_\infty$ the number of cusps and $\nu_2,\nu_3$ the numbers of elliptic points of order $2$ and $3$; for $\Gamma_0(N)$ the number of cusps is $\sum_{d\mid N}\varphi(\gcd(d,N/d))$, which is $2$ for prime $N$, and the elliptic-point counts are computed by counting the solutions of the congruence conditions in $SL_2(\mathbb{Z}/N\mathbb{Z})$ that fix $i$ and $\rho$. The displayed values follow by substitution.

Corollary. For a congruence subgroup $\Gamma$ the quotient $X_\Gamma$ is a compact Riemann surface of the genus above, the cusps being the points added at the rational boundary; the natural map $X_\Gamma\to X_{\Gamma(1)}$ is a covering of degree $[\bar\Gamma(1):\bar\Gamma]$, ramified at the elliptic points and the cusps.

Modular Forms and Cusp Forms

The Automorphy Factor and the Definition

Definition. For $k\in\mathbb{Z}$ and $\gamma = \begin{pmatrix}a&b\\c&d\end{pmatrix}$ define the automorphy factor and the weight-$k$ action $$ j(\gamma,z) = cz+d, \qquad (f\mid_k\gamma)(z) = j(\gamma,z)^{-k}f(\gamma z) = (cz+d)^{-k}f\Bigl(\frac{az+b}{cz+d}\Bigr). $$ Proposition (the cocycle relation). $j(\gamma\delta,z) = j(\gamma,\delta z)\,j(\delta,z)$, and consequently $\mid_k$ is a right action of $SL_2(\mathbb{R})$ on functions on $\mathbb{H}$; the scalar matrix $-I$ acts by $(-1)^k$, so that $\mid_k$ factors through $\bar\Gamma(1)$ exactly when $k$ is even.

Proof. Both statements are the chain rule for the derivative of a Möbius transformation, $(\gamma\delta)'(z) = \gamma'(\delta z)\delta'(z)$, the derivative of $z\mapsto(az+b)/(cz+d)$ being $j(\gamma,z)^{-2}$.

Definition. A modular form of weight $k$ for $\Gamma$ is a holomorphic function $f:\mathbb{H}\to\mathbb{C}$ such that $f\mid_k\gamma = f$ for every $\gamma\in\Gamma$ and such that $f$ is holomorphic at every cusp: for each $\sigma\in SL_2(\mathbb{Z})$ the function $f\mid_k\sigma$ has a Fourier expansion $\sum_{n\geq0}a_nq^{n/h}$ with $a_n = 0$ for $n<0$, where $h$ is the width of the cusp. It is a cusp form if $a_0 = 0$ at every cusp, equivalently if $f$ vanishes at every cusp. The spaces are $M_k(\Gamma)\supseteq S_k(\Gamma)$, and $M_k(\Gamma) = 0$ for odd $k$ when $-I\in\Gamma$.

Definition. A modular form $f$ is a modular function if $k=0$ and the only pole is at the cusps, and $f$ is entire if $k=0$ and holomorphic everywhere including the cusps; the entire modular functions are the constants, by the maximum principle applied on the compact quotient $X_\Gamma$.

Theorem (finite dimensionality; the growth of coefficients). $M_k(\Gamma)$ and $S_k(\Gamma)$ are finite-dimensional, of dimension bounded by the Riemann–Roch formula below, and if $f = \sum a_nq^n$ is a cusp form of weight $k$ for a congruence subgroup then $a_n = O(n^{k/2+\epsilon})$ for every $\epsilon>0$, with the sharper bound $a_n = O(n^{(k-1)/2}\sigma_1(n))$ in the classical form.

Proof sketch. The finite dimensionality is the Riemann–Roch statement of the next subsection. For the coefficients, the Petersson formula expresses $a_n$ as an inner product of $f$ with a Poincaré series, and the estimate of that inner product uses the decay of $f$ at the cusps; the sharp bound $a_n = O(n^{(k-1)/2+\epsilon})$ is the consequence of Deligne's theorem on the absolute values of the eigenvalues of Frobenius, stated below.

The Fourier Expansion and Examples

Definition. The Fourier expansion (or $q$-expansion) of a modular form for $\Gamma$ at the cusp $\infty$ is $$ f(z) = \sum_{n\geq0} a_n q^n, \qquad q = e^{2\pi iz}, $$ valid for $\operatorname{Im}z$ large, the sum being finite in negative powers by the holomorphy at the cusp; a normalised form has $a_1 = 1$.

Definition. For even $k\geq4$ the Eisenstein series of weight $k$ for $\Gamma(1)$ is $$ G_k(z) = \sum_{(m,n)\neq(0,0)}\frac{1}{(mz+n)^k} = 2\zeta(k)E_k(z), \qquad E_k(z) = 1 - \frac{2k}{B_k}\sum_{n\geq1}\sigma_{k-1}(n)q^n, $$ with $\sigma_{r}(n) = \sum_{d\mid n}d^r$ and $B_k$ the Bernoulli number of Zeta Functions; the series converges absolutely for $k\geq4$ and defines a modular form of weight $k$ for $\Gamma(1)$.

Proof sketch. Absolute convergence and the automorphy are immediate from the definition — the sum is over a lattice and $mz+n$ transforms by $j(\gamma,z)$ — and the Fourier expansion is the Lipschitz formula, obtained by Poisson summation on the $n$ variable, which produces the divisor sums and the Bernoulli factors. The two normalisations differ by $2\zeta(k)$, so that $E_k$ has constant term $1$.

Remark (the weight two exception). For $k=2$ the Eisenstein series converges only conditionally and transforms by an extra term: with $E_2(z) = 1-24\sum\sigma_1(n)q^n$ one has $$ E_2(\gamma z) = j(\gamma,z)^2E_2(z) + \frac{6}{\pi i}c\,j(\gamma,z), $$ so that $E_2$ is a quasimodular form and not a modular form; the non-holomorphic completion $E_2(z)-\frac{3}{\pi y}$ is modular and is the standard substitute in the theory of almost holomorphic forms.

Definition. The discriminant and the Dedekind eta function are $$ \Delta = \frac{E_4^3-E_6^2}{1728}, \qquad \eta(z) = q^{1/24}\prod_{n\geq1}(1-q^n), \qquad \Delta = \eta^{24}. $$ Proposition. $\Delta$ is a cusp form of weight $12$ for $\Gamma(1)$ with $\Delta = q\prod_{n\geq1}(1-q^n)^{24}$, and writing $\Delta = \sum_{n\geq1}\tau(n)q^n$ for the Ramanujan function, $$ \tau(1)\dots\tau(12) = 1,\ -24,\ 252,\ -1472,\ 4830,\ -6048,\ -16744,\ 84480,\ -113643,\ -115920,\ 534612,\ -370944 . $$ Proof sketch. $E_4^3$ and $E_6^2$ are both forms of weight $12$ with constant term $1$, so their difference is a cusp form; the value of the constant $1728$ and the product formula are checked by comparing the $q$-expansions, the product being the classical Jacobi triple-product identity in the form $\eta^{24} = \Delta$. The coefficients are obtained by expanding the product, and the numbers have been checked by exact integer computation from the definition $E_4 = 1+240\sum\sigma_3(n)q^n$, $E_6 = 1-504\sum\sigma_5(n)q^n$.

The Valence Formula and the Dimension

Theorem (the valence formula). Let $\Gamma$ contain $-I$ and let $k$ be even. For a nonzero $f\in M_k(\Gamma)$, $$ \sum_{P\in\bar\Gamma\backslash\mathbb{H}^*}\frac{1}{\lvert\bar\Gamma_P\rvert}\operatorname{ord}_P(f) = \frac{k\,[\bar\Gamma(1):\bar\Gamma]}{12}, $$ the sum over the orbits of the points of $\mathbb{H}^*$, where $\operatorname{ord}_P$ is the order of vanishing at $P$ and $\lvert\bar\Gamma_P\rvert$ the order of the stabiliser.

Proof sketch. The form $f$ has a $\Gamma$-invariant divisor on the compact curve $X_\Gamma$; the divisor of the differential $f(z)(dz)^{k/2}$ is computed and its degree is compared with the degree of $dz^{k/2}$, which is $\frac{k}{12}[\bar\Gamma(1):\bar\Gamma]$ by the Gauss–Bonnet computation on the orbifold quotient, the contributions of the elliptic points and cusps being the reciprocal stabiliser factors.

Corollary (the case of level one). For $f\in M_k(\Gamma(1))$ nonzero and $k$ even, $$ \operatorname{ord}_\infty f + \tfrac12\operatorname{ord}_if + \tfrac13\operatorname{ord}_\rho f + \sum_{P\neq i,\rho}\operatorname{ord}_Pf = \frac{k}{12}, $$ and consequently $M_k(\Gamma(1))\neq0$ only for even $k$, with $M_0$ the constants, $M_2 = 0$, and $\dim M_k = \lfloor k/12\rfloor+1$ for even $k\geq4$ with $k\not\equiv2\pmod{12}$, while $\dim M_k = \lfloor k/12\rfloor$ for $k\equiv2\pmod{12}$.

Proof. The valence formula is the statement; the dimension is obtained by counting the possible orders of vanishing of a form of weight $k$: a cusp form has $\operatorname{ord}_\infty f\geq1$, a form vanishing at $i$ contributes $1/2$ and one vanishing at $\rho$ contributes $1/3$, and the integrality of the resulting count forces the exceptional congruence. The vanishing of $M_2$ follows because every positive contribution to the valency sum is at least $\frac13$, while the sum required for a form of weight $2$ is $\frac{2}{12} = \frac16$; no nonzero form of weight $2$ can therefore exist.

Corollary (the dimension in general). For a congruence subgroup $\Gamma$ and even $k\geq4$ the Eisenstein subspace of $M_k(\Gamma)$ has dimension $\nu_\infty$, one form for each cusp, obtained by averaging the Eisenstein series of the cusps over the cosets of $\Gamma$; hence $$ \dim M_k(\Gamma) = \dim S_k(\Gamma) + \nu_\infty , $$ while $\dim S_k(\Gamma)$ is the number of holomorphic sections of the $(k/2)$-th power of the canonical bundle of $X_\Gamma$ vanishing at the cusps, computed from the genus, the elliptic points and the cusps by the Riemann–Roch theorem for the orbifold $X_\Gamma$. For $\Gamma = \Gamma_0(N)$ the computation reduces to the data of the previous section, and in particular $\dim S_2(\Gamma_0(N)) = g(\Gamma_0(N))$, the weight-two cusp forms being the holomorphic differentials of the first kind.

The Ring of Level One Forms

The Structure Theorem

Theorem. The graded ring of modular forms of level one is $$ \bigoplus_{k}M_k(\Gamma(1)) = \mathbb{C}[E_4,E_6], $$ the two generators being algebraically independent, of weights $4$ and $6$; the ideal of cusp forms is $\Delta\cdot\mathbb{C}[E_4,E_6]$.

Proof. By the dimension formula, $M_k$ is spanned by monomials $E_4^aE_6^b$ with $4a+6b = k$: the number of such monomials equals the value of $\dim M_k$ in each weight, so the monomials span by the dimension count; the algebraic independence is the statement that the map $\mathbb{C}[X,Y]\to\bigoplus M_k$, $X\mapsto E_4$, $Y\mapsto E_6$, is injective, which follows by writing a relation of minimal weight in the form $E_4^3 = E_6^2+\Delta$ and comparing the orders of vanishing at the elliptic points, or, more directly, by using the valence formula to show that a form of weight $k$ vanishing to the highest possible order at $\rho$ and $i$ is a multiple of the appropriate monomial. The cusp forms are the multiples of $\Delta$ because $\Delta$ is a cusp form of minimal weight $12$ and multiplication by $\Delta$ maps $M_{k-12}$ isomorphically onto $S_k$ for every $k\geq12$, by the valence formula.

Corollary. $\dim M_k$ is the number of solutions of $4a+6b = k$ in nonnegative integers, whence the formula of the previous section; and every form of level one with rational Fourier coefficients is a polynomial with rational coefficients in $E_4$ and $E_6$, since $E_4,E_6,\Delta$ have rational coefficients and the map is defined over $\mathbb{Q}$.

The Discriminant and the Ramanujan Function

Theorem (Ramanujan). The coefficients $\tau(n)$ of $\Delta$ satisfy

(a) $\tau(mn) = \tau(m)\tau(n)$ whenever $\gcd(m,n)=1$;

(b) $\tau(p^{r+1}) = \tau(p)\tau(p^r) - p^{11}\tau(p^{r-1})$ for every prime $p$ and $r\geq1$;

(c) $\lvert\tau(p)\rvert\leq 2p^{11/2}$ (Deligne's theorem, formerly the Ramanujan conjecture);

(d) $\tau(n)\equiv\sigma_{11}(n)\pmod{691}$.

Proof sketch. Parts (a) and (b) are the multiplicativity of the Hecke operators, of which $\Delta$ is an eigenform: the space $S_{12}(\Gamma(1))$ is one-dimensional, so $\Delta$ is an eigenvector for every Hecke operator and its eigenvalue is the coefficient of $q$, namely $\tau(n)$; the relations follow from the composition law of the $T_n$. Part (c) is the bound of Deligne, obtained from the Weil conjectures: the eigenvalues of the Frobenius acting on the relevant cohomology of a surface have absolute value $p$ in the appropriate normalisation. Part (d) is the congruence between the Eisenstein series and the cusp form, proved by comparing $E_{12}$ with $E_4^3$ and using the Clausen–von Staudt congruence for the Bernoulli numbers.

Remark (the arithmetic of $\tau$). The multiplicativity and the recursion determine $\tau(n)$ from the values at the primes; the values have been checked against the expansion of $E_4^3-E_6^2$ by exact computation, which gives $\tau(4) = \tau(2)^2-2^{11}$, $\tau(6) = \tau(2)\tau(3)$, $\tau(8) = \tau(2)\tau(4)-2^{11}\tau(2)$ and $\tau(9) = \tau(3)^2-3^{11}$, all four identities holding for the coefficients displayed above. The function $\tau$ is the prototypical multiplicative arithmetic function arising from a modular form.

The $j$-Invariant

Definition. The modular invariant is $$ j(z) = \frac{E_4(z)^3}{\Delta(z)} = q^{-1}+744+196884q+21493760q^2+\dots $$ Theorem. $j$ is a modular function of weight $0$ for $\Gamma(1)$, holomorphic on $\mathbb{H}$ with a simple pole at the cusp and no other pole, invariant under $\Gamma(1)$, and it induces a bijection $X_{\Gamma(1)}\to\mathbb{P}^1(\mathbb{C})$; consequently the field of modular functions of level one is $\mathbb{C}(j)$, and every modular function is a rational function of $j$. Its values at the elliptic points are $j(i) = 1728$ and $j(\rho) = 0$.

Proof sketch. $E_4^3$ and $\Delta$ are forms of weights $12$ with $\Delta$ vanishing to order $1$ at the cusp and $E_4$ not vanishing there, so the quotient is a modular function with a simple pole and no other pole; the valence formula applied to $j-c$ shows that it has exactly one zero, so the map to $\mathbb{P}^1$ is bijective, and the bijectivity gives $\mathbb{C}(j)$ for the field of modular functions. The values at $i$ and $\rho$ follow from the valence formula applied to $E_4$ and to $E_6$, or directly from the definitions.

Remark (the arithmetic of $j$). The Fourier coefficients of $j$ are integers; the value $j(\tau)$ at a quadratic irrationality $\tau$ of discriminant $d$ is an algebraic integer, and $\mathbb{Q}(j(\tau))$ is the Hilbert class field of the imaginary quadratic field of discriminant $d$, of degree the class number — the theory of complex multiplication, whose arithmetic side belongs to Galois Theory and Algebraic Number Theory. The $q$-expansion displayed has been checked by exact power series computation from $j = E_4^3/\Delta$.

Hecke Operators and $L$-Functions

The Hecke Operators

Definition. For $f\in M_k(\Gamma)$ with $\Gamma\supseteq\Gamma(N)$ and $n\geq1$, the Hecke operator is $$ (T_nf)(z) = n^{k-1}\sum_{d\mid n}d^{-k}\sum_{b=0}^{d-1}f\Bigl(\frac{nz+b}{d}\Bigr), $$ and on the Fourier expansion it acts by $$ T_n\Bigl(\sum_{m\geq0}a_mq^m\Bigr) = \sum_{m\geq0}\Bigl(\sum_{d\mid\gcd(m,n)}d^{k-1}a_{mn/d^2}\Bigr)q^m . $$ Theorem. $T_n$ maps $M_k(\Gamma)$ to itself and $S_k(\Gamma)$ to $S_k(\Gamma)$; the operators commute with one another and with the action of the diamond operators $\langle d\rangle$ of $\Gamma_0(N)/\Gamma_1(N)$; the $T_n$ with $\gcd(n,N)=1$ are normal with respect to the Petersson inner product and pairwise commute, and they are self-adjoint on the space of forms of trivial nebentypus, so that $S_k(\Gamma_0(N))$ has a basis of simultaneous eigenvectors for them.

Proof sketch. The action of $T_n$ averages over the coset representatives of the double coset $\Gamma\begin{pmatrix}1&0\\0&n\end{pmatrix}\Gamma$, so it preserves the space and commutes with $\Gamma$; the coefficient formula is read off from the $q$-expansion; the commutation relations follow from the composition law of the double cosets, and the normality for $\gcd(n,N)=1$ from the invariance of the measure $y^{-2}dx\,dy$, which makes $T_n$ self-adjoint as soon as the diamond operator $\langle n\rangle$ acts trivially.

Definition. A Hecke eigenform is a nonzero $f\in M_k(\Gamma_0(N))$ which is an eigenvector for every $T_n$ with $\gcd(n,N)=1$ and for every diamond operator; it is normalised if $a_1 = 1$.

Theorem. A normalised Hecke eigenform with $a_1=1$ has $T_nf = a_nf$ for every $\gcd(n,N)=1$, and its Fourier coefficients satisfy $a_{mn} = a_ma_n$ for $\gcd(m,n)=1$ and $a_{p^{r+1}} = a_pa_{p^r}-\langle p\rangle p^{k-1}a_{p^{r-1}}$ for $p\nmid N$.

Proof. The eigenvalue of $T_n$ is computed by applying the operator to the normalised form and reading the coefficient of $q$; the relations are the coefficient form of the commutation relations.

Newforms and the Petersson Inner Product

Definition. The Petersson inner product on $S_k(\Gamma)$ is $$ \langle f,g\rangle = \int_{\Gamma\backslash\mathbb{H}}f(z)\overline{g(z)}\,y^k\,\frac{dx\,dy}{y^2}, $$ absolutely convergent because both forms vanish at the cusps.

Theorem (the Petersson and Eichler–Selberg formulae). The inner product is positive definite and Hermitian and $T_n$ is self-adjoint for $\gcd(n,N)=1$; the inner product of two Poincaré series is given by an explicit formula whose main term involves the Kloosterman sums, and consequently the eigenvalues of the $T_n$ on the cusp forms are real algebraic integers.

Definition. For $\Gamma_0(N)$ the old subspace of $S_k(\Gamma_0(N))$ is the span of the images of $S_k(\Gamma_0(M))$ for the proper divisors $M\mid N$, $Mnew subspace is its orthogonal complement with respect to the Petersson inner product, and a newform of level $N$ is a normalised form in the new subspace which is an eigenform for the full Hecke algebra, including the operators at the primes dividing $N$.

Theorem (Atkin–Lehner; multiplicity one). $S_k(\Gamma_0(N))$ is the direct sum of the old subspace and the new subspace; the new subspace is spanned by the newforms, which are eigenforms for the full Hecke algebra including the operators at the primes dividing $N$, and the space spanned by the Galois conjugates of a newform is irreducible. The newforms of level $N$ and weight $k$ with rational coefficients correspond bijectively to the normalised eigenforms with rational coefficients.

Proof sketch. The old subspace is the image of the sum of the two degeneracy maps from the levels $M$ and $N/M$, and the new subspace is its orthogonal complement; the Atkin–Lehner operators $W_Q$ for $Q\parallel N$ refine the decomposition, and multiplicity one is proved by comparing two eigenforms through the recursions at two primes.

The $L$-Function of a Modular Form

Definition. For a normalised Hecke eigenform $f = \sum a_nq^n$ of weight $k$ and level $N$ (that is, on $\Gamma_0(N)$, eigen at the primes not dividing $N$), the $L$-function is $$ L(f,s) = \sum_{n\geq1}\frac{a_n}{n^s} = \prod_{p\nmid N}\frac{1}{1-a_pp^{-s}+p^{k-1-2s}}\prod_{p\mid N}\frac{1}{1-a_pp^{-s}}, $$ the product converging for $\Re s > (k+1)/2$ by the bound on the coefficients.

Theorem (Hecke; the functional equation). With the completed function $$ \Lambda(f,s) = N^{s/2}(2\pi)^{-s}\Gamma(s)L(f,s), $$ the product converges absolutely for $\Re s>(k+1)/2$, $L(f,s)$ has an analytic continuation to an entire function, and $$ \Lambda(f,s) = \varepsilon\,\Lambda(f,k-s), \qquad \varepsilon = \pm1 , $$ the sign $\varepsilon$ being the root number; the critical line is $\Re s = k/2$ and the nontrivial zeros lie in the critical strip $(k-1)/2<\Re s<(k+1)/2$, since the Euler product is nonvanishing for $\Re s>(k+1)/2$ and the functional equation excludes the zeros for $\Re s<(k-1)/2$.

Proof sketch. The analytic continuation and the functional equation are proved by the Mellin transform of $f$: for a cusp form of weight $k$, $$ \Lambda(f,s) = \int_0^\infty f(iy)\,y^{s-1}\,dy , $$ so that the modular transformation law $f\bigl(i/(Ny)\bigr) = \varepsilon\,(i\sqrt{N}y)^kf(iy)$ — the $(-1/(Nz))$-invariance of the Fricke involution $W_N$ — turns the reflection $y\mapsto1/(Ny)$ into the reflection $s\mapsto k-s$. The Euler product is the eigenvalue relations of the Hecke operators.

Remark (the zeros and the generalised hypothesis). The zeros of $L(f,s)$ in the critical strip are conjectured to satisfy $\Re s = k/2$; this is the generalised Riemann hypothesis for a modular form, treated in The Riemann Hypothesis as one instance of the grand generalised hypothesis, where the completed function above is an $L$-function in the sense of L-Functions with conductor $N$, degree $2$ and spectral parameter $k/2$. The analytic properties established here — the Euler product, the continuation and the functional equation — are exactly the axioms of that article, so that the entire analytic theory of $L$-functions applies to the modular case.

Theta Series and the Half-Integral Weight Theory

Definition. For a positive definite quadratic form $Q$ on $\mathbb{Z}^n$ the theta series is $$ \theta_Q(z) = \sum_{x\in\mathbb{Z}^n}q^{Q(x)} = \sum_{m\geq0}r_Q(m)q^m, \qquad r_Q(m) = \#\{x : Q(x) = m\}, $$ and for the standard form $Q(x) = x^2$ the series is $\theta(z) = \sum_{n\in\mathbb{Z}}q^{n^2}$.

Theorem. If $Q$ is integral of even rank $n$ and its associated bilinear form is even unimodular, then $\theta_Q$ is a modular form of weight $n/2$ for $\Gamma(1)$. In particular $\theta\in M_{1/2}(\Gamma_0(4),\chi)$ with the multiplier $\chi$ of the theta function, and $\theta^4\in M_2(\Gamma_0(4))$ is the Eisenstein series of weight $2$ with $$ \theta(z)^4 = \sum_{n\geq0}r_4(n)q^n, \qquad r_4(n) = 8\sum_{\substack{d\mid n\\ 4\nmid d}}d . $$ Proof sketch. The transformation law of the theta series is the Poisson summation formula of Zeta Functions, which gives $\theta(-1/z) = (z/i)^{1/2}\theta(z)$ with the principal branch of the square root; that law is the automorphy with the multiplier $\chi$. The fourth power kills the multiplier, and $\dim M_2(\Gamma_0(4)) = 2$ forces it to be an Eisenstein series; comparing the first coefficients gives the four-square formula.

Example. The formula gives $r_4(1) = 8$, $r_4(2) = 24$, $r_4(3) = 32$, $r_4(4) = 24$: for $n=2$ the representations are the permutations of $(\pm1,\pm1,0,0)$, of which there are $24$, and for $n=4$ the representations are the eight of $(\pm2,0,0,0)$ and the sixteen of $(\pm1,\pm1,\pm1,\pm1)$, again $24$. The agreement with the formula is the classical theorem of Jacobi on the four-square numbers.

Remark (weight one-half and the multiplier). The half-integral weight theory requires a multiplier system, since $j(\gamma,z)^{1/2}$ is not a function of $\gamma$ alone; the resulting spaces $M_{k/2}(\Gamma,\chi)$ carry a Hecke theory of their own, due to Shimura, whose operators raise the weight by $2$, and the Waldspurger formula relates the Fourier coefficients of a half-integral weight form to the central values of the $L$-functions of the corresponding integral weight forms.

Modular Forms and Arithmetic

Modular Curves and the Eichler–Shimura Isomorphism

Theorem (Eichler–Shimura). For a congruence subgroup $\Gamma$ and $k\geq2$ there is an isomorphism, compatible with the Hecke action, $$ S_k(\Gamma)\ \cong\ H^1(X_\Gamma,\ \operatorname{Sym}^{k-2}\mathcal{V}), $$ where $\mathcal{V}$ is the local system attached to the standard representation of $\Gamma$ and the cohomology is the de Rham cohomology of the modular curve with coefficients in the corresponding local system; for $k=2$ this is the statement that $S_2(\Gamma)$ is the space of holomorphic differentials of the first kind on $X_\Gamma$, so that $\dim S_2(\Gamma) = g(\Gamma)$, and $H^1(X_\Gamma,\mathbb{C})$ is the direct sum of $S_2$ and its complex conjugate.

Proof sketch. A cusp form $f$ of weight $k$ gives the holomorphic section $f(z)(dz)^{k/2}$ of the $(k/2)$-th power of the canonical bundle, and $f\mapsto$ the cohomology class of the corresponding invariant form is $\Gamma$-equivariant and injective; the surjectivity is the Hodge decomposition, and the dimension comparison is the Riemann–Roch computation of $S_k$.

Corollary (the arithmetic of the differentials). The space $S_2(\Gamma_0(N))$ has dimension the genus of $X_0(N)$, and the Hecke operators act on it as correspondences on the Jacobian $J_0(N)$ of the modular curve; the abelian variety generated by a newform is a quotient of $J_0(N)$ defined over $\mathbb{Q}$, whose dimension is the degree of the field of coefficients of the newform. This is the mechanism that attaches an abelian variety to a modular form on the arithmetic side.

Galois Representations and Modularity

Theorem (Deligne). Let $f$ be a normalised newform of weight $k$ and level $N$ with coefficients in a number field $E$. For every prime $\lambda$ of $E$ there is a continuous representation $$ \rho_{f,\lambda}:\operatorname{Gal}(\overline{\mathbb{Q}}/\mathbb{Q})\to GL_2(E_\lambda), $$ unramified outside $N$ and $\lambda$, such that the characteristic polynomial of the Frobenius at an unramified prime $p$ is $X^2-a_pX+\langle p\rangle p^{k-1}$, and the eigenvalues of the Frobenius have absolute value $p^{(k-1)/2}$ at every unramified prime. The representation is irreducible and its determinant is the product of the cyclotomic character with the $(k-1)$-st power of $\langle\cdot\rangle$.

Proof sketch. The newform contributes a two-dimensional piece of the cohomology of the modular curve with coefficients in a local system, carrying an action of the Galois group defined over a number field; the Frobenius traces are computed by the Eichler–Shimura congruence relation, which identifies $T_p$ with the Frobenius on the reduction modulo $p$, and the absolute values of the eigenvalues follow from the purity of the cohomology, so that the Ramanujan conjecture becomes a theorem about a variety.

Theorem (modularity). Every elliptic curve $E$ over $\mathbb{Q}$ is modular: there is a surjective morphism $X_0(N)\to E$ defined over $\mathbb{Q}$, with $N$ the conductor of $E$, and the $L$-function of $E$ equals the $L$-function of a normalised newform of weight $2$ and level $N$. Consequently $L(E,s)$ has an analytic continuation to an entire function and satisfies the functional equation with the sign equal to the parity of the order of vanishing at $s=1$.

Proof sketch. The theorem was proved by Wiles and Taylor for semistable $E$ and by Breuil, Conrad, Diamond and Taylor in general: a deformation ring of the residual representation of the Galois group is identified with a Hecke ring by a numerical criterion together with the automorphic theory of $GL_2$, and the $L$-functions agree because their Frobenius traces agree, on the modular side by definition and on the arithmetic side by the Lefschetz trace formula. The analytic continuation on the modular side is the theorem of the previous subsection.

Remark (the seam with the automorphic theory). Every object of this article has an adelic translation. The upper half-plane is $GL_2(\mathbb{R})/O(2)$, a modular form of weight $k$ corresponds to a function on $GL_2(\mathbb{A})/GL_2(\mathbb{Q})$ of central character $\lvert\cdot\rvert^{k}$ transforming as a highest weight vector at the Archimedean place, the Hecke operators become the convolution operators at the finite places, the Fourier coefficients become the Whittaker coefficients, and the $L$-function is a zeta integral. In that form the definitions make sense for $GL_n$ and the theory becomes the theory of automorphic representations, whose development — the adeles of Adeles and Ideles, the Fourier analysis of Adelic Analysis, the representation theory of Part II — belongs to Automorphic Forms, in the neighbouring category of this Part and written in parallel. The classical theory of this article is the case $GL_2$ and its arithmetic input is over $\mathbb{Q}$.

Summary

A modular form of weight $k$ for a congruence subgroup $\Gamma$ of $SL_2(\mathbb{Z})$ is a holomorphic function on the upper half-plane satisfying $f((az+b)/(cz+d)) = (cz+d)^kf(z)$ and holomorphic at the cusps; it is a cusp form if it vanishes at every cusp. The group $\bar\Gamma(1)$ is the free product of $\mathbb{Z}/2$ and $\mathbb{Z}/3$ generated by $S:z\mapsto-1/z$ and $T:z\mapsto z+1$, its fundamental domain has one cusp and the elliptic points $i$ and $\rho$, and the congruence subgroups $\Gamma(N)\subseteq\Gamma_1(N)\subseteq\Gamma_0(N)$ have indices $N^3\prod(1-p^{-2})$, $N\varphi(N)\prod(1+p^{-1})$ and $N\prod(1+p^{-1})$ in $\Gamma(1)$. The valence formula $\sum_P\lvert\bar\Gamma_P\rvert^{-1}\operatorname{ord}_Pf = k[\bar\Gamma(1):\bar\Gamma]/12$ makes $M_k$ and $S_k$ finite-dimensional and gives $\dim M_k = \lfloor k/12\rfloor+1$ at level one for $k\not\equiv2\pmod{12}$ and $\lfloor k/12\rfloor$ for $k\equiv2\pmod{12}$; the level-one ring is $\mathbb{C}[E_4,E_6]$, its cusp ideal is $\Delta\mathbb{C}[E_4,E_6]$, and the discriminant $\Delta = (E_4^3-E_6^2)/1728 = q\prod(1-q^n)^{24}$ has coefficients $\tau(n)$ that are multiplicative, satisfy $\tau(p^{r+1}) = \tau(p)\tau(p^r)-p^{11}\tau(p^{r-1})$, obey $\lvert\tau(p)\rvert\leq2p^{11/2}$ by Deligne, and satisfy $\tau(n)\equiv\sigma_{11}(n)\pmod{691}$. The $j$-invariant $j = E_4^3/\Delta = q^{-1}+744+196884q+\dots$ generates the field of modular functions of level one.

The Hecke operators $T_n$ act on the spaces, commute, and are normal for $\gcd(n,N)=1$ — self-adjoint on the forms of trivial nebentypus — with respect to the Petersson inner product; a normalised eigenform has $T_nf = a_nf$ with multiplicative eigenvalues and a recursion at the prime powers, and its $L$-function has a degree-two Euler product and the functional equation $\Lambda(f,s) = \varepsilon\Lambda(f,k-s)$ for $\Lambda(f,s) = N^{s/2}(2\pi)^{-s}\Gamma(s)L(f,s)$, proved by the Mellin transform and the invariance of the form under $z\mapsto-1/(Nz)$. The new subspace is the orthogonal complement of the oldforms and is spanned by the newforms, by Atkin–Lehner and multiplicity one; theta series of even unimodular lattices are modular forms of half-integral weight and Jacobi's formula $r_4(n) = 8\sum_{d\mid n,4\nmid d}d$ is the weight-two case. The Eichler–Shimura isomorphism identifies $S_k(\Gamma)$ with a piece of the cohomology of the modular curve $X_\Gamma$, so that $\dim S_2(\Gamma) = g(\Gamma)$ and the Hecke action is a correspondence on the Jacobian; Deligne attaches to each newform a two-dimensional Galois representation with Frobenius characteristic polynomial $X^2-a_pX+\langle p\rangle p^{k-1}$ and eigenvalues of absolute value $p^{(k-1)/2}$; and the modularity theorem identifies every elliptic curve over $\mathbb{Q}$ with a weight-two newform of its conductor, giving the analytic continuation of its $L$-function.

Summary of Notation

Symbol Meaning
$\mathbb{H}$, $q$ Upper half-plane; $q = e^{2\pi iz}$
$\Gamma(1)$, $\bar\Gamma(1)$ $SL_2(\mathbb{Z})$ and its image $PSL_2(\mathbb{Z})$
$S$, $T$ Generators $z\mapsto-1/z$ and $z\mapsto z+1$
$\mathcal{F}$, $i$, $\rho$ Fundamental domain; elliptic points $i$ and $e^{2\pi i/3}$
$\Gamma(N)$, $\Gamma_0(N)$, $\Gamma_1(N)$ Principal, Hecke and unipotent congruence subgroups of level $N$
$\varphi(N)$ Euler function; order of $(\mathbb{Z}/N\mathbb{Z})^\times$
$X_\Gamma$, $Y_\Gamma$, $g$ Modular curve, its affine part, its genus
$\nu_\infty,\nu_2,\nu_3$ Numbers of cusps and of elliptic points of order $2$ and $3$
$j(\gamma,z)$, $\mid_k$ Automorphy factor $cz+d$ and weight-$k$ action
$M_k(\Gamma)$, $S_k(\Gamma)$ Modular forms, cusp forms of weight $k$
$E_k$, $G_k$ Normalised Eisenstein series, the unnormalised sum $=2\zeta(k)E_k$
$\sigma_r(n)$, $B_k$ Divisor sum, Bernoulli numbers of Zeta Functions
$\Delta$, $\eta$, $\tau(n)$ Discriminant, Dedekind eta, Ramanujan function
$j$ Modular invariant $E_4^3/\Delta$
$T_n$, $\langle d\rangle$ Hecke operators, diamond operators
$\langle f,g\rangle$ Petersson inner product $\int f\bar g y^k dx\,dy/y^2$
$L(f,s)$, $\Lambda(f,s)$, $\varepsilon$ $L$-function, completed $L$-function, root number
$k/2$, $k$ Critical line and right edge of the critical strip of $L(f,s)$
$\theta$, $r_Q(m)$, $r_4(n)$ Theta series, representation numbers
$\rho_{f,\lambda}$ $\lambda$-adic Galois representation of a newform
$N$, $a_p$, $\varepsilon$ Level/conductor, Hecke eigenvalue, root number

Further Reading

  • Jean-Pierre Serre, A Course in Arithmetic (Springer, 1973), for the level-one theory, the structure theorem and the four-square theorem.
  • Goro Shimura, Introduction to the Arithmetic Theory of Automorphic Functions (Princeton University Press, 1971), for the congruence subgroups, the dimension formulas, the Hecke operators and the newform theory.
  • Robert A. Rankin, Modular Forms and Functions (Cambridge University Press, 1977), for the analytic theory, the Petersson formula and the growth of the coefficients.
  • Neal Koblitz, Introduction to Elliptic Curves and Modular Forms (2nd ed., Springer, 1993), for the modular parametrisation of an elliptic curve and the arithmetic applications.
  • Fred Diamond and Jerry Shurman, A First Course in Modular Forms (Springer, 2005), for a modern treatment of the Hecke theory, the new subspace and the modularity theorem.
  • Martin Eichler and Don Zagier, The Theory of Jacobi Forms (Birkhäuser, 1985), for the half-integral weight theory and the theta correspondence in its classical form.
  • Andrew Wiles, Modular elliptic curves and Fermat's last theorem (Annals of Mathematics 141, 1995), and Richard Taylor and Andrew Wiles, Ring-theoretic properties of certain Hecke algebras (Annals of Mathematics 141, 1995), for the modularity theorem in the semistable case.
  • Christophe Breuil, Brian Conrad, Fred Diamond and Richard Taylor, On the modularity of elliptic curves over $\mathbb{Q}$ (Journal of the American Mathematical Society 14, 2001), for the general modularity theorem.
  • Pierre Deligne, Formes modulaires et représentations $\ell$-adiques (Séminaire Bourbaki 355, 1969), and La conjecture de Weil I (Publications Mathématiques de l'IHÉS 43, 1974), for the Galois representations and the bound on the coefficients.
  • Jean-Pierre Serre and André Weil, on the modularity of the Ramanujan function and the congruence $\tau(n)\equiv\sigma_{11}(n)\pmod{691}$, in Séminaire Delange–Pisot–Poitou (1968), for the arithmetic of $\Delta$.
  • Jun-ichi Igusa, Theta Functions (Springer, 1972), for the theta series of lattices and the transformation laws of half-integral weight.