Eigenvalues and Diagonalisation
Introduction
A linear operator is at its simplest when it preserves a line. The scalar by which it stretches that line is an eigenvalue, the line an eigenvector, and an operator that has enough eigenvectors to give a basis of the whole space is diagonalisable: in that basis it is nothing but a list of scalars. Everything in this article is organised around that idea and around its failure. The failure is quantitative — an eigenspace can be too small — and it is measured by comparing two multiplicities attached to each eigenvalue, one read from the characteristic polynomial and one read from the operator's kernel.
Throughout, $F$ is a field, $V$ is a finite-dimensional $F$-vector space of dimension $n$, and $T:V \to V$ is $F$-linear. The base is a field and not a ring, because the characteristic polynomial is a polynomial with coefficients in the scalars and because its roots must be compared with elements of the scalars; the module case over a principal ideal domain is treated in the applications article of this category on modules over $k[x]$ and the Jordan form. Matrices of linear maps, change of basis, trace and determinant were fixed in the preceding article of the category, and the notation $c_T(x)=\det(xI-A)$ is used throughout.
The article proceeds from eigen-elements to the characteristic polynomial, then to the diagonalisation criterion, then to the finer structure that replaces a diagonal form when diagonalisation is impossible: invariant subspaces, triangular forms and the minimal polynomial. Simultaneous diagonalisation of commuting operators closes the theory with the statement that the spectral decompositions are compatible.
Eigenvalues and Eigenvectors
Definitions
Definition. Let $T:V \to V$ be linear. A scalar $\lambda \in F$ is an eigenvalue of $T$ if there exists a nonzero $v \in V$ with
$$ T(v)=\lambda v . $$
Such a $v$ is an eigenvector for $\lambda$, and the eigenspace is
$$ V_\lambda=\ker(T-\lambda \operatorname{id}_V)=\{v \in V : T(v)=\lambda v\}, $$
a subspace of $V$ containing the zero vector, which is not itself an eigenvector. The nonzero elements of $V_\lambda$ are the eigenvectors for $\lambda$.
The eigenvalue equation is equivalent to $\ker(T-\lambda\,\mathrm{id}) \neq 0$, that is, to $T-\lambda \operatorname{id}$ being non-invertible. If $\lambda$ is an eigenvalue, the set of its eigenvectors together with $0$ is a subspace by linearity of $T-\lambda\operatorname{id}$; conversely, any nonzero element of $\ker(T-\lambda\operatorname{id})$ is an eigenvector. The eigenvalues of a matrix $A \in M_n(F)$ are those of the linear map $x \mapsto Ax$ on $F^n$.
Example. The matrix $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}$ has eigenvalues $3$ and $1$: the vector $(1,1)$ satisfies $A(1,1)^{\mathsf{T}}=(3,3)^{\mathsf{T}}=3(1,1)^{\mathsf{T}}$, and $(1,-1)$ is fixed, $A(1,-1)^{\mathsf{T}}=(1,-1)^{\mathsf{T}}$. Hence $V_3=\operatorname{span}\{(1,1)\}$ and $V_1=\operatorname{span}\{(1,-1)\}$, and $F^2=V_3\oplus V_1$.
Eigenvectors and Lines
An eigenvector determines a line, and the eigenvalue is the factor by which $T$ stretches it. Two eigenvectors for the same eigenvalue need not be multiples of one another, and the set of all of them for a fixed $\lambda$ is the eigenspace.
Proposition. Eigenvectors for distinct eigenvalues are linearly independent; more generally, if $\lambda_1,\dots,\lambda_k$ are distinct then
$$ V_{\lambda_1}+\cdots+V_{\lambda_k}=V_{\lambda_1}\oplus\cdots\oplus V_{\lambda_k}, $$
so the sum of the eigenspaces is direct.
Proof. Suppose $v_1+\cdots+v_k=0$ with $v_i \in V_{\lambda_i}$, and choose such a relation with the fewest nonzero terms; reindex so that $v_1,\dots,v_m \neq 0$. Applying $T$ gives $\lambda_1 v_1+\cdots+\lambda_m v_m=0$, and subtracting $\lambda_1$ times the original relation gives $(\lambda_2-\lambda_1)v_2+\cdots+(\lambda_m-\lambda_1)v_m=0$, a shorter relation with nonzero coefficients; this contradicts minimality unless $m=1$, in which case $v_1=0$, again a contradiction. Hence no nontrivial relation exists and the sum is direct.
A crucial point is that eigenvalues depend on the field. The real matrix $J=\begin{pmatrix}0&-1\\1&0\end{pmatrix}$ has no real eigenvalue, because $J(x,y)=(-y,x)$ rotates the plane and fixes no line; over $\mathbb{C}$ its eigenvalues are $i$ and $-i$, with eigenvectors $(1,-i)$ and $(1,i)$. The characteristic polynomial below makes this precise: it has real coefficients and its real roots are the real eigenvalues.
The Characteristic Polynomial
Definition and Basic Properties
Definition. Let $A \in M_n(F)$. The characteristic polynomial of $A$ is
$$ c_A(x)=\det(xI_n-A) \in F[x], $$
and the characteristic polynomial of a linear operator $T$ is that of its matrix in any basis.
The definition is legitimate for $T$ because similar matrices have equal determinants: if $B=P^{-1}AP$ then $xI-B=P^{-1}(xI-A)P$, so $c_B(x)=\det(P^{-1})c_A(x)\det(P)=c_A(x)$. Thus $c_T$ is well defined.
Theorem. $c_A(x)=x^n-(\operatorname{tr}A)x^{n-1}+\cdots+(-1)^n\det A$, a monic polynomial of degree $n$; and $\lambda \in F$ is an eigenvalue of $A$ if and only if $c_A(\lambda)=0$.
Proof. The determinant is a sum over permutations, and the only permutations contributing to the coefficient of $x^{n-k}$ involve at least $n-k$ diagonal factors; expanding identifies the leading coefficient as $1$, the coefficient of $x^{n-1}$ as $-\sum_i a_{ii}$, and the constant term as $\det(-A)=(-1)^n\det A$. For the second statement, $\lambda$ is an eigenvalue exactly when $\ker(\lambda I-A) \neq 0$, which over a field is equivalent to $\det(\lambda I-A)=0$.
Definition. The algebraic multiplicity of an eigenvalue $\lambda$ is its multiplicity as a root of $c_T$, and the geometric multiplicity is $\dim_F V_\lambda$.
Proposition. For every eigenvalue, the geometric multiplicity is at most the algebraic multiplicity.
Proof. Extend a basis $v_1,\dots,v_g$ of $V_\lambda$ to a basis of $V$. In that basis the matrix of $T$ is block upper triangular,
$$ \begin{pmatrix} \lambda I_g & * \\ 0 & B \end{pmatrix}, $$
so $c_T(x)=(x-\lambda)^g c_B(x)$ and $\lambda$ occurs to the power at least $g$.
The inequality can be strict, and the gap is exactly what prevents diagonalisation.
Eigenvalues of Special Matrices
Proposition. (i) The eigenvalues of an upper triangular matrix are its diagonal entries, and $c_A(x)=\prod_i(x-a_{ii})$. (ii) If $A$ is invertible with eigenvalues $\lambda_1,\dots,\lambda_n$ then $A^{-1}$ has eigenvalues $\lambda_i^{-1}$, and $A^k$ has eigenvalues $\lambda_i^k$. (iii) $A$ and $A^{\mathsf{T}}$ have the same characteristic polynomial.
Proof. (i) $xI-A$ is upper triangular with diagonal $x-a_{ii}$, so its determinant is the product. (ii) $Av=\lambda v$ gives $A^{-1}v=\lambda^{-1}v$ for $\lambda \neq 0$, and $A^kv=\lambda^k v$ by induction. (iii) $\det(xI-A^{\mathsf{T}})=\det((xI-A)^{\mathsf{T}})=\det(xI-A)$.
Diagonalisation
The Criterion
Definition. $T$ is diagonalisable if $V$ has a basis consisting of eigenvectors of $T$, equivalently if there is a basis in which the matrix of $T$ is diagonal. A matrix $A$ is diagonalisable if it is similar to a diagonal matrix.
Theorem. For a linear operator $T$ on an $n$-dimensional space the following are equivalent: (i) $T$ is diagonalisable; (ii) $V$ is the direct sum of the eigenspaces of $T$; (iii) the sum of the geometric multiplicities of the distinct eigenvalues equals $n$; (iv) the characteristic polynomial splits over $F$ and the geometric multiplicity equals the algebraic multiplicity for every eigenvalue.
Proof. (i) $\Leftrightarrow$ (ii): a basis of eigenvectors is exactly a union of bases of the eigenspaces, and the union is a basis precisely when the sum of the eigenspaces is direct and equal to $V$. (ii) $\Leftrightarrow$ (iii): the sum of the eigenspaces is always direct, and it equals $V$ exactly when its dimension, the sum of the geometric multiplicities, is $n$. (iii) $\Leftrightarrow$ (iv): the sum of the algebraic multiplicities is $n$ when $c_T$ splits, and the sum of the geometric multiplicities is at most that sum with equality exactly when each geometric multiplicity equals the corresponding algebraic multiplicity; so a sum of geometric multiplicities equal to $n$ forces both conditions, and they clearly imply the sum is $n$.
Corollary. If $T$ has $n$ distinct eigenvalues then $T$ is diagonalisable.
Proof. Each eigenspace is nonzero and the sum is direct, so the sum has dimension at least the number of eigenvalues, which is $n$; hence the sum is $V$.
The converse fails: the identity has a single eigenvalue and is diagonalisable. Distinctness of eigenvalues is sufficient, not necessary.
Computing the Decomposition
For a matrix $A$ the eigenspace for $\lambda$ is $\ker(\lambda I-A)$, so its dimension is computed by rank–nullity:
$$ \dim_F V_\lambda=n-\operatorname{rk}(\lambda I-A), $$
and a basis of the eigenspace is a basis of the null space of $\lambda I-A$, obtained by Gaussian elimination. Diagonalisation, when possible, is the change of basis $P$ whose columns are bases of the eigenspaces, giving $P^{-1}AP=\operatorname{diag}(\lambda_1,\dots,\lambda_n)$, with each eigenvalue repeated according to its geometric multiplicity.
Examples. The matrix $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}$ has $\dim V_3=\dim V_1=1$, the sum is $2=n$, and $A$ is diagonalisable with $P=\begin{pmatrix}1&1\\1&-1\end{pmatrix}$ and $P^{-1}AP=\operatorname{diag}(3,1)$. The matrix $N=\begin{pmatrix}1&1\\0&1\end{pmatrix}$ has $c_N(x)=(x-1)^2$, algebraic multiplicity $2$; but $N-I=\begin{pmatrix}0&1\\0&0\end{pmatrix}$ has rank $1$, so $\dim V_1=2-1=1$ and the geometric multiplicity is $1$. The criterion fails, and $N$ is not diagonalisable. The rotation $J=\begin{pmatrix}0&-1\\1&0\end{pmatrix}$ has $c_J(x)=x^2+1$, which does not split over $\mathbb{R}$; over $\mathbb{C}$ it splits with distinct roots $\pm i$, so $J$ is not diagonalisable over $\mathbb{R}$ but is over $\mathbb{C}$, with $P=\begin{pmatrix}1&1\\-i&i\end{pmatrix}$ (up to column order).
Invariant Subspaces
Definition and Algebraicity
Definition. A subspace $W \subseteq V$ is $T$-invariant if $T(W) \subseteq W$. Then $T$ restricts to a linear operator $T|_W$ on $W$, and $T$ induces an operator on the quotient $V/W$.
Invariance is preserved by the operators that are polynomials in $T$: if $W$ is $T$-invariant and $p \in F[x]$, then $W$ is $p(T)$-invariant, because $T^k(W) \subseteq W$ for every $k$. In particular the kernel and image of any polynomial in $T$ are $T$-invariant.
Every eigenspace is invariant, and so is the sum of the eigenspaces for a set of eigenvalues, and the generalised eigenspace
$$ V^{(\lambda)}=\{v \in V : (T-\lambda\operatorname{id})^k v=0 \text{ for some } k \ge 1\}=\bigcup_{k \ge 1}\ker(T-\lambda\operatorname{id})^k . $$
The generalised eigenspaces are the correct replacement for the eigenspaces when diagonalisation fails: they are the primary components of $V$ as a module over $F[x]$ under the action $x \cdot v=T(v)$, and they are treated in the applications article on the Jordan form.
Proposition. If $V=W_1 \oplus \cdots \oplus W_k$ with each $W_i$ invariant under $T$, then in a basis adapted to the decomposition the matrix of $T$ is block diagonal, with blocks the matrices of the restrictions $T|_{W_i}$; consequently $c_T=\prod_i c_{T|_{W_i}}$.
Proof. The image of $W_i$ lies in $W_i$, so the matrix has no components carrying $W_i$ into $W_j$ for $j \neq i$, and the determinant of a block diagonal matrix is the product of the determinants of its blocks.
Triangularisation
Definition. $T$ is triangularisable if $V$ has a basis in which the matrix of $T$ is upper triangular.
Theorem. $T$ is triangularisable if and only if $c_T$ splits over $F$ into linear factors.
Proof. If the matrix is upper triangular then $c_T$ is the product of the diagonal entries $x-a_{ii}$ by the proposition on triangular matrices, so it splits. Conversely, if $c_T$ splits, it has a root $\lambda_1 \in F$, hence an eigenvector $v_1$; the quotient $V/\langle v_1\rangle$ has dimension $n-1$ and the induced operator has characteristic polynomial $c_T(x)/(x-\lambda_1)$, which splits, so by induction the quotient has a basis of the required form, and lifting it gives a basis $v_1,\dots,v_n$ of $V$ in which the matrix is upper triangular with the roots on the diagonal.
Triangularisation is always available after extending scalars to a splitting field, and over $\mathbb{C}$ every matrix is triangularisable; diagonalisation is not, because the diagonal entries may repeat with a single eigenvector.
The Minimal Polynomial
Definition
Definition. The minimal polynomial $m_T \in F[x]$ is the monic polynomial of least degree with $m_T(T)=0$.
The set $\{p \in F[x] : p(T)=0\}$ is a nonzero ideal of $F[x]$: it is nonzero because the powers $I,T,T^2,\dots,T^{n^2}$ are linearly dependent in the $n^2$-dimensional space $\operatorname{End}_F(V)$, so some nonzero polynomial annihilates $T$; and it is an ideal because $p(T)=0$ implies $(pq)(T)=p(T)q(T)=0$. Since $F[x]$ is a principal ideal domain, the ideal is generated by a unique monic polynomial, which is $m_T$; hence $p(T)=0$ if and only if $m_T \mid p$.
Theorem (Cayley–Hamilton). $c_T(T)=0$, so $m_T$ divides $c_T$; in particular $\deg m_T \le n$.
This is the standard Cayley–Hamilton theorem, proved by either of the usual routes: reduction to a triangular matrix after extending scalars, or the adjugate identity for $xI-A$ over the ring $F[x]$. It is quoted here as standard.
Minimal Polynomial and Diagonalisability
Theorem. $T$ is diagonalisable over $F$ if and only if $m_T$ is a product of distinct linear factors in $F[x]$.
Proof. If $T$ is diagonal, let $\mu_1,\dots,\mu_r$ be its distinct eigenvalues; then $\prod_{j=1}^{r}(x-\mu_j)$ annihilates $T$, since it vanishes at every diagonal entry, so $m_T$ divides it, hence $m_T$ is a product of distinct linear factors. Conversely, suppose $m_T=\prod_{i=1}^{k}(x-\lambda_i)$ with the $\lambda_i$ distinct. The polynomials $q_i=\prod_{j \neq i}(x-\lambda_j)$ have no common factor, so there exist $a_i \in F[x]$ with $\sum_i a_i q_i=1$ by Bézout. For $v \in V$, write $v=\sum_i a_i(T)q_i(T)v$, and $q_i(T)v \in \ker(T-\lambda_i\operatorname{id})$ because $(x-\lambda_i)q_i=m_T$ annihilates $T$. Hence $V$ is the sum of the eigenspaces, and it is a direct sum because eigenspaces for distinct eigenvalues are independent.
Theorem. Let $K$ be a splitting field for $m_T$. The roots of $m_T$ in $K$ are exactly the eigenvalues of $T$ in $K$, and they coincide with the roots of $c_T$ in $K$.
Proof. If $\lambda$ is an eigenvalue with eigenvector $v$, then $0=m_T(T)v=m_T(\lambda)v$ and $v \neq 0$, so $m_T(\lambda)=0$. Conversely, if $\mu$ is a root of $m_T$, write $m_T=(x-\mu)^k g$ with $g(\mu) \neq 0$ and $k \ge 1$. If $\mu$ were not an eigenvalue then $T-\mu\operatorname{id}$ would be invertible, hence so would $(T-\mu\operatorname{id})^k$, and multiplying $0=m_T(T)=(T-\mu\operatorname{id})^kg(T)$ by the inverse would give $g(T)=0$ with $\deg g<\deg m_T$, contradicting minimality. Hence $\mu$ is an eigenvalue. Since $m_T \mid c_T$ by Cayley–Hamilton, every root of $m_T$ is a root of $c_T$; and every root of $c_T$ is an eigenvalue, hence a root of $m_T$ by the first part. So the two polynomials have the same set of roots.
The minimal polynomial is a finer invariant than the characteristic polynomial and detects diagonalisability exactly. For the matrix $N=\begin{pmatrix}1&1\\0&1\end{pmatrix}$ one has $c_N=(x-1)^2$, and also $m_N=(x-1)^2$, since $N-I \neq 0$ but $(N-I)^2=0$; the repeated root occurs in both, so $N$ is not diagonalisable. For $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}$ one has $c_A=(x-3)(x-1)=m_A$.
Simultaneous Diagonalisation
Theorem. Let $T_1,\dots,T_k$ be diagonalisable operators on $V$ that commute pairwise, and suppose each $T_i$ splits over $F$. Then there is a basis of $V$ in which every $T_i$ is diagonal.
Proof. For $k=1$ this is the definition. Inductively, let $T_1$ be diagonalisable and decompose $V=\bigoplus_\lambda V_\lambda$ into its eigenspaces. Each $V_\lambda$ is invariant under every $T_i$, because if $v \in V_\lambda$ then $T_1T_i v=T_iT_1v=\lambda T_iv$, so $T_iv \in V_\lambda$. The restrictions $T_2|_{V_\lambda},\dots,T_k|_{V_\lambda}$ commute pairwise and remain diagonalisable with splitting characteristic polynomial on the invariant subspace $V_\lambda$; by induction they are simultaneously diagonal on $V_\lambda$. Combining bases over the finitely many $\lambda$ gives a basis of $V$ diagonal for all $T_i$.
Corollary. Over an algebraically closed field the hypothesis that each $T_i$ splits is automatic, so every finite commuting family of diagonalisable operators is simultaneously diagonalisable. In particular a single diagonalisable operator lies in a maximal commutative subalgebra consisting entirely of simultaneously diagonalisable operators, namely the algebra of all operators that are diagonal in its eigenbasis; the algebra of polynomials in the operator and the scalars is the subalgebra of that one generated by the operator.
Powers of an Operator
Polynomials in a Diagonalisable Operator
Let $T=SDS^{-1}$ with $D=\operatorname{diag}(\lambda_1,\dots,\lambda_n)$. Then $T^k=SD^kS^{-1}=S\operatorname{diag}(\lambda_1^k,\dots,\lambda_n^k)S^{-1}$ for every $k \ge 0$, and more generally
$$ p(T)=S\,p(D)\,S^{-1}=S\operatorname{diag}\bigl(p(\lambda_1),\dots,p(\lambda_n)\bigr)S^{-1} $$
for every polynomial $p \in F[x]$. So on a diagonalisable operator every polynomial is computed eigenvalue by eigenvalue, and the eigenbasis is the basis in which the computation is trivial. This is the practical content of diagonalisation: it turns the algebra of polynomials in a single operator into the algebra of diagonal matrices, evaluated pointwise on the spectrum.
Example (a linear recurrence). Let $M=\begin{pmatrix}1&1\\1&0\end{pmatrix}$ and let $F_0=0$, $F_1=1$, $F_{k}=F_{k-1}+F_{k-2}$. Induction gives
$$ M^k=\begin{pmatrix}F_{k+1}&F_k\\ F_k&F_{k-1}\end{pmatrix}. $$
The characteristic polynomial is $c_M(x)=x^2-x-1$, with distinct real roots $\varphi=\tfrac12(1+\sqrt5)$ and $\psi=\tfrac12(1-\sqrt5)$, so $M$ is diagonalisable over $\mathbb{R}$ and $F_k=\dfrac{\varphi^k-\psi^k}{\varphi-\psi}$. Taking determinants in the displayed identity gives Cassini's identity $F_{k+1}F_{k-1}-F_k^2=(-1)^k$, because $\det M^k=(\det M)^k=(-1)^k$.
Further Examples
A cycle over two fields. Let $C=\begin{pmatrix}0&1&0\\0&0&1\\1&0&0\end{pmatrix}$, the matrix of the $3$-cycle. Since $C^3=I$ and $C \neq I$, the minimal polynomial divides $x^3-1$ and is not $x-1$; in fact $x^3-1=(x-1)(x^2+x+1)$ has distinct roots over $\mathbb{C}$, namely $1,\omega,\omega^2$ with $\omega=e^{2\pi i/3}$. Over $\mathbb{C}$ the matrix has three distinct eigenvalues and is diagonalisable; over $\mathbb{R}$ its only eigenvalue is $1$, with eigenspace spanned by $(1,1,1)$, so the sum of the geometric multiplicities is $1<3$ and $C$ is not diagonalisable over $\mathbb{R}$. The same matrix is a different object over the two fields.
A Jordan block. Let $N=\begin{pmatrix}1&1&0\\0&1&1\\0&0&1\end{pmatrix}$. Then $c_N(x)=(x-1)^3$ and $N-I$ has rank $2$, so $V_1$ has dimension $3-2=1$ and the geometric multiplicity is strictly less than the algebraic multiplicity. The minimal polynomial is $(x-1)^3$, the generalised eigenspace $V^{(1)}$ is all of $F^3$, and $N$ is not diagonalisable. This is the simplest matrix for which the generalised eigenspace is needed.
Orthogonal Diagonalisation Deferred to Part II
Requiring the diagonalising basis to be orthogonal is a demand that uses more than the field and the linear structure, so it is not a purely algebraic question; that refinement of diagonalisation belongs to Part II, where it is treated after the geometric structure it needs has been introduced.
The Spectrum and the Spectral Radius
Definition. The spectrum of $T$ is the set of its eigenvalues in $F$. Over a splitting field $K$ it is the multiset of roots of $c_T$ in $K$, each counted with its algebraic multiplicity, and it is written $\sigma(T)$. When $F \subseteq \mathbb{C}$ the spectral radius is
$$ \rho(T)=\max_{\lambda \in \sigma(T)}|\lambda| . $$
The Jordan Decomposition in General
Diagonalisation is the special case of a finer canonical form that always exists after a mild hypothesis on the field.
Definition. A Jordan block of size $e$ with eigenvalue $\lambda$ is the $e \times e$ matrix $J_e(\lambda)=\lambda I_e+N_e$, where $N_e$ is the matrix with ones on the superdiagonal and zeros elsewhere; $N_e$ is nilpotent with $N_e^e=0$ and $N_e^{e-1} \neq 0$.
Theorem (Jordan). Suppose $c_T$ splits over $F$, which holds automatically when $F$ is algebraically closed. Then $T$ is similar to a block diagonal matrix whose blocks are Jordan blocks, and the multiset of blocks is determined by $T$ up to reordering. Moreover $T$ is diagonalisable if and only if every block has size $1$, equivalently if and only if $m_T$ is a product of distinct linear factors, equivalently if and only if each generalised eigenspace equals the corresponding eigenspace.
The theorem is proved in the applications article of this category on modules over $k[x]$ and the Jordan form, where it is read off from the structure theorem for finitely generated modules over the principal ideal domain $F[x]$ applied to $V$ with $x$ acting as $T$. It is quoted here because it is the precise sense in which diagonalisation fails: the failure is measured by the sizes of the Jordan blocks of size at least $2$, one for each repeated root of $m_T$, and the minimal polynomial is the product of $(x-\lambda)^{e_\lambda}$ over the eigenvalues, where $e_\lambda$ is the largest Jordan block at $\lambda$.
Summary
An eigenvalue of $T:V \to V$ is a scalar $\lambda$ with $T(v)=\lambda v$ for some nonzero $v$, the eigenvector; the eigenvectors for $\lambda$ together with $0$ form the eigenspace $V_\lambda=\ker(T-\lambda\operatorname{id})$. Eigenvectors for distinct eigenvalues are independent, and the sum of the eigenspaces is always direct. The characteristic polynomial $c_T(x)=\det(xI-A)$ is a similarity invariant, monic of degree $n$, with $c_T=x^n-(\operatorname{tr}T)x^{n-1}+\cdots+(-1)^n\det T$, and its roots in $F$ are exactly the eigenvalues.
$T$ is diagonalisable when $V$ has a basis of eigenvectors, equivalently when $V$ is the direct sum of the eigenspaces, equivalently when the geometric multiplicities sum to $n$, equivalently when $c_T$ splits over $F$ and every geometric multiplicity equals the corresponding algebraic multiplicity. Distinct eigenvalues always suffice but are not necessary. The eigenspace dimensions are computed as $n-\operatorname{rk}(\lambda I-A)$, so diagonalisation is decided by rank computations on the matrices $\lambda I-A$.
When diagonalisation fails one uses invariant subspaces: a generalised eigenspace collects the vectors killed by a power of $T-\lambda\operatorname{id}$, the space decomposes into generalised eigenspaces whenever $c_T$ splits, and triangularisation is available over any field over which $c_T$ splits. The minimal polynomial $m_T$ generates the ideal of polynomials annihilating $T$, divides $c_T$ by Cayley–Hamilton, has the same roots as $c_T$, and decides diagonalisability: $T$ is diagonalisable exactly when $m_T$ is a product of distinct linear factors. Commuting diagonalisable operators are simultaneously diagonalisable, because each preserves the eigenspaces of the others. On a diagonalisable operator every polynomial is evaluated eigenvalue by eigenvalue, $p(T)=S\operatorname{diag}(p(\lambda_1),\dots,p(\lambda_n))S^{-1}$.
Summary of Notation
| Symbol | Meaning |
|---|---|
| $F$ | a field |
| $V$ | finite-dimensional $F$-vector space, $n=\dim_F V$ |
| $T:V \to V$ | linear operator |
| $A$, $B$ | matrices in $M_n(F)$ |
| $\lambda$ | eigenvalue |
| $V_\lambda=\ker(T-\lambda\operatorname{id})$ | eigenspace |
| $c_T(x)=\det(xI-A)$ | characteristic polynomial |
| $\operatorname{tr}T$, $\det T$ | trace and determinant, similarity invariants |
| $m_T(x)$ | minimal polynomial |
| $T|_W$ | restriction to a $T$-invariant subspace $W$ |
| $V^{(\lambda)}$ | generalised eigenspace, $\bigcup_k\ker(T-\lambda\operatorname{id})^k$ |
| $\operatorname{id}$, $I$ | identity operator, identity matrix |
| $P^{-1}AP$ | similar matrices |
| $\operatorname{diag}(\lambda_1,\dots,\lambda_n)$ | diagonal matrix |
| $\varphi,\psi$ | the roots $\tfrac12(1\pm\sqrt5)$ of $x^2-x-1$ |
| $\omega$ | primitive cube root of unity $e^{2\pi i/3}$ |
| $\sigma(T)$ | spectrum: eigenvalues of $T$, with algebraic multiplicities over a splitting field |
| $\rho(T)$ | spectral radius, $\max_{\lambda\in\sigma(T)}|\lambda|$ |
| $J_e(\lambda)=\lambda I_e+N_e$ | Jordan block of size $e$ |
| $N_e$ | nilpotent superdiagonal shift, $N_e^e=0 \neq N_e^{e-1}$ |
| $\mathbb{R}$, $\mathbb{C}$ | real and complex numbers; splitting fields for real matrices |
Further Reading
- Sheldon Axler, Linear Algebra Done Right (Springer, 3rd ed. 2015), for the determinant-free treatment of eigenvalues and diagonalisation.
- Nicolas Bourbaki, Algebra I: Chapters 1–3 (Springer, 1998), for the module-theoretic view of the minimal polynomial.
- Paul M. Cohn, Basic Algebra: Groups, Rings and Fields (Springer, 2003), for the rational canonical form and invariant subspaces.
- David S. Dummit and Richard M. Foote, Abstract Algebra (Wiley, 3rd ed. 2004), for worked diagonalisation examples and the minimal polynomial.
- Kenneth Hoffman and Ray Kunze, Linear Algebra (Prentice Hall, 2nd ed. 1971), for the characteristic and minimal polynomials with full proofs.
- Serge Lang, Linear Algebra (Springer, 3rd ed. 1987), for triangularisation and simultaneous diagonalisation.
- Steven Roman, Advanced Linear Algebra (Springer, 3rd ed. 2008), for the structure theory of a single operator.