Math Core

Lesson 7.1 · Symmetric Matrices

Diagonalizing symmetric matrices

Many of the matrices that show up in practice are symmetric: covariance matrices in statistics, Hessians in calculus, stiffness matrices in engineering. Symmetric matrices are the best-behaved matrices there are. Their eigenvalues are always real, and they can always be diagonalized using an orthonormal basis of eigenvectors, so the messy P−1P^{-1} in A=PDP−1A = PDP^{-1} becomes a simple transpose.

Symmetric matrices

Definition

Symmetric matrix

A square matrix AA is symmetric if AT=AA^T = A. Equivalently, aij=ajia_{ij} = a_{ji} for every ii and jj: the entries are mirror images across the main diagonal.

For example,

[3113],[2−40−47505−1]\begin{bmatrix} 3 & 1 \\ 1 & 3 \end{bmatrix}, \qquad \begin{bmatrix} 2 & -4 & 0 \\ -4 & 7 & 5 \\ 0 & 5 & -1 \end{bmatrix}

are symmetric, while [1231]\begin{bmatrix} 1 & 2 \\ 3 & 1 \end{bmatrix} is not, and a non-square matrix never is. The diagonal entries can be anything; only the off-diagonal pairs must match.

A handy source of symmetric matrices: for any matrix BB (any size), BTBB^TB is symmetric, because (BTB)T=BT(BT)T=BTB(B^TB)^T = B^T(B^T)^T = B^TB. You will use this fact in the lesson on the singular value decomposition.

Eigenvectors of a symmetric matrix are orthogonal

Recall from the last unit that eigenvectors for different eigenvalues are always linearly independent. For a symmetric matrix, something much stronger is true.

Orthogonal eigenvectors

If AA is symmetric, then any two eigenvectors from different eigenspaces are orthogonal.

The proof is short and worth knowing. Suppose Av1=λ1v1A\mathbf{v}_1 = \lambda_1\mathbf{v}_1 and Av2=λ2v2A\mathbf{v}_2 = \lambda_2\mathbf{v}_2 with λ1≠λ2\lambda_1 \ne \lambda_2. Using x⋅y=xTy\mathbf{x}\cdot\mathbf{y} = \mathbf{x}^T\mathbf{y} and AT=AA^T = A,

λ1(v1⋅v2)=(Av1)Tv2=v1TATv2=v1T(Av2)=λ2(v1⋅v2).\lambda_1(\mathbf{v}_1 \cdot \mathbf{v}_2) = (A\mathbf{v}_1)^T\mathbf{v}_2 = \mathbf{v}_1^TA^T\mathbf{v}_2 = \mathbf{v}_1^T(A\mathbf{v}_2) = \lambda_2(\mathbf{v}_1 \cdot \mathbf{v}_2).

So (λ1−λ2)(v1⋅v2)=0(\lambda_1 - \lambda_2)(\mathbf{v}_1 \cdot \mathbf{v}_2) = 0. Since λ1−λ2≠0\lambda_1 - \lambda_2 \ne 0, the dot product v1⋅v2\mathbf{v}_1 \cdot \mathbf{v}_2 must be 00.

Orthogonal diagonalization

If you normalize orthogonal eigenvectors, they become orthonormal, and a square matrix PP with orthonormal columns is an orthogonal matrix: it satisfies PTP=IP^TP = I, so P−1=PTP^{-1} = P^T. That leads to the following definition.

Definition

Orthogonally diagonalizable

A square matrix AA is orthogonally diagonalizable if there are an orthogonal matrix PP and a diagonal matrix DD with

A=PDPT=PDP−1.A = PDP^T = PDP^{-1}.

If A=PDPTA = PDP^T, then AT=(PDPT)T=(PT)TDTPT=PDPT=AA^T = (PDP^T)^T = (P^T)^TD^TP^T = PDP^T = A, since a diagonal matrix equals its own transpose. So only symmetric matrices can be orthogonally diagonalizable. The remarkable fact is that the converse is also true.

The spectral theorem for symmetric matrices

An n×nn \times n matrix AA is orthogonally diagonalizable if and only if it is symmetric. In that case:

  1. AA has nn real eigenvalues, counted with multiplicity.
  2. The dimension of each eigenspace equals the multiplicity of its eigenvalue.
  3. Eigenspaces for different eigenvalues are orthogonal.
  4. A=PDPTA = PDP^T, where the columns of PP are an orthonormal basis of Rn\mathbb{R}^n made of eigenvectors of AA.

The set of eigenvalues of a matrix is called its spectrum, which is where the name comes from. Part 2 is the big surprise: a symmetric matrix never has a deficient eigenspace, so it is always diagonalizable, even with repeated eigenvalues.

The recipe

  1. Find the eigenvalues.
  2. For each eigenvalue, find a basis of its eigenspace.
  3. Make each basis orthonormal. If an eigenspace is one-dimensional, just normalize. If it has dimension 2 or more, run Gram-Schmidt on its basis, then normalize.
  4. Put the orthonormal eigenvectors into the columns of PP and the matching eigenvalues on the diagonal of DD.

You never need Gram-Schmidt between different eigenspaces: they are automatically orthogonal.

Worked example: A 2 × 2 matrix

Orthogonally diagonalize A=[3113]A = \begin{bmatrix} 3 & 1 \\ 1 & 3 \end{bmatrix}.

The characteristic polynomial is (3−λ)2−1=λ2−6λ+8=(λ−4)(λ−2)(3 - \lambda)^2 - 1 = \lambda^2 - 6\lambda + 8 = (\lambda - 4)(\lambda - 2).

For λ=4\lambda = 4: A−4I=[−111−1]A - 4I = \begin{bmatrix} -1 & 1 \\ 1 & -1 \end{bmatrix}, so x1=x2x_1 = x_2 and v1=(1,1)\mathbf{v}_1 = (1, 1).

For λ=2\lambda = 2: A−2I=[1111]A - 2I = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}, so x1=−x2x_1 = -x_2 and v2=(1,−1)\mathbf{v}_2 = (1, -1).

As promised, v1⋅v2=1−1=0\mathbf{v}_1 \cdot \mathbf{v}_2 = 1 - 1 = 0. Both have length 2\sqrt{2}, so

P=[1/21/21/2−1/2],D=[4002],A=PDPT.P = \begin{bmatrix} 1/\sqrt{2} & 1/\sqrt{2} \\ 1/\sqrt{2} & -1/\sqrt{2} \end{bmatrix}, \qquad D = \begin{bmatrix} 4 & 0 \\ 0 & 2 \end{bmatrix}, \qquad A = PDP^T.

Worked example: A repeated eigenvalue

Orthogonally diagonalize A=[211121112]A = \begin{bmatrix} 2 & 1 & 1 \\ 1 & 2 & 1 \\ 1 & 1 & 2 \end{bmatrix}.

Every row sums to 44, so A(1,1,1)=(4,4,4)A(1, 1, 1) = (4, 4, 4) and λ=4\lambda = 4 is an eigenvalue with eigenvector (1,1,1)(1, 1, 1). Next, A−I=[111111111]A - I = \begin{bmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \end{bmatrix} has rank 11, so λ=1\lambda = 1 is an eigenvalue with a two-dimensional eigenspace: the plane x1+x2+x3=0x_1 + x_2 + x_3 = 0. (Check with the trace: 4+1+1=6=2+2+24 + 1 + 1 = 6 = 2 + 2 + 2.)

A basis for the plane is w1=(−1,1,0)\mathbf{w}_1 = (-1, 1, 0), w2=(−1,0,1)\mathbf{w}_2 = (-1, 0, 1), but these are not orthogonal (w1⋅w2=1\mathbf{w}_1 \cdot \mathbf{w}_2 = 1). Apply Gram-Schmidt:

w2−w2⋅w1w1⋅w1w1=(−1,0,1)−12(−1,1,0)=(−12,−12,1).\mathbf{w}_2 - \frac{\mathbf{w}_2 \cdot \mathbf{w}_1}{\mathbf{w}_1 \cdot \mathbf{w}_1}\mathbf{w}_1 = (-1, 0, 1) - \tfrac{1}{2}(-1, 1, 0) = \left(-\tfrac{1}{2}, -\tfrac{1}{2}, 1\right).

Scale this to (−1,−1,2)(-1, -1, 2). Now normalize all three eigenvectors:

P=[1/3−1/2−1/61/31/2−1/61/302/6],D=[400010001].P = \begin{bmatrix} 1/\sqrt{3} & -1/\sqrt{2} & -1/\sqrt{6} \\ 1/\sqrt{3} & 1/\sqrt{2} & -1/\sqrt{6} \\ 1/\sqrt{3} & 0 & 2/\sqrt{6} \end{bmatrix}, \qquad D = \begin{bmatrix} 4 & 0 & 0 \\ 0 & 1 & 0 \\ 0 & 0 & 1 \end{bmatrix}.

Notice that (1,1,1)(1, 1, 1) was already orthogonal to both (−1,1,0)(-1, 1, 0) and (−1,−1,2)(-1, -1, 2), exactly as the theorem guarantees.

Common mistake

Inside an eigenspace of dimension 2 or more, a basis you find by row reduction is usually not orthogonal. If you just normalize those vectors and put them in PP, then PP is not an orthogonal matrix and PT≠P−1P^T \ne P^{-1}. Always run Gram-Schmidt within each multi-dimensional eigenspace.

The spectral decomposition

Write P=[u1⋯un]P = \begin{bmatrix} \mathbf{u}_1 & \cdots & \mathbf{u}_n \end{bmatrix}. Multiplying out PDPTPDP^T column-by-row gives

A=λ1u1u1T+λ2u2u2T+⋯+λnununT.A = \lambda_1\mathbf{u}_1\mathbf{u}_1^T + \lambda_2\mathbf{u}_2\mathbf{u}_2^T + \cdots + \lambda_n\mathbf{u}_n\mathbf{u}_n^T.

This is the spectral decomposition of AA. Each ujujT\mathbf{u}_j\mathbf{u}_j^T is an n×nn \times n matrix of rank 11, and it is exactly the matrix of orthogonal projection onto the line spanned by uj\mathbf{u}_j. So a symmetric matrix acts by splitting a vector into perpendicular pieces along its eigenvectors and scaling each piece by its eigenvalue.

Worked example: Building a matrix from its spectrum

Let A=[7224]A = \begin{bmatrix} 7 & 2 \\ 2 & 4 \end{bmatrix}. Find the spectral decomposition of AA.

The characteristic polynomial is λ2−11λ+24=(λ−8)(λ−3)\lambda^2 - 11\lambda + 24 = (\lambda - 8)(\lambda - 3).

For λ=8\lambda = 8: A−8I=[−122−4]A - 8I = \begin{bmatrix} -1 & 2 \\ 2 & -4 \end{bmatrix} gives x1=2x2x_1 = 2x_2, so u1=15(2,1)\mathbf{u}_1 = \frac{1}{\sqrt{5}}(2, 1).

For λ=3\lambda = 3: the eigenvector must be orthogonal to (2,1)(2, 1), so u2=15(1,−2)\mathbf{u}_2 = \frac{1}{\sqrt{5}}(1, -2).

u1u1T=15[4221],u2u2T=15[1−2−24].\mathbf{u}_1\mathbf{u}_1^T = \frac{1}{5}\begin{bmatrix} 4 & 2 \\ 2 & 1 \end{bmatrix}, \qquad \mathbf{u}_2\mathbf{u}_2^T = \frac{1}{5}\begin{bmatrix} 1 & -2 \\ -2 & 4 \end{bmatrix}.

Check: 8u1u1T+3u2u2T=15[32+316−616−68+12]=15[35101020]=A8\mathbf{u}_1\mathbf{u}_1^T + 3\mathbf{u}_2\mathbf{u}_2^T = \frac{1}{5}\begin{bmatrix} 32 + 3 & 16 - 6 \\ 16 - 6 & 8 + 12 \end{bmatrix} = \frac{1}{5}\begin{bmatrix} 35 & 10 \\ 10 & 20 \end{bmatrix} = A.

Tip

For a 2×22 \times 2 symmetric matrix, once you have one eigenvector (a,b)(a, b), the other eigenvector is (b,−a)(b, -a) (or (−b,a)(-b, a)). No second row reduction is needed.

Practice

Practice 1

Which matrix is guaranteed to be orthogonally diagonalizable?

Practice 2

Find the eigenvalues of A=[5222]A = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}.

Separate answers with commas, e.g. 2, -5

Practice 3

The matrix P=[3/5−4/54/53/5]P = \begin{bmatrix} 3/5 & -4/5 \\ 4/5 & 3/5 \end{bmatrix} is orthogonal. Find the (1,2)(1, 2) entry of P−1P^{-1}.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

Let A=[102030201]A = \begin{bmatrix} 1 & 0 & 2 \\ 0 & 3 & 0 \\ 2 & 0 & 1 \end{bmatrix}. Its eigenvalues are 33 and −1-1. What is the dimension of the eigenspace for λ=3\lambda = 3?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A symmetric 2×22 \times 2 matrix AA has eigenvalues 33 and −2-2, and (1,2)(1, 2) is an eigenvector for λ=3\lambda = 3. Find the (1,1)(1, 1) entry of AA.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

Let u1=12(1,1)\mathbf{u}_1 = \frac{1}{\sqrt{2}}(1, 1) and u2=12(1,−1)\mathbf{u}_2 = \frac{1}{\sqrt{2}}(1, -1), and let A=6u1u1T−2u2u2TA = 6\mathbf{u}_1\mathbf{u}_1^T - 2\mathbf{u}_2\mathbf{u}_2^T. Find the (1,2)(1, 2) entry of AA.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

Find the eigenvalues of A=[2−1−1−12−1−1−12]A = \begin{bmatrix} 2 & -1 & -1 \\ -1 & 2 & -1 \\ -1 & -1 & 2 \end{bmatrix}, then describe an orthonormal eigenbasis.

Separate answers with commas, e.g. 2, -5

Practice 8

Let AA be an n×nn \times n symmetric matrix. Which statement is false?