Given a subspace W and a vector y that is not in it, which vector in W is closest to y? The answer is the orthogonal projection of y onto W: drop a perpendicular from y to W. This one idea powers least-squares data fitting, the Gram-Schmidt process, and much of approximation theory.
Projecting onto a line
Start with a line L=Span{u}, where u=0. You want to split y into two pieces,
y=y^+z,
with y^=αu on the line and z orthogonal to u. The condition z⋅u=0 reads (y−αu)⋅u=0, so y⋅u=α(u⋅u), which determines α.
Projection onto a line
The orthogonal projection of y onto L=Span{u} is
y^=projLy=u⋅uy⋅uu.
The component of y orthogonal to u is z=y−y^.
The projection depends only on the line, not on which nonzero u you pick: replacing u by cu multiplies the fraction by c1 and the vector by c. Notice the fraction is exactly the weight formula from the previous lesson.
Worked example: Projection in the plane
Let y=(7,6) and u=(4,2). Find the projection of y onto Span{u}, write y as a sum of a vector on the line and a vector orthogonal to it, and find the distance from y to the line.
y⋅u=28+12=40 and u⋅u=16+4=20, so
y^=2040(4,2)=(8,4),z=y−y^=(−1,2).
Check: z⋅u=−4+4=0. So (7,6)=(8,4)+(−1,2).
The closest point on the line to y is y^, so the distance from y to the line is ∥z∥=1+4=5.
ŷ = (8, 4) is the foot of the perpendicular from y = (7, 6) to the line through u. The dashed segment z = (−1, 2) is orthogonal to the line.Open in grapher →
Projecting onto a subspace
The same splitting works for any subspace W of Rn, as long as you have an orthogonal basis for it.
The orthogonal decomposition theorem
Let W be a subspace of Rn. Every y in Rn can be written uniquely as
y=y^+z,y^ in W,z in W⊥.
If {u1,…,up} is any orthogonal basis of W, then
y^=projWy=u1⋅u1y⋅u1u1+⋯+up⋅upy⋅upup.
In words: projWy is the sum of the projections of y onto the separate basis lines. To see that z=y−y^ is orthogonal to W, dot it with u1. Orthogonality kills the cross terms, leaving y⋅u1−u1⋅u1y⋅u1(u1⋅u1)=0, and likewise for each uj.
Two quick consequences: if y is already in W, then projWy=y (this is the weight formula from the last lesson), and if y is in W⊥, then projWy=0.
Worked example: Projection onto a plane in R³
Let u1=(1,1,0), u2=(1,−1,2) and y=(3,1,5). Write y as a vector in W=Span{u1,u2} plus a vector orthogonal to W.
First check the basis is orthogonal: u1⋅u2=1−1+0=0. Then
So z=y−y^=(−1,1,1). Check: z⋅u1=−1+1=0 and z⋅u2=−1−1+2=0. The decomposition is
(3,1,5)=(4,0,4)+(−1,1,1).
Common mistake
The projection formula requires an orthogonal basis. If u1⋅u2=0, adding up the one-line projections gives the wrong answer, and y−y^ won't be orthogonal to W. Always check orthogonality first; if the basis isn't orthogonal, make it so with Gram-Schmidt (next lesson).
The best approximation theorem
Best approximation
Let W be a subspace of Rn and y^=projWy. Then y^ is the point of W closest to y:
∥y−y^∥<∥y−v∥for every v in W with v=y^.
The proof is the Pythagorean theorem. For v in W, write y−v=(y−y^)+(y^−v). The first piece is in W⊥ and the second is in W, so they are orthogonal and
∥y−v∥2=∥y−y^∥2+∥y^−v∥2,
which is larger than ∥y−y^∥2 unless v=y^.
So the distance from y to W is ∥y−y^∥=∥z∥. In the plane example above, the distance from (3,1,5) to W is ∥(−1,1,1)∥=3.
Projection matrices
With an orthonormal basis {u1,…,up} the denominators are all 1, so projWy=(y⋅u1)u1+⋯+(y⋅up)up. With U=[u1⋯up], the vector of weights is UTy, and combining the columns with those weights gives:
Projection with orthonormal columns
If the columns of U are an orthonormal basis of W, then projWy=UUTy for every y.
Worked example: A projection matrix
Let W be the line spanned by u=(53,54), a unit vector. Find the matrix P of projW and use it to project y=(5,0).