Projections and least squares: find the closest fit

Project a vector onto a direction, measure its orthogonal residual, and connect that geometry to least squares, regression, and nonunique coefficients.

By 10 min read

What you will learn

  • Find the closest multiple of a nonzero vector to a target.
  • Check that a least-squares residual is orthogonal to every fitted direction.
  • Distinguish a unique fitted vector from unique coefficients.
  • Connect projection geometry with linear regression and numerical least-squares solvers.

Before you start

A projection finds the closest vector you can make from allowed directions. Least squares asks the same question through an error function: which coefficients make the sum of squared residuals smallest?

Start with one direction and two coordinates. The calculation will carry over to fitting many measurements with a matrix.

Find the closest allowed vector

Suppose a simplified robot can move along the direction a = (1, 2). A controller requests velocity b = (3, 1), measured in meters per second on both axes. Allow any real coefficient c, so the available velocities are c a.

Those velocities form the line through a and the origin: span(a). The closest one to b is its orthogonal projection, written p. The residual r = b − p records the part the fit cannot reproduce.

For a nonzero direction, the coefficient and projection are:

c* = (a · b) / (a · a)
p = c* a
r = b − p

The dot product measures alignment. Dividing by a · a accounts for the direction's length. The coefficient c* equals a signed projection length only when a has unit length.

This example uses a dimensionless direction and a velocity coefficient. It models no speed bounds or one-way motion; those would add constraints and could change the answer. Measuring both target coordinates in the same units also makes the Euclidean distance meaningful.

Work through a two-coordinate fit

For a = (1, 2) and b = (3, 1), calculate two dot products:

a · b = 1 × 3 + 2 × 1 = 5
a · a = 1² + 2² = 5
c* = 5 / 5 = 1

The fitted velocity is p = (1, 2). Subtraction gives r = (3, 1) − (1, 2) = (2, −1).

Check the residual's direction: a · r = 1 × 2 + 2 × (−1) = 0. It is perpendicular to a, and therefore to every vector in its span. The squared error is 2² + (−1)² = 5; the residual length is √5.

Compare a candidate with the best fit

The amber point marks the projection. The square marks the candidate c a, which you can move along the span. Compare their squared errors.

Projection and least squares

Find the closest vector in the span of a

Compare the best projection p with a candidate c a. The dashed residual runs from p to the target b.

Projection of target (3.000, 1.000) onto the span of direction (1.000, 2.000)The projection is (1.000, 2.000) and residual is (2.000, -1.000). The candidate coefficient 1.000 gives (1.000, 2.000). Both axes run from -4 to 4.-40440-4xy
Target bProjection pCandidate c aSpan of a
Coincident points share nested markers. Both axes use the same scale and resize to include the candidate.

Changing a or b resets c to zero. Use the fit button to choose the computed optimum.

Best coefficient 1.000 gives squared error 5.000. Candidate 1.000 gives squared error 5.000.

Fitted coefficient
1.000
Projection p
(1.000, 2.000)
Residual r
(2.000, -1.000)
Minimum squared error
5.000
Candidate squared error
5.000
a · r
0.000

Keep the default vectors and try c = 0 or c = 2. Both give squared error 10. Use least-squares fit returns c to 1 and lowers the error to 5.

Changing a or b resets the candidate coefficient to zero. The fitted readouts always show the optimum. Axis limits change together to keep the candidate visible without distorting angles.

Try Reverse direction. The coefficient changes sign, while the span, projection, and minimum error stay the same. Then try Perpendicular target: the best fit is the origin, even though a is nonzero.

Explain why the residual is smallest

For any candidate c, ordinary least squares uses the squared Euclidean residual:

F(c) = ‖b − c a‖₂²
F(c) = (3 − c)² + (1 − 2c)²
F(c) = 5(c − 1)² + 5

The square cannot be negative. Setting c = 1 makes it zero, so the smallest possible error is 5. Minimizing residual length gives the same answer because squaring preserves the order of nonnegative lengths.

The geometry explains the general case. Write b − c a = r + (c* − c)a. The two terms are perpendicular, so the Pythagorean theorem gives:

‖b − c a‖₂² = ‖r‖₂² + (c* − c)² ‖a‖₂²

Every move away from c* adds a nonnegative term. For nonzero a, that added term vanishes only at c*. This proves both the minimum and the coefficient's uniqueness.

Handle the zero direction explicitly

For a = (0, 0), the coefficient formula divides by zero. It does not define a fitted coefficient.

The set of allowed vectors still makes sense: span((0, 0)) contains only the origin. Its projection is p = (0, 0), with r = b. Every coefficient produces that same fitted vector, so the coefficient is not unique.

In the zero-direction preset, b = (3, 1) gives squared error 3² + 1² = 10. Moving the candidate control cannot change it.

Fit a combination of several columns

Let A have m rows and n columns. Each column supplies an allowed direction in ℝᵐ. A coefficient vector x combines them into A x.

Least squares finds a fitted vector p = A x* closest to b. Its residual must be orthogonal to every column of A. A transpose collects those dot products into one equation:

Aᵀr = 0
Aᵀ(b − A x*) = 0
AᵀA x* = Aᵀb

These are the normal equations. Here “normal” refers to perpendicularity. They express the projection condition for any real matrix A. MIT's projection and least-squares material develops this connection.

If A has full column rank, its columns are independent and x* is unique. With dependent columns, different coefficients can produce the same fitted vector. Rank and null space explain why: adding a null-space vector z changes x* but leaves A(x* + z) = A x*.

The closest fitted vector p remains unique, even in this rank-deficient case. A pseudoinverse selects the least-squares coefficient vector with the smallest Euclidean norm. NumPy's lstsq documentation states that minimum-norm choice.

Connect the geometry to regression

To fit a line prediction = intercept + slope × t, put one observation in each row of A. Row i is (1, tᵢ), x holds the intercept and slope, and b contains all measured responses.

For m observations, b and A x live in ℝᵐ, the observation space. Projecting b onto A's column space minimizes the sum of squared prediction errors. The linear regression lesson works through that fit with plotted measurements.

On the usual two-axis scatter plot, these errors are vertical differences between observed and predicted responses. Their orthogonality belongs to the column-space picture in ℝᵐ. Ordinary least squares does not minimize perpendicular distances from scatter-plot points to the drawn regression line.

The objective also encodes a choice: every squared residual receives equal weight. Different units, measurement reliabilities, or error costs can call for scaling or a different objective. A low training residual alone does not establish good predictions on new data.

Choose a numerical least-squares solver

The normal equations describe the solution mathematically. For numerical work, use a least-squares solver built around suitable matrix factorizations. Forming an explicit inverse of AᵀA is unnecessary, and dependent columns make that inverse unavailable.

LAPACK provides QR-based solvers for full-rank problems and SVD-based solvers that handle rank deficiency. Its least-squares guide describes the assumptions of each family. The singular value decomposition lesson explains the directions and scales behind that approach.

In floating-point work, deciding which very small singular values count as zero also affects the reported rank and coefficients. NumPy exposes that cutoff through the rcond argument of lstsq. Choose it with the scale and precision of the problem in mind.

Reproduce the fit in Python

This standard-library example uses exact fractions for the one-direction calculation. It assumes two-coordinate vectors and explicitly rejects the zero direction because the coefficient is not unique.

from fractions import Fraction

def dot(u, v):
    return sum(x * y for x, y in zip(u, v))

def fit(a, b):
    denominator = dot(a, a)
    if denominator == 0:
        raise ValueError("The zero direction has no unique coefficient")
    c = Fraction(dot(a, b), denominator)
    p = tuple(c * x for x in a)
    r = tuple(y - x for x, y in zip(p, b))
    return c, p, r

def vector_text(v):
    return "(" + ", ".join(str(x) for x in v) + ")"

a = (1, 2)
b = (3, 1)
c, p, r = fit(a, b)
print("Coefficient:", c)
print("Projection:", vector_text(p))
print("Residual:", vector_text(r))
print("Squared error:", dot(r, r))
print("a dot residual:", dot(a, r))
for candidate in (0, 1, 2):
    residual = tuple(y - candidate * x for x, y in zip(a, b))
    print(f"c={candidate}: squared error={dot(residual, residual)}")

Expected output:

Coefficient: 1
Projection: (1, 2)
Residual: (2, -1)
Squared error: 5
a dot residual: 0
c=0: squared error=10
c=1: squared error=5
c=2: squared error=10

Try it yourself

Exercise 1. Fit b = (3, 4) with a multiple of a = (2, 0). Find the coefficient, projection, residual, and squared error. Why does the coefficient differ from the projection's first coordinate?

Show solution: account for the direction's length

The coefficient is c* = 6/4 = 1.5. Multiplying by a gives p = (3, 0), so r = (0, 4) and the squared error is 16.

The direction's first coordinate is 2. A coefficient of 1.5 produces a first coordinate of 3. The residual is perpendicular to a because (2, 0) · (0, 4) = 0.

Exercise 2. A has columns a = (1, 2) and 2a = (2, 4). Fit b = (3, 1). Show that coefficient vectors (1, 0) and (0, 0.5) give the same best fit. Which quantities are unique?

Show solution: separate coefficients from fitted vectors

Both choices give A x = (1, 2): the first uses a once, and the second uses 2a half as much. Every coefficient pair satisfying x₁ + 2x₂ = 1 produces that projection.

The fitted vector p = (1, 2) and residual r = (2, −1) are unique. The minimum squared error is 5. The coefficients are not unique because the second column adds no new direction.

Sources and further study