Singular value decomposition: directions, gains, and low-rank approximation

Build an SVD from orthogonal directions and nonnegative gains. See a circle become an ellipse, identify lost directions, and measure the error from keeping one singular component.

By 12 min read

What you will learn

  • Interpret a singular value as the gain along a paired input and output direction.
  • Track the shapes of full and reduced SVD factors for a rectangular matrix.
  • Connect singular values to rank, lost directions, and sensitivity to input errors.
  • Calculate the spectral and Frobenius errors of a truncated SVD.

Before you start

A matrix can stretch one direction threefold while leaving a perpendicular direction unchanged. If those directions are tilted, the entries alone may hide that simple behavior. Singular value decomposition, or SVD, makes the directions and their gains explicit.

You will construct a matrix from those pieces, inspect what happens when one gain reaches zero, and calculate what an approximation loses when it drops the weaker component.

Separate directions from gains

Every real m × n matrix A has a full SVD:

A = UΣVᵀ

The columns of V are orthonormal input directions. The columns of U are orthonormal output directions. Orthonormal means unit length and mutually perpendicular. The rectangular diagonal matrix Σ contains the singular values, conventionally ordered σ₁ ≥ σ₂ ≥ … ≥ 0.

With the column-vector convention used here, read the operations from right to left. Vᵀ expresses an input in the V basis. Σ scales those coordinates. U combines them along the output basis directions. Review matrix composition if the order is unfamiliar.

The square U and V in the full SVD are orthogonal matrices: UᵀU = I and VᵀV = I. They preserve Euclidean lengths. An orthogonal factor can include a reflection as well as a rotation, so singular values do not carry a negative sign to represent a flip.

Follow a pair of singular directions

For a paired right singular vector vᵢ and left singular vector uᵢ:

Avᵢ = σᵢuᵢ

This equation supplies the geometric meaning. A unit input vᵢ becomes an output of length σᵢ, pointing along uᵢ when σᵢ is positive.

Take V = I, gains 3 and 1, and U as a 90° counterclockwise rotation. Using semicolons to separate rows:

U = [0, −1; 1, 0]
Σ = [3, 0; 0, 1]
A = [0, −1; 3, 0]

The input v₁ = (1, 0) becomes (0, 3), while v₂ = (0, 1) becomes (−1, 0). Any input x = (a, b) becomes Ax = (−b, 3a). For the unit input (0.6, 0.8), the output is (−0.8, 1.8).

The unit circle becomes an ellipse with semiaxis lengths 3 and 1. The directions of those semiaxes are the columns of U. Changing V changes which circle points arrive at particular output points, while the ellipse as a set stays the same when U and Σ stay fixed.

Turn a circle into an ellipse

The experiment starts with U = V = I and gains 3 and 1, so A = [3, 0; 0, 1]. The 45° unit probe produces approximately (2.121, 0.707). The dashed segment shows the image after retaining only the first singular component.

Interactive experiment

Build a matrix from its singular directions

Choose the gains and two orthonormal bases. Compare A with the approximation that keeps only its first singular component.

3
1
0°
0°
45°

σ₂ stays at or below σ₁. Reflecting flips the second column of U. All coordinates use fixed perpendicular unit axes; gains are dimensionless.

U
1.0000.000
0.0001.000
Σ
3.0000.000
0.0001.000
Vᵀ
1.0000.000
0.0001.000
A = UΣVᵀ
3.0000.000
0.0001.000

Input: the unit circle

Input: the unit circleThe filled dot is the unit input (0.707, 0.707). The open circle marks v₁; the open square marks v₂. Input ticks run from −1 to 1.-1-1011xy
The filled dot is the unit input (0.707, 0.707). The open circle marks v₁; the open square marks v₂. Input ticks run from −1 to 1.

Output: the circle's image

Output: the circle's imageThe filled dot is A times the input: (2.121, 0.707). The large open ring is the approximate output. The dashed blue segment is the rank-at-most-one image. Output ticks run from −4 to 4.-4-4044xy
The filled dot is A times the input: (2.121, 0.707). The large open ring is the approximate output. The dashed blue segment is the rank-at-most-one image. Output ticks run from −4 to 4.
Each singular direction has one gain: Avᵢ = σᵢuᵢ
Input directionGainOutput image
v1(1.000, 0.000)3.000(3.000, 0.000)
v2(0.000, 1.000)1.000(0.000, 1.000)
Rank
2
2-norm condition
3.000
Rank ≤1 spectral error
1.000
Rank ≤1 Frobenius error
1.000
Probe output
(2.121, 0.707)
Approximate output
(2.121, 0.000)
Probe error
0.707

Both gains are positive. The ellipse has two independent directions, and the discarded gain σ₂ is the approximation's largest error over all unit inputs.

The two plots use different scales. Changing V changes which inputs reach each output without changing the ellipse as a set. This experiment constructs known factors; rank comes from their exact gains. Displayed coordinates are rounded to three decimals.

Set Output basis angle to 90° to reproduce the worked matrix. Change Input basis angle while watching both the direction table and the probe: the output ellipse stays fixed, but the input directions and probe output change.

Select Rank one. The second gain is zero, so the ellipse collapses to a segment. Zero map collapses everything to the origin. Reflected output basis flips the second output direction while retaining the same positive gains.

Both plots show their coordinates, but their axis scales differ. The model uses dimensionless gains and fixed perpendicular unit axes. It constructs known factors within bounded controls; it does not run a decomposition algorithm on arbitrary matrix entries. Rank follows the chosen gains exactly, while the displayed trigonometric coordinates are rounded.

Check full and reduced factor shapes

SVD also applies when inputs and outputs have different dimensions. Set k = min(m, n) for an m × n matrix.

FactorFull SVDReduced SVD
Um × mm × k
Σm × nk × k
Vᵀn × nk × n

Both products reconstruct A. Reduced U and V have orthonormal columns; they may be rectangular. Reduced SVD retains k singular values, including any zeros. It is not automatically a truncation to the nonzero values. NumPy's SVD documentation specifies these shapes and its sorted singular-value output.

For a 3 × 5 matrix, reduced factors have shapes 3 × 3, 3 × 3, and 3 × 5. The full V has two additional columns. Those extra input directions map to zero even when all three listed singular values are positive. A wide matrix can therefore have full row rank while still losing input directions.

Find preserved and lost directions

In exact arithmetic, the rank is the number of positive singular values. A zero gain sends its input direction to zero. The rank and null-space lesson explains why an m × n matrix of rank r has an input null space of dimension n − r.

In the full SVD, the last n − r columns of V form an orthonormal basis for that null space. For wide matrices, include the extra columns just described, not only columns paired with a listed zero singular value.

The transpose connects SVD to eigenvalues:

AᵀAvᵢ = σᵢ²vᵢ

Thus the right singular directions are eigenvectors of AᵀA, with eigenvalues equal to squared singular values; additional null directions have eigenvalue zero. This follows by substituting A = UΣVᵀ and using UᵀU = I. Singular values of A are generally different from eigenvalues of A.

For measured data, deciding whether a tiny singular value counts as zero requires a tolerance and a scale. A floating-point SVD does not decide which physical effects matter to your task.

For a robot arm, manipulability turns the position Jacobian’s singular values into a velocity ellipse. Change the arm’s pose to compare its available motion directions, ellipse area, and condition number.

Recognize when directions are not unique

The ordered singular values are determined by A, but singular vectors need not be unique. Negating both uᵢ and vᵢ leaves their contribution σᵢuᵢvᵢᵀ unchanged.

Repeated positive singular values allow more freedom. For A = 2I₂, both gains are 2. Choose any orthonormal basis for V and use the same basis for U: U(2I₂)Vᵀ = 2I₂. No preferred longest direction exists because every unit input has gain 2.

Try Repeated gains, then set both basis angles to 45°. A stays 2I₂, while its displayed singular directions change. The rank-at-most-one approximation also changes because it keeps a different direction, with the same approximation error.

For a zero singular value, Avᵢ = 0 does not determine uᵢ. Dividing Avᵢ by σᵢ would be invalid. Null-space bases have freedom too; for the zero matrix, any orthogonal U and V work.

Compare the largest and smallest gains

For an invertible square matrix, its 2-norm condition number is:

κ₂(A) = σmax / σmin

The default matrix has condition number 3/1 = 3. Gains 3 and 0.25 give 12: errors along the weakest output direction are amplified more when solving backward. A condition number describes worst-case relative sensitivity; a particular right-hand side and perturbation may behave better. Stanford's SVD applications notes derive this ratio.

A singular square matrix has no ordinary inverse. The widget reports its condition number as Infinite, including the zero matrix. It does not evaluate 0/0 to reach that result. Conversely, a small nonzero scaled identity εI₂ has condition number 1: all directions shrink equally.

For rectangular problems, full row or column rank can give a finite ratio among the k listed singular values. That does not establish a two-sided inverse. A wide matrix can still have an input null space, and solving a least-squares problem brings additional questions about residuals and the requested solution.

Keep the strongest singular component

The SVD also writes a matrix as a sum of simple pieces:

A = Σᵢ σᵢuᵢvᵢᵀ
A₁ = σ₁u₁v₁ᵀ

Each nonzero term has rank one. A₁ has rank at most one, since it becomes zero if σ₁ = 0. For our two-by-two model, discarding the second term gives:

  • Spectral error: ‖A − A₁‖₂ = σ₂. This is the largest output error over all unit inputs.
  • Frobenius error: ‖A − A₁‖F = σ₂. This is the square root of the sum of squared errors in the matrix entries.

Those two norms happen to agree here because only one singular component is discarded. The error for an individual probe can be smaller. In the rotated example, input (0.6, 0.8) gives approximate output (0, 1.8), an error of 0.8, while the maximum unit-input error is 1.

More generally, keep the first q components, with 0 ≤ q < k = min(m, n). This gives a best approximation of rank at most q in either norm. Its spectral error is σq₊₁; its Frobenius error is the square root of the sum of squares of all discarded singular values.

Keeping all k components gives zero error. These are the Eckart–Young results taught in MIT's matrix methods course.

A best approximation need not be unique. Equal singular values at the cutoff can give different, equally good choices of retained directions. Also, a small matrix error does not guarantee a useful model: feature units, measurement noise, and the downstream task determine whether discarding a component is sensible.

Reproduce a known SVD in Python

This standard-library example constructs the quarter-turn case from known factors. It does not implement a general SVD solver. Multiplying the factors and measuring the residual lets you check the geometric claims independently.

from math import sqrt


def multiply(a, b):
    return [
        [sum(x * y for x, y in zip(row, column))
         for column in zip(*b)]
        for row in a
    ]


def apply(a, x):
    return tuple(sum(v * w for v, w in zip(row, x)) for row in a)


def show(x):
    return "(" + ", ".join(f"{value:.3f}" for value in x) + ")"


u = [[0, -1], [1, 0]]
sigma = [[3, 0], [0, 1]]
vt = [[1, 0], [0, 1]]
a = multiply(multiply(u, sigma), vt)
a1 = multiply(multiply(u, [[3, 0], [0, 0]]), vt)
x = (0.6, 0.8)  # Unit length: 0.6**2 + 0.8**2 = 1.
residual = [a[i][j] - a1[i][j] for i in range(2) for j in range(2)]

print("A:", a)
print("A v1:", show(apply(a, (1, 0))))
print("A v2:", show(apply(a, (0, 1))))
print("A x:", show(apply(a, x)))
print("A1 x:", show(apply(a1, x)))
print(f"Frobenius error: {sqrt(sum(r * r for r in residual)):.3f}")

Expected output:

A: [[0, -1], [3, 0]]
A v1: (0.000, 3.000)
A v2: (-1.000, 0.000)
A x: (-0.800, 1.800)
A1 x: (0.000, 1.800)
Frobenius error: 1.000

For general matrices, use a tested numerical library's SVD routine. For real matrices, NumPy returns a factor named Vh; it contains Vᵀ.

Try it yourself

Exercise 1. Use V = I, a 90° counterclockwise rotation for U, and singular values 4 and 2. Find A, its rank, its condition number, and both matrix-error norms after dropping the second component.

Show solution 1

Multiplying U by the diagonal gains gives A = [0, −2; 4, 0]. Both gains are positive, so its rank is 2. Its condition number is 4/2 = 2.

The approximation is A₁ = [0, 0; 4, 0]. The residual [0, −2; 0, 0] has spectral norm 2 and Frobenius norm 2. Input (0, 1) attains that output error; input (1, 0) has no error.

Exercise 2. A 3 × 4 matrix has singular values 5, 2, and 0. Give its reduced factor shapes, rank, input nullity, and the spectral and Frobenius errors of retaining only the first component.

Show solution 2

Here k = 3. Reduced U is 3 × 3, Σ is 3 × 3, and Vᵀ is 3 × 4. There are two positive singular values, so the rank is 2 and input nullity is 4 − 2 = 2.

One null direction is paired with the listed zero singular value; full V supplies the additional null direction. Dropping the gains 2 and 0 gives spectral error 2 and Frobenius error √(2² + 0²) = 2. No ordinary inverse exists for this rectangular matrix.

Use projections and least squares to study the closest output a matrix can produce. The pseudoinverse then uses singular directions to reverse nonzero gains and select a solution when an ordinary inverse is unavailable.

Sources and further study