Vector norms and normalization: L1, L2, and L∞

Measure vectors with L1, L2, and infinity norms. Explore unit boundaries, normalize a robot displacement, and distinguish zero vectors from tiny nonzero inputs.

By 9 min read

What you will learn

  • Calculate L1, L2, and infinity norms and explain what each measures.
  • Normalize a nonzero vector using a specified norm.
  • Explain why zero-vector normalization is undefined and the L0 count fails a norm rule.
  • Distinguish vector normalization from standardizing features across a dataset.

Before you start

Moving three meters east and four meters north gives a displacement vector v = (3, 4). Its straight-line length is five meters. Its two coordinate movements add up to seven meters.

Both numbers measure something useful. A norm makes the choice of size measure explicit. Normalization then rescales a nonzero vector until that chosen measure equals one.

Define a consistent measure of size

For real vectors, a norm follows three rules:

  • Its value is nonnegative and equals zero only for the zero vector.
  • Scaling a vector by a multiplies its norm by |a|.
  • The norm of u + v is at most the norm of u plus the norm of v.

The last rule is the triangle inequality. It keeps a direct displacement no larger than the combined size of two successive displacements, under the same norm. Nick Higham gives these defining properties.

Write a norm with double bars, such as ‖v‖. A subscript identifies the measure. The vectors still use the addition and scaling rules from vector spaces.

Measure a vector three ways

For v with real coordinates v₁ through vₙ:

‖v‖₁ = Σᵢ |vᵢ|
‖v‖₂ = √(Σᵢ vᵢ²)
‖v‖∞ = maxᵢ |vᵢ|

L1, also called the Manhattan norm, adds absolute coordinate sizes. L2 gives Euclidean length in orthonormal coordinates. L∞, read “L infinity,” takes the largest absolute coordinate. The infinity symbol names the norm; a finite vector has a finite mathematical value.

These are vector norms. Passing a matrix to a library can select different definitions, so check the input shape and axis. NumPy documents its vector and matrix conventions separately.

You may also see L0, meaning the number of nonzero coordinates. It fails the scaling rule: (3, 0) and twice that vector, (6, 0), both have one nonzero entry. A norm would double. L0 is a useful count, but it is not a mathematical norm; NumPy's vector ord=0 returns this count.

Separate a robot's direction from its distance

Use perpendicular east and north axes, with both coordinates in meters. For v = (3, 4):

MeasureCalculationResult
L13 + 47 m
L2√(3² + 4²)5 m
L∞max(3, 4)4 m

Five meters is the direct displacement length. Seven meters is the path length if the robot travels three meters along one axis, then four along the other. Four meters is the larger coordinate displacement. These interpretations assume the stated axes and units.

To describe direction independently of distance, divide by the L2 norm:

u = v / ‖v‖₂ = (3/5, 4/5) = (0.6, 0.8)
‖u‖₂ = √(0.6² + 0.8²) = 1

The unit vector u is dimensionless. Multiplying it by a speed of 2 m/s gives a velocity of (1.2, 1.6) m/s, whose Euclidean speed is 2 m/s. This specifies a velocity direction and magnitude; it does not establish whether a particular robot can execute the motion.

Compare the unit boundaries

A unit boundary contains vectors with norm exactly one. The region on and inside it is the unit ball.

  • L1 gives a diamond: the absolute coordinates add to one.
  • L2 gives a circle: Euclidean length equals one.
  • L∞ gives a square: the largest absolute coordinate equals one.

Interactive experiment

How does the norm change the unit vector?

Measure one vector three ways, then divide it by the selected norm. The amber boundary contains vectors whose selected norm equals one.

Enter coordinates from −6 to 6. Press Enter or leave a field to apply it; Escape restores its value. Small nonzero entries keep their size and direction.

Original vector v

Original vector (3.000, 4.000)The vector runs from the origin to x 3, y 4. Axis tick marks show plus and minus 4. The axes rescale for each vector.-4-4440xyv
The original axes rescale to keep tiny vectors visible. Read the tick values before comparing arrow lengths across presets.

Unit boundaries and result u

L2 unit boundary and normalized vector (0.600, 0.800)L1 has a diamond boundary, L2 a circle, and L infinity a square. Amber highlights L2. The normalized vector ends at (0.600, 0.800) on that boundary.-1-1110xyu
Dividing by L2 puts u on the circle. This panel always uses the same scale.
L1: diamondL2: circle (selected)L∞: square
L1 norm
7
L2 norm
5
L∞ norm
4
Nonzero entries
2
Normalized vector
(0.600, 0.800)
Result norm
1.000

Vector (3.000, 4.000) normalized by L2 gives (0.600, 0.800). Its L2 norm is 1.000.

Nonzero entries is a count, often called L0. It is not a mathematical norm and is not offered as a normalization choice.

Start with the 3–4–5 vector and change Normalization norm. The normalized endpoint moves along the same ray to meet the selected boundary. The original panel rescales its axes for visibility; the unit panel keeps a fixed scale.

Reverse direction selects (−3, −4). All three norms stay unchanged while the normalized arrow reverses. The presets preserve your selected norm. Coordinate edits apply on Enter or when you leave the field.

State which norm becomes one

For any selected norm and nonzero v, define u = v / ‖v‖. The denominator is positive, so the direction stays the same. The scaling rule gives ‖u‖ = 1.

For the same (3, 4) input, the three results differ:

Divide byNormalized resultEuclidean length of result
L1 = 7(3/7, 4/7)5/7
L2 = 5(0.6, 0.8)1
L∞ = 4(0.75, 1)1.25

“Unit” must refer to a particular norm. In ordinary geometric discussions, a unit vector usually means L2 length one.

For real vectors, ‖v‖₂ = √(v · v). Two L2-normalized vectors therefore have a dot product equal to the cosine of their angle. L1 or L∞ normalization does not establish that identity. A nonzero cross product can also be L2-normalized to specify a unit perpendicular direction.

Separate zero from small

The zero vector has norm zero. Dividing by it is undefined, and multiplying zero by a scalar cannot produce a unit vector. The experiment reports Undefined and draws no normalized arrow.

Small nonzero vectors are different. The Tiny vector preset uses (3 × 10⁻¹², 4 × 10⁻¹²). Its L2 norm is 5 × 10⁻¹², and its L2-normalized result is still (0.6, 0.8).

Computer arithmetic needs care. Squaring extremely small coordinates can round them to zero before taking the square root. The helper first rescales the vector by its largest absolute coordinate, then normalizes that scaled vector. Higham describes this scaling approach to avoid damaging underflow and overflow.

A numerical threshold is an application policy, not a change to the definition. If measurement noise is comparable to the vector's size, its direction may be unreliable. Decide how to handle that uncertainty explicitly; returning zero does not produce a unit vector.

Choose coordinates and units deliberately

For an error vector with comparable coordinate units, L∞ measures the worst coordinate error. A requirement ‖error‖∞ ≤ 0.1 means every coordinate's absolute error is at most 0.1. L2 instead measures the combined Euclidean error.

Our robot example uses meters on both axes. Combining a distance in meters with a motor temperature in degrees gives numbers that need a justified scaling before their norm has a useful interpretation. Changing meters to millimeters can otherwise change which feature dominates the result.

Norms also depend on coordinate geometry. A rotation in an orthonormal frame preserves L2 length; a general matrix transformation can stretch it. State the coordinates and the measure when comparing transformed vectors.

Distinguish normalization from standardization

Vector normalization divides each nonzero sample by its own norm. For L2, (3, 4) and (6, 8) both become (0.6, 0.8), so their original lengths disappear. Scikit-learn's normalization reference describes this operation and its selectable axis.

Feature standardization works across samples. It subtracts each feature's training mean and divides by that feature's training standard deviation, when nonzero. It can change directions and does not generally make each row's norm one.

StandardScaler documents those per-feature statistics. Estimate them from training data and reuse them on later data. Choose preprocessing based on which information the task needs; removing magnitude can discard a useful signal.

Reproduce the calculation in Python

This standard-library example uses nonempty finite vectors. Scaling before division keeps tiny nonzero inputs from becoming an artificial zero case.

from math import fsum, hypot

def norm(v, kind):
    if kind == "L1":
        return fsum(abs(x) for x in v)
    if kind == "L2":
        return hypot(*v)
    if kind == "Linf":
        return max(abs(x) for x in v)
    raise ValueError("Unknown norm")

def normalize(v, kind):
    scale = max(abs(x) for x in v)
    if scale == 0:
        return None
    scaled = tuple(x / scale for x in v)
    denominator = norm(scaled, kind)
    return tuple(x / denominator for x in scaled)

def pair(v):
    return f"({v[0]:.3f}, {v[1]:.3f})"

v = (3, 4)
for kind in ["L1", "L2", "Linf"]:
    u = normalize(v, kind)
    print(f"{kind}: norm={norm(v, kind):.3f}; "
          f"u={pair(u)}; result={norm(u, kind):.3f}")
print("zero:", normalize((0, 0), "L2"))
print("tiny L2:", pair(normalize((3e-200, 4e-200), "L2")))

Expected output:

L1: norm=7.000; u=(0.429, 0.571); result=1.000
L2: norm=5.000; u=(0.600, 0.800); result=1.000
Linf: norm=4.000; u=(0.750, 1.000); result=1.000
zero: None
tiny L2: (0.600, 0.800)

Try it yourself

Exercise 1. For v = (−3, 4), calculate all three norms, its nonzero-entry count, and its L2-normalized vector. What changes if you double v?

Show solution: scale the size and preserve the direction

The norms are L1 = 7, L2 = 5, and L∞ = 4. There are two nonzero entries. Dividing by five gives (−0.6, 0.8).

Doubling gives (−6, 8), with norms 14, 10, and 8. Its normalized vector stays (−0.6, 0.8). Its nonzero-entry count stays two, demonstrating why that count fails norm homogeneity.

Exercise 2. A displacement error is (0.08, −0.08) meters. Does it meet a requirement that each coordinate error be at most 0.1 meters in absolute value? Does it meet a requirement that the Euclidean error be at most 0.1 meters?

Show solution: match the norm to the requirement

Its L∞ norm is 0.08 meters, so both coordinate errors meet the first requirement.

Its L2 norm is √(0.08² + 0.08²) ≈ 0.1131 meters, which exceeds 0.1 meters. It fails the Euclidean requirement. The two requirements define different acceptable regions.

Sources and further study