explainer
Vector norms and normalization: L1, L2, and L∞
Measure vectors with L1, L2, and infinity norms. Explore unit boundaries, normalize a robot displacement, and distinguish zero vectors from tiny nonzero inputs.
What you will learn
- Calculate L1, L2, and infinity norms and explain what each measures.
- Normalize a nonzero vector using a specified norm.
- Explain why zero-vector normalization is undefined and the L0 count fails a norm rule.
- Distinguish vector normalization from standardizing features across a dataset.
Before you start
Moving three meters east and four meters north gives a displacement vector v = (3, 4). Its straight-line length is five meters. Its two coordinate movements add up to seven meters.
Both numbers measure something useful. A norm makes the choice of size measure explicit. Normalization then rescales a nonzero vector until that chosen measure equals one.
Define a consistent measure of size
For real vectors, a norm follows three rules:
- Its value is nonnegative and equals zero only for the zero vector.
- Scaling a vector by a multiplies its norm by |a|.
- The norm of u + v is at most the norm of u plus the norm of v.
The last rule is the triangle inequality. It keeps a direct displacement no larger than the combined size of two successive displacements, under the same norm. Nick Higham gives these defining properties.
Write a norm with double bars, such as ‖v‖. A subscript identifies the measure. The vectors still use the addition and scaling rules from vector spaces.
Measure a vector three ways
For v with real coordinates v₁ through vₙ:
‖v‖₁ = Σᵢ |vᵢ|
‖v‖₂ = √(Σᵢ vᵢ²)
‖v‖∞ = maxᵢ |vᵢ|
L1, also called the Manhattan norm, adds absolute coordinate sizes. L2 gives Euclidean length in orthonormal coordinates. L∞, read “L infinity,” takes the largest absolute coordinate. The infinity symbol names the norm; a finite vector has a finite mathematical value.
These are vector norms. Passing a matrix to a library can select different definitions, so check the input shape and axis. NumPy documents its vector and matrix conventions separately.
You may also see L0, meaning the number of nonzero coordinates. It fails the scaling rule: (3, 0) and twice that vector, (6, 0), both have one nonzero entry. A norm would double. L0 is a useful count, but it is not a mathematical norm; NumPy's vector ord=0 returns this count.
Separate a robot's direction from its distance
Use perpendicular east and north axes, with both coordinates in meters. For v = (3, 4):
| Measure | Calculation | Result |
|---|---|---|
| L1 | 3 + 4 | 7 m |
| L2 | √(3² + 4²) | 5 m |
| L∞ | max(3, 4) | 4 m |
Five meters is the direct displacement length. Seven meters is the path length if the robot travels three meters along one axis, then four along the other. Four meters is the larger coordinate displacement. These interpretations assume the stated axes and units.
To describe direction independently of distance, divide by the L2 norm:
u = v / ‖v‖₂ = (3/5, 4/5) = (0.6, 0.8)
‖u‖₂ = √(0.6² + 0.8²) = 1
The unit vector u is dimensionless. Multiplying it by a speed of 2 m/s gives a velocity of (1.2, 1.6) m/s, whose Euclidean speed is 2 m/s. This specifies a velocity direction and magnitude; it does not establish whether a particular robot can execute the motion.
Compare the unit boundaries
A unit boundary contains vectors with norm exactly one. The region on and inside it is the unit ball.
- L1 gives a diamond: the absolute coordinates add to one.
- L2 gives a circle: Euclidean length equals one.
- L∞ gives a square: the largest absolute coordinate equals one.
Start with the 3–4–5 vector and change Normalization norm. The normalized endpoint moves along the same ray to meet the selected boundary. The original panel rescales its axes for visibility; the unit panel keeps a fixed scale.
Reverse direction selects (−3, −4). All three norms stay unchanged while the normalized arrow reverses. The presets preserve your selected norm. Coordinate edits apply on Enter or when you leave the field.
State which norm becomes one
For any selected norm and nonzero v, define u = v / ‖v‖. The denominator is positive, so the direction stays the same. The scaling rule gives ‖u‖ = 1.
For the same (3, 4) input, the three results differ:
| Divide by | Normalized result | Euclidean length of result |
|---|---|---|
| L1 = 7 | (3/7, 4/7) | 5/7 |
| L2 = 5 | (0.6, 0.8) | 1 |
| L∞ = 4 | (0.75, 1) | 1.25 |
“Unit” must refer to a particular norm. In ordinary geometric discussions, a unit vector usually means L2 length one.
For real vectors, ‖v‖₂ = √(v · v). Two L2-normalized vectors therefore have a dot product equal to the cosine of their angle. L1 or L∞ normalization does not establish that identity. A nonzero cross product can also be L2-normalized to specify a unit perpendicular direction.
Separate zero from small
The zero vector has norm zero. Dividing by it is undefined, and multiplying zero by a scalar cannot produce a unit vector. The experiment reports Undefined and draws no normalized arrow.
Small nonzero vectors are different. The Tiny vector preset uses (3 × 10⁻¹², 4 × 10⁻¹²). Its L2 norm is 5 × 10⁻¹², and its L2-normalized result is still (0.6, 0.8).
Computer arithmetic needs care. Squaring extremely small coordinates can round them to zero before taking the square root. The helper first rescales the vector by its largest absolute coordinate, then normalizes that scaled vector. Higham describes this scaling approach to avoid damaging underflow and overflow.
A numerical threshold is an application policy, not a change to the definition. If measurement noise is comparable to the vector's size, its direction may be unreliable. Decide how to handle that uncertainty explicitly; returning zero does not produce a unit vector.
Choose coordinates and units deliberately
For an error vector with comparable coordinate units, L∞ measures the worst coordinate error. A requirement ‖error‖∞ ≤ 0.1 means every coordinate's absolute error is at most 0.1. L2 instead measures the combined Euclidean error.
Our robot example uses meters on both axes. Combining a distance in meters with a motor temperature in degrees gives numbers that need a justified scaling before their norm has a useful interpretation. Changing meters to millimeters can otherwise change which feature dominates the result.
Norms also depend on coordinate geometry. A rotation in an orthonormal frame preserves L2 length; a general matrix transformation can stretch it. State the coordinates and the measure when comparing transformed vectors.
Distinguish normalization from standardization
Vector normalization divides each nonzero sample by its own norm. For L2, (3, 4) and (6, 8) both become (0.6, 0.8), so their original lengths disappear. Scikit-learn's normalization reference describes this operation and its selectable axis.
Feature standardization works across samples. It subtracts each feature's training mean and divides by that feature's training standard deviation, when nonzero. It can change directions and does not generally make each row's norm one.
StandardScaler documents those per-feature statistics. Estimate them from training data and reuse them on later data. Choose preprocessing based on which information the task needs; removing magnitude can discard a useful signal.
Reproduce the calculation in Python
This standard-library example uses nonempty finite vectors. Scaling before division keeps tiny nonzero inputs from becoming an artificial zero case.
from math import fsum, hypot
def norm(v, kind):
if kind == "L1":
return fsum(abs(x) for x in v)
if kind == "L2":
return hypot(*v)
if kind == "Linf":
return max(abs(x) for x in v)
raise ValueError("Unknown norm")
def normalize(v, kind):
scale = max(abs(x) for x in v)
if scale == 0:
return None
scaled = tuple(x / scale for x in v)
denominator = norm(scaled, kind)
return tuple(x / denominator for x in scaled)
def pair(v):
return f"({v[0]:.3f}, {v[1]:.3f})"
v = (3, 4)
for kind in ["L1", "L2", "Linf"]:
u = normalize(v, kind)
print(f"{kind}: norm={norm(v, kind):.3f}; "
f"u={pair(u)}; result={norm(u, kind):.3f}")
print("zero:", normalize((0, 0), "L2"))
print("tiny L2:", pair(normalize((3e-200, 4e-200), "L2")))
Expected output:
L1: norm=7.000; u=(0.429, 0.571); result=1.000
L2: norm=5.000; u=(0.600, 0.800); result=1.000
Linf: norm=4.000; u=(0.750, 1.000); result=1.000
zero: None
tiny L2: (0.600, 0.800)
Try it yourself
Exercise 1. For v = (−3, 4), calculate all three norms, its nonzero-entry count, and its L2-normalized vector. What changes if you double v?
Show solution: scale the size and preserve the direction
The norms are L1 = 7, L2 = 5, and L∞ = 4. There are two nonzero entries. Dividing by five gives (−0.6, 0.8).
Doubling gives (−6, 8), with norms 14, 10, and 8. Its normalized vector stays (−0.6, 0.8). Its nonzero-entry count stays two, demonstrating why that count fails norm homogeneity.
Exercise 2. A displacement error is (0.08, −0.08) meters. Does it meet a requirement that each coordinate error be at most 0.1 meters in absolute value? Does it meet a requirement that the Euclidean error be at most 0.1 meters?
Show solution: match the norm to the requirement
Its L∞ norm is 0.08 meters, so both coordinate errors meet the first requirement.
Its L2 norm is √(0.08² + 0.08²) ≈ 0.1131 meters, which exceeds 0.1 meters. It fails the Euclidean requirement. The two requirements define different acceptable regions.
Sources and further study
- Nick Higham: What Is a Vector Norm?, for norm properties and scaled numerical evaluation.
- NumPy: numpy.linalg.norm, for vector-norm definitions, the nonzero-entry count, and matrix conventions.
- Scikit-learn: normalize, for normalizing individual samples or features by a selected norm.
- Scikit-learn: StandardScaler, for per-feature means and standard deviations learned from training data.