explainer
Partial derivatives and gradients: predict a multivariable change
Hold one input fixed to find a partial derivative, combine the partials into a gradient, and compare a directional derivative with the actual change from a finite step.
What you will learn
- Compute partial derivatives while holding the other inputs fixed.
- Use a gradient and a displacement to predict a small cost change.
- Compare unit directional derivatives with finite-step changes.
- Explain the role of coordinate scales and the limits of a zero gradient.
Before you start
A cost can change when several inputs move together. Partial derivatives measure each input’s local effect while the others stay fixed. The gradient collects those effects so you can predict a small combined change.
For the cost below, the gradient predicts an increase of 1.5 for one chosen step. The actual increase is 1.75. Working through that gap shows what a derivative can tell you about a finite move.
Start with a cost that couples two inputs
Use this smooth, two-input function:
f(x, y) = x² + xy + 2y²
Think of x and y as two normalized calibration errors in a synthetic robot example. Each coordinate uses a chosen reference scale, so both coordinates and the cost are dimensionless. This is a teaching model, with no claim that it fits a measured robot.
The xy term couples the inputs. Changing y changes how sensitive the cost is to x. A model with only x² + 2y² would miss that interaction.
At (1, 1), the cost is 1 + 1 + 2 = 4. Start there and keep track of three separate quantities: the starting cost, a local rate of change, and the change after a chosen step.
Hold the other input fixed
To find the partial derivative with respect to x, treat y as a constant. The symbol ∂ signals that the function has other inputs:
∂f/∂x = limₕ→₀ [f(x + h, y) − f(x, y)] / h
∂f/∂x = 2x + y
The x² term contributes 2x. The xy term contributes y, and 2y² contributes zero because y stays fixed. For the y partial, hold x fixed:
∂f/∂y = x + 4y
At (1, 1), these rates are 3 along x and 5 along y. A very small positive y change has the larger local effect when you compare equal coordinate increments. The single-variable derivative rules apply to each slice.
For example, fixing y = 1 gives the curve f(x, 1) = x² + x + 2. Its slope at x = 1 is 3. Fixing x = 1 gives f(1, y) = 1 + y + 2y², whose slope at y = 1 is 5.
Collect the partials into a column
In this lesson, input vectors and gradients are columns. For two inputs, write:
∇f(x, y) = [∂f/∂x; ∂f/∂y] = [2x + y; x + 4y]
∇f(1, 1) = [3; 5]
The semicolon separates rows. Readouts use the compact pair (3, 5) for the same column. The symbol ∇ is pronounced “nabla,” and ∇f is “the gradient of f.”
The gradient depends on the point. At (−1, 2), it becomes (0, 7), so the x slice is locally flat there while the y slice has positive slope. One fixed gradient cannot describe this whole cost surface.
For a differentiable scalar function of n inputs, the gradient contains n partial derivatives. The total derivative maps an input change to one scalar change; with column inputs, its matrix is the row ∇fᵀ. The Jacobian lesson extends that row to functions with several outputs.
Predict a small change in both inputs
Let p = (x, y) and let Δ = (Δx, Δy) be a displacement. For a differentiable function, the local linear model is:
f(p + Δ) ≈ f(p) + ∇f(p) · Δ
Δf_predicted = (∂f/∂x)Δx + (∂f/∂y)Δy
This sum is a dot product. Each partial derivative multiplies the corresponding coordinate change. At (1, 1), a displacement (0.1, −0.2) predicts 3(0.1) + 5(−0.2) = −0.7.
For our polynomial, expanding the expression gives an exact identity:
f(p + Δ) = f(p) + ∇f(p) · Δ
- Δx² + ΔxΔy + 2Δy²
The last three terms form the prediction error. For (0.1, −0.2), they total 0.01 − 0.02 + 0.08 = 0.07. The actual new cost is 4 − 0.7 + 0.07 = 3.37.
Halving this displacement divides its linear prediction by two and its error by four. That squared scaling belongs to this quadratic model. General differentiability says the error divided by ‖Δ‖ approaches zero as Δ approaches zero.
Measure change along a unit direction
Choose a direction u with Euclidean length one. Move along the path p + su, where s measures signed distance in the chosen coordinates. The directional derivative measures the rate at s = 0:
Dᵤf(p) = d/ds f(p + su) at s = 0
Dᵤf(p) = ∇f(p) · u, with ‖u‖₂ = 1
At (1, 1), choosing u = (1, 0) gives 3, and u = (0, 1) gives 5. Choosing u = (3/5, 4/5) gives 3(3/5) + 5(4/5) = 5.8. A direction can change both inputs together.
A non-unit vector v defines a path p + tv too. Its derivative is ∇f · v, measured per unit of t; the vector’s length changes the speed along that path. Normalize v to compare rates per unit distance across directions. MIT derives this unit-direction dot product.
For a step of length h in unit direction u, the predicted change is h Dᵤf. The derivative itself stays a local rate even when h is large.
Compare the prediction with a finite step
The initial point is (1, 1), with direction (1, 0) and step length 0.5. The endpoint is (1.5, 1), and the predicted increase is 0.5 × 3 = 1.5. Direct evaluation gives 2.25 + 1.5 + 2 = 5.75, an increase of 1.75 and an error of 0.25.
The contour plot uses equal scales on both axes. Each ellipse joins points with the same cost, and the readouts give the complete coordinates. The blue arrow shows the gradient’s direction at a fixed display length; Gradient length gives its actual magnitude.
Try these comparisons:
- Along y: the local rate is 5. The half-unit step predicts 2.5 and produces an increase of 3.
- Against gradient: the local rate is −√34, about −5.831. The half-unit step lowers the cost from 4 to about 1.629.
- Perpendicular to gradient: the local rate is zero. The half-unit straight step still raises the cost by 7/34, about 0.206.
The perpendicular direction is tangent to the level curve at the starting point. A straight step leaves that curved level set, which explains its positive finite change. Set Step length to 0.25 and the error becomes one quarter of its value at 0.5.
The controls bound starting coordinates to [−2, 2] and step length to [0, 1]. They explore this exact function with no simulated sensor readings or fitted parameters. Three-decimal readouts round small values, while the math keeps its precision.
State the metric behind steepest
When the gradient is nonzero, its unit direction gives the largest directional derivative among all Euclidean unit directions. The dot-product bound explains why:
|∇f · u| ≤ ‖∇f‖₂
u_up = ∇f / ‖∇f‖₂, with rate +‖∇f‖₂
u_down = −∇f / ‖∇f‖₂, with rate −‖∇f‖₂
At (1, 1), the steepest unit uphill direction is (3, 5)/√34. Its rate, √34, exceeds the x and y rates. “Steepest” compares local changes for equal Euclidean step lengths.
Coordinate scales affect that comparison. If x measures meters and you replace it with X = 100x in centimeters, then ∂f/∂X = (∂f/∂x)/100. Treating one unit of X as equivalent to one unchanged y unit changes the steps you compare.
Our normalized coordinates make the chosen scales explicit. Real models that mix position, angle, or other units need meaningful reference scales or a metric that weights displacements. Boyd and Vandenberghe’s treatment of steepest descent shows how the norm used to measure a step affects its steepest direction.
A downhill directional derivative guarantees a decrease for sufficiently small positive steps. It does not choose a safe finite step size for you. Gradient descent turns the downhill direction into an update rule and examines that choice.
Check what a zero gradient tells you
At the origin, our cost has gradient (0, 0). All unit directional derivatives are zero, so no direction has a strict local first-order advantage. The experiment cannot normalize that zero vector and asks you to choose an independent direction.
For this particular function, completing the square proves that the origin is the unique global minimum:
f(x, y) = (x + y/2)² + 7y²/4 ≥ 0
f(x, y) = 0 only at (0, 0)
The zero gradient alone does not prove a minimum. For g(x, y) = x² − y², the gradient also vanishes at the origin. Nearby points along x have positive values, while nearby points along y have negative values, making the origin a saddle point.
The Hessian lesson measures this difference through second derivatives. It shows how curvature in different directions can distinguish a minimum from a saddle, and when that test leaves the answer unresolved.
For a differentiable function at an unconstrained interior local minimum, the gradient must vanish. Boundaries and constraints change the test. The function f(x) = x on the domain x ≥ 0 has its minimum at the boundary x = 0, even though its derivative is 1.
Reproduce the calculation in Python
This standard-library example computes the cost, gradient, and two directions at the same point. The unit-direction check prevents a scaled vector from silently changing the meaning of the rate. The printed cases also compare two step lengths.
from math import hypot, isclose
def cost(point):
x, y = point
return x*x + x*y + 2*y*y
def gradient(point):
x, y = point
return (2*x + y, x + 4*y)
def inspect(point, direction, step):
if not isclose(hypot(*direction), 1.0, rel_tol=1e-12):
raise ValueError("Use a unit direction")
rate = sum(a*b for a, b in zip(gradient(point), direction))
delta = tuple(step*x for x in direction)
endpoint = tuple(a+b for a, b in zip(point, delta))
predicted = step*rate
actual = cost(endpoint) - cost(point)
return rate, predicted, actual, cost(delta)
point = (1.0, 1.0)
gx, gy = gradient(point)
length = hypot(gx, gy)
tangent = (-gy/length, gx/length)
print(f"Start cost: {cost(point):.3f}")
print(f"Gradient: ({gx:.3f}, {gy:.3f})")
for label, direction, step in [
("x, h=0.50", (1.0, 0.0), 0.50),
("x, h=0.25", (1.0, 0.0), 0.25),
("tangent, h=0.50", tangent, 0.50),
]:
rate, predicted, actual, error = inspect(point, direction, step)
rate = 0.0 if abs(rate) < 1e-12 else rate
predicted = 0.0 if abs(predicted) < 1e-12 else predicted
print(f"{label}: rate={rate:.3f}, predicted={predicted:.3f}, "
f"actual={actual:.3f}, error={error:.3f}")
Expected output:
Start cost: 4.000
Gradient: (3.000, 5.000)
x, h=0.50: rate=3.000, predicted=1.500, actual=1.750, error=0.250
x, h=0.25: rate=3.000, predicted=0.750, actual=0.812, error=0.062
tangent, h=0.50: rate=0.000, predicted=0.000, actual=0.206, error=0.206
Python rounds the exact ties 0.8125 and 0.0625 to the nearest even final digit here. Their full values preserve the fourfold error ratio. The remainder calculation cost(delta) follows this polynomial’s exact expansion; other functions need their own error analysis.
Try it yourself
Exercise 1. At p = (1, −1), use unit direction u = (3/5, 4/5) and step length h = 0.5. Find the gradient, directional derivative, predicted change, endpoint, actual new cost, and prediction error.
Show solution 1
The starting cost is 2, and the gradient is (1, −3). The directional derivative is 1(3/5) − 3(4/5) = −9/5 = −1.8. The predicted change is −0.9.
The displacement is (0.3, 0.4), giving endpoint (1.3, −0.6). Direct evaluation gives 1.69 − 0.78 + 0.72 = 1.63, so the actual change is −0.37. The error is −0.37 − (−0.9) = 0.53, also equal to 0.3² + 0.3(0.4) + 2(0.4²).
Exercise 2. For g(x, y) = x² − y², find the gradient at the origin. Compare the local directional derivatives along x and y with the finite changes at endpoints (h, 0) and (0, h), for nonzero h. Does a zero gradient establish a minimum?
Show solution 2
The gradient is (2x, −2y), which becomes (0, 0) at the origin. Both unit directional derivatives are zero there. The finite changes are +h² along x and −h² along y.
Every neighborhood contains both higher and lower values. The origin is a saddle point, and the vanishing first-order terms do not distinguish it from a minimum. The signs of the remaining changes reveal the difference.
Continue with Jacobian matrices when one input change affects several outputs. Each output contributes its own row of partial derivatives.
Sources and further study
- MIT: The Tangent Plane and the Gradient Vector. Differentiability and the local linear model of a scalar function.
- MIT: The Gradient and Directional Derivatives. Partials with other inputs fixed and the dot product with a unit direction.
- Boyd and Vandenberghe: Unconstrained Minimization. Steepest descent under a chosen norm and its relationship to coordinate changes.