Hessians: measure curvature in every direction

Differentiate a gradient to build the Hessian, calculate directional curvature, and classify stationary points. Explore coupled quadratics, saddles, and the limits of zero eigenvalues.

By 11 min read

What you will learn

  • Build a Hessian by differentiating each component of a scalar function's gradient.
  • Calculate how a step changes the gradient and contributes a quadratic output change.
  • Read directional curvature and stationary-point classifications from a symmetric Hessian.
  • Explain why zero eigenvalues can leave the second-derivative test inconclusive.

Before you start

A gradient predicts a small change in a function. The Hessian predicts how that gradient changes as you move. Its eigenvalues can also reveal whether a stationary point sits at a minimum, maximum, or saddle.

For f(x, y) = x² + xy + 2y², a step from (1, 1) to (1.5, 1) changes the gradient from (3, 5) to (4, 5.5). We will calculate that change and the extra output change that a linear prediction misses.

Measure how the gradient changes

Write the input as a column p and its displacement as Δp. The gradient ∇f(p) is also a column. It contains the first partial derivatives of the scalar output f.

The Hessian H(p) is the derivative of that gradient map. It gives the local prediction:

∇f(p + Δp) ≈ ∇f(p) + H(p)Δp

This is a Jacobian whose output happens to be a gradient. For n scalar input coordinates, H has n rows and n columns. Cornell's calculus notes place the Hessian in this second-derivative framework.

Our examples use dimensionless, scaled coordinates. A Hessian entry has units of output divided by the product of its two input units. For a cost in joules and two coordinates in meters, its entries have units J/m².

In robot calibration, the inputs might be joint offsets and link lengths, with a cost measuring disagreement between predicted and observed positions. The Hessian describes how the cost's gradient changes when those parameters move, including interactions between parameters. For a smooth machine-learning loss, the same calculation applies to model weights. Both uses require deliberate parameter scaling. A Hessian evaluated at one point supplies a local model; it does not describe every part of a nonlinear cost.

Differentiate every gradient component

For two inputs, arrange the second partial derivatives as:

H =
[fₓₓ, fₓᵧ;
fᵧₓ, fᵧᵧ]

Here the first row records changes in ∂f/∂x, and the second records changes in ∂f/∂y. The columns correspond to changing x or y. We use functions with continuous second partial derivatives, so the two mixed entries agree.

Start with f(x, y) = x² + xy + 2y², the coupled cost from the gradient lesson. Its gradient is (2x + y, x + 4y)ᵀ, giving:

Gradient componentDifferentiate with respect to xDifferentiate with respect to y
2x + y21
x + 4y14

Thus H = [2, 1; 1, 4]. This quadratic has a constant Hessian. For a general function, the entries can change with the input point.

Read the mixed terms and symmetry

The off-diagonal value 1 means changing y changes the x-gradient component, and changing x changes the y-gradient component. The xy term couples the two inputs. Looking only at fₓₓ and fᵧᵧ would miss that interaction.

Continuous mixed second partials agree by Clairaut's theorem, making H symmetric. OpenStax states the continuity condition. Mere existence of second partials at one point does not guarantee that equality.

A symmetric H defines the quadratic form ΔpᵀHΔp. If H = [a, b; b, c], expanding it gives aΔx² + 2bΔxΔy + cΔy². Both off-diagonal entries contribute to the mixed term.

For our exact quadratic, f(p) = ½pᵀHp. The one-half factor turns H's diagonal entries 2 and 4 into the coefficients 1 and 2, and its two mixed contributions into the coefficient 1 on xy.

Take a slice in a chosen direction

Choose a unit vector u and follow the straight line p + su. The scalar parameter s measures signed distance in the chosen coordinate scale. Applying the chain rule twice gives:

d/ds f(p + su) at s = 0: ∇f(p) · u
d²/ds² f(p + su) at s = 0: uᵀH(p)u

We call uᵀHu directional curvature here. It is the second derivative along a straight input-space slice. This definition concerns how the slope changes; it is not the arc-length curvature of a drawn curve.

Positive curvature means that slice's slope increases locally. Negative curvature means it decreases. A function can still decrease along a direction with positive curvature if its first derivative in that direction is negative.

Unit length matters when comparing directions. Doubling a direction vector multiplies its second directional derivative by four. Changing coordinate units also changes numerical curvature values, so comparisons need a stated scaling.

Calculate a complete quadratic change

Set p = (1, 1) and Δp = (0.5, 0). The starting gradient is (3, 5), and the Hessian remains [2, 1; 1, 4].

Multiplying H by the displacement gives HΔp = (1, 0.5). Adding this to the old gradient produces (4, 5.5). Direct differentiation at the new point (1.5, 1) gives the same result.

The scalar output change has two contributions:

Δf = ∇f(p) · Δp + ½ΔpᵀHΔp
= 1.5 + 0.25 = 1.75

The first term is 3(0.5) + 5(0) = 1.5. The quadratic term is ½(0.5, 0) · (1, 0.5) = 0.25. The output rises from 4 to 5.75.

The equality is exact for this quadratic, apart from computer rounding. A smooth nonlinear function generally has a higher-order remainder. The Taylor expansion lesson develops that distinction.

Compare bowls, a saddle, and a flat direction

The experiment starts at the origin of the coupled minimum. At that point, the linear contribution is zero and the selected step's entire change comes from the quadratic term.

Interactive experiment

Inspect curvature and the changing gradient

Choose a quadratic, then compare its gradient change and directional curvature. Every example has a stationary origin.

0.000
0.000
0°
0.500

f(x, y) = x² + xy + 2y². Directions use 0° along +x and positive angles toward +y.

Contours and the chosen step

f(x, y) = x² + xy + 2y². Start (0.000, 0.000), endpoint (0.500, 0.000).The x and y axes have equal scales. Solid blue contours have value 1 or 4, dashed amber contours −1 or −4, and dotted charcoal contours zero. A solid dot marks the start and an open circle the endpoint. The line connects them.-2-2022xy
Equal coordinate scales. Blue solid levels: 1 and 4. Amber dashed levels: −1 and −4. Charcoal dotted level: 0. Only levels the function reaches appear; an isolated zero-level origin has no contour line.

Slice along the unit direction

Output change along p + s u, with u = (1.000, 0.000).The solid curve is f(p + s u) minus f(p). The dashed line is the linear change. At s = 0.500, the actual change is 0.250. The directional second derivative is 2.000. Vertical range adjusts; horizontal and vertical axes have different units.Δf0.00.51.0-101s
Solid black: actual output change along the line. Dashed blue: linear prediction. Amber dot: the selected step. The gap comes from the quadratic term. The vertical scale adjusts between examples.
Hessian H = derivative of the column gradient
Gradient outputChange xChange y
∂f/∂x2.0001.000
∂f/∂y1.0004.000
Eigenvalues
(1.586, 4.414)
Test at the origin
Strict minimum
Gradient at start
(0.000, 0.000)
Unit direction
(1.000, 0.000)
Directional curvature
2.000
Gradient change
(1.000, 0.500)
Linear change
0.000
Quadratic change
0.250
Actual change
0.250
Step endpoint
(0.500, 0.000)

The stationary origin is a strict minimum. At start (0.000, 0.000), the chosen direction has curvature 2.000. Linear plus quadratic change equals 0.250.

The origin test stays at (0, 0) when you move the start. All four Hessians are constant. For these exact quadratics, Δ∇f = HΔp and Δf = ∇f · Δp + ½ΔpᵀHΔp. Coordinates are dimensionless. Calculations retain precision; displays round to three decimals.

Choose Use worked point to reproduce the calculation above. Moving the start changes the gradient but leaves these quadratic Hessians fixed. The readout Test at the origin always concerns (0, 0), even when your selected start moves elsewhere.

The Coupled maximum reverses every sign. The Saddle has f(x, y) = x² − y²: a slice along x curves upward, and a slice along y curves downward. Change the direction angle to see both signs.

At 45° in the saddle, uᵀHu is zero because the positive and negative contributions cancel. Its Hessian still has no zero eigenvalue. A zero-curvature direction alone does not establish a null direction of H.

The Flat direction uses f(x, y) = x². Along y, both the function and gradient stay unchanged. Its zero eigenvalue calls for the separate analysis below.

Classify a stationary point with eigenvalues

First locate a stationary point, where ∇f = 0. For an interior point with continuous second partials nearby, the eigenvalues of its symmetric Hessian give these conclusions:

Eigenvalue signsSecond-derivative conclusion
All positiveStrict local minimum
All negativeStrict local maximum
At least one positive and one negativeSaddle
Zero eigenvalues, with all remaining values of one signInconclusive
All zeroInconclusive

Cornell describes this test through the signs of the Hessian's eigenvalues. Positive and negative definite mean that the quadratic form has the respective strict sign for every nonzero direction. An indefinite Hessian admits both signs.

In an orthonormal eigenvector basis, uᵀHu = Σᵢ λᵢcᵢ², where u has coordinates cᵢ and Σᵢcᵢ² = 1. The smallest and largest eigenvalues bound the directional curvature, and their eigenvectors attain those bounds.

Our minimum has eigenvalues 3 − √2 and 3 + √2, about 1.586 and 4.414. Both are positive. The exact quadratic has its unique global minimum at the origin; the general test establishes a local conclusion.

A positive definite Hessian at a point with a nonzero gradient does not make that point a minimum. Constraints also change which directions are available. Apply this unconstrained test only under its stated assumptions.

Recognize an inconclusive test

For f(x, y) = x², H = [2, 0; 0, 0]. The Hessian is positive semidefinite: uᵀHu is nonnegative, but some nonzero directions give zero. The second-derivative test is inconclusive.

The full formula settles this example. Every point with x = 0 has value zero, and no point has a smaller value. Each is a non-strict global minimum because nearby points on that line share its value.

Now compare two functions with the same Hessian at the origin:

  • x² + y⁴ has a strict minimum there, since every other point has positive value.
  • x² − y⁴ has a saddle there, with positive values along x and negative values along y.

The fourth-order terms decide what the identical Hessians cannot. A strict minimum therefore does not require a positive definite Hessian. OpenStax marks the zero-determinant case of its two-variable test as inconclusive.

Near-zero computed eigenvalues raise an extra numerical question: are they zero in the model or only small? This lab uses four fixed, analytically known spectra. It does not apply a numerical threshold to classify an arbitrary measured Hessian.

Reproduce the calculation in Python

This standard-library example evaluates the quadratic form and its gradient. The eigenvalue formula applies to a symmetric two-by-two matrix.

from math import hypot


def dot(a, b):
    return sum(x * y for x, y in zip(a, b))


def multiply(matrix, vector):
    return tuple(dot(row, vector) for row in matrix)


def value(matrix, point):
    return 0.5 * dot(point, multiply(matrix, point))


def eigenvalues(matrix):
    a, b = matrix[0]
    c = matrix[1][1]
    middle = (a + c) / 2
    radius = hypot((a - c) / 2, b)
    return (middle - radius, middle + radius)


def pair(values):
    return "(" + ", ".join(f"{v:.3f}" for v in values) + ")"


H = ((2, 1), (1, 4))
p, delta = (1, 1), (0.5, 0)
end = tuple(a + b for a, b in zip(p, delta))
gradient = multiply(H, p)
gradient_change = multiply(H, delta)
linear = dot(gradient, delta)
quadratic = 0.5 * dot(delta, gradient_change)
actual = value(H, end) - value(H, p)
print("Eigenvalues:", pair(eigenvalues(H)))
print("Start gradient:", pair(gradient))
print("Gradient change:", pair(gradient_change))
print("End gradient:", pair(multiply(H, end)))
print(f"Linear={linear:.3f}, quadratic={quadratic:.3f}, actual={actual:.3f}")

for name, matrix in [("saddle", ((2, 0), (0, -2))),
                     ("flat", ((2, 0), (0, 0)))]:
    u = (0, 1)
    curvature = dot(u, multiply(matrix, u))
    print(f"{name}: eigenvalues={pair(eigenvalues(matrix))}, "
          f"y curvature={curvature:.3f}")

Expected output:

Eigenvalues: (1.586, 4.414)
Start gradient: (3.000, 5.000)
Gradient change: (1.000, 0.500)
End gradient: (4.000, 5.500)
Linear=1.500, quadratic=0.250, actual=1.750
saddle: eigenvalues=(-2.000, 2.000), y curvature=-2.000
flat: eigenvalues=(0.000, 2.000), y curvature=0.000

Try it yourself

Exercise 1. For f(x, y) = x² + xy + 2y² at the origin, use the unit direction u = (1, 1)/√2 and step length 0.5. Find the directional curvature, gradient change, and output change.

Show solution 1

The Hessian is [2, 1; 1, 4]. Multiplication gives Hu = (3, 5)/√2, so uᵀHu = (3 + 5)/2 = 4.

The displacement is Δp = (0.5, 0.5)/√2. The gradient change is HΔp = (1.5, 2.5)/√2 ≈ (1.061, 1.768).

The initial gradient is zero, so the linear change vanishes. The output change is ½(0.5)²(4) = 0.5, exactly for this quadratic.

Exercise 2. At the origin, compare f(x, y) = x² + y⁴ and g(x, y) = x² − y⁴. Find both gradients and Hessians there. What conclusions require information beyond the Hessians?

Show solution 2

Both gradients are (0, 0). The y derivatives are ±4y³ and ±12y², which vanish at zero. Both Hessians are therefore [2, 0; 0, 0], with eigenvalues 0 and 2.

The second-derivative test is inconclusive for each. Directly inspecting f gives a strict minimum: it is positive at every nonzero point. Inspecting g gives a saddle: g(x, 0) = x² is positive, while g(0, y) = −y⁴ is negative away from zero.

Continue with Taylor expansion and linearization to compare first- and second-order approximations when the Hessian changes across the domain.

Sources and further study