Derivatives and the chain rule: predict a small change

Understand derivatives as local rates, compare secant and tangent slopes, and multiply the correct factors through nested functions. Test a smooth curve and a corner with an interactive experiment.

By 10 min read

What you will learn

  • Explain a derivative as a limit of finite-step slopes and attach the correct units.
  • Use a derivative to predict a nearby output and measure its approximation error.
  • Differentiate a nested function by evaluating and multiplying its inner and outer factors.
  • Recognize a corner where the ordinary derivative does not exist.

Before you start

  • Function notation, substitution, and basic algebra
  • A line's slope as change in output divided by change in input

A derivative predicts how an output responds to a small input change. For f(x) = (x² + 1)², the derivative at x = 1 is 8. A step of 0.01 therefore predicts an output increase of 0.08; the actual increase is 0.08080401.

That small gap is part of the lesson. You will calculate the local rate, follow it through nested functions, and see where the prediction works.

Give the rate a meaning and units

Suppose a robot moves along a straight rail with position s(t) = k t², where k = 1 m/s². At t = 2 seconds, its position is 4 meters. At t = 2.1 seconds, its position is 4.41 meters.

The average velocity over that interval is 0.41 m / 0.1 s = 4.1 m/s. Its instantaneous velocity at t = 2 is s′(2) = 4 m/s. The prime mark means “derivative”; the next section shows how that value follows from a limit.

Derivative units are output units divided by input units. A position derivative with respect to time has velocity units. A motor-current derivative with respect to voltage has units A/V; time need not be the input.

The sign also carries meaning. A negative velocity means position decreases along the chosen axis. Speed is the magnitude of velocity, so the two words describe different quantities.

Bring two points together

Choose a base input x and a nonzero step h. The line through (x, f(x)) and (x + h, f(x + h)) is a secant line. Its slope is the difference quotient:

secant slope = [f(x + h) − f(x)] / h, h ≠ 0

The derivative is the finite value these slopes approach as h approaches zero from both signs:

f′(x) = limₕ→₀ [f(x + h) − f(x)] / h

The quotient uses nonzero steps throughout the limit. Substituting h = 0 directly gives division by zero. OpenStax defines the derivative through this limiting secant slope.

For f(x) = x², expand the numerator:

[(x + h)² − x²] / h
= (2xh + h²) / h = 2x + h
f′(x) = 2x

Multiplying x² by a constant k multiplies the derivative by k. This gives the rail model's velocity s′(t) = 2kt, including its units.

Predict a nearby output

At a differentiable point, the derivative supplies the tangent line's slope. A step h along that line predicts:

f(x + h) ≈ f(x) + f′(x)h

The term f′(x)h predicts the change; adding f(x) gives the predicted output. OpenStax calls this the linear approximation or linearization.

For the rail example, the prediction at t = 2.1 is 4 + 4(0.1) = 4.4 meters. The actual position is 4.41 meters, leaving a 0.01-meter error. Curvature creates that gap.

Differentiability means the error divided by |h| approaches zero as h approaches zero. It does not promise a small error for every step. A local prediction also cannot establish how well a motion model matches a real robot.

The Taylor expansion and linearization lesson adds a curvature term to this prediction and compares the error as the step grows. That comparison helps you decide how far a local approximation is useful.

Follow the inner and outer functions

A nested function sends x through an intermediate value u:

u = g(x), y = F(u)
y = F(g(x))

If g is differentiable at x and F is differentiable at g(x), the chain rule gives:

dy/dx = (dy/du)(du/dx)
= F′(g(x))g′(x)

Evaluate the outer derivative at the inner output g(x). Then multiply by the inner derivative at x. OpenStax states the chain rule with these differentiability conditions.

The local predictions explain the product: Δu ≈ g′(x)Δx, then Δy ≈ F′(u)Δu. Substituting the first into the second multiplies their rate factors. Their units combine as (y units/u units) × (u units/x units).

The fraction-like notation helps track those units. The justification comes from differentiability; treating the symbols as ordinary fractions can hide missing assumptions.

Calculate both chain-rule factors

For y = (x² + 1)², use u = x² + 1 and F(u) = u². The inner derivative is 2x, because adding a constant changes no output differences. The outer derivative is 2u.

At x = 1:

QuantityCalculationValue
Inner value u1² + 12
Output y2²4
Inner derivative du/dx2(1)2
Outer derivative dy/du2(2)4
Complete derivative dy/dx4 × 28

For general x, dy/dx = 4x(x² + 1). Dropping the inner factor would give 4 at x = 1, only half the correct rate.

Take h = 0.25. The tangent predicts 4 + 8(0.25) = 6. Direct evaluation gives f(1.25) = 6.56640625, so the secant slope is 2.56640625 / 0.25 = 10.265625.

These values answer separate questions: 8 is the local rate at the base point, while 10.265625 measures the average change across this interval.

Compare the derivative with a finite step

The experiment starts with that calculation. Its nested-square example uses dimensionless inputs and outputs.

Interactive experiment

Compare a tangent with a finite step

Follow two chain-rule factors, then test how a tangent predicts the output after a finite input change.

1.000
0.250

u = x² + 1; y = u². Chain rule: dy/dx = (2u)(2x).

  • Solid black: function
  • Long blue dashes: tangent
  • Short amber dashes: secant
Function, tangent, and finite step for f(x) = (x² + 1)²f(x) = (x² + 1)². Base (1.000, 4.000), actual endpoint (1.250, 6.566). Derivative 8.000; secant slope 10.266; tangent prediction 6.000. The view follows x and rescales the vertical axis. Horizontal and vertical units have different screen scales.f(x)-1.26.313.90.51.01.5x
The view follows the base point and rescales vertically. Read the numerical slopes; axis units use different screen scales. An open circle marks the base, an amber dot the actual endpoint, and an open square the prediction. Coincident markers nest.
Base value f(x)
4.000
Inner value u
2.000
Inner derivative du/dx
2.000
Outer derivative dy/du
4.000
Derivative dy/dx
8.000
Secant slope
10.266
Actual value f(x + h)
6.566
Linear prediction
6.000
Prediction error
0.566

At x = 1.000, the derivative is 8.000. The finite-step secant slope is 10.266; actual minus predicted output is 0.566.

Prediction = f(x) + f′(x)h. Error = actual value minus prediction. The mathematical examples use dimensionless x and y. Values display three decimals; calculations retain their precision.

Choose Smaller step to set h = 0.01. The secant slope becomes 8.080401, closer to the derivative 8. The actual output becomes 4.08080401, close to the prediction 4.08.

Choose Zero inner derivative to move the base to x = 0. The inner rate becomes zero, so the full rate is zero even though the outer factor is 2. At h = 0.25 the function still increases from 1 to 1.12890625.

A zero derivative describes a first-order change at one point. It does not make a function constant nearby. In optimization, a zero derivative alone also does not identify a minimum: x³ has derivative zero at the origin and keeps increasing through it.

The graph follows the base point and rescales vertically. Compare the numerical slopes when moving between states. At h = 0, the secant readout becomes Undefined while any existing derivative remains available.

Check where a derivative exists

Choose At the corner for f(x) = |x| at x = 0. Positive steps give secant slope +1; negative steps give −1. The two sides approach different values, so the ordinary two-sided derivative does not exist.

The function remains continuous there. Differentiability implies continuity, but continuity alone does not imply differentiability. OpenStax uses absolute value to explain this distinction.

The lab draws no tangent or derivative-based prediction at that corner. This differs from the horizontal tangent of the smooth nested function at zero.

Choose Across the corner to start at x = 0.1 with h = −0.25. The derivative at the base exists and equals 1. Extending its tangent across zero predicts −0.15, while the actual absolute value is +0.15.

That step crosses a change in the function's rule. Local differentiability at the starting point gives no promise that its tangent will remain accurate across the corner.

If a chain-rule factor is undefined, analyze the composite separately. The stated theorem requires both derivatives.

Use finite differences with care

A finite difference estimates a derivative by evaluating a function at nearby inputs. Our secant uses a forward or backward step depending on the sign of h. The central difference uses [f(x + h) − f(x − h)] / (2h).

Neither calculation alone proves differentiability. For |x| at zero, the central difference equals zero for every nonzero h, even though the two-sided derivative is undefined. Checking both one-sided slopes exposes the corner.

Very small steps also meet the limits of floating-point arithmetic. Adding h may leave the stored x unchanged, and subtracting nearby outputs can lose precision. SciPy documents the finite-precision limits of numerical differentiation.

Check several step sizes and use an analytic derivative when available. A derivative can guide gradient descent, while the chosen step still determines whether a finite update lowers the loss.

Reproduce the calculation in Python

This example uses ordinary Python arithmetic. It calculates the chain-rule factors analytically and compares them with finite differences.

def f(x):
    return (x * x + 1) ** 2


def chain_factors(x):
    u = x * x + 1
    inner = 2 * x
    outer = 2 * u
    return u, inner, outer, inner * outer


x = 1.0
u, inner, outer, derivative = chain_factors(x)
print(f"u={u:.3f}, inner={inner:.3f}, outer={outer:.3f}")
print(f"derivative={derivative:.3f}")
for h in [0.25, 0.01, -0.01]:
    actual = f(x + h)
    prediction = f(x) + derivative * h
    secant = (actual - f(x)) / h
    print(f"h={h:+.2f}: secant={secant:.6f}, "
          f"actual={actual:.6f}, prediction={prediction:.6f}")

print(f"derivative at zero={chain_factors(0.0)[3]:.3f}")
for h in [-0.1, 0.1]:
    print(f"abs at zero, h={h:+.1f}: secant={abs(h) / h:+.1f}")

Expected output:

u=2.000, inner=2.000, outer=4.000
derivative=8.000
h=+0.25: secant=10.265625, actual=6.566406, prediction=6.000000
h=+0.01: secant=8.080401, actual=4.080804, prediction=4.080000
h=-0.01: secant=7.920399, actual=3.920796, prediction=3.920000
derivative at zero=0.000
abs at zero, h=-0.1: secant=-1.0
abs at zero, h=+0.1: secant=+1.0

Try it yourself

Exercise 1. Let y = (3x + 1)². Find dy/dx at x = 1, predict y after a step h = 0.1, and calculate the actual output and prediction error. Which factor would disappear if you forgot the inner derivative?

Show solution 1

The inner value is u = 3(1) + 1 = 4, so y = 16. Its inner derivative is 3, and the outer derivative is 2u = 8. Multiplying gives dy/dx = 24.

The prediction is 16 + 24(0.1) = 18.4. Direct evaluation gives (3(1.1) + 1)² = 4.3² = 18.49, so actual minus predicted output is 0.09. Omitting the inner factor 3 would give an incorrect derivative of 8.

Exercise 2. For f(x) = |x| at x = 0, calculate the forward, backward, and central difference estimates with step magnitude 0.01. Does the central estimate establish a derivative?

Show solution 2

The forward estimate is 0.01 / 0.01 = +1. The backward estimate is 0.01 / (−0.01) = −1. Their disagreement persists as the step magnitude shrinks.

The central estimate is (0.01 − 0.01) / 0.02 = 0. Symmetry cancels the numerator, but the two-sided derivative does not exist. A central estimate cannot settle that question by itself.

Continue with partial derivatives and gradients when an output depends on several inputs. For several outputs as well, Jacobian matrices organize the local rates and extend the chain rule to matrix multiplication.

Sources and further study