explainer
Taylor expansion and linearization: predict locally and check the error
Build constant, linear, and quadratic approximations around a chosen center. Compare their errors, calculate a Taylor remainder bound, and connect the same idea to gradients, Hessians, and Jacobians.
What you will learn
- Build constant, linear, and quadratic Taylor models at a chosen center.
- Compare a signed approximation error with a bound on its magnitude.
- Explain why a higher-order polynomial can still fail far from its center.
- Recognize the gradient, Hessian, and Jacobian versions of a local approximation.
Before you start
A linear approximation to exp(0.5) gives 1.5. Adding one quadratic term gives 1.625, closer to the actual value of about 1.648721. Taylor expansion builds these local models from derivatives at a chosen point.
You will calculate the models, test their errors, and find a target where adding that quadratic term makes the prediction worse. The useful question is how accurate a particular approximation is over the range you need.
Choose the point where the model must agree
Let a be the expansion center, and write the target as x = a + h. The displacement h can be positive, negative, or zero. All coefficients in the approximation come from derivatives at a.
For example, a = 0 and h = 0.5 target x = 0.5. Changing h tests the same local model at another input. Changing a rebuilds the model around a different point.
This lesson uses f(x) = eˣ with dimensionless x and output. Every derivative of this function also equals eˣ, which keeps the arithmetic focused on the approximation. The experiment evaluates a known mathematical function, with no fitted sensor data.
You can picture f as a synthetic nonlinear response to a normalized input. In a real application, you would first choose meaningful input scales and establish whether that function describes the system.
Match value, slope, and second derivative
The constant, linear, and quadratic Taylor polynomials are:
T₀(a + h) = f(a)
T₁(a + h) = f(a) + f′(a)h
T₂(a + h) = f(a) + f′(a)h + ½f″(a)h²
Each model retains another local property:
- Constant: matches the function value at a.
- Linear: also matches its first derivative at a.
- Quadratic: also matches its second derivative at a.
The factor ½ makes the second derivative agree. Differentiating ch² twice gives 2c, so the required coefficient is c = f″(a)/2. More generally, Taylor coefficients divide the kth derivative by k!, the product 1 × 2 × ⋯ × k.
The linear model is the tangent-line approximation from the derivative lesson. As a function of the original x, it generally includes a constant offset and is affine. The predicted increment f′(a)h is linear in h.
At a = 0, these are also called Maclaurin polynomials. An order can exceed the polynomial’s actual degree when a coefficient vanishes. For example, sin x has second derivative zero at the origin, so its first- and second-order Taylor polynomials there both equal x. OpenStax defines these polynomials and their derivative coefficients.
Approximate the exponential at one target
At a = 0, the exponential and its first two derivatives all equal 1. This gives:
T₀(h) = 1
T₁(h) = 1 + h
T₂(h) = 1 + h + h²/2
At h = 0.5, the quadratic correction is 0.5²/2 = 0.125. Add it to the linear value 1.5 to get 1.625. Direct evaluation gives exp(0.5) ≈ 1.648721.
Define the signed error as Rₙ = f − Tₙ. A positive error means the model predicts too low; a negative error means it predicts too high.
| Model | Predicted value | Signed error |
|---|---|---|
| Constant T₀ | 1.000000 | +0.648721 |
| Linear T₁ | 1.500000 | +0.148721 |
| Quadratic T₂ | 1.625000 | +0.023721 |
All three agree with the function exactly at h = 0. At this positive target, the quadratic model has the smallest absolute error. That comparison concerns this function, center, and displacement.
Move the target and rebuild the center
The experiment starts with the worked example. Its graph includes the actual function and all three polynomial models, while the readouts compare their target values.
Use the presets to make these comparisons:
- Small step: h = 0.1 gives actual value 1.105171, linear value 1.1, and quadratic value 1.105. The quadratic error falls to about 0.000171.
- Zero step: all models equal f(a), and both the error and its bound are zero.
- Shifted center: a = 1 and h = 0.5 target x = 1.5. Every derivative coefficient now uses e¹, about 2.718282.
At the shifted center, T₁ = 1.5e ≈ 4.077423 and T₂ = 1.625e ≈ 4.417208. The actual value is exp(1.5) ≈ 4.481689. Reusing the coefficients from a = 0 would create a different polynomial.
The plot follows the interval from center to target and rescales both axes. Judge errors from the numerical values when comparing states; the same gap on the screen can represent a different output difference. Six-decimal readouts can also round a small nonzero error to zero.
The Closer at this target readout compares the linear and quadratic models. The constant model remains visible as a separate baseline. You can explore centers from −1 to 1 and displacements from −3 to 3.
Put a bound on the omitted terms
Subtracting the approximation from the exact function gives the observed error. Taylor’s theorem can also bound that error using derivative information over the interval. For a function with continuous derivatives through order three on an interval containing a and a + h:
f(a + h) = T₂(a + h) + f‴(ξ)h³/6
for some ξ between a and a + h, when h ≠ 0
The intermediate point ξ is generally unknown. If |f‴(z)| ≤ M throughout that interval, the Lagrange remainder bound gives:
|R₂| ≤ M|h|³/6
For our exponential model, f‴(z) = eᶻ increases with z. The largest derivative on the segment occurs at its larger endpoint, giving M = exp(max(a, a + h)). This rule works for either sign of h.
At a = 0 and h = 0.5, the bound is exp(0.5)(0.5³)/6 ≈ 0.034348. The actual absolute error, about 0.023721, lies below it. A bound supplies a ceiling; equality is not required.
You can choose a uniform M for a whole planned range. For a = 0 and |h| ≤ 0.1, using M = exp(0.1) proves |R₂| ≤ exp(0.1)(0.1³)/6 ≈ 0.000184 across that range. This guarantees error below 0.001 for the exact function without checking each target individually.
With the same M, halving |h| divides the cubic bound by eight. Actual errors need not follow an exact eightfold ratio. Replacing M with a tighter interval-specific value changes the bound’s ratio too. OpenStax states Taylor’s remainder theorem and its interval requirement.
Check the range before trusting the order
Choose Far-left step, which keeps a = 0 and sets h = −3. The actual value remains positive, e⁻³ ≈ 0.049787, while the predictions are:
T₁(−3) = 1 − 3 = −2
T₂(−3) = 1 − 3 + 9/2 = 2.5
The linear absolute error is about 2.049787. The quadratic absolute error is about 2.450213, which is larger. At this target, even the constant prediction 1 has a smaller error than either of them.
The quadratic error bound is still valid: M = 1 on [−3, 0], so |R₂| ≤ 27/6 = 4.5. That ceiling permits an error too large for many uses. Mathematical validity and practical accuracy answer separate questions.
Moving the expansion center nearer the target can improve a low-order model. Adding terms is another option, but you must check its error. The first omitted term is an estimate unless a theorem supplies the needed derivative bound.
A finite Taylor polynomial also differs from an infinite Taylor series. For eˣ, the full series converges to the function at every finite real x. A fixed low-order polynomial can still be poor there; other functions need their own convergence analysis.
Smoothness alone does not guarantee that an infinite Taylor series reproduces its function. For local numerical work, state the approximation order and a usable error range. A formula at a corner also needs care: |x| has no ordinary first derivative at zero, so its tangent linearization there does not exist.
Connect the scalar formula to vectors
For a scalar function f of several inputs, use a column displacement Δ. The gradient supplies the first-order term, and the Hessian collects the second partial derivatives:
f(p + Δ) ≈ f(p) + ∇f(p)ᵀΔ + ½ΔᵀH(p)Δ
For n inputs, ∇f has shape n × 1 and H has shape n × n. Both correction terms produce a scalar. Continuous second derivatives near p give this second-order expansion with a remainder that becomes negligible relative to ‖Δ‖² as Δ shrinks. MIT’s optimization notes use this gradient-and-Hessian model.
Take the earlier cost f(x, y) = x² + xy + 2y² at p = (1, 1). Its gradient is (3, 5), and its Hessian has rows (2, 1) and (1, 4). For Δ = (0.1, −0.2), the linear change is −0.7 and the quadratic correction is 0.07.
The predicted value is 4 − 0.7 + 0.07 = 3.37. This prediction is exact because the original function is already quadratic. Its expansion has no higher-order remainder.
For a vector output F, Jacobian linearization gives F(p + Δ) ≈ F(p) + J(p)Δ. If F maps n inputs to m outputs, J has shape m × n. Each row predicts the change in one output.
For states constrained to a curved set, a local direction also has to respect that constraint. The manifolds and tangent spaces lesson shows why a tangent vector can be valid even though adding a finite multiple of it leaves the set.
Keep approximation error separate from model error
Taylor error measures the gap between a function and its local polynomial. It does not measure whether the original function describes a real sensor, actuator, or loss surface well. More Taylor terms cannot repair a wrong physical assumption.
Units must also agree across the expansion. If x has input units, f′ has output units per input unit, and f″ has output units per input unit squared. Multiplying by h and h² makes both corrections compatible with f(a).
In robot dynamics, an ordinary differential equation describes how a state changes over time. Linearizing its rate function around a fixed equilibrium uses Jacobians with respect to state and input. The equilibrium makes the constant rate term zero; expansion around a general fixed point retains that constant term.
The University of Illinois control-systems notes derive this local model for small deviations around an equilibrium. A local rate approximation does not guarantee that a long simulated trajectory stays accurate. The state can leave the region where the approximation works.
Reproduce the calculation in Python
This standard-library example builds the two polynomial predictions and checks their signed errors. Its bound uses the maximum third derivative along the complete segment, including for a negative displacement.
from math import exp
def inspect(a, h):
base = exp(a)
target = a + h
actual = exp(target)
linear = base * (1 + h)
quadratic = base * (1 + h + h*h/2)
maximum_third_derivative = exp(max(a, target))
bound = maximum_third_derivative * abs(h)**3 / 6
return actual, linear, quadratic, bound
for a, h in [(0.0, 0.5), (0.0, 0.1), (0.0, -3.0), (1.0, 0.5)]:
actual, linear, quadratic, bound = inspect(a, h)
print(f"a={a:.1f}, h={h:.1f}: actual={actual:.6f}, "
f"linear={linear:.6f}, quadratic={quadratic:.6f}")
print(f" errors: linear={actual-linear:.6f}, "
f"quadratic={actual-quadratic:.6f}; bound={bound:.6f}")
Expected output:
a=0.0, h=0.5: actual=1.648721, linear=1.500000, quadratic=1.625000
errors: linear=0.148721, quadratic=0.023721; bound=0.034348
a=0.0, h=0.1: actual=1.105171, linear=1.100000, quadratic=1.105000
errors: linear=0.005171, quadratic=0.000171; bound=0.000184
a=0.0, h=-3.0: actual=0.049787, linear=-2.000000, quadratic=2.500000
errors: linear=2.049787, quadratic=-2.450213; bound=4.500000
a=1.0, h=0.5: actual=4.481689, linear=4.077423, quadratic=4.417208
errors: linear=0.404266, quadratic=0.064481; bound=0.093369
For extremely small h, subtracting nearly equal floating-point values can erase an error numerically. The experiment computes small exponential remainders without that subtraction; the short Python example uses direct subtraction at the moderate steps shown. A printed zero error does not prove exact agreement.
Try it yourself
Exercise 1. Approximate exp(−0.2) with constant, linear, and quadratic polynomials centered at a = 0. Find their signed errors using the actual value 0.8187307531. Then bound the absolute quadratic error using the largest third derivative on [−0.2, 0].
Show solution 1
The predictions are T₀ = 1, T₁ = 0.8, and T₂ = 1 − 0.2 + 0.04/2 = 0.82. Their signed errors are approximately −0.181269, +0.018731, and −0.001269.
The maximum third derivative is M = e⁰ = 1, so the bound is 0.2³/6 ≈ 0.001333. The absolute quadratic error lies below it. Using exp(−0.2) as M would miss the larger derivative values nearer zero.
Exercise 2. Let f(x) = x² + 3x + 2. Build its linear and quadratic approximations at a = 1, then evaluate them at x = 1.4. Which prediction is exact, and what does the third derivative tell you about the remainder?
Show solution 2
At a = 1, f(a) = 6, f′(a) = 5, and f″(a) = 2. The displacement is h = 0.4. The linear prediction is 6 + 5(0.4) = 8, and the quadratic prediction adds ½(2)(0.4²) = 0.16 to give 8.16.
Direct evaluation gives 1.4² + 3(1.4) + 2 = 8.16. The quadratic model is exact for every target because f‴ is zero everywhere. The remainder bound is zero, while the linear model omits h².
Continue with Hessians to see how the quadratic correction changes with direction. Keep the expansion center, input scales, and permitted error alongside the formula.
Sources and further study
- OpenStax: Taylor and Maclaurin Series. Taylor coefficients, polynomial approximations, remainder bounds, and the separate question of series convergence.
- MIT: Hessians, Preconditioning, and Newton’s Method. The multivariable second-order model built from a gradient and Hessian.
- University of Illinois: Linearization of State-Space Models. Jacobian linearization around an equilibrium and its local range of use.