explainer
Linear regression: fit a line and inspect its errors
Fit linear regression to robot sensor readings. Work through least squares, residuals, and MSE, then test an outlier with Python and exercises.
What you will learn
- Predict a continuous value with a slope and intercept.
- Calculate residuals and mean squared error, then check a least squares fit.
- Explain how an outlier and new operating conditions limit a fitted line.
Before you start
- Arithmetic and reading an x-y graph
- Dot products, for the connection to several input features
Linear regression predicts a numeric value by fitting a line to examples. The line gives a simple rule you can inspect. Its errors show where that rule needs closer attention.
Imagine comparing a robot's distance sensor with reference measurements. For each reported reading, record the reference distance. A fitted line can correct an offset and a scale difference over the range you measured.
This lesson uses five synthetic readings. They illustrate the calculation; they do not establish the performance of a real sensor.
Turn a reading into a prediction
Let x be the sensor's reported distance and y the reference distance. Both use metres. Predict y with:
ŷ = mx + b
The hat in ŷ marks a prediction. The slope m sets how much the prediction changes when x increases by one. The intercept b gives the prediction at x = 0.
For m = 0.84 and b = 0.32, a reading of 2.5 m gives 0.84 × 2.5 + 0.32 = 2.42 m. Here the slope is a ratio of metres to metres, so it has no units. The intercept has units of metres.
With several input features, the prediction becomes ŷ = w · x + b. The dot product combines the features with their weights. This lesson keeps one input so every calculation fits on a graph.
Measure the errors with MSE
A residual is the observed value minus the prediction: e = y − ŷ. A positive residual means the model predicts too little. A negative residual means it predicts too much.
Adding residuals alone can hide mistakes: +2 and −2 cancel. Mean squared error, or MSE, squares each residual, adds the squares, and divides by the number of observations:
MSE = (1/n) ∑(yi − ŷi)²
The index i identifies one observation; n counts them. With distances in metres, MSE has units of square metres. Taking its square root gives RMSE, in metres. Google's loss guide explains these two measures.
Squaring gives large errors more influence. A residual of 2 contributes four times as much squared error as a residual of 1. Choose the loss with the cost of your real prediction errors in mind.
The Laplace distribution lesson connects absolute error to a model for measurement errors. For a constant location estimate under that model, minimizing absolute error gives a sample median.
Adjust the calibration line
Start with “Calibration readings.” The initial line is ŷ = x, which treats each sensor reading as the reference distance. Its training MSE is 0.1240 m².
Move the slope and intercept controls. Try to lower the MSE before pressing “Fit least squares.” Then compare your line with the computed fit.
Switch to “One changed reading” after fitting. The last observation moves from (4, 3.4) to (4, 5.4). Your current line stays in place, so you can see the new error before fitting again.
The new fit has m = 1.24 and b = −0.08. One changed observation moves predictions across the entire input range. The negative prediction at x = 0 also shows that this model imposes no physical distance constraint.
An unusual observation could reflect a recording error or a real condition. Investigate it before deciding how to handle it. NIST describes this sensitivity to outliers.
Calculate the least squares fit
Ordinary least squares chooses the line with the smallest sum of squared residuals. Dividing that sum by a fixed n leaves the minimizing line unchanged, so it also minimizes MSE.
For a straight line with an intercept, calculate the input mean x̄ and output mean ȳ. Then use:
m = ∑[(xi − x̄)(yi − ȳ)] / ∑(xi − x̄)²
b = ȳ − mx̄
These formulas follow from setting both partial derivatives of squared error to zero. NIST gives the derivation and estimators.
For the original five readings, x̄ = 2 and ȳ = 2. The numerator is 8.4, and the denominator is 10. Therefore m = 8.4 / 10 = 0.84 and b = 2 − 0.84 × 2 = 0.32.
Check each prediction and residual:
| x | y | ŷ = 0.84x + 0.32 | y − ŷ |
|---|---|---|---|
| 0 | 0.2 | 0.32 | −0.12 |
| 1 | 1.3 | 1.16 | 0.14 |
| 2 | 1.8 | 2.00 | −0.20 |
| 3 | 3.3 | 2.84 | 0.46 |
| 4 | 3.4 | 3.68 | −0.28 |
The squared residuals sum to 0.364. Divide by five to get MSE = 0.0728 m². The RMSE is approximately 0.2698 m.
Gradient descent can minimize this loss through repeated updates. The formula above solves this particular fit directly. For several features, numerical libraries commonly solve the least squares problem using matrix factorizations. The projections and least squares lesson explains the geometry, and singular value decomposition exposes independent directions in the data.
Inspect what the line misses
The residual plot puts each error against its input. An upward trend suggests an incorrect slope. A curve can suggest that a straight line misses a systematic relationship. A changing spread can indicate that the error scale varies with the input.
Five points provide little evidence about those patterns. Use more measurements before drawing strong conclusions. NIST's model validation guide shows why residual plots complement a single score.
Fitting a line requires no assumption of normally distributed residuals. Statistical confidence intervals require additional assumptions about the data and errors. The widget reports a point fit and training loss; it does not calculate uncertainty intervals.
Test on new readings
Every displayed observation helps choose the fitted coefficients. A smaller training MSE alone cannot establish accuracy on another robot run.
Fit coefficients using training data. Use validation data to choose between modeling approaches, then assess the chosen approach on untouched test data. Estimate any preprocessing values from the training data, too. Scikit-learn explains how information leaking across that boundary inflates evaluation results.
The pandas lesson shows how to select sensor readings, handle missing values, and check joins before those rows reach a model. Preserving the run labels helps you split the data deliberately.
Nearby sensor readings from one run can share conditions and errors. If the goal is performance on new runs, keep whole runs together when splitting data. Choose a split that reflects how the model will face new observations.
The fitted line predicts 5.36 m at x = 6. Our observations only cover x from 0 to 4, so that prediction is extrapolation. The arithmetic still works, but the data do not establish that the relationship continues there.
Even within the measured range, a different surface or temperature can change performance. A low loss under one set of conditions needs fresh checks under another.
For an outcome such as “obstacle present” or “obstacle absent,” logistic regression maps a linear score to a probability. Its lesson also separates fitting that probability model from choosing a decision threshold.
Reproduce the fit in Python
This example uses only the Python standard library. It fits the original training readings and reports their loss.
from statistics import mean
from math import sqrt
data = [(0, 0.2), (1, 1.3), (2, 1.8), (3, 3.3), (4, 3.4)]
x_bar = mean(x for x, y in data)
y_bar = mean(y for x, y in data)
sxx = sum((x - x_bar) ** 2 for x, y in data)
if sxx == 0:
raise ValueError("At least two distinct x values are required")
m = sum((x - x_bar) * (y - y_bar) for x, y in data) / sxx
b = y_bar - m * x_bar
residuals = [y - (m * x + b) for x, y in data]
mse = mean(error ** 2 for error in residuals)
print(f"slope={m:.2f}, intercept={b:.2f}")
print(f"training MSE={mse:.4f}, RMSE={sqrt(mse):.4f}")
print(f"prediction at x=2.5: {m * 2.5 + b:.2f}")
Expected output:
slope=0.84, intercept=0.32
training MSE=0.0728, RMSE=0.2698
prediction at x=2.5: 2.42
Try it yourself
Exercise 1. Keep the original slope of 0.84, but increase the intercept from 0.32 to 0.42. Predict how every residual changes. Calculate the new MSE.
Show solution: shift every prediction
Every prediction increases by 0.10 m, so every residual decreases by 0.10 m. The new residuals are −0.22, 0.04, −0.30, 0.36, −0.38.
Their squares sum to 0.414. The MSE is 0.0828 m², higher than the fitted minimum of 0.0728 m². Check the result with the intercept control.
Exercise 2. The changed dataset has a fitted MSE of 0.1688 m². Does fitting it prove that its new slope is more accurate for future readings? Explain what evidence you still need.
Show solution: separate fitting from evaluation
The fit minimizes squared error for those five observations. It does not establish which dataset better represents future readings. Investigate the changed observation, then compare candidate approaches using measurements that did not help fit or select them.
Choose test conditions that reflect the intended use. Also inspect individual errors, since an average can hide a large mistake on one reading.
Sources and further study
- NIST: Least Squares, for the slope and intercept formulas.
- Google Machine Learning Crash Course: Loss, for MSE and RMSE.
- NIST: Linear Least Squares Regression, for model scope, outliers, and extrapolation.
- NIST: How can I tell if a model fits my data?, for residual analysis.
- Scikit-learn: Common pitfalls, for separating training and evaluation data.