Linear regression: fit a line and inspect its errors

Fit linear regression to robot sensor readings. Work through least squares, residuals, and MSE, then test an outlier with Python and exercises.

By 9 min read

What you will learn

  • Predict a continuous value with a slope and intercept.
  • Calculate residuals and mean squared error, then check a least squares fit.
  • Explain how an outlier and new operating conditions limit a fitted line.

Before you start

Linear regression predicts a numeric value by fitting a line to examples. The line gives a simple rule you can inspect. Its errors show where that rule needs closer attention.

Imagine comparing a robot's distance sensor with reference measurements. For each reported reading, record the reference distance. A fitted line can correct an offset and a scale difference over the range you measured.

This lesson uses five synthetic readings. They illustrate the calculation; they do not establish the performance of a real sensor.

Turn a reading into a prediction

Let x be the sensor's reported distance and y the reference distance. Both use metres. Predict y with:

ŷ = mx + b

The hat in ŷ marks a prediction. The slope m sets how much the prediction changes when x increases by one. The intercept b gives the prediction at x = 0.

For m = 0.84 and b = 0.32, a reading of 2.5 m gives 0.84 × 2.5 + 0.32 = 2.42 m. Here the slope is a ratio of metres to metres, so it has no units. The intercept has units of metres.

With several input features, the prediction becomes ŷ = w · x + b. The dot product combines the features with their weights. This lesson keeps one input so every calculation fits on a graph.

Measure the errors with MSE

A residual is the observed value minus the prediction: e = y − ŷ. A positive residual means the model predicts too little. A negative residual means it predicts too much.

Adding residuals alone can hide mistakes: +2 and −2 cancel. Mean squared error, or MSE, squares each residual, adds the squares, and divides by the number of observations:

MSE = (1/n) ∑(yi − ŷi)²

The index i identifies one observation; n counts them. With distances in metres, MSE has units of square metres. Taking its square root gives RMSE, in metres. Google's loss guide explains these two measures.

Squaring gives large errors more influence. A residual of 2 contributes four times as much squared error as a residual of 1. Choose the loss with the cost of your real prediction errors in mind.

The Laplace distribution lesson connects absolute error to a model for measurement errors. For a constant location estimate under that model, minimizing absolute error gives a sample median.

Adjust the calibration line

Start with “Calibration readings.” The initial line is ŷ = x, which treats each sensor reading as the reference distance. Its training MSE is 0.1240 m².

Move the slope and intercept controls. Try to lower the MSE before pressing “Fit least squares.” Then compare your line with the computed fit.

Interactive experiment

Fit a sensor calibration line

Change the slope and intercept. Compare the line with five synthetic calibration readings.

Calibration observations and the prediction lineFive sensor readings from 0 to 4 metres. The prediction is 1.00 times the reading, plus 0.00 metres. Dashed vertical segments connect predictions to measured distances. Current training MSE is 0.1240 square metres.Measured distance (m)-1.02.05.001234Sensor reading x (m)
The amber line predicts distance. Dark marks show observations; dashed segments show residuals. Each circle is one observation. The vertical scale adjusts to the displayed values.
Read observations and residuals
Distances and signed residuals, in metres
Sensor xActual yModel ŷy − ŷ
0.000.200.000.20
1.001.301.000.30
2.001.802.00-0.20
3.003.303.000.30
4.003.404.00-0.60
Slope m
1.00
Intercept b (m)
0.00
Training MSE (m²)
0.1240
Training RMSE (m)
0.3521

Current training MSE: 0.1240 m². The least squares line reaches 0.0728 m² on this dataset.

Residuals against sensor readingsResidual means measured minus predicted distance. Positive values are underpredictions. The five residuals in metres are 0.20, 0.30, -0.20, 0.30, -0.60. The vertical range is minus 1.0 to plus 1.0 metres.Residual (m)-1.00.01.001234Sensor reading x (m)
Above zero, the line predicts too little. Below zero, it predicts too much. The vertical scale adjusts; use the numbers to compare fits.

Changing the dataset keeps your line in place. Fit again to see how the changed observation moves it. All displayed losses use the five training readings.

Switch to “One changed reading” after fitting. The last observation moves from (4, 3.4) to (4, 5.4). Your current line stays in place, so you can see the new error before fitting again.

The new fit has m = 1.24 and b = −0.08. One changed observation moves predictions across the entire input range. The negative prediction at x = 0 also shows that this model imposes no physical distance constraint.

An unusual observation could reflect a recording error or a real condition. Investigate it before deciding how to handle it. NIST describes this sensitivity to outliers.

Calculate the least squares fit

Ordinary least squares chooses the line with the smallest sum of squared residuals. Dividing that sum by a fixed n leaves the minimizing line unchanged, so it also minimizes MSE.

For a straight line with an intercept, calculate the input mean x̄ and output mean ȳ. Then use:

m = ∑[(xi − x̄)(yi − ȳ)] / ∑(xi − x̄)²


b = ȳ − mx̄

These formulas follow from setting both partial derivatives of squared error to zero. NIST gives the derivation and estimators.

For the original five readings, x̄ = 2 and ȳ = 2. The numerator is 8.4, and the denominator is 10. Therefore m = 8.4 / 10 = 0.84 and b = 2 − 0.84 × 2 = 0.32.

Check each prediction and residual:

xyŷ = 0.84x + 0.32y − ŷ
00.20.32−0.12
11.31.160.14
21.82.00−0.20
33.32.840.46
43.43.68−0.28

The squared residuals sum to 0.364. Divide by five to get MSE = 0.0728 m². The RMSE is approximately 0.2698 m.

Gradient descent can minimize this loss through repeated updates. The formula above solves this particular fit directly. For several features, numerical libraries commonly solve the least squares problem using matrix factorizations. The projections and least squares lesson explains the geometry, and singular value decomposition exposes independent directions in the data.

Inspect what the line misses

The residual plot puts each error against its input. An upward trend suggests an incorrect slope. A curve can suggest that a straight line misses a systematic relationship. A changing spread can indicate that the error scale varies with the input.

Five points provide little evidence about those patterns. Use more measurements before drawing strong conclusions. NIST's model validation guide shows why residual plots complement a single score.

Fitting a line requires no assumption of normally distributed residuals. Statistical confidence intervals require additional assumptions about the data and errors. The widget reports a point fit and training loss; it does not calculate uncertainty intervals.

Test on new readings

Every displayed observation helps choose the fitted coefficients. A smaller training MSE alone cannot establish accuracy on another robot run.

Fit coefficients using training data. Use validation data to choose between modeling approaches, then assess the chosen approach on untouched test data. Estimate any preprocessing values from the training data, too. Scikit-learn explains how information leaking across that boundary inflates evaluation results.

The pandas lesson shows how to select sensor readings, handle missing values, and check joins before those rows reach a model. Preserving the run labels helps you split the data deliberately.

Nearby sensor readings from one run can share conditions and errors. If the goal is performance on new runs, keep whole runs together when splitting data. Choose a split that reflects how the model will face new observations.

The fitted line predicts 5.36 m at x = 6. Our observations only cover x from 0 to 4, so that prediction is extrapolation. The arithmetic still works, but the data do not establish that the relationship continues there.

Even within the measured range, a different surface or temperature can change performance. A low loss under one set of conditions needs fresh checks under another.

For an outcome such as “obstacle present” or “obstacle absent,” logistic regression maps a linear score to a probability. Its lesson also separates fitting that probability model from choosing a decision threshold.

Reproduce the fit in Python

This example uses only the Python standard library. It fits the original training readings and reports their loss.

from statistics import mean
from math import sqrt

data = [(0, 0.2), (1, 1.3), (2, 1.8), (3, 3.3), (4, 3.4)]
x_bar = mean(x for x, y in data)
y_bar = mean(y for x, y in data)
sxx = sum((x - x_bar) ** 2 for x, y in data)

if sxx == 0:
    raise ValueError("At least two distinct x values are required")

m = sum((x - x_bar) * (y - y_bar) for x, y in data) / sxx
b = y_bar - m * x_bar
residuals = [y - (m * x + b) for x, y in data]
mse = mean(error ** 2 for error in residuals)

print(f"slope={m:.2f}, intercept={b:.2f}")
print(f"training MSE={mse:.4f}, RMSE={sqrt(mse):.4f}")
print(f"prediction at x=2.5: {m * 2.5 + b:.2f}")

Expected output:

slope=0.84, intercept=0.32
training MSE=0.0728, RMSE=0.2698
prediction at x=2.5: 2.42

Try it yourself

Exercise 1. Keep the original slope of 0.84, but increase the intercept from 0.32 to 0.42. Predict how every residual changes. Calculate the new MSE.

Show solution: shift every prediction

Every prediction increases by 0.10 m, so every residual decreases by 0.10 m. The new residuals are −0.22, 0.04, −0.30, 0.36, −0.38.

Their squares sum to 0.414. The MSE is 0.0828 m², higher than the fitted minimum of 0.0728 m². Check the result with the intercept control.

Exercise 2. The changed dataset has a fitted MSE of 0.1688 m². Does fitting it prove that its new slope is more accurate for future readings? Explain what evidence you still need.

Show solution: separate fitting from evaluation

The fit minimizes squared error for those five observations. It does not establish which dataset better represents future readings. Investigate the changed observation, then compare candidate approaches using measurements that did not help fit or select them.

Choose test conditions that reflect the intended use. Also inspect individual errors, since an average can hide a large mistake on one reading.

Sources and further study