Logistic regression: turn a score into a probability

Explore logistic regression with synthetic robot observations. Calculate sigmoid probabilities and log loss, then change a threshold and inspect the confusion matrix.

By 10 min read

Logistic regression estimates the probability of a class from input features. A decision threshold turns that probability into a label. Keeping those two steps separate helps you understand both training and classification errors.

Imagine a robot checking whether a sensor return comes from an obstacle. Use a scaled return-strength feature x, with label y = 1 for an obstacle and y = 0 for its absence. The eight observations below are synthetic; they illustrate the calculation without establishing a real detector's performance.

From a sensor feature to a probability

Start with the same slope-and-intercept form used in linear regression. Call its output the score z:

z = mx + b
p = σ(z) = 1 / (1 + e−z)

The sigmoid function σ maps the score to a probability. Here p estimates P(y = 1 given x), a conditional probability. The constant e is about 2.718; a score of zero gives p = 0.5.

Positive scores give probabilities above 0.5. Negative scores give probabilities below 0.5. For finite scores, the mathematical function stays strictly between zero and one; computer arithmetic can round extreme values to an endpoint.

The score also equals ln(p / (1 − p)), the log-odds. Slope m changes log-odds per unit of x. Its effect on probability varies along the curve. Google's sigmoid guide explains this connection.

Our x values have no units. A negative value means a return below a chosen reference level. It does not mean a negative physical distance.

Choose a decision threshold

Let t be a probability cutoff. This lesson predicts class 1 when p ≥ t, including an exact tie. Otherwise it predicts class 0.

At t = 0.5, the rule predicts an obstacle whenever z ≥ 0. For a nonzero slope, the crossing occurs at x = −b/m. With m = 0, every input receives the same probability σ(b).

A lower threshold allows more positive predictions. On a fixed dataset, that can catch additional obstacles and produce additional false alarms. A higher threshold can remove false alarms while missing more obstacles.

Choose a cutoff based on error costs and validation results. A detector that requests another observation may need a different rule from one that triggers an immediate stop. Scikit-learn's threshold guide describes this decision step.

Calculate one prediction

Use m = 1, b = 0, and the observation x = 1, y = 0:

z = 1 × 1 + 0 = 1
p = 1 / (1 + e−1) ≈ 0.7311

At threshold 0.5, the decision is 1. The actual label is 0, so this observation is a false positive. At threshold 0.8, the decision becomes 0 and the probability remains about 73.1%.

The model can assign a high probability to an event that does not occur. One such observation cannot tell you whether its probabilities are reliable across many cases.

Move the curve and the cutoff

The initial curve uses m = 1 and b = 0. Four examples have each label. Some labels overlap along x, so one cutoff on this feature cannot separate them perfectly.

Interactive experiment

Separate probability from the decision

Inspect eight synthetic robot observations. Class 1 means an obstacle is present; class 0 means absent.

Probability and observed class

Logistic probability curve, observed classes, and decision thresholdThe amber curve is sigmoid of 1.00 times x plus 0.00. Dark circles mark eight actual labels at 0 or 1. The dashed horizontal line marks threshold 0.50. The outlined square marks the selected probability: 73.1% at x = 1. Predictions at or above the threshold are positive. Exact numbers follow in the observations table.0.00.51.0-2-1012

Scaled sensor feature x

Amber curve: model probabilities. Dark circles: actual labels. Dashed line: decision cutoff. The outlined square follows your selected example.
Read all observations and probabilities
Eight training examples. Decisions use the full probability before rounding.
Feature xActual yP(y = 1)Predict
-2.000.1190
-1.500.1820
-1.010.2690
-0.500.3780
0.510.6221
1.000.7311
1.510.8181
2.010.8811
Selected score z
1.00
P(obstacle), selected
73.1%
Selected decision
Positive (1)
Training log loss
0.5289

At x = 1, the actual label is 0. The model predicts 1 at threshold 0.50. This example contributes 1.3133 log loss.

Confusion matrix: all eight training examples
ActualPredict 1Predict 0
Obstacle 1TP 3FN 1
Absent 0FP 1TN 3

Precision: 75.0%. Recall: 75.0%. Accuracy: 75.0%.

Changing the threshold keeps the curve and log loss fixed. Changing a coefficient changes the probabilities. All results here describe the training examples.

First move only the threshold from 0.50 to 0.80. The confusion matrix changes from TP 3, FN 1, FP 1, TN 3 to TP 2, FN 2, FP 0, TN 4. The curve and training log loss stay fixed.

The matrix compares predictions with actual labels:

  • True positive (TP): predicts an obstacle that is present.
  • False negative (FN): misses an obstacle that is present.
  • False positive (FP): predicts an obstacle that is absent.
  • True negative (TN): predicts absence correctly.

Reset, then set the slope to 2. All decisions at threshold 0.5 stay the same, but log loss rises from 0.5289 to 0.6267. The curve makes more confident errors at x = −1 and x = 1.

Measure probability errors with log loss

Binary cross entropy, also called log loss, scores the probability assigned to the observed label. For one observation:

ℓ = −y ln(p) − (1 − y) ln(1 − p)

When y = 1, the loss is −ln(p). When y = 0, it is −ln(1 − p). Use the natural logarithm and average the losses over all examples. Google's loss guide gives the full mean-loss expression.

For the worked example, y = 0 and p ≈ 0.7311. Its loss is −ln(0.2689) ≈ 1.3133. If a model assigns p = 0.99 to that same absent obstacle, the loss grows to about 4.6052.

The calculation needs care near zero and one. Rounded probabilities can make log(0) appear even when the original score is finite. Compute directly from z using this equivalent form:

ℓ = max(z, 0) − yz + ln(1 + e−|z|)

The exponential now has a nonpositive input. The code uses log1p(a) to calculate ln(1 + a) accurately for small a. For z = 1000 and y = 0, this gives a loss of 1000 without overflowing. TensorFlow documents this stable formulation.

Train the coefficients

Training searches for coefficients that lower mean log loss. Under a model of conditionally independent binary labels, minimizing the unpenalized loss matches maximizing their likelihood. The threshold plays no part in this probability calculation.

For one input feature, the mean-loss gradients are:

∂L/∂m = (1/n) ∑(pi − yi)xi
∂L/∂b = (1/n) ∑(pi − yi)

Here n is the number of examples. Gradient descent subtracts a learning rate times each gradient from its coefficient. A suitable step lowers the loss; an oversized step can make it rise.

Training often adds a penalty on coefficient size, called regularization. It can limit extreme fits and help generalization, with its strength chosen through validation. Scikit-learn documents the logistic objective and its penalties.

The widget lets you choose coefficients by hand. Those controls illustrate the model family. They do not claim to find the best fit.

Random forests take another approach to classification: learn split rules from training examples, then combine predictions from several trees. That lesson lets you inspect a small trained forest and see where its trees disagree.

Check performance on new observations

The widget's eight examples all contribute to its displayed training metrics. A real evaluation needs fresh observations. Fit preprocessing and coefficients on training data, choose settings and the threshold using validation data, then evaluate the fixed system on a held-out test set.

For robot logs, keep related readings from the same run together when splitting data. Nearby frames can share almost the same scene. Test conditions should also cover the surfaces, lighting, and sensor behavior you expect to encounter.

Accuracy counts all correct decisions. Precision is TP / (TP + FP), the fraction of positive predictions that are correct. Recall is TP / (TP + FN), the fraction of actual positives you catch.

An always-negative rule gets 99% accuracy when only 1% of examples are positive. It catches none of them. Inspect error counts and class frequencies alongside any single metric; Google explains these measures and their limits.

Calibration asks whether probabilities match observed frequencies across many comparable predictions. Among many cases assigned probabilities near 0.7, about 70% should be positive for a well-calibrated model. Eight points cannot establish that pattern, and a sigmoid output alone does not guarantee it.

Use calibration curves on data separate from model fitting. Log loss reflects several aspects of prediction quality, so a lower value alone cannot prove better calibration. Scikit-learn's calibration guide explains this distinction.

A learned positive coefficient describes an association within the fitted model. It does not establish that changing the sensor feature would cause an obstacle to appear. Scikit-learn's coefficient guide discusses causal interpretation errors.

For uncertainty about the coefficients themselves, a probability for one event is only part of the picture. The Bayesian inference lesson introduces a distribution over an unknown parameter.

Reproduce the numbers in Python

This standard-library script evaluates the same chosen coefficients and eight observations. It computes loss from scores and decisions from unrounded probabilities. Change threshold to 0.8 to reproduce the second confusion matrix.

from math import exp, log1p

data = [(-2, 0), (-1.5, 0), (-1, 1), (-0.5, 0),
        (0.5, 1), (1, 0), (1.5, 1), (2, 1)]
m, b, threshold = 1.0, 0.0, 0.5

def sigmoid(z):
    if z >= 0:
        return 1 / (1 + exp(-z))
    ez = exp(z)
    return ez / (1 + ez)

def log_loss(z, y):
    return max(z, 0) - y * z + log1p(exp(-abs(z)))

scores = [m * x + b for x, y in data]
probabilities = [sigmoid(z) for z in scores]
predictions = [int(p >= threshold) for p in probabilities]
loss = sum(log_loss(z, y) for z, (x, y) in zip(scores, data)) / len(data)

tp = sum(y == 1 and pred == 1 for (x, y), pred in zip(data, predictions))
fn = sum(y == 1 and pred == 0 for (x, y), pred in zip(data, predictions))
fp = sum(y == 0 and pred == 1 for (x, y), pred in zip(data, predictions))
tn = sum(y == 0 and pred == 0 for (x, y), pred in zip(data, predictions))

print(f"At x=1: z={scores[5]:.2f}, p={probabilities[5]:.4f}")
print(f"Mean log loss: {loss:.4f}")
print(f"TP={tp}, FN={fn}, FP={fp}, TN={tn}")
print(f"Loss at z=1000, y=0: {log_loss(1000, 0):.1f}")

Expected output:

At x=1: z=1.00, p=0.7311
Mean log loss: 0.5289
TP=3, FN=1, FP=1, TN=3
Loss at z=1000, y=0: 1000.0

Try it yourself

Exercise 1. Set m = 0 and b = 0. Predict the probabilities, mean log loss, and confusion matrix at threshold 0.5. Then raise only the threshold to 0.51.

Show solution: set every score to zero

Solution: Every probability equals 0.5, and every loss equals ln(2) ≈ 0.6931. At threshold 0.5, the inclusive tie rule predicts all eight examples as positive: TP 4, FN 0, FP 4, TN 0.

At 0.51, all predictions become negative: TP 0, FN 4, FP 0, TN 4. The probabilities and loss stay fixed. Precision is undefined because its denominator, the number of positive predictions, is zero.

Exercise 2. Reset the experiment. At threshold 0.5, accuracy is 75%. Raise the threshold to 0.8 and calculate accuracy, precision, and recall again. Does unchanged accuracy mean an unchanged detector?

Show solution: compare the decisions

Solution: The new counts are TP 2, FN 2, FP 0, TN 4. Accuracy remains 6/8 = 75%. Precision becomes 2/2 = 100%, and recall becomes 2/4 = 50%.

The detector now misses two of the four obstacles and makes no false alarms on these examples. Decide which errors matter for the intended action, then test that choice on new observations.

Sources and further study