Random forests: train different trees and combine their votes

Build a small random forest from synthetic robot observations. Inspect learned splits, bootstrap samples, random feature choices, and individual tree votes.

By 11 min read

What you will learn

  • Trace a decision tree and calculate how a split changes Gini impurity.
  • Explain how bootstrap rows and random candidate features produce different trees.
  • Combine hard class votes and identify limits of agreement and feature importance.

Before you start

Random forests combine decision trees trained with different random choices. Each tree follows a sequence of questions to predict a class. The forest aggregates those predictions.

Imagine a robot classifying sensor returns as an obstacle or its absence. Each observation has return strength x₁ in scaled units and apparent height x₂ in centimetres. The label is 1 for obstacle present and 0 for absent.

This lesson uses 12 synthetic observations. Their fitted models illustrate the algorithm; they do not establish a detector's performance on a real robot.

On the map of machine learning fields, this setup combines a robotics task, tabular inputs, supervised labels, and a tree ensemble. Those descriptions answer different questions about the same model.

Follow a decision tree

A decision tree asks threshold questions such as “Is return strength at most 6.5?” Each answer selects a branch. A final leaf holds the prediction.

In this lesson, a leaf votes for the most common class among its training rows. If the counts tie, it votes 0. A path can use both features, which lets later questions depend on earlier answers.

Training searches for splits that separate the labels. One common measure of their mixture is Gini impurity. If a node contains a fraction p of class-1 rows, its binary impurity is:

G = 1 − p² − (1 − p)² = 2p(1 − p)

A pure node has impurity 0. An equal mix of both classes has impurity 0.5. A useful split lowers the weighted average impurity of its children. Scikit-learn's tree guide describes this split criterion.

Work through a learned split

Consider four separate training rows with one feature:

Return strengthObserved class
10
20
71
81

The parent contains two rows of each class, so G = 0.5. Test a threshold halfway between 2 and 7: x₁ ≤ 4.5. Both children become pure, giving weighted impurity (2/4) × 0 + (2/4) × 0 = 0.

The impurity decrease is 0.5 − 0 = 0.5. A competing threshold of 1.5 leaves one pure row on the left and labels [0, 1, 1] on the right. Its weighted impurity is (1/4) × 0 + (3/4) × (4/9) = 1/3, so 4.5 wins.

The trained question now routes x₁ = 4.5 left and x₁ = 4.6 right. Exact threshold ties go left throughout the experiment. Training learns the threshold from rows; a query only follows the resulting rules.

Give trees different training choices

A forest introduces variation in two places:

  • Bootstrap the rows. Draw n rows from the n training rows, with replacement. A row can appear several times, and another can be absent. Fit one tree to that sampled dataset.
  • Sample candidate features at each split. Choose a random subset of features, then search their thresholds for the best available split. Repeat this choice as the tree grows.

Each tree trains separately. Its fit does not depend on another tree's predictions. That allows training trees in parallel, although their errors can still be correlated because they share the same source data.

Using every feature at each split keeps the bootstrap variation and produces bagged trees. Restricting the candidates adds another source of variation. Scikit-learn's forest guide explains these two mechanisms.

For example, a strong height feature might win most root splits when both inputs are available. Sampling only return strength forces a tree to search that feature. This can reduce shared mistakes, but restricting the choices can also weaken individual trees.

Train and inspect a small forest

The experiment trains actual trees from the listed rows. Each bootstrap contains 12 draws. Trees have at most three splits along a path and search midpoint thresholds on the sampled features.

Interactive experiment

Train small trees and inspect their votes

Each tree learns from a seeded bootstrap of 12 synthetic observations. Class 1 means obstacle present; class 0 means absent.

Height x₂ (cm)

Forest decisions across two sensor featuresTwelve training observations appear as open circles for class zero and filled triangles for class one. Amber background cells predict one; pale cells predict zero. The cross marks return strength 5.0 and height 5.0 centimetres. The forest gives 1 obstacle votes out of 5, predicting class 0. Exact votes and the training data follow in tables.00551010

Return strength x₁ (scaled units)

Open circles: actual 0. Filled triangles: actual 1. The cross is your query. Background cells sample the forest decision on a 0.5-unit grid; the readouts evaluate your exact point.
Read the 12 training observations
Synthetic sensor features and observed class
Rowx₁x₂ (cm)Class
1110
2230
3181
4370
5420
6450
7620
8661
9791
10831
11961
12991
Forest decision
Absent (0)
Obstacle votes
1 / 5
Vote fraction
20.0%
Trees
5

At (5.0, 5.0), 1 of 5 trees vote obstacle. This fraction measures tree agreement; it is not a calibrated event probability.

Individual tree votes
TreeVoteLeaf 1s / rows
100 / 6
200 / 6
300 / 4
416 / 6
500 / 6

Leaf counts include repeated bootstrap rows. Each tree casts one vote for its leaf's majority class. A tied leaf votes 0.

Inspect tree 1's sample and path

Bootstrap row IDs, in draw order: 3, 7, 1, 6, 2, 9, 8, 5, 9, 2, 12, 11.

Rows omitted from this tree: 4, 10.

  1. Candidate feature: Return strength x₁. Return strength x₁ = 5.0 is ≤ 6.5; follow the left branch.
  2. Candidate feature: Height x₂. Height x₂ = 5.0 is ≤ 5.5; follow the left branch.

The reached leaf contains 0 class-1 rows out of 6. Tree 1 votes 0.

Trees use at most three splits along any path and stop when the sampled candidate features cannot reduce Gini impurity. A seed repeats the same training choices. Moving the query changes predictions without retraining.

The initial settings use five trees, one candidate feature, seed 17, and query (5, 5). The votes are [0, 0, 0, 1, 0]. Four trees choose absence, so the forest predicts class 0.

Open the first tree's sample and path. Its bootstrap row IDs are 3, 7, 1, 6, 2, 9, 8, 5, 9, 2, 12, 11. Rows 2 and 9 occur twice; rows 4 and 10 never appear.

For query (5, 5), tree 1 asks x₁ ≤ 6.5, then x₂ ≤ 5.5. Both answers go left. The leaf contains six sampled rows, all class 0.

Move the query to (10, 10). All five trees now vote 1. Moving the query leaves the trained trees intact; changing the seed or candidate-feature setting trains a different forest.

Increase the tree count to add trees while preserving the earlier ones. A fixed seed makes the choices repeatable. Changing a seed explores the effect of sampling; a seed with attractive results on these rows does not establish better future performance.

This small implementation stops at a pure node, its depth cap, or a lack of positive impurity improvement on the sampled features. A constant sampled feature can therefore stop a branch even when another feature could split it. These limits keep the rules inspectable.

Combine votes carefully

This lesson uses hard voting: each tree contributes one class label. For T trees, the class-1 vote fraction is:

v = (h₁(x) + h₂(x) + ⋯ + hT(x)) / T

Each h is either 0 or 1. A strict majority of class-1 votes produces class 1; a tied forest produces class 0. The widget offers odd tree counts, which avoid a forest tie.

At the initial point, v = 1/5 = 20%. This measures agreement among fitted trees. It does not establish a 20% chance that a real obstacle is present.

The rule follows the class-voting definition in Breiman's original random forests paper. Implementations can aggregate differently: scikit-learn averages the trees' class-probability estimates. Check the rule before comparing outputs.

Shared errors limit the benefit of aggregation. If every tree learns the same misleading cue, their votes can reinforce it. More trees can stabilize the voting result, but an added tree can still change a finite forest's decision in either direction.

Reliable event probabilities need calibration checks on suitable data. Scikit-learn's calibration guide explains how observed frequencies can test probability estimates. The conditional probability lesson provides the foundation for interpreting those frequencies.

Validate on observations the forest has not seen

A tree can fit accidental patterns in a small dataset. A forest can also overfit or carry systematic errors into new settings. Choose depth, minimum leaf size, candidate-feature count, and other settings using validation data or cross-validation.

The pandas lesson covers the labeled tables that can hold these observations. Its selection, missing-data, and join examples help you check which rows a training pipeline actually uses.

Keep a final test set for evaluating the completed choice. For robot logs, split by run or scene when nearby observations share conditions. Randomly scattering adjacent frames across sets can make evaluation look easier than deployment. Scikit-learn's validation guide covers separate test data and grouped splits.

An out-of-bag observation is a row omitted from a particular tree's bootstrap. An out-of-bag prediction uses only trees that omitted that row. Breiman and Cutler describe how to combine those votes; a very small forest may leave some rows with few or no eligible trees.

Out-of-bag checks still need an appropriate sampling setup. If another frame from the same scene enters a tree's training sample, omitting one row does not remove that shared context. The widget displays omissions for inspection and makes no out-of-bag accuracy claim.

Inspect false alarms and missed obstacles, especially when one class is rare. The logistic regression lesson explains precision, recall, and the limits of accuracy. Agreement among trees cannot replace those checks.

A forest can also predict continuous values by averaging numeric tree outputs. The linear regression lesson introduces that prediction goal and its error measures. The experiment here uses binary labels throughout.

Read feature importance with care

Impurity importance credits features for the training impurity reductions their splits achieve. Features with many possible values get more chances to find appealing splits, including accidental ones. A large score can reflect training fit without proving useful performance on fresh observations.

Permutation importance shuffles one feature's values and measures the resulting change in a chosen model score. Using held-out data connects the calculation to predictions on fresh examples. Repeat the shuffle to see how much its result varies.

Correlated features complicate this measure. If one feature supplies much of the same information as another, shuffling either can leave a useful substitute. Both importance scores can look small even when the shared information matters. Scikit-learn's permutation guide explains these limitations.

Treat importance as evidence about a fitted model's predictive behavior. It does not, by itself, establish that changing a feature would cause the outcome to change.

Fit three small trees in Python

This standard-library example trains three stumps, trees with at most one split, on the four-row worked dataset. It learns thresholds, draws bootstrap samples, and aggregates the resulting class votes. With one feature, it demonstrates bagging; the widget adds feature subsampling and deeper trees.

from random import Random

data = [(1, 0), (2, 0), (7, 1), (8, 1)]
rng = Random(17)

def gini(rows):
    p = sum(y for x, y in rows) / len(rows)
    return 2 * p * (1 - p)

def majority(rows):
    return int(sum(y for x, y in rows) > len(rows) / 2)

def fit_stump(rows):
    best = (None, majority(rows), majority(rows))
    best_loss = gini(rows)
    values = sorted(set(x for x, y in rows))
    for a, b in zip(values, values[1:]):
        cut = (a + b) / 2
        left = [row for row in rows if row[0] <= cut]
        right = [row for row in rows if row[0] > cut]
        loss = (len(left) * gini(left) + len(right) * gini(right)) / len(rows)
        if loss < best_loss - 1e-12:
            best_loss = loss
            best = (cut, majority(left), majority(right))
    return best

samples = [[rng.randrange(len(data)) for _ in data] for _ in range(3)]
trees = [fit_stump([data[i] for i in sample]) for sample in samples]
query = 4.5
votes = [left if cut is None or query <= cut else right
         for cut, left, right in trees]
print('Bootstrap row IDs:', [[i + 1 for i in sample] for sample in samples])
print(f'Votes at x={query}: {votes}')
print(f'Class-1 vote fraction: {sum(votes) / len(votes):.1%}')
print('Forest decision:', int(sum(votes) > len(votes) / 2))

Expected output:

Bootstrap row IDs: [[4, 3, 3, 3], [2, 3, 1, 1], [2, 4, 4, 3]]
Votes at x=4.5: [1, 0, 0]
Class-1 vote fraction: 33.3%
Forest decision: 0

The first sample contains only positive labels, so its stump predicts 1 everywhere. The other two learn splits and vote 0 at the query. Inspecting the samples explains the disagreement.

Try it yourself

Exercise 1. A node contains labels [0, 1, 1, 1]. A proposed split sends the single zero left and all three ones right. Calculate the parent impurity, weighted child impurity, and improvement.

Show solution: calculate the split improvement

The positive fraction is p = 3/4. Parent impurity is 2 × (3/4) × (1/4) = 0.375.

Both children are pure, so their weighted impurity is (1/4) × 0 + (3/4) × 0 = 0. The improvement is 0.375. The left leaf votes 0 and the right leaf votes 1.

Exercise 2. Four trees vote [1, 0, 1, 0]. Calculate the vote fraction and forest decision under this lesson's tie rule. Then explain what you know about the actual event probability.

Show solution: interpret a tied vote

The class-1 vote fraction is 2/4 = 50%. The vote ties, so the forest predicts 0 under the stated rule.

The trees split evenly on this query. Their votes alone cannot establish a 50% event probability. Check predicted scores against actual outcomes on suitable new observations before making a calibration claim.

Sources and further study