explainer
Conditional probability: reading a robot sensor
Learn conditional probability with an interactive robot sensor example. Work through the formula, Bayes' rule, base rates, and common mistakes using counts.
What you will learn
- Identify the reference group in a conditional probability.
- Calculate an obstacle's probability after a positive sensor reading.
- Explain how prevalence changes the meaning of a detection.
- Distinguish independent events from mutually exclusive events.
Before you start
- Fractions and percentages
- Probability as a number between zero and one
A sensor that detects 90% of obstacles can still produce as many false alarms as real detections. The missing detail is how often obstacles occur.
Conditional probability gives you a way to reason through that detail. We'll build a hypothetical robot sensor model, count its outcomes, and see why reversing a probability changes the question.
Choose the reference group
Conditional probability is the probability of one event given another event. Write P(A | B) and read it as “the probability of A, given B.” The event after the bar sets your reference group.
Ask: “Among outcomes where B happens, what fraction also have A?” The definition captures that fraction. Penn State's conditional probability lesson develops it from a count table.
P(A | B) = P(A ∩ B) / P(B), with P(B) > 0
Here, ∩ means “and.” The numerator counts the overlap. The denominator counts the group you kept.
For a robot, let O mean “an obstacle occupies the path” and + mean “the sensor reports an obstacle.” These questions use different reference groups:
- P(+ | O): Among real obstacles, how often does the sensor report one?
- P(O | +): Among positive readings, how often does an obstacle exist?
The first describes the sensor's detection rate. The second describes what a positive reading tells you.
Try the sensor experiment
Start with 10% obstacle prevalence, 90% detection, and 10% false alarms. “Prevalence” means the share of readings that encounter an actual obstacle. The false alarm rate applies to clear paths.
In the upper rectangle, column width shows how common each actual state is. Filled regions show positive readings within each state. The lower bar collects those positive readings into one reference group.
Before changing prevalence, predict whether rarer obstacles will raise or lower P(O | +). Then choose Rare obstacles (1%). Keep the two sensor rates fixed while comparing the results.
Work through 1,000 readings
Reset the experiment. Imagine 1,000 readings under the starting assumptions:
- 100 readings encounter obstacles: 10% of 1,000.
- 90 produce true positives: 90% of those 100 obstacles.
- 10 miss an obstacle: the remaining obstacle readings.
- 900 readings encounter clear paths: the other 90% of the population.
- 90 produce false positives: 10% of those 900 clear paths.
- 810 correctly report clear: the remaining clear-path readings.
The sensor reports positive 180 times. Only 90 of those readings correspond to an obstacle.
P(O | +) = 90 / (90 + 90) = 50%
Now reverse the question. Among the 100 obstacle readings, 90 produce a positive result, so P(+ | O) = 90 / 100 = 90%.
Both calculations use the same 90 true positives. Their denominators differ because they ask about different groups. Saying “the sensor is 90% accurate” hides that distinction; name the metric and its denominator.
The same care applies to vector matching: a cosine similarity score needs its own interpretation before you use it as evidence.
Joint and marginal probabilities
A joint probability describes two events together. Our 90 true positives out of 1,000 give P(O ∩ +) = 0.09. Multiplying along the obstacle branch gives the same result: 0.10 × 0.90.
A marginal probability describes one event while ignoring the other. The table's row total gives P(O) = 100 / 1,000 = 0.10. Its positive-column total gives P(+) = 180 / 1,000 = 0.18.
To find P(+), add the two ways a positive reading can happen: a detected obstacle or a false alarm. This is the law of total probability. Substituting it into the conditional probability definition gives Bayes' rule for our sensor. Penn State's Bayes' theorem lesson explains this reversal.
P(O | +) = (d × p) / (d × p + f × (1 − p))
Here p is prevalence, d is detection rate, and f is false alarm rate. With our starting values, the calculation is 0.09 / (0.09 + 0.09) = 0.5. The formula scales the same counting argument to any population size.
Change the base rate
Suppose obstacles occur in just 1% of readings. In 1,000 readings, you now expect 10 obstacles and 990 clear paths. The same sensor produces 9 true positives and 99 false positives.
That gives P(O | +) = 9 / 108 ≈ 8.3%. A positive reading still raises the obstacle probability above its starting 1%. The large pool of clear paths also generates many false alarms.
At 50% prevalence, the counts shift again: 450 true positives and 50 false positives. P(O | +) becomes 90%. These changes follow from the population even while the detection and false alarm rates stay fixed.
Independence and exclusivity
Two events are independent when knowing one leaves the other's probability unchanged. Equivalently, P(A ∩ B) = P(A)P(B). If P(B) is positive, this means P(A | B) = P(A). Penn State's independent events lesson connects these definitions.
Choose Uninformative sensor. It reports positive on 50% of obstacle readings and 50% of clear readings. At 10% prevalence, you get 50 true positives and 450 false positives, so P(O | +) stays at 10%.
Mutually exclusive events cannot happen together. An obstacle and a clear path are mutually exclusive in this model. Learning that the path is clear rules out an obstacle.
When both events have positive probability, mutually exclusive events are dependent: their joint probability is zero, while the product of their probabilities is positive.
What the model assumes
This example uses invented rates to teach the calculation. Applying it to real sensor data requires a defined population and evidence for those rates.
- Define “obstacle,” the detection distance, and what counts as one reading.
- Estimate detection and false alarms under the conditions where the robot operates.
- Treat the display as expected counts. A finite run will vary, and some settings produce fractional averages.
- Check dependence before combining repeated readings. Several frames can share the same reflection or missed object.
The calculation concerns one reading. It does not require successive readings to be independent. A model that multiplies evidence across readings needs further assumptions.
Markov chains use conditional probabilities to model a sequence of states. Their key assumption is that the current state contains the information from the past needed to predict the next state.
What if the success rate itself is unknown? The Bayesian inference lesson starts with a distribution over that rate, then updates it using observed outcomes. It also explains the assumptions needed to combine the evidence.
A detector can also estimate an obstacle probability from a measured feature. Logistic regression shows that model and lets you change the decision threshold to compare false alarms with missed detections.
If the sensor never reports positive, P(O | +) has a zero denominator and is undefined. The experiment shows that case explicitly. MIT's conditioning and Bayes' rule notes develop the general rules behind this model.
Check your understanding
Exercise 1. Return to 10% prevalence and 90% detection. Reduce the false alarm rate to 1%. Out of 1,000 readings, how many positives do you expect, and what is P(O | +)?
Show the worked solution
There are 100 obstacles, so 90 readings are true positives. There are 900 clear paths, so 9 readings are false positives. The positive group contains 99 readings, and P(O | +) = 90 / 99 ≈ 90.9%.
The smaller false alarm rate removes 81 false positives from the starting example. The number of true positives stays at 90.
Exercise 2. At the starting settings, the sensor reports a negative reading. What is the probability that the path still contains an obstacle? Identify the denominator before calculating.
Show the worked solution
The reference group contains all negative readings: 10 missed obstacles plus 810 correctly identified clear paths. Thus P(O | negative) = 10 / 820 ≈ 1.22%.
A negative reading lowers the modeled obstacle probability from 10% to about 1.22%. To choose an action, the robot also needs a decision rule that accounts for the consequences of a missed obstacle.
Sources and further study
- Penn State STAT 414: Conditional Probability, for the definition and count-table approach.
- Penn State STAT 414: Bayes' Theorem, for total probability and reversing a conditional.
- Penn State STAT 414: Independent Events, for the independence conditions.
- MIT: Conditioning and Bayes' Rule, for a second treatment with radar examples.
For your next probability claim, write the reference group in plain words before reaching for a formula.