Conditional probability:
how new information changes the odds
A plain-language path from a two-way table to P(A given B), with a spam-filter example that shows why order matters.
Calcylator Editorial Team
Updated · 4 min read
When extra information changes the probability
Conditional probability is the chance that something happens once you know that something else already has. A coin toss gives 50 percent heads regardless of anything else, but the chance that a card is a king rises from 4 in 52 to 4 in 12 once you learn it is a face card. The condition shrinks the world you are drawing from.
That shrinking is the core idea. You stop looking at everyone and look only at the group where the condition holds, then ask what fraction of that group also has the outcome you care about.
The formula and its pieces
- P(A | B):
- probability of A given that B has happened
- P(A and B):
- probability that both happen
- P(B):
- probability of the condition, must be above zero
The denominator is the probability of the condition, not of the outcome. Once B is known to have occurred, only the part of A that overlaps with B is still possible. Dividing by P(B) rescales that overlap so the probabilities inside B add to 1.
Rearranged, the same relationship gives the multiplication rule: P(A and B) = P(B) × P(A | B). It is how you build probabilities of chains of events, like two draws without replacement.
A useful habit is to say the sentence out loud: among all the people who are B, what fraction are also A? If the answer to that phrasing matches the number you computed, the denominator is right. If you find yourself dividing by the grand total, you have found the joint probability instead.
Worked example from a two-way table
A school has 120 students. 40 play a sport and 12 of those also play a musical instrument. What is the chance a randomly chosen sports player also plays music?
Total students
120
Play sport (B)
40, so P(B) = 40 ÷ 120 = 0.3333
Play sport and music (A and B)
12, so P(A and B) = 12 ÷ 120 = 0.10
Divide
0.10 ÷ 0.3333
P(music given sport)
0.30, or 30%
The shortcut of 12 ÷ 40 gives the same answer.
The shortcut matters. When the table is given in counts, divide the count in the overlap by the count in the condition group. The 120 cancels out and you never need to turn anything into a decimal.
Counts make this kind of problem easy to follow because you can see the groups. Writing the numbers into a small two-way table, with sport against music across the rows and columns, shows 12 in the overlap, 28 who play only sport, 18 who play only music and 62 who do neither. The totals check: 12 + 28 + 18 + 62 = 120.
Order matters: P(A | B) is not P(B | A)
Reversing the condition is the most common error in everyday reasoning. Suppose 30 students play music. Then the chance that a musician also plays a sport is 12 ÷ 30 = 40%, which is not the 30% from before. The overlap is the same, yet the group it is measured against changed.
| Question | Group being conditioned on | Calculation | Result |
|---|---|---|---|
| Sport players who play music | 40 sport players | 12 ÷ 40 | 30% |
| Musicians who play a sport | 30 musicians | 12 ÷ 30 | 40% |
| Students who do both | All 120 | 12 ÷ 120 | 10% |
Media reports often blur this difference, for example by quoting how many criminals had a trait as though it were how many people with that trait become criminals. Always ask which group is the denominator.
Chains of events and the multiplication rule
The product form of the rule is what you use when events happen one after another and each step changes the odds of the next. Drawing two cards from a deck without putting the first back is the standard case.
First card is an ace
4 ÷ 52
Second card is an ace, given the first was
3 ÷ 51
Multiply
4 ÷ 52 × 3 ÷ 51
P(two aces)
12 ÷ 2,652 = 0.0045, or 1 in 221
With replacement the chance would be (4 ÷ 52)² = 0.0059, so the condition lowers it.
A tree diagram lays the same thing out visually. Each branch carries a conditional probability, and the probability of any complete path is the product of the branches along it. Adding the paths that end in the same outcome gives the total probability of that outcome, which is exactly the denominator used in Bayes' rule.
Flipping the condition with Bayes' rule
- P(B | A):
- the reversed conditional
- P(A):
- overall probability of A, found by adding across all cases
A spam filter makes this concrete. Suppose 20% of mail is spam, 60% of spam contains the word free, and 5% of ordinary mail does too.
P(spam)
0.20
P(free given spam)
0.60
P(free given not spam)
0.05
P(free)
0.20 × 0.60 + 0.80 × 0.05 = 0.16
P(spam given free)
0.12 ÷ 0.16 = 0.75, or 75%
A message with the word is three times as likely to be spam as the 20% baseline suggests, not certain.
Checking independence
Two events are independent when learning one tells you nothing about the other. In that case P(A | B) equals P(A). If the school above had 30% of all students playing music and 30% of sport players doing so as well, the activities would be independent. Because the overall music rate is 25% (30 of 120) and the rate among athletes is 30%, they are mildly connected.
Independence is useful but rare in data. A calculator can help you compare P(A | B) with P(A) quickly; if they differ meaningfully, the events are related and any model that treats them as independent will mislead.
Real-world conditioning also hides selection effects. If you only ever see data from people who walked into a clinic, a shop or a website, the group you are conditioning on was already filtered, and the percentages you compute describe that filtered group. Ask who is missing from the table before you generalise the answer to everyone.
Pitfalls and limits
- The condition must have a probability above zero. You cannot condition on something impossible.
- Small groups give unstable percentages. 3 of 10 is also 30%, but with far more uncertainty than 300 of 1,000.
- Conditional probability describes association, not cause. Sport players playing music more often does not mean one produces the other.
- Base rates matter. A test that is right 95% of the time can still be mostly wrong on a positive result if the condition is rare.
Common questions
What is the formula for conditional probability?
P(A | B) = P(A and B) ÷ P(B), provided P(B) is above zero. Divide the probability that both events occur by the probability of the condition. With counts, divide the number in both by the number in B.
How is conditional probability different from joint probability?
Joint probability P(A and B) measures both events across the whole population. Conditional probability P(A | B) measures A inside only the group where B is true. In the school example, joint is 10 percent and conditional is 30 percent.
Is P(A|B) the same as P(B|A)?
No. They share the same overlap but use different denominators. If 12 of 40 sport players play music, P(music | sport) is 30 percent, while with 30 musicians P(sport | music) is 12 ÷ 30 = 40 percent.
What does it mean for two events to be independent?
Events are independent when P(A | B) equals P(A), so knowing B changes nothing about A. Equivalently, P(A and B) = P(A) × P(B). Rolling two dice is independent; drawing two cards without replacement is not.
When do I need Bayes' theorem?
Use it when you know P(A | B) but want P(B | A), such as the chance a message is spam given it contains a word. It needs the base rate P(B) and the overall P(A) to flip the condition correctly.
Was this guide helpful?
Continue reading
View all blogsCoefficient of Variation: Compare Spread Fairly
The coefficient of variation is SD ÷ mean. A 50 g SD on 5 kg bags (1%) is steadier than a 4 g SD on 200 g packs (2%), despite the larger SD.
6 min read
How to Find Quartiles and the IQR
Q1, Q3 and the IQR explained with a 12-value set: Q1 = 12.5, Q3 = 25, IQR = 12.5, and the 1.5 × IQR fences flag 58 as an outlier.
5 min read
Odds to Probability: Convert 3:1, 2.5 and More
Odds of 3 to 1 against are a 25% chance, 3 to 2 in favour is 60%, and decimal 2.50 is 40%. Convert any odds format with one short rule.
5 min read




