Bayes' rule

Frequentist and Bayesian probability

Two main approaches to probability:

The Frequentist approach

A frequentist distrusts any specific measurements, treating them as approximations of a true value.

If a frequentist wants to know the height of a mountain:

The Bayesian approach

A Bayesian trusts every measurement, though it may vary slightly.

If a Bayesian wants to know the height of a mountain:

Bayesian coin flipping

Consider a game with two coins:

  1. one fair, with heads and tails equally likely
  2. one rigged, with a bias of two-thirds, so it comes up heads two-thirds of the time

We pick a coin at random and want to figure out which one we got. It has a 50% chance of being fair and a 50% chance of being rigged.

What is the probability that the fair coin was picked?

We flip the coin once and observe Heads, so we can make a quantitative statement about which coin we picked.

Step Explanation Image

Picturing the coin probabilities

Represent the prior odds of choosing the fair coin (50:50) as a wall divided in half.

Prior odds of choosing the fair coin

Split up the fair and rigged regions

Split each region by the probability of heads or tails.

Now imagine throwing a dart at this wall.

Wherever it lands represents the outcome of flipping a coin.

Split each region by the probability of heads or tails

We've already flipped and observed heads

Since we've already flipped and observed heads, we know the dart must have landed in one of two zones:

  • Fair Heads
  • Rigged Heads

The probability that the coin is fair given that we observed heads is the area of "Fair Heads" divided by the total area that produces heads.

Because the "Rigged Heads" region is larger than the "Fair Heads" region:

  • it's more likely that the observed Heads came from the rigged coin
  • in other words, after seeing heads, we should update our belief: it's now slightly more likely that we flipped the rigged coin
Probability that the coin is fair given that we observed Heads

Let's rephrase the last diagram using probability notation:

Coin flips as probabilities

We can rewrite the joint probabilities by their expanded version:

\[\begin{aligned} \\ P(F|H) &= \frac{ {\color{Goldenrod}P(H,F)} }{ {\color{Goldenrod}P(H,F)} + {\color{crimson}P(H,R)} } \\ \\ &= \frac{ {\color{Goldenrod}P(H|F) \times P(F)} }{ {\color{Goldenrod}P(H|F) \times P(F)} + {\color{crimson}P(H|R) \times P(R)} } \\ \\ \end{aligned}\]

Let's assign a number to each term:

\[\begin{aligned} \\ P(F|H) &= \frac{ \frac{1}{2} \times \frac{1}{2} }{ (\frac{1}{2} \times \frac{1}{2}) + (\frac{2}{3} \times \frac{1}{2}) } \\ \\ &= \frac{ \frac{1}{4} }{ \frac{1}{4} + \frac{1}{3} } = \frac{ \frac{3}{12} }{ \frac{3}{12} + \frac{4}{12} } = \frac{3}{7} = 0.43 \\ \\ \end{aligned}\]

After just one flip:

Note:

Bayes' Rule

The bottom part of the ratio, P(H,F) + P(H,R) represents all the possible ways we could have gotten heads:

\[\begin{aligned} \\ P(F|H) &= \frac{ P(H|F) \times P(F) }{ P(H,F) + P(H,R) } \\ \\ &= \frac{ P(H|F) \times P(F) }{ P(H) } \\ \\ \end{aligned}\]

This gives us the well-known Bayes' Rule (or Bayes' Theorem):

\[ \fbox{$ P(F|H) = \frac{ P(H|F) \ P(F)}{P(H)} $} \]

The left-hand side is always a conditional probability: "the probability of F given H". This is why the question we ask of Bayes' Rule needs to be in the form of a conditional probability:

What is the probability that (something1) is true, given that (something2) is true?

If we can't express our problem in that form, then Bayes' Rule isn't the right tool for answering it.

Bayes' Rule is tricky to remember, but it can be re-derived:

\[\begin{aligned} \\ P(F,H) &= P(H,F)\\ \\ P(F|H) P(H) &= P(H|F) P(F)\\ \\ \frac{P(F|H) P(H)}{\relax{\color{ForestGreen}P(H)}\relax} &= \frac{P(H|F) P(F)}{\relax{\color{ForestGreen}P(H)}\relax}\\ \\ P(F|H) &= \frac{P(H|F) P(F)}{P(H)}\\ \\ \end{aligned}\]

Each term has a conventional name, and the traditional letters are A and B:

\[ \fbox{$ { {\color{BlueViolet}P(A|B)} \atop {\color{BlueViolet}\text{Posterior}\relax} } = \frac{ { {\color{crimson}\text{Likelihood}\relax} \atop {\color{crimson}P(B|A)} } \times { {\color{DarkOrchid}\text{Prior}\relax} \atop {\color{DarkOrchid}P(A)} } } { { {\color{ForestGreen}P(B)} \atop {\color{ForestGreen}\text{Evidence}\relax} } } $} \]

Term Definition Example
Prior P(A) Belief about a hypothesis before observing data. The chance the coin is fair before flipping it.
Likelihood P(B|A) If the hypothesis were true, how likely would this observation be? Getting heads if the coin is fair.
Evidence P(B) The probability of the observed outcome. Getting heads.
Posterior P(A|B) The result of Bayes' Rule: the updated belief in the hypothesis after observing data. The probability the coin is fair given that we saw heads.

A key strength of Bayesian reasoning is that it makes assumptions explicit through the prior P(A):

While the likelihood P(B|A) and evidence P(B) usually come from the experimental setup, we have to guess the prior P(A).

In simple problems, such as a coin flip, choosing a prior is easy. In more complex settings, it is not.

Priors may be chosen by the analyst (subjective Bayes) or selected algorithmically (automatic Bayes).

Bayes' Rule and confusion matrices

Confusion matrices are useful when combined with Bayes' Rule to assess real-world probabilities.

Example:

So the question is:

Is there life on this planet?

To find out, you send a probe, which reports: "No life detected".

Because no probe is perfect, we must ask the question:

What is the probability that the planet contains life, given that the probe detected nothing?

Defining the events:

The probe is sent to 1,000 known planets, of which 101 were known to contain life.

That gives us our prior:

Probe testing results:

Probe testing results

From the confusion matrix, we can compute:

Using Bayes' Rule, \( P(L|not-D) \) is the probability that there actually is life given that our probe says there isn't:

\[\begin{aligned} \\ P(L|not-D) &= \frac{P(not-D|L) \times P(L)}{P(not-D)} \\ \\ &= \frac{ \frac{1}{101} \times \frac{101}{1,000} }{ \frac{870}{1,000} }\\ \\ &= \frac{ \frac{1}{1,000} }{ \frac{870}{1,000} } = \frac{1}{870} = 0.001 \\ \\ \end{aligned}\]

If the probe reports "no life", there's only a 0.1% chance the planet actually has life. That's a lot of confidence.

Now, what if the probe does detect life?

How confident can we be that there really is life on that planet?

\[\begin{aligned} \\ P(L|D) &= \frac{P(D|L) \times P(L)}{P(D)} \\ \\ &= \frac{ \frac{100}{101} \times \frac{101}{1,000} }{ \frac{130}{1,000} }\\ \\ &= \frac{ \frac{100}{1,000} }{ \frac{130}{1,000} } = \frac{100}{130} = 0.77 \\ \\ \end{aligned}\]

If the probe says it found life, we can be about 77% confident that life is truly there, just from this one probe.

To be even more sure:

Repeating Bayes' rule

One event (or measurement) isn't much evidence.

If we apply Bayes' Rule repeatedly in a loop, we can refine our beliefs every time new data arrives.

Each round of Bayes' Rule works like this:

  1. start with a prior P(A) or your initial belief
  2. collect new evidence P(B)
  3. apply Bayes' Rule to calculate the posterior P(A|B)
  4. use this posterior as your new prior and repeat with the next piece of evidence

The Bayes Loop:

The Bayes Loop

Each pass through the loop improves our prior, moving from a rough guess toward a more confident estimate.

Bayes' Rule accumulates evidence:

The Bayes Loop provides a systematic way to update our beliefs based on evidence, getting closer to the truth with each piece of new information.

Multiple hypotheses

Bayes' Rule can handle multiple hypotheses simultaneously — e.g., "the coin is fair" and "the coin is rigged".

Each hypothesis gets updated independently after every observation, allowing us to compare their relative probabilities.

Imagine a box containing coins with 5 different biases:

After each coin flip, the posterior becomes the new prior for the next flip.

Result of the first flip → heads:

The most likely hypothesis emerges gradually:

With many hypotheses, posteriors resemble a Gaussian curve:

By increasing the number of hypotheses toward infinity, Bayesian updating naturally leads to continuous probability distributions, allowing arbitrarily precise estimates.

Robustness to poor priors:

Previous Measuring performance All ⏎ Next Curves and surfaces

A Kemar Joint