← back to EEE 485
Week 2110 min full read
7 concepts20 worked examples30 exercises4 exam-level7 figures
What are you here for?

02 Bayesian and frequentist machine learning: posteriors, point estimates, credible and confidence intervals

Start with this

One question before you read anything. Getting it wrong is the point: it shows you what this section is for.

§02.1 — what three clicks in three views support

A new banner is shown to $3$ visitors and all $3$ click. Before reading on, pick the statement you find easiest to defend.

Find(a) Which statement do these data support?
Given
  • $3$ clicks in $3$ views

  • views independent, each clicked with the same unknown rate $\theta$

Hint 1/4

Ask how often different rates would produce exactly this result.

Hint 2/4

With independent views, $\Pr(3\text{ clicks in }3\mid\theta)=\theta^3$.

Hint 3/4

$\theta=0.6$ gives $0.6^3=0.216$; $\theta=0.9$ gives $0.729$; $\theta=0.1$ gives $0.001$.

Hint 4/4

Only the statement about $0.6$ holds up: that rate gives this result about $1$ time in $5$.

Show solution

Testing each claim against the probability $\theta^3$ is faster than arguing about it in words.

Probability of the data under a few rates

$$0.1^3=0.001,\quad 0.6^3=0.216,\quad 0.9^3=0.729,\quad 1^3=1$$

Three independent clicks, so the rate is cubed.

Judge each claim

$$0.216\approx\tfrac15$$

A modest rate explains the data about one time in five, so neither 'the rate is 1' nor 'certainly above 0.9' follows; 'nothing at all' fails too, since $0.1$ is almost ruled out.

Answer $$\boxed{\theta=0.6\ \text{gives this result about 1 time in 5}}$$
Check

The ratio $0.729/0.216\approx 3.4$ says $0.9$ fits better than $0.6$, but by a factor of about three, not by certainty.

Three observations are weak evidence, not zero evidence and not proof; this section is about measuring exactly how weak.

A music app shows a new playlist suggestion to four listeners, and all four accept it. The dashboard divides $4$ by $4$, prints an acceptance rate of 100%, and the product team wants to announce that listeners always accept. How much should four out of four really convince you?

By the end you can turn those four acceptances into a full distribution over the true rate, read off that a rate above $0.9$ has probability only about $0.41$, and attach a credible or a to any rate estimated from counts.

In 60 seconds

Bayesian learning puts a distribution on the unknown rate and updates it by counting, Beta prior in and Beta posterior out; frequentist learning keeps the rate fixed and judges a recipe, the fraction of heads and its $\bar\theta\pm 2\,\mathrm{SE}$ interval, by how often it works.

Beta update
$$\mathrm{Beta}(a,b)\ \xrightarrow{\ N_1,\ N_0\ }\ \mathrm{Beta}(N_1+a,\,N_0+b)$$

any yes/no data with a Beta prior

MLE, MAP, posterior mean
$$\frac{N_1}{N},\qquad \frac{a+N_1-1}{a+b+N-2},\qquad \frac{a+N_1}{a+b+N}$$

one number: no prior, the posterior's peak, the posterior's average

Credible interval
$$\Pr(\theta_l\le\Theta\le\theta_u\mid D)=1-\alpha$$

a probability statement about the rate, given your data

95% confidence interval
$$\bar\theta\pm 2\sqrt{\bar\theta(1-\bar\theta)/n}$$

a frequentist interval for a proportion, $n$ large

Three most common mistakes
  1. Reading a 95% confidence interval as 'the true rate is in this interval with probability $0.95$'. The $0.95$ belongs to the recipe over repeated samples; that sentence needs a credible interval.

  2. Using the posterior-mean formula for the MAP, or the reverse: the MAP has the $-1$ and $-2$, the mean does not.

  3. Adding the total $N$ instead of the tails $N_0$ to the second Beta parameter.

Two course documents disagree on the weights. The STARS syllabus for Fall 2026-27 (printed 21 September) gives Midterm 30%, Final 30%, Problem sets + Quiz 20%, project 20%; the Chapter 1 slides (issued 15 and again 24 September) give 25%, 25%, 20% and 30% in the same order. Confirm with the course which split applies.

How much time do you have?
10 minutes

The update rule that answers most Bayesian questions, and the interval recipe with the one sentence that reads it correctly.

The 60-second card · The Beta prior · Confidence intervals · Formula card
45 minutes

Every estimate and interval of the chapter once, each with a worked example, then one ladder from a full solution to a bare problem.

The 60-second card · Likelihood and posterior · The Beta prior · Maximum likelihood · The MAP estimate · The posterior mean · Credible intervals · Confidence intervals · Scaffolding comes off · Formula card
full read

Where each formula comes from, when each estimate is the right one, and enough mixed practice to decide the method yourself.

The opening pages · Recall first · Likelihood and posterior · The Beta prior · Maximum likelihood · The MAP estimate · The posterior mean · Credible intervals · Confidence intervals · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
  1. Write the likelihood of coin-flip data and turn it into a posterior density with Bayes rule, normalizing when needed.

  2. Update a Beta prior on yes/no data to the posterior $\mathrm{Beta}(N_1+a,\,N_0+b)$, in one batch or one observation at a time.

  3. Derive the MLE $N_1/N$, including the endpoint cases, and apply it to raw 0/1 records and to pooled batches.

  4. Compute the MAP estimate under a Beta prior and explain why it wins when only an exact hit counts.

  5. Compute the posterior mean, write it as a weighted average of the prior mean and the MLE, and justify it under squared loss.

  6. Construct a central or one-sided credible interval from a posterior CDF, or from mean $\pm$ 2 sd when the counts are large.

  7. Construct the 95% confidence interval $\bar\theta\pm 2\,\mathrm{SE}$ for a proportion and state correctly what its 95% means.

Syllabus coverage

Bayesian and frequentist machine learning — covered

  • Likelihood and posterior for coin-flip data
  • the Beta prior and its conjugate update
  • the MLE
  • the MAP estimate and the posterior mean with the loss each one minimizes
  • why a hides uncertainty
  • the frequentist view, in which the parameter is fixed and the data are random

These are the lecture's sub-topics for the chapter, in the lecture's order; they fill the first five blocks and the start of the last one.

credible intervals — covered

The definition, the central interval with $\alpha/2$ in each tail, exact intervals from a closed-form posterior CDF, and one-sided bounds when the posterior peaks at $0$ or $1$.

confidence intervals — covered

The central limit theorem applied to the sample mean, the interval $\bar\theta\pm 2\sqrt{\bar\theta(1-\bar\theta)/n}$, and its reading as a statement about repeated experiments.

their use in data science — covered

  • Click-through, conversion and open rates
  • A/B test intervals
  • a classifier's error rate on a test set
  • sensor packet loss
  • planning a sample size

The applications run through the examples of every block rather than sitting in one place.

normal approximation to a Beta posterior — off syllabus

Mean $\pm$ 2 sd as a hand-calculated credible interval.

Further reading, not on the lecture slides. It is used here only as a calculator-free shortcut, always next to the exact interval.

actual coverage of the two-standard-error recipe — off syllabus

How often the $\pm 2$ SE interval really covers the true rate for small samples and rates near $0$.

Further reading. The lecture states the recipe for large $n$; the table in that block shows by exact calculation how far the coverage falls when $n$ is small.

Recall first
Bayes rule for random variables

$p(y\mid x)=\dfrac{p(x\mid y)\,p(y)}{p(x)}$, with $p(x)=\int p(x\mid y')\,p(y')\,dy'$, or a sum when $y$ is discrete.

The posterior of this section is this formula with the unknown rate in the role of $y$.

Bernoulli and binomial distributions

$X\sim\mathrm{Ber}(\theta)$ takes $1$ with probability $\theta$. The number of ones in $n$ independent trials has $\Pr(S=k)=\binom{n}{k}\theta^k(1-\theta)^{n-k}$, mean $n\theta$ and variance $n\theta(1-\theta)$.

It is the likelihood of every data set on this page.

$X\sim\mathrm{Beta}(x,y)$ has density $\dfrac{1}{B(x,y)}t^{x-1}(1-t)^{y-1}$ on $[0,1]$, mean $\dfrac{x}{x+y}$ and variance $\dfrac{xy}{(x+y)^2(x+y+1)}$.

Priors and posteriors here are Betas; the mean and variance give the posterior mean and the approximate credible interval.

Expectation and variance rules

$E[aX+b]=aE[X]+b$, $\mathrm{Var}(aX)=a^2\mathrm{Var}(X)$, and variances of independent variables add. For an event $A$, $E[I(X\in A)]=\Pr(X\in A)$.

They give the of $\bar\theta$ and the expected losses that single out the MAP and the posterior mean.

Law of large numbers and central limit theorem

For i.i.d. $X_i$ with mean $\mu$ and variance $\sigma^2$, $\bar X_n\to\mu$, and $\dfrac{\bar X_n-\mu}{\sigma/\sqrt n}$ is approximately $\mathcal N(0,1)$ for large $n$; a standard normal lies in $[-2,2]$ with probability about $0.954$.

The confidence interval is built from exactly these two facts.

Hoeffding's inequality

For independent $Z_i\in[0,1]$ with average $\bar Z$: $\Pr(\vert\bar Z-E[\bar Z]\vert\ge\varepsilon)\le 2e^{-2n\varepsilon^2}$.

One practice question compares it with the central-limit interval.

Maximizing on a closed interval

A differentiable function on $[0,1]$ attains its maximum at a critical point or at an endpoint; since $\ln$ is increasing, $f$ and $\ln f$ peak at the same place.

The MLE and the MAP are maximizers, and the all-heads case is decided at an endpoint.

Try it yourself first (2 questions)
1§02.1 — Bayes rule with two coins

A drawer holds two coins: a fair one and a bent one that lands heads with probability $0.8$. You pick one at random and flip it twice: heads, heads.

Find(a) What is the probability that you picked the bent coin?
Given
  • each coin is picked with probability $0.5$

  • $\Pr(H\mid\text{fair})=0.5$, $\Pr(H\mid\text{bent})=0.8$

  • flips independent once the coin is fixed

Hint 1/4

We want a probability about the coin after seeing the flips, so this is Bayes rule with two hypotheses.

Hint 2/4

Bayes rule: $\Pr(\text{bent}\mid HH)\propto\Pr(HH\mid\text{bent})\Pr(\text{bent})$, then divide by the same product summed over both coins.

Hint 3/4

$\Pr(HH\mid\text{bent})=0.8^2=0.64$, $\Pr(HH\mid\text{fair})=0.5^2=0.25$, and both priors are $0.5$.

Hint 4/4

$\Pr(\text{bent}\mid HH)=0.32/0.445\approx 0.719$.

Show solution

Two hypotheses make the denominator a two-term sum, so we compute it directly.

Likelihoods

$$\Pr(HH\mid\text{bent})=0.64,\qquad \Pr(HH\mid\text{fair})=0.25$$

Given the coin, the two flips are independent, so their probabilities multiply: $0.8^2$ and $0.5^2$.

Bayes rule

$$\frac{0.5(0.64)}{0.5(0.64)+0.5(0.25)}=\frac{0.32}{0.445}\approx 0.719$$

Bayes rule: each coin's prior times its likelihood, divided by the total over both coins so the answers add to one.

Answer $$\boxed{\Pr(\text{bent}\mid HH)\approx 0.719}$$
Check

In odds form: prior odds $1$ times likelihood ratio $0.64/0.25=2.56$ gives posterior odds $2.56$, and $2.56/3.56\approx 0.719$. Two heads make the bent coin more likely than not, but far from certain.

Replace the two coins by a whole interval of possible rates and the sum becomes an integral: that is the posterior of this section.

2§02.7 — reading a confidence interval

A report on a survey says that the 95% confidence interval for the true rate is $[0.42,\,0.58]$. A classmate rewrites the sentence in her notes.

Find(a) Is her note a correct reading of the report? Answer true or false.
Given
  • 95% confidence interval $[0.42,\,0.58]$, from one sample

  • her note: there is a $0.95$ probability that the true rate lies in $[0.42,\,0.58]$

Hint 1/4

Ask what is random in the frequentist model: the true rate, or the interval?

Hint 2/4

A confidence interval's $0.95$ is the probability, over repeated samples, that the recipe's interval covers the fixed true rate.

Hint 3/4

Here the recipe produced $[0.42,\,0.58]$ from one sample; the true rate is a fixed number either inside or outside it.

Hint 4/4

So the note is false: the $0.95$ belongs to the recipe, not to this interval.

Show solution

Naming the random quantity settles the question faster than any calculation.

Identify what is random

$$\theta_{\mathrm{true}}\ \text{fixed},\qquad [\bar\theta-2\,\mathrm{SE},\ \bar\theta+2\,\mathrm{SE}]\ \text{random}$$

In the frequentist model only the data, and so the interval, vary from sample to sample.

Read the 95%

$$\Pr_{\text{samples}}(\text{interval covers }\theta_{\mathrm{true}})\approx 0.95$$

The probability is about the recipe; once the numbers $0.42$ and $0.58$ are in, nothing random is left.

Answer $$\boxed{\text{False}}$$
Check

Try a fixed true rate: if it is $0.5$, the statement $0.42\le 0.5\le 0.58$ is simply true, probability $1$; if it is $0.6$, it is false, probability $0$. For a fixed number the answer is never $0.95$.

A probability statement about the parameter itself needs a posterior; that is what a credible interval gives.

Notation
symbolreads asmeanswatch out
$\Theta$

capital theta

the unknown rate treated as a random variable with values in $[0,1]$

Only in the Bayesian view; the frequentist writes the same unknown as a fixed $\theta_{\mathrm{true}}$.

$\theta$

theta

one particular value of the rate, for example $0.3$

Densities are functions of $\theta$; probabilities are statements about $\Theta$.

$N_1,\ N_0,\ N$

N one, N zero, N

numbers of heads (successes), tails (failures) and flips, with $N=N_0+N_1$

$N_1$ always feeds the first Beta parameter.

$D$

D

the observed data, $\{N_1\text{ heads},\,N_0\text{ tails}\}$

For coin-flip data only the counts matter, not the order.

$p_{D\mid\Theta}(D\mid\theta)$

the likelihood of D given theta

probability of the observed data if the rate were $\theta$

As a function of $\theta$ it does not integrate to $1$.

$p_{\Theta\mid D}(\theta\mid D)$

the posterior density

the density of the rate after seeing $D$

This one does integrate to $1$ over $\theta$.

$\mathrm{Beta}(a,b)$

Beta a b

the density $\theta^{a-1}(1-\theta)^{b-1}/B(a,b)$ on $[0,1]$

$B(a,b)$ is only the constant that makes the area $1$.

$\hat\theta_{\mathrm{MLE}},\ \hat\theta_{\mathrm{MAP}}$

theta hat MLE, theta hat MAP

the maximizers of the likelihood and of the posterior

The hat marks a number computed from data.

$E[\Theta\mid D]$

the posterior mean

the average of $\Theta$ under the posterior

Equal to the MAP only when the posterior is symmetric.

$\mathrm{Cred}_\alpha(D)$

the credible interval at level one minus alpha

an interval $[\theta_l,\theta_u]$ holding posterior probability $1-\alpha$

$\alpha=0.05$ gives a 95% interval, with $0.025$ in each tail if it is central.

$\theta_{\mathrm{true}},\ S_n,\ \bar\theta$

theta true, S n, theta bar

the fixed unknown rate, the number of successes in $n$ trials, and $S_n/n$

$\bar\theta$ changes from sample to sample; $\theta_{\mathrm{true}}$ does not.

$\mathrm{SE}$

standard error

$\sqrt{\bar\theta(1-\bar\theta)/n}$, the estimated standard deviation of $\bar\theta$

It shrinks like $1/\sqrt n$, not like $1/n$.

Conventions used here
Capital and small theta.

$\Theta$ is the unknown rate as a random variable (the Bayesian view) and $\theta$ is a value in $[0,1]$. In the frequentist blocks the unknown rate is a fixed number, $\theta_{\mathrm{true}}$, and nothing about it is random.

Most interval mistakes come from forgetting which symbol carries the randomness.

Which count feeds which parameter.

$N_1$ counts the ones (heads, clicks, defects, whatever the question counts) and goes into the first Beta parameter; $N_0$ counts the zeros and goes into the second.

Defects or losses can be the ones; the rule follows what is counted, not whether it is good news.

The 95% multiplier.

Following the lecture, a 95% confidence interval steps $2$ standard errors each side. The exact normal value is $1.96$, which gives probability $0.950$ instead of $0.954$. If a question fixes the multiplier use it; otherwise say which one you used.

Answers built with $2$ and with $1.96$ differ in the third decimal, and both appear in textbooks.

Intervals are closed and live in [0, 1].

All intervals here are closed, $[\theta_l,\theta_u]$. If an approximate interval pokes outside $[0,1]$, the approximation is being used outside its safe range, and we say so instead of quietly clipping it.

A lower end of $-0.006$ is a warning sign about the method, not a rate.

What the probability refers to.

In a credible interval the probability is over $\Theta$, given the data you have. In a confidence interval it is over repeated data sets, with $\theta_{\mathrm{true}}$ held fixed.

The two can give almost the same numbers and still mean different things.

Rounding.

Intermediate steps keep at least four significant figures; final answers are rounded to three decimals unless an exact fraction is asked for.

Rounding a standard error early can move an interval end by one unit in the third decimal.

2.1Likelihood and posterior: one formula, read two ways

Treats the unknown rate as a random variable and uses Bayes rule to turn counts into a distribution over that rate.

The previous section gave us Bayes rule for random variables; here we point it at the unknown probability itself.

Solvable with what we have
  • Compute the chance of the data for a given rate: if $\theta=0.7$, four acceptances in four happen with probability $0.7^4\approx 0.24$.

  • Compare two named rates with Bayes rule, say $0.9$ against $0.5$.

  • Say where the fraction of successes is heading as the sample grows, by the law of large numbers.

Not solvable yet
  • Say how probable it is that the true rate exceeds $0.9$, given only four listeners.

  • Give an estimate after four out of four that is not the absurd value $1$.

  • Attach any error bar to a rate measured on four people.

Take the fraction as the truth: $\hat\theta=4/4=1$. Then the next $100$ listeners all accept with probability $1^{100}=1$, and the team can promise it.

Why it fails

A rate of $0.8$ produces four out of four with probability $0.8^4\approx 0.41$, which is hardly rare. The fraction makes a rate above $0.9$ certain; a distribution over $\theta$ will put that chance near $0.41$. One number throws the whole range of plausible rates away.

RuleBayes rule for a parameter
Conditions
  • $\Theta$ takes values in $[0,1]$ and has prior density $p_\Theta(\theta)$

  • given $\Theta=\theta$, the flips $X_1,\dots,X_N$ are i.i.d. $\mathrm{Ber}(\theta)$

  • the data $D$ record $N_1$ heads and $N_0$ tails, $N=N_1+N_0$

$$\boxed{\;\textcolor{#d1690a}{p_{\Theta\mid D}(\theta\mid D)}=\frac{\textcolor{#1f6feb}{p_{D\mid\Theta}(D\mid\theta)}\,\textcolor{#8250df}{p_\Theta(\theta)}}{p_D(D)}\;\propto\;\textcolor{#1f6feb}{\theta^{N_1}(1-\theta)^{N_0}}\,\textcolor{#8250df}{p_\Theta(\theta)}\;}$$

After the data, each candidate rate gets a weight equal to its prior weight times how well it explains the data. The denominator $p_D(D)$ does not involve $\theta$; it only rescales the curve so that its area is $1$.

Where the formula comes from

Take one particular sequence, say heads, heads, tails. Independence lets us multiply: $\theta\cdot\theta\cdot(1-\theta)$. Any sequence with $N_1$ heads and $N_0$ tails has probability $\theta^{N_1}(1-\theta)^{N_0}$, whatever the order.

The data record only the counts, and $\binom{N_1+N_0}{N_1}$ orders give the same counts, so $$p_{D\mid\Theta}(D\mid\theta)=\binom{N_1+N_0}{N_1}\theta^{N_1}(1-\theta)^{N_0}.$$

Bayes rule for random variables, with $\Theta$ as the unknown and $D$ as the observation, gives $$p_{\Theta\mid D}(\theta\mid D)=\frac{p_{D\mid\Theta}(D\mid\theta)\,p_\Theta(\theta)}{p_D(D)},\qquad p_D(D)=\int_0^1 p_{D\mid\Theta}(D\mid t)\,p_\Theta(t)\,dt.$$

Neither the binomial coefficient nor $p_D(D)$ contains $\theta$. Dropping them changes the height of the curve, not its shape, which is why we write $\propto$ and normalize at the end.

Looks like this, but is not

The likelihood $\theta^{4}$ from four out of four is a nonnegative curve on $[0,1]$, so it looks like a density for $\theta$.

Its area is $\int_0^1\theta^4\,d\theta=\tfrac15$, not $1$. For each fixed $\theta$ the likelihood is a distribution over data; as a function of $\theta$ it is only a weight. Times the prior and divided by $p_D(D)$, it becomes the density $5\theta^4$.

Four acceptances out of four: the chance that the rate exceeds 0.9

A playlist suggestion is shown to $4$ listeners and all $4$ accept. Take a flat prior, $p_\Theta(\theta)=1$ on $[0,1]$, and find the posterior probability that the acceptance rate is above $0.9$.

Find$\Pr(\Theta>0.9\mid D)$
Given
  • $N_1=4$ acceptances, $N_0=0$ refusals

  • prior $p_\Theta(\theta)=1$ for $0\le\theta\le 1$

Solution

We write the posterior up to a constant and fix the constant with one easy integral; computing $p_D(D)$ separately would cost the same integral twice.

Write the posterior up to a constant

$$p(\theta\mid D)\propto\theta^{4}(1-\theta)^{0}\cdot 1=\theta^{4}$$

Likelihood times prior. The binomial coefficient is $1$ here and does not involve $\theta$ anyway.

Fix the constant

$$\int_0^1 c\,\theta^4\,d\theta=\frac{c}{5}=1\ \Rightarrow\ c=5$$

A density must have total area $1$, and $c$ is the only unknown left.

$$\textcolor{#d1690a}{p(\theta\mid D)=5\theta^{4}},\quad 0\le\theta\le 1$$

With $c=5$ the area is exactly $1$; this is the orange curve in the figure.

Integrate over the region asked for

$$\Pr(\Theta>0.9\mid D)=\int_{0.9}^{1}5\theta^4\,d\theta=\Big[\theta^5\Big]_{0.9}^{1}$$

A probability for a continuous variable is an area under its density.

$$=1-0.9^{5}=1-0.59049=0.40951$$

$0.9^5=0.59049$ by repeated multiplication.

Answer $$\boxed{\Pr(\Theta>0.9\mid D)=1-0.9^{5}\approx 0.41}$$
Check

The posterior median solves $\theta^5=0.5$, so it is $0.5^{1/5}\approx 0.871<0.9$; the region above $0.9$ must therefore hold less than half the area, and $0.41$ does. Under the flat prior the same region held $0.1$, so four acceptances roughly quadruple it.

The dashboard's 100% was the wrong kind of answer: a rate above $0.9$ gets about 41 chances in 100. A normalized posterior turns vague plausibility into numbers you can report.

Two candidate rates: Bayes rule with a sum instead of an integral

A product lead believes a new suggestion is either average, $\theta=0.5$, or excellent, $\theta=0.9$, and puts probability $0.8$ on average. Then $4$ listeners out of $4$ accept. How probable is excellent now?

Find$\Pr(\Theta=0.9\mid D)$
Given
  • two candidate rates: $0.5$ with prior $0.8$, and $0.9$ with prior $0.2$

  • $4$ acceptances in $4$

Solution

With only two candidates the evidence $p_D(D)$ is a two-term sum, so we compute it directly instead of normalizing a curve.

Likelihood of each candidate

$$p(D\mid 0.9)=0.9^4=0.6561,\qquad p(D\mid 0.5)=0.5^4=0.0625$$

Four independent acceptances, so the rate is raised to the fourth power.

Prior times likelihood, then normalize

$$0.2\times 0.6561=0.13122,\qquad 0.8\times 0.0625=0.05$$

Bayes rule multiplies each candidate's prior by its likelihood; these weights are not yet probabilities.

$$\Pr(\Theta=0.9\mid D)=\frac{0.13122}{0.13122+0.05}\approx 0.724$$

The denominator is $p_D(D)$, the total probability of the data.

Answer $$\boxed{\Pr(\Theta=0.9\mid D)\approx 0.724}$$
Check

In odds form: prior odds $0.2/0.8=0.25$ times likelihood ratio $0.6561/0.0625=10.4976$ gives posterior odds $2.624$, and $2.624/3.624\approx 0.724$. The data were enough to overturn a prior that was $4$ to $1$ against excellent.

Continuous or discrete, the recipe is the same: prior times likelihood, then divide by the total so the probabilities add to one.

Checkpoint
§02.1 — posterior shape from prior and likelihood

A colleague's prior for a coin's heads rate is tilted toward heads: $p_\Theta(\theta)=2\theta$ on $[0,1]$. The coin is then flipped $5$ times.

Find(a) Which expression is proportional to the posterior density $p_{\Theta\mid D}(\theta\mid D)$?
Given
  • $2$ heads and $3$ tails

  • prior $p_\Theta(\theta)=2\theta$ for $0\le\theta\le 1$

Hint 1/4

We need the shape of the posterior as a function of $\theta$, so any factor without $\theta$ can be dropped.

Hint 2/4

Posterior $\propto$ likelihood $\times$ prior: $p(\theta\mid D)\propto\theta^{N_1}(1-\theta)^{N_0}\,p_\Theta(\theta)$.

Hint 3/4

Here $N_1=2$, $N_0=3$ and $p_\Theta(\theta)=2\theta$, so the product is $\theta^{2}(1-\theta)^{3}\cdot 2\theta$.

Hint 4/4

Dropping the constant $2$ leaves $\theta^{3}(1-\theta)^{3}$.

Show solution

Multiplying first and simplifying second avoids losing the prior's factor of $\theta$, which is the whole point of the question.

Multiply likelihood and prior

$$\theta^{2}(1-\theta)^{3}\cdot 2\theta=2\,\theta^{3}(1-\theta)^{3}$$

Bayes rule multiplies; the prior's $\theta$ raises the heads exponent by one.

Drop what does not depend on theta

$$p(\theta\mid D)\propto\theta^{3}(1-\theta)^{3}$$

The factor $2$ disappears when we normalize anyway.

Answer $$\boxed{p(\theta\mid D)\propto\theta^{3}(1-\theta)^{3}}$$
Check

The result is symmetric about $0.5$: two heads plus a prior worth one extra head balance three tails, so a symmetric posterior is exactly what we should expect.

A prior of the form $\theta^{k}$ acts like $k$ extra heads; keep an eye out for that pattern in the next block.

⚠ Treating the likelihood as a density over the rate

it is a nonnegative curve in $\theta$, so it looks like one

wrong$$\int_0^1\theta^{N_1}(1-\theta)^{N_0}\,d\theta=1$$
right$$\int_0^1\theta^{4}\,d\theta=\tfrac15\neq 1\ \Rightarrow\ \text{normalize the posterior, not the likelihood}$$
⚠ Dropping a prior that is not flat

with the flat prior the posterior is just the normalized likelihood, and the habit sticks

wrong$$p(\theta\mid D)\propto\theta^{2}(1-\theta)^{3}\ \ \text{under the prior }2\theta$$
right$$p(\theta\mid D)\propto\theta^{2}(1-\theta)^{3}\cdot 2\theta\propto\theta^{3}(1-\theta)^{3}$$

2.2The Beta prior: updating a belief by adding counts

A Beta prior makes the posterior another Beta, so updating on data means adding the head and tail counts to its parameters.

The flat prior made the integral easy; to keep it easy for any prior belief, we choose a prior with the same shape as the likelihood.

TheoremConjugate update: Beta prior, Bernoulli data
Conditions
  • prior $\Theta\sim\mathrm{Beta}(a,b)$ with $a>0$ and $b>0$

  • given $\Theta=\theta$, the flips are i.i.d. $\mathrm{Ber}(\theta)$, with $N_1$ heads and $N_0$ tails

$$\boxed{\;\textcolor{#8250df}{\mathrm{Beta}(a,b)}\ \xrightarrow{\ N_1\text{ heads},\ N_0\text{ tails}\ }\ \textcolor{#d1690a}{\Theta\mid D\sim\mathrm{Beta}(N_1+a,\;N_0+b)}\;}$$

Add the heads to the first parameter and the tails to the second. For the posterior mean, $\mathrm{Beta}(a,b)$ acts exactly like $a$ earlier heads and $b$ earlier tails, so a large $a+b$ means a stubborn prior.

Why the posterior stays a Beta

The prior density is $p_\Theta(\theta)=\dfrac{1}{B(a,b)}\theta^{a-1}(1-\theta)^{b-1}$ on $[0,1]$.

Multiply by the likelihood and keep only what depends on $\theta$: $$p(\theta\mid D)\propto\theta^{N_1}(1-\theta)^{N_0}\cdot\theta^{a-1}(1-\theta)^{b-1}=\theta^{N_1+a-1}(1-\theta)^{N_0+b-1}.$$

The right side has the $\mathrm{Beta}(N_1+a,\,N_0+b)$ shape. A density with that shape can have only one constant in front, $1/B(N_1+a,\,N_0+b)$, because its area must be $1$. So the posterior is exactly that Beta.

A prior whose posterior stays in the same family is called conjugate to the likelihood. The Beta family is conjugate to Bernoulli and binomial data.

Looks like this, but is not

A normal prior on the rate, say $\mathcal N(0.5,\,0.1^2)$, looks as sensible as a Beta: smooth, one peak, centred where we want it.

Multiplied by $\theta^{N_1}(1-\theta)^{N_0}$ it gives neither a normal nor a Beta, so there is no closed form and every summary needs numerical integration. It also puts a little probability below $0$ and above $1$, where a rate cannot be.

Email open rate: a Beta(2, 6) prior, then 7 opens in 10

Past campaigns suggest open rates around $0.25$, encoded as the prior $\mathrm{Beta}(2,6)$. A new subject line goes to $10$ test users and $7$ open it. Find the posterior and the posterior probability that the new open rate exceeds $0.5$.

FindThe posterior and $\Pr(\Theta>0.5\mid D)$.
Given
  • prior $\mathrm{Beta}(2,6)$, mean $2/8=0.25$

  • $N_1=7$ opens, $N_0=3$ non-opens

Solution

The conjugate rule gives the posterior in one line, so no integral is needed for the first part; for the second, the posterior turns out symmetric, and symmetry answers without integrating.

Update the parameters

$$\Theta\mid D\sim\mathrm{Beta}(7+2,\;3+6)=\mathrm{Beta}(9,9)$$

Opens go to the first parameter, non-opens to the second.

Use the symmetry of the posterior

$$p(\theta\mid D)\propto\theta^{8}(1-\theta)^{8}$$

Swapping $\theta$ and $1-\theta$ leaves this unchanged, so the density is symmetric about $0.5$.

$$\Pr(\Theta>0.5\mid D)=\Pr(\Theta<0.5\mid D)=\tfrac12$$

The two halves have equal area and together make $1$.

Answer $$\boxed{\Theta\mid D\sim\mathrm{Beta}(9,9),\qquad \Pr(\Theta>0.5\mid D)=0.5}$$
Check

The posterior mean $9/18=0.5$ lies between the prior mean $0.25$ and the data fraction $0.7$, as a compromise must. Ten real users against a prior worth $8$ earlier ones should land near the middle, and it does.

Seven opens in ten did not prove the new line beats one half; they pulled a pessimistic prior up to even odds.

Two days of clicks: updating twice gives the same posterior as updating once

Start from the flat prior $\mathrm{Beta}(1,1)$ for a click rate. Monday brings $3$ clicks in $8$ views; Tuesday brings $5$ clicks in $12$ views. Compare updating after each day with updating once on the pooled counts.

FindThe posterior after both days, computed both ways.
Given
  • prior $\mathrm{Beta}(1,1)$

  • Monday: $3$ clicks, $5$ skips

  • Tuesday: $5$ clicks, $7$ skips

Solution

Monday's posterior becomes Tuesday's prior; doing it both ways is the quickest proof that batching the data does not matter.

Update day by day

$$\mathrm{Beta}(1,1)\to\mathrm{Beta}(1+3,\,1+5)=\mathrm{Beta}(4,6)$$

Monday's $3$ clicks join the first parameter and its $5$ skips the second, starting from the flat prior.

$$\mathrm{Beta}(4,6)\to\mathrm{Beta}(4+5,\,6+7)=\mathrm{Beta}(9,13)$$

Monday's posterior serves as Tuesday's prior.

Update once on the totals

$$\mathrm{Beta}(1+8,\,1+12)=\mathrm{Beta}(9,13)$$

Eight clicks and twelve skips over the two days.

Answer $$\boxed{\Theta\mid D\sim\mathrm{Beta}(9,13),\qquad E[\Theta\mid D]=\tfrac{9}{22}\approx 0.409}$$
Check

Both routes add the same numbers to the same parameters in a different order, so they must agree. The pooled fraction $8/20=0.4$ sits close to the posterior mean $0.409$, as it should with a prior worth only two views.

A running posterior is all a system has to store: two numbers, updated one observation at a time.

Checkpoint
§02.2 — conjugate Beta update

A coin collector believes her coins are close to fair and encodes that as the prior $\mathrm{Beta}(4,4)$. She flips one coin $20$ times.

Find(a) Which distribution is the posterior for the heads rate?
Given
  • prior $\mathrm{Beta}(4,4)$

  • $15$ heads in $20$ flips

Hint 1/4

Split the $20$ flips into heads and tails before touching the prior.

Hint 2/4

$\mathrm{Beta}(a,b)$ prior with $N_1$ heads and $N_0$ tails gives $\mathrm{Beta}(N_1+a,\,N_0+b)$.

Hint 3/4

Here $a=b=4$, $N_1=15$ and $N_0=20-15=5$.

Hint 4/4

The posterior is $\mathrm{Beta}(19,9)$.

Show solution

The conjugate rule needs the tails count, which the question hides inside the total; getting it first prevents the most common slip.

Get both counts

$$N_0=20-15=5$$

The second parameter wants tails, not flips.

Add them to the prior

$$\mathrm{Beta}(15+4,\;5+4)=\mathrm{Beta}(19,9)$$

A Beta prior with Bernoulli data stays Beta: heads join $a$, tails join $b$.

Answer $$\boxed{\mathrm{Beta}(19,9)}$$
Check

The parameters add to $28=8+20$: the prior's $8$ pseudo-flips plus the $20$ real ones. A total that does not match signals a slip.

Always check that the posterior parameters add up to $a+b+N$.

⚠ Adding all the flips to the second parameter

the second number is easy to read as the total

wrong$$\mathrm{Beta}(N_1+a,\;N+b)$$
right$$\mathrm{Beta}(N_1+a,\;N_0+b)$$
⚠ Swapping the roles of heads and tails

the tails count is the second number in the data and feels like it belongs first somewhere

wrong$$\mathrm{Beta}(N_0+a,\;N_1+b)$$
right$$\mathrm{Beta}(N_1+a,\;N_0+b)$$
00.20.40.60.810123θ, the click ratedensityafter 8 views: 5 clicks, 3 skipsposterior Beta(6, 4)

Eight views arrive one at a time: click, skip, click, click, skip, click, skip, click. Step back to the flat prior and forward again. Each click adds $1$ to the first parameter, each skip adds $1$ to the second, and the curve narrows as the counts grow.

At the edges
0 views Beta(1, 1)

Nothing seen yet, so the posterior is the flat prior.

8 views Beta(6, 4)

Five clicks and three skips: the same curve you get by updating once on the totals.

2.3Maximum likelihood: the rate that makes the data most probable

Picks the value of $\theta$ that makes the observed data most probable; for coin flips, the fraction of heads.

A posterior needs a prior; maximum likelihood drops the prior and asks only which $\theta$ explains the data best.

Theorem for Bernoulli data
Conditions
  • i.i.d. $\mathrm{Ber}(\theta)$ flips with $N_1$ heads and $N_0$ tails, $N=N_1+N_0\ge 1$

  • $\theta$ ranges over the closed interval $[0,1]$

$$\boxed{\;\textcolor{#d1690a}{\hat\theta_{\mathrm{MLE}}}:=\arg\max_{\theta\in[0,1]}\textcolor{#1f6feb}{p(D\mid\theta)}=\frac{N_1}{N_0+N_1}\;}$$

Of all the rates, the one that makes the observed counts most probable is the plain fraction of heads. No prior enters; only the data and the form of the likelihood do.

Derivation, with the endpoint check

Maximize $L(\theta)=\theta^{N_1}(1-\theta)^{N_0}$ on $[0,1]$; the binomial coefficient does not move the maximizer. Since $\ln$ is increasing, $L$ and its logarithm peak at the same place: $$\ell(\theta)=N_1\ln\theta+N_0\ln(1-\theta),\qquad 0<\theta<1.$$

Set the derivative to zero: $$\ell'(\theta)=\frac{N_1}{\theta}-\frac{N_0}{1-\theta}=0\ \Longrightarrow\ N_1(1-\theta)=N_0\,\theta\ \Longrightarrow\ \theta=\frac{N_1}{N_0+N_1}.$$

It is a maximum: $\ell''(\theta)=-\dfrac{N_1}{\theta^2}-\dfrac{N_0}{(1-\theta)^2}<0$ on $(0,1)$ when both counts are positive.

Endpoints: with $N_1\ge 1$ and $N_0\ge 1$, $L(0)=L(1)=0$, so the interior point wins. With $N_0=0$, $L(\theta)=\theta^{N_1}$ rises all the way and the maximum is $\theta=1=N_1/N$; $N_1=0$ mirrors it at $0$. The formula holds in every case.

Looks like this, but is not

With all heads, say $N_1=3$ and $N_0=0$, it seems we can still find the MLE by solving $\ell'(\theta)=0$.

Now $\ell'(\theta)=3/\theta$ is positive on all of $(0,1)$ and never zero. The likelihood $\theta^3$ keeps rising, so the maximum is the endpoint $\theta=1$. The answer $N_1/N=1$ survives, but only the endpoint check finds it.

θL(θ) = θ³(1 − θ)⁷

$0.1$

$0.000478$

$0.2$

$0.001678$

$0.3$

$0.002224$

$0.4$

$0.001792$

$0.5$

$0.000977$

The column rises to $\theta=0.3$ and falls after it: $0.002224$ beats both neighbours, $0.001678$ and $0.001792$. The derivative found the same point without a grid.

Packet loss on a sensor link: the MLE from a 0/1 record

A sensor node sends $12$ packets and logs $1$ for a lost packet and $0$ for a delivered one: $$0,0,1,0,0,0,0,1,0,0,0,0.$$ Model the losses as i.i.d. $\mathrm{Ber}(\theta)$ and find the maximum likelihood estimate of the loss rate.

Find$\hat\theta_{\mathrm{MLE}}$
Given
  • the 12-packet record above

  • losses i.i.d. $\mathrm{Ber}(\theta)$

Solution

Counting first and applying the formula is faster than differentiating a 12-factor product; we still test the answer against its neighbours, because a miscount fails silently.

Count

$$N_1=2\ \text{lost},\quad N_0=10\ \text{delivered},\quad N=12$$

Only the counts enter the likelihood, not where the ones sit.

Maximize

$$\hat\theta_{\mathrm{MLE}}=\frac{N_1}{N}=\frac{2}{12}=\frac16\approx 0.167$$

Both counts are positive, so the interior critical point is the maximum.

Check against neighbours

$$L(\tfrac16)=(\tfrac16)^2(\tfrac56)^{10}\approx 0.00449$$

We evaluate the likelihood at the candidate first, so the neighbours have a number to beat.

$$L(0.1)\approx 0.00349,\qquad L(0.25)\approx 0.00352$$

Both are smaller, as they must be if $1/6$ is the maximum.

Answer $$\boxed{\hat\theta_{\mathrm{MLE}}=\tfrac16\approx 0.167}$$
Check

The log-likelihood $2\ln\theta+10\ln(1-\theta)$ has derivative $2/\theta-10/(1-\theta)$, and at $\theta=1/6$ this is $12-12=0$, so $1/6$ is the critical point.

With Bernoulli data the MLE never needs the full product: count, divide, and make sure the result is a fraction between $0$ and $1$.

Three days of one banner: pool the counts, do not average the rates

An online shop runs the same banner on three days: $4$ clicks in $20$ views, $7$ in $30$, and $1$ in $10$. Assuming one click rate $\theta$ for all views, find its MLE.

Find$\hat\theta_{\mathrm{MLE}}$ for the common rate.
Given
  • day 1: $4$ of $20$

  • day 2: $7$ of $30$

  • day 3: $1$ of $10$

  • all views i.i.d. $\mathrm{Ber}(\theta)$

Solution

Independent days multiply, so their counts add; averaging the three daily fractions would give a 10-view day the same say as a 30-view day.

Multiply the three likelihoods

$$L(\theta)=\theta^{4}(1-\theta)^{16}\cdot\theta^{7}(1-\theta)^{23}\cdot\theta^{1}(1-\theta)^{9}$$

The days are independent, so their probabilities multiply.

$$=\theta^{12}(1-\theta)^{48}$$

Exponents add: $4+7+1$ clicks and $16+23+9$ skips.

Apply the formula to the pooled counts

$$\hat\theta_{\mathrm{MLE}}=\frac{12}{12+48}=0.2$$

Same maximization as before with $N_1=12$ and $N_0=48$.

Answer $$\boxed{\hat\theta_{\mathrm{MLE}}=0.2}$$
Check

At $\theta=0.2$ the expected number of clicks in $60$ views is $12$, exactly the total observed. The average of the daily fractions, $\tfrac13\big(\tfrac{4}{20}+\tfrac{7}{30}+\tfrac{1}{10}\big)\approx 0.178$, would predict only about $10.7$.

Pool counts, not fractions: the MLE weights every observation equally, whatever batch it came in.

Checkpoint
§02.3 — MLE from counts

A quality engineer tests $20$ solder joints and $5$ of them fail. She models failures as i.i.d. $\mathrm{Ber}(\theta)$.

Find(a) Which value of $\theta$ maximizes $\theta^{5}(1-\theta)^{15}$?
Given$N_1=5$ failures, $N_0=15$ good joints
Hint 1/4

We want the peak of the likelihood, and for Bernoulli data the peak has a closed form.

Hint 2/4

$\hat\theta_{\mathrm{MLE}}=N_1/(N_0+N_1)$.

Hint 3/4

Here $N_1=5$ and $N_0=15$, so $N=20$.

Hint 4/4

The maximizer is $5/20=0.25$.

Show solution

The closed form saves differentiating; a one-line check with the log-derivative confirms it.

Apply the formula

$$\hat\theta=\frac{5}{5+15}=0.25$$

The MLE of a Bernoulli rate is the count of ones over the total: failures over all joints.

Confirm with the derivative

$$\frac{5}{0.25}-\frac{15}{0.75}=20-20=0$$

The log-likelihood's slope is zero there.

Answer $$\boxed{\hat\theta_{\mathrm{MLE}}=0.25}$$
Check

The neighbours do worse: $L(0.2)\approx 1.13\times 10^{-5}$ and $L(0.3)\approx 1.15\times 10^{-5}$, against $L(0.25)\approx 1.31\times 10^{-5}$.

Divide by the total, never by the other count.

⚠ Dividing heads by tails instead of by the total

the ratio of the counts is what the data look like

wrong$$\hat\theta=\frac{N_1}{N_0}=\frac{5}{15}$$
right$$\hat\theta=\frac{N_1}{N_0+N_1}=\frac{5}{20}$$
⚠ Averaging batch fractions instead of pooling counts

each batch already comes with a fraction, and averaging feels fair

wrong$$\tfrac13\big(\tfrac{4}{20}+\tfrac{7}{30}+\tfrac{1}{10}\big)\approx 0.178$$
right$$\frac{4+7+1}{20+30+10}=0.2$$

2.4The MAP estimate: the peak of the posterior and the best exact guess

The posterior's peak: an MLE on counts inflated by the prior, and the best guess when only an exact hit scores.

The posterior is a whole curve, but a spam filter or a pricing rule needs one number; the first candidate is the curve's highest point.

TheoremMAP estimate under a Beta prior
Conditions
  • prior $\mathrm{Beta}(a,b)$ and data with $N_1$ heads and $N_0$ tails

  • $a+N_1>1$ and $b+N_0>1$, so the peak lies inside $(0,1)$

$$\boxed{\;\textcolor{#d1690a}{\hat\theta_{\mathrm{MAP}}}:=\arg\max_{\theta}p(\theta\mid D)=\frac{a+N_1-1}{a+b+N_0+N_1-2}\;}$$

The MAP estimate is the MLE computed on inflated counts: $a-1$ extra heads and $b-1$ extra tails. The flat prior $\mathrm{Beta}(1,1)$ adds no extras, so there MAP equals MLE.

Derivation, and why the peak is the right guess for exact hits

The posterior is $\mathrm{Beta}(\alpha,\beta)$ with $\alpha=a+N_1$ and $\beta=b+N_0$, so we maximize $\theta^{\alpha-1}(1-\theta)^{\beta-1}$. That is the MLE problem with $N_1$ replaced by $\alpha-1$ and $N_0$ by $\beta-1$: $$\hat\theta_{\mathrm{MAP}}=\frac{\alpha-1}{\alpha+\beta-2}=\frac{a+N_1-1}{a+b+N_0+N_1-2}.$$

Suppose $\Theta$ can take only a few values and you score only by naming it exactly. Under the loss $I(\hat\theta\neq\Theta)$ your expected loss is $$E[I(\hat\theta\neq\Theta)\mid D]=\Pr(\Theta\neq\hat\theta\mid D)=1-\Pr(\Theta=\hat\theta\mid D),$$ which is smallest for the most probable value.

For a continuous $\Theta$ every single point has probability $0$, so exact hits are impossible. The honest version counts a miss when $\vert\hat\theta-\Theta\vert>\varepsilon$; the best guess is the centre of the heaviest window of width $2\varepsilon$, and as $\varepsilon$ shrinks that centre approaches the peak.

Looks like this, but is not

The posterior mean $\dfrac{a+N_1}{a+b+N}$ looks like the MAP formula with the $-1$ and $-2$ tidied away, and for the posterior $\mathrm{Beta}(9,9)$ both give $0.5$.

They agree only when the posterior is symmetric. For $\mathrm{Beta}(2,5)$ the peak is $\tfrac{1}{5}=0.2$ but the mean is $\tfrac27\approx 0.286$. The subtractions are what move you from the average to the peak.

Click rate with a Beta(3, 7) prior after 12 clicks in 40 views

A recommender team encodes earlier experience as the prior $\mathrm{Beta}(3,7)$ for a new item's click rate. The item then gets $12$ clicks in $40$ views. Find the MAP estimate and compare it with the MLE and with the prior's own peak.

Find$\hat\theta_{\mathrm{MAP}}$, with $\hat\theta_{\mathrm{MLE}}$ and the prior mode for comparison.
Given
  • prior $\mathrm{Beta}(3,7)$

  • $N_1=12$ clicks, $N_0=28$ skips

Solution

Updating first and taking the peak second keeps each step checkable; one long substitution into the formula works too but hides where an error came from.

Update

$$\Theta\mid D\sim\mathrm{Beta}(12+3,\;28+7)=\mathrm{Beta}(15,35)$$

Clicks to the first parameter, skips to the second.

Take the peak

$$\hat\theta_{\mathrm{MAP}}=\frac{15-1}{15+35-2}=\frac{14}{48}\approx 0.292$$

Mode of $\mathrm{Beta}(\alpha,\beta)$; both parameters exceed $1$, so the peak is interior.

Compare with the two ingredients

$$\hat\theta_{\mathrm{MLE}}=\frac{12}{40}=0.3,\qquad \text{prior mode}=\frac{3-1}{3+7-2}=0.25$$

The MAP pools the prior's with the data, so its two ingredients are the prior's peak and the MLE.

$$\frac{8}{48}(0.25)+\frac{40}{48}(0.3)\approx 0.2917$$

The prior's peak counts as $8$ pseudo-views and the data as $40$ real ones.

Answer $$\boxed{\hat\theta_{\mathrm{MAP}}=\tfrac{14}{48}\approx 0.292}$$
Check

It sits between the prior mode $0.25$ and the MLE $0.3$, much closer to the MLE because $40$ real views outweigh $8$ pseudo-views; the weighted-average line reproduces it independently.

Read a MAP as data plus pseudo-counts: once the real counts dwarf $a+b-2$, MAP and MLE are nearly the same number.

Three possible defect rates: the MAP minimizes the chance of being wrong

A supplier's machine runs at one of three defect rates, $0.1$, $0.3$ or $0.5$, with prior probabilities $0.6$, $0.3$ and $0.1$. A sample of $5$ items has $2$ defectives. You must name the rate and are scored only on naming it exactly. Which do you name, and how often will you be wrong?

FindThe rate to name and the probability that it is wrong.
Given
  • candidates $0.1,\ 0.3,\ 0.5$ with prior $0.6,\ 0.3,\ 0.1$

  • $2$ defectives in $5$ items

Solution

With three candidates Bayes rule is a three-term sum, so we build the whole posterior table; the binomial coefficient is common to all rows and can be left out.

Likelihood of each candidate

$$\theta^2(1-\theta)^3:\quad 0.00729,\ \ 0.03087,\ \ 0.03125$$

Two defectives and three good items, for $\theta=0.1,\ 0.3,\ 0.5$; the binomial coefficient is the same for all three and is left out.

Prior times likelihood, then normalize

$$0.004374,\ \ 0.009261,\ \ 0.003125$$

Each likelihood times its prior: $0.6$, $0.3$ and $0.1$.

$$\text{sum}=0.01676\ \Rightarrow\ \Pr(\Theta=\theta\mid D)=0.261,\ 0.553,\ 0.186$$

Divide each weight by the sum so the three add to $1$.

Pick the most probable value

$$\hat\theta_{\mathrm{MAP}}=0.3,\qquad \Pr(\text{wrong}\mid D)=1-0.553=0.447$$

Under the exact-hit loss the expected loss of a guess is one minus its posterior probability.

Answer $$\boxed{\text{name }0.3;\ \ \Pr(\text{wrong})\approx 0.447}$$
Check

Naming $0.5$ would be wrong with probability $1-0.186=0.814$ and naming $0.1$ with $1-0.261=0.739$, both worse. By a hair, $0.5$ has the larger likelihood, $0.03125$ against $0.03087$, so the MLE among the three would be $0.5$: the prior is what moves the answer.

When only an exact hit counts, name the value with the highest posterior probability, even if another value explains the data slightly better.

Checkpoint
§02.4 — MAP with a Beta prior

A coach's prior for a player's free-throw rate is $\mathrm{Beta}(2,2)$, mild and centred at one half. In practice the player makes $9$ of $10$.

Find(a) What is the MAP estimate of the free-throw rate?
Given
  • prior $\mathrm{Beta}(2,2)$

  • $9$ made, $1$ missed

Hint 1/4

The MAP is the peak of the posterior, so get the posterior parameters first.

Hint 2/4

$\hat\theta_{\mathrm{MAP}}=\dfrac{a+N_1-1}{a+b+N-2}$.

Hint 3/4

With $a=b=2$, $N_1=9$ and $N=10$ this is $\dfrac{2+9-1}{2+2+10-2}$.

Hint 4/4

$\hat\theta_{\mathrm{MAP}}=10/12\approx 0.833$.

Show solution

Writing the posterior $\mathrm{Beta}(11,3)$ first makes the $-1$ and $-2$ easy to place.

Update

$$\mathrm{Beta}(2+9,\;2+1)=\mathrm{Beta}(11,3)$$

Makes join the first parameter and misses the second; the Beta shape survives the update.

Peak

$$\frac{11-1}{11+3-2}=\frac{10}{12}\approx 0.833$$

Mode of a Beta with both parameters above $1$.

Answer $$\boxed{\hat\theta_{\mathrm{MAP}}\approx 0.833}$$
Check

It lies between the prior mode $0.5$ and the MLE $0.9$, closer to the MLE because $10$ real attempts outweigh $2$ pseudo-attempts.

Update, then subtract one from the top and two from the bottom.

⚠ Using the posterior-mean formula for the MAP

the two formulas differ only by the $-1$ and $-2$

wrong$$\hat\theta_{\mathrm{MAP}}=\frac{a+N_1}{a+b+N}$$
right$$\hat\theta_{\mathrm{MAP}}=\frac{a+N_1-1}{a+b+N-2}$$
⚠ Reading the MAP as the whole answer

a single number feels like a conclusion

wrong$$\hat\theta_{\mathrm{MAP}}=0\ \Rightarrow\ \text{the click rate is }0$$
right$$\hat\theta_{\mathrm{MAP}}=0\ \text{but}\ \Pr(\Theta>0.1\mid D)=0.9^{5}\approx 0.59$$

2.5The posterior mean: the average over the posterior

Averages $\theta$ over the posterior: the best guess when errors cost their square, pulling the MLE toward the prior mean.

The peak uses one point of the posterior; the mean uses all of it.

TheoremPosterior mean under a Beta prior
Conditions
  • prior $\mathrm{Beta}(a,b)$, data with $N_1$ heads and $N_0$ tails, $N=N_0+N_1$

$$\boxed{\;\textcolor{#d1690a}{E[\Theta\mid D]}=\frac{a+N_1}{a+b+N_0+N_1}\;}$$

Average every candidate rate, weighted by its posterior density. For a Beta posterior the result is a weighted average of the prior mean and the fraction of heads, and the weight on the prior shrinks as $N$ grows.

Derivation, the weighted average, and why squared loss picks the mean

A $\mathrm{Beta}(x,y)$ variable has mean $x/(x+y)$. The posterior is $\mathrm{Beta}(a+N_1,\,b+N_0)$, which gives the boxed formula at once.

Split it to see the weights: $$\frac{a+N_1}{a+b+N}=\underbrace{\frac{a+b}{a+b+N}}_{w}\cdot\frac{a}{a+b}+\underbrace{\frac{N}{a+b+N}}_{1-w}\cdot\frac{N_1}{N}.$$

Under squared loss the expected loss of a guess $\hat\theta$ splits in two, with $m=E[\Theta\mid D]$: $$E[(\hat\theta-\Theta)^2\mid D]=(\hat\theta-m)^2+\mathrm{Var}(\Theta\mid D).$$ Only the first part depends on $\hat\theta$, and it vanishes exactly at $\hat\theta=m$.

The split comes from writing $\hat\theta-\Theta=(\hat\theta-m)+(m-\Theta)$ and expanding the square; the cross term has expectation $2(\hat\theta-m)\,E[m-\Theta\mid D]=0$.

Looks like this, but is not

The posterior mean averages everything the posterior knows, so it looks like the safest guess in any game.

With the three possible defect rates of the previous block, the posterior mean is $0.1(0.261)+0.3(0.553)+0.5(0.186)\approx 0.285$, which is not one of the three rates. If only an exact hit scores, it loses every time. The best estimate depends on the loss.

The posterior mean as a weighted average: a Beta(4, 6) prior and 30 heads in 40

The prior $\mathrm{Beta}(4,6)$ has mean $0.4$. The data are $30$ heads in $40$ flips. Find the posterior mean directly, then again as a weighted average of the prior mean and the MLE.

Find$E[\Theta\mid D]$, computed both ways.
Given
  • prior $\mathrm{Beta}(4,6)$

  • $N_1=30$, $N_0=10$, $N=40$

Solution

Computing it twice costs one extra line and catches the most common slip, feeding a count into the wrong parameter.

Direct formula

$$E[\Theta\mid D]=\frac{4+30}{4+6+40}=\frac{34}{50}=0.68$$

Mean of the posterior $\mathrm{Beta}(34,16)$.

As a weighted average

$$w=\frac{a+b}{a+b+N}=\frac{10}{50}=0.2$$

The prior counts as $10$ pseudo-flips against $40$ real ones.

$$0.2(0.4)+0.8(0.75)=0.08+0.6=0.68$$

The weight $w$ multiplies the prior mean $4/10=0.4$, and the rest goes to the MLE $30/40=0.75$.

Answer $$\boxed{E[\Theta\mid D]=0.68}$$
Check

The two routes agree, and $0.68$ lies between $0.4$ and $0.75$: $0.07$ from the MLE and $0.28$ from the prior mean, four times closer to the data because the data carry four times the weight.

Prior strength is measured in flips: $a+b$ pseudo-flips against $N$ real ones decides how far the estimate moves.

Zero clicks in four views: MAP or mean under squared loss?

A new ad gets $0$ clicks in $4$ views, and the prior is flat, so the posterior is $\mathrm{Beta}(1,5)$ with MAP $0$ and mean $1/6$. The team pays $(\hat\theta-\theta)^2$ for its estimate. Compare the expected cost of reporting the MAP with that of reporting the mean.

Find$E[(\hat\theta-\Theta)^2\mid D]$ for both estimates.
Given
  • posterior $\mathrm{Beta}(1,5)$

  • $\hat\theta_{\mathrm{MAP}}=0$, $E[\Theta\mid D]=1/6$

Solution

The split into squared distance from the mean plus the posterior variance gives both costs without a single integral.

Posterior variance

$$\mathrm{Var}(\Theta\mid D)=\frac{xy}{(x+y)^2(x+y+1)}=\frac{1\cdot 5}{36\cdot 7}=\frac{5}{252}\approx 0.0198$$

Beta variance with $x=1$ and $y=5$.

Cost of each report

$$\text{mean: }\ 0^2+0.0198=0.0198$$

Reporting the mean removes the first term entirely.

$$\text{MAP: }\ \big(0-\tfrac16\big)^2+0.0198=0.0278+0.0198=0.0476$$

The MAP pays the squared distance to the mean on top of the shared variance.

Answer $$\boxed{0.0198\ \text{(mean)}\quad\text{vs}\quad 0.0476\ \text{(MAP)}}$$
Check

The MAP is $0$, so its cost is just $E[\Theta^2\mid D]$, which the Beta moment formula gives directly: $\frac{1\cdot 2}{6\cdot 7}=\frac{2}{42}\approx 0.0476$, the same number as the split.

Under squared loss, report the mean; for a skewed posterior it beats the peak by a wide margin.

Checkpoint
§02.5 — posterior mean

A coin is believed to be roughly fair, with prior $\mathrm{Beta}(3,3)$. It is flipped $8$ times and lands heads once.

Find(a) What is the posterior mean of the heads rate?
Given
  • prior $\mathrm{Beta}(3,3)$

  • $1$ head, $7$ tails

Hint 1/4

The mean of the posterior is asked, so find the posterior's parameters first.

Hint 2/4

$E[\Theta\mid D]=\dfrac{a+N_1}{a+b+N}$.

Hint 3/4

With $a=b=3$, $N_1=1$ and $N=8$: $\dfrac{3+1}{3+3+8}$.

Hint 4/4

$E[\Theta\mid D]=4/14\approx 0.286$.

Show solution

Update, then use the Beta mean; the weighted-average form doubles as a check.

Update

$$\mathrm{Beta}(3+1,\;3+7)=\mathrm{Beta}(4,10)$$

One head joins the first parameter, seven tails the second.

Mean

$$\frac{4}{4+10}=\frac{4}{14}\approx 0.286$$

The posterior is a Beta, and a Beta's mean is its first parameter over the sum.

Answer $$\boxed{E[\Theta\mid D]\approx 0.286}$$
Check

As a weighted average: $\tfrac{6}{14}(0.5)+\tfrac{8}{14}(0.125)=\tfrac{3+1}{14}\approx 0.286$, the same number.

A fair-leaning prior keeps one head in eight from dragging the estimate all the way down to $0.125$.

⚠ Letting the data speak alone once they arrive

the MLE is the familiar number and the prior feels like a starting guess only

wrong$$E[\Theta\mid D]=\frac{N_1}{N}$$
right$$E[\Theta\mid D]=\frac{a+N_1}{a+b+N}$$
⚠ Borrowing the MAP's pseudo-counts for the mean

the $-2$ from the peak formula leaks into the weight

wrong$$w=\frac{a+b-2}{a+b+N-2}\ \ \text{for the mean}$$
right$$w=\frac{a+b}{a+b+N}\ \ \text{for the mean}$$

2.6Credible intervals: where the posterior puts its probability

Reports an interval that holds the parameter with a stated posterior probability, usually cutting equal probability from each tail.

A point estimate hides the spread of the posterior; an interval reports it.

DefinitionCredible interval
Conditions
  • a posterior density $p(\theta\mid D)$ is available

  • $0<\alpha<1$; $\alpha=0.05$ gives a 95% interval

$$\boxed{\begin{aligned}&\mathrm{Cred}_\alpha(D)=[\theta_l,\theta_u]:\ \ \Pr(\theta_l\le\Theta\le\theta_u\mid D)=\int_{\theta_l}^{\theta_u}p(\theta\mid D)\,d\theta=1-\alpha\\&\text{central interval: }\ \Pr(\Theta<\theta_l\mid D)=\Pr(\Theta>\theta_u\mid D)=\tfrac{\alpha}{2}\end{aligned}}$$

Given the data you actually saw, $\Theta$ lies in $[\theta_l,\theta_u]$ with probability $1-\alpha$. Many intervals have that property; the central one cuts $\alpha/2$ of the posterior probability from each end.

Three ways to get the numbers

Closed-form CDF. When the posterior is $\mathrm{Beta}(a,1)$ or $\mathrm{Beta}(1,b)$, its CDF is $\theta^{a}$ or $1-(1-\theta)^{b}$, and each end point solves a one-line equation.

Integrate a polynomial. For small whole-number parameters the density is a polynomial, so a tail probability such as $\Pr(\Theta>0.5\mid D)$ is an exact integral.

Normal approximation (further reading, not on the lecture slides). For large counts a Beta posterior is close to a normal curve with the same mean and variance, so $$[\theta_l,\theta_u]\approx E[\Theta\mid D]\pm 2\sqrt{\mathrm{Var}(\Theta\mid D)}.$$ With few observations, or a mean near $0$ or $1$, compare it with the exact interval.

Looks like this, but is not

It seems the central 95% interval must contain the most probable value, since that is where the density is highest.

For the posterior $\mathrm{Beta}(1,5)$ the density is highest at $\theta=0$, but the central interval is about $[0.005,\,0.522]$ and leaves $0$ out, because it must cut $0.025$ from the left end too. When the peak sits on the boundary, a one-sided interval $[0,\theta_u]$ is the better report.

posteriorexact (numerical)mean ± 2 sd

$\mathrm{Beta}(5,1)$

$[0.478,\ 0.995]$

$[0.552,\ 1.115]$

$\mathrm{Beta}(3,60)$

$[0.010,\ 0.112]$

$[-0.006,\ 0.101]$

$\mathrm{Beta}(47,155)$

$[0.177,\ 0.293]$

$[0.173,\ 0.292]$

For $\mathrm{Beta}(5,1)$, four observations under a flat prior, the shortcut runs past $1$. With a mean near $0$ it dips below $0$. With about two hundred observations and a mean near $0.23$ it is off by less than $0.005$.

Four out of four: a 90% central credible interval from an exact CDF

Back to the four listeners who all accepted, with the flat prior. The posterior is $\mathrm{Beta}(5,1)$ with density $5\theta^4$. Find the central 90% credible interval for the acceptance rate.

Find$[\theta_l,\theta_u]$
Given
  • posterior density $5\theta^4$ on $[0,1]$

  • $\alpha=0.10$, so $0.05$ in each tail

Solution

The CDF of this posterior is a single power, so each end point is one root and no approximation is needed.

Write the CDF

$$F(\theta)=\int_0^{\theta}5t^4\,dt=\theta^{5}$$

The CDF is the area to the left of $\theta$, and a central interval is defined through it.

Cut 0.05 from each end

$$F(\theta_l)=0.05\ \Rightarrow\ \theta_l=0.05^{1/5}\approx 0.549$$

The left tail holds $\alpha/2=0.05$.

$$F(\theta_u)=0.95\ \Rightarrow\ \theta_u=0.95^{1/5}\approx 0.990$$

The right tail also holds $0.05$, so $0.95$ lies to the left of $\theta_u$.

Answer $$\boxed{\mathrm{Cred}_{0.10}(D)=[0.549,\;0.990]}$$
Check

$F(0.990)-F(0.549)=0.95-0.05=0.90$. The interval also contains the posterior mean $5/6\approx 0.833$, a quick plausibility check.

Given these four listeners, the rate lies between $0.55$ and $0.99$ with probability $0.9$. That is the sentence the dashboard should have printed.

Two hundred users: an approximate 95% credible interval

With a flat prior, $46$ of $200$ users click a new button, so the posterior is $\mathrm{Beta}(47,155)$. Find an approximate central 95% credible interval from the posterior mean and standard deviation.

FindAn approximate $[\theta_l,\theta_u]$.
Given
  • posterior $\mathrm{Beta}(47,155)$

  • $\alpha=0.05$

Solution

With 200 observations the posterior is close to a normal curve, and mean $\pm$ 2 sd takes three lines; the exact Beta quantiles need a calculator or software.

Posterior mean

$$m=\frac{47}{47+155}=\frac{47}{202}\approx 0.2327$$

The shortcut interval is centred at the posterior mean.

Posterior standard deviation

$$\mathrm{Var}=\frac{47\cdot 155}{202^2\cdot 203}\approx 0.000879$$

The width of the shortcut comes from the posterior variance, here with $x=47$ and $y=155$.

$$s=\sqrt{0.000879}\approx 0.0297$$

The interval steps in standard deviations, the square root of the variance.

Two standard deviations each side

$$0.2327\pm 2(0.0297)=[0.173,\;0.292]$$

A normal curve keeps about $0.95$ of its area within two standard deviations of its mean.

Answer $$\boxed{\mathrm{Cred}_{0.05}(D)\approx[0.173,\;0.292]}$$
Check

A numerical calculation of the exact $\mathrm{Beta}(47,155)$ quantiles gives $[0.177,\,0.293]$, so the shortcut is off by about $0.004$ at the left end. The interval also brackets the data fraction $46/200=0.23$.

For a few hundred observations and a rate away from $0$ and $1$, mean $\pm$ 2 sd is a safe shortcut; with a handful of observations, go back to the exact CDF.

Checkpoint
§02.6 — one-sided credible bound

A new classifier is tested on $8$ images and makes no mistakes. With a flat prior on its error rate, the posterior is $\mathrm{Beta}(1,9)$.

Find(a) What is $\theta_u$ in the one-sided 95% credible interval $[0,\theta_u]$?
Given
  • posterior $\mathrm{Beta}(1,9)$

  • its CDF: $F(\theta)=1-(1-\theta)^{9}$

Hint 1/4

We need the point that cuts off the top tail of the posterior, leaving the stated probability to its left.

Hint 2/4

Solve $F(\theta_u)=0.95$.

Hint 3/4

$1-(1-\theta_u)^9=0.95$ with $F(\theta)=1-(1-\theta)^9$, so $(1-\theta_u)^9=0.05$.

Hint 4/4

$\theta_u=1-0.05^{1/9}\approx 0.283$.

Show solution

A one-sided interval starting at $0$ is the natural report here because the posterior peaks at $0$.

Set up the tail

$$1-(1-\theta_u)^9=0.95\ \Rightarrow\ (1-\theta_u)^9=0.05$$

Probability $0.95$ below $\theta_u$.

Solve

$$\theta_u=1-0.05^{1/9}=1-e^{\ln 0.05/9}\approx 1-0.717=0.283$$

Take the ninth root; $\ln 0.05\approx -2.996$.

Answer $$\boxed{\theta_u\approx 0.283}$$
Check

Plug back in: $1-(1-0.283)^9=1-0.717^9\approx 1-0.050=0.950$.

Zero errors in eight tests still leaves error rates up to about $0.28$ plausible at the 95% level.

⚠ Putting all of alpha into each tail

the level is remembered, the split is not

wrong$$\Pr(\Theta<\theta_l\mid D)=0.05\ \ \text{for a central 95\% interval}$$
right$$\Pr(\Theta<\theta_l\mid D)=0.025\ \ \text{and}\ \ \Pr(\Theta>\theta_u\mid D)=0.025$$
⚠ Using mean plus or minus 2 sd on a tiny sample

it worked for two hundred users

wrong$$\mathrm{Beta}(5,1):\ 0.833\pm 2(0.141)=[0.552,\,1.115]$$
right$$\text{exact: }[0.025^{1/5},\,0.975^{1/5}]=[0.478,\,0.995]$$

2.7Confidence intervals: the frequentist guarantee and how to read it

Builds an interval from the sample mean whose recipe traps the fixed true rate in about 95 of every 100 repeated experiments.

Everything so far put a distribution on $\theta$; the frequentist view refuses to, holding $\theta$ fixed and treating the data as the random part.

TheoremApproximate 95% confidence interval for a proportion
Conditions
  • $X_1,\dots,X_n$ i.i.d. $\mathrm{Ber}(\theta_{\mathrm{true}})$, with $\theta_{\mathrm{true}}$ fixed but unknown

  • $n$ large and $\bar\theta$ not close to $0$ or $1$, so the central limit theorem applies

$$\boxed{\;\textcolor{#d1690a}{\bar\theta\pm 2\sqrt{\frac{\bar\theta(1-\bar\theta)}{n}}},\qquad \bar\theta=\frac{S_n}{n}=\frac{X_1+\cdots+X_n}{n}\;}$$

Take the fraction of successes and step two standard errors to each side. The 95% belongs to this recipe: over many repetitions of the experiment, about 95 of every 100 intervals it produces contain $\theta_{\mathrm{true}}$.

From the central limit theorem to the interval

Each $X_i$ has mean $\theta_{\mathrm{true}}$ and variance $\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})$, so $S_n$ has mean $n\theta_{\mathrm{true}}$ and variance $n\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})$.

By the central limit theorem, for large $n$ $$\frac{S_n-n\theta_{\mathrm{true}}}{\sqrt{n\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})}}\approx\mathcal N(0,1),\quad\text{so}\quad \bar\theta\approx\mathcal N\Big(\theta_{\mathrm{true}},\,\frac{\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})}{n}\Big).$$

A standard normal lies within $\pm 2$ with probability $0.954$. With $\sigma_n=\sqrt{\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})/n}$ this reads $$\Pr\big(\vert\bar\theta-\theta_{\mathrm{true}}\vert\le 2\sigma_n\big)\approx 0.95.$$

The event $\vert\bar\theta-\theta_{\mathrm{true}}\vert\le 2\sigma_n$ is the same as $\bar\theta-2\sigma_n\le\theta_{\mathrm{true}}\le\bar\theta+2\sigma_n$. Last, $\sigma_n$ contains the unknown $\theta_{\mathrm{true}}$; for large $n$ the law of large numbers puts $\bar\theta$ close to it, so we plug in $\bar\theta$.

Looks like this, but is not

Our data gave the 95% confidence interval $[0.18,\,0.26]$, so it seems there is probability $0.95$ that $\theta_{\mathrm{true}}$ lies between $0.18$ and $0.26$.

In the frequentist model $\theta_{\mathrm{true}}$ is a fixed number: it is in $[0.18,\,0.26]$ or it is not, and no probability is left to assign. The $0.95$ describes the recipe over repeated samples. A probability about this one interval needs a posterior, which is what a credible interval supplies.

nθtrue = 0.5θtrue = 0.1

$10$

$89$

$65$

$30$

$96$

$81$

$100$

$94$

$93$

$400$

$95$

$95$

With $10$ samples and a rate of $0.1$, only about $65$ intervals in $100$ catch the truth, far from the promised $95$. The recipe earns its name once both the success count and the failure count are comfortably large.

Click-through rate: 88 clicks in 400 views

A recommendation widget is shown $400$ times and clicked $88$ times. Treat the views as i.i.d. $\mathrm{Ber}(\theta_{\mathrm{true}})$ and give a 95% confidence interval for the click-through rate.

FindThe interval $\bar\theta\pm 2\,\mathrm{SE}$.
Given$n=400$ views, $S_n=88$ clicks
Solution

With $88$ clicks and $312$ non-clicks both counts are large, so the central-limit recipe is safe and needs no prior.

Point estimate

$$\bar\theta=\frac{88}{400}=0.22$$

The mean of the 0/1 outcomes is the click fraction.

Standard error

$$\mathrm{SE}=\sqrt{\frac{0.22\times 0.78}{400}}=\sqrt{0.000429}\approx 0.0207$$

Plug the estimate into the Bernoulli variance $\theta(1-\theta)/n$.

Two standard errors each side

$$0.22\pm 2(0.0207)=0.22\pm 0.0414$$

The central limit theorem puts about $0.95$ of the sampling distribution within two standard errors.

$$[0.179,\;0.261]$$

The two ends, rounded to three decimals.

Answer $$\boxed{[0.179,\;0.261]}$$
Check

Scale check: the widest possible half-width at $n=400$ is $2\sqrt{0.25/400}=0.05$, and ours, $0.041$, is below it because $0.22$ is away from $0.5$.

Say it this way: the recipe that produced $[0.179,\,0.261]$ catches the true rate about 95 times in 100. Do not say the true rate is in this interval with probability $0.95$.

Twenty honest experiments: how many intervals miss?

A lab repeats the same $100$-flip experiment $20$ times and builds a 95% interval each time, as in the figure. Assume each interval covers $\theta_{\mathrm{true}}$ with probability $0.95$, independently of the others. How many misses should they expect, and how likely is at least one?

FindThe expected number of misses and $\Pr(\text{at least one miss})$.
Given
  • $K=20$ independent intervals

  • $0.95$ each

Solution

The number of misses counts independent yes/no events, so the binomial model from the previous section answers both parts.

Expected misses

$$E[\text{misses}]=20\times 0.05=1$$

A binomial count has mean $n$ times $p$.

At least one miss

$$\Pr(\text{no miss})=0.95^{20}\approx 0.358$$

All twenty independent intervals must cover.

$$\Pr(\text{at least one miss})=1-0.358=0.642$$

At least one miss is the opposite event of no miss.

Answer $$\boxed{E[\text{misses}]=1,\qquad \Pr(\ge 1\ \text{miss})\approx 0.64}$$
Check

Order of magnitude: with mean $1$ and small miss probability, the miss count is close to Poisson$(1)$, which gives $\Pr(\ge 1)\approx 1-e^{-1}\approx 0.632$, close to $0.642$. Exactly one miss has probability $20(0.05)(0.95)^{19}\approx 0.377$, the most likely count, so the figure's single miss is typical.

A 95% recipe guarantees misses in the long run; the guarantee is about how often they happen, never about which interval is one.

Checkpoint
§02.7 — 95% interval for a proportion

An exit poll asks $100$ voters a yes/no question, and exactly half say yes.

Find(a) Which is the 95% confidence interval $\bar\theta\pm 2\,\mathrm{SE}$?
Given
  • $n=100$

  • $\bar\theta=0.5$

Hint 1/4

Build the standard error first; the interval is the estimate plus or minus two of them.

Hint 2/4

$\mathrm{SE}=\sqrt{\bar\theta(1-\bar\theta)/n}$ and the interval is $\bar\theta\pm 2\,\mathrm{SE}$.

Hint 3/4

$\bar\theta=0.5$ and $n=100$ give $\mathrm{SE}=\sqrt{0.25/100}=0.05$.

Hint 4/4

The interval is $0.5\pm 0.1=[0.40,\,0.60]$.

Show solution

At $\bar\theta=0.5$ the standard error is as large as it gets, which makes the arithmetic clean.

Standard error

$$\sqrt{\frac{0.5\times 0.5}{100}}=0.05$$

Bernoulli variance over $n$, square-rooted.

Interval

$$0.5\pm 2(0.05)=[0.40,\;0.60]$$

Two standard errors, because a normal curve keeps about $0.95$ of its area within two standard deviations.

Answer $$\boxed{[0.40,\;0.60]}$$
Check

With $n=100$ the half-width $0.1$ equals $1/\sqrt{n}$, the known worst case for the two-standard-error recipe.

A poll of $100$ pins a proportion down to about $\pm 0.1$; for $\pm 0.05$ you need four times as many people.

⚠ Giving one interval a probability it does not have

it is the sentence everyone wants to say

wrong$$\Pr(0.18\le\theta_{\mathrm{true}}\le 0.26)=0.95$$
right$$\Pr_{\text{repeated samples}}\big(\bar\theta-2\,\mathrm{SE}\le\theta_{\mathrm{true}}\le\bar\theta+2\,\mathrm{SE}\big)\approx 0.95$$
⚠ Dividing by n instead of taking the square root

the variance of the sample mean has $n$ in its denominator, and the root gets lost

wrong$$\bar\theta\pm 2\,\frac{\bar\theta(1-\bar\theta)}{n}$$
right$$\bar\theta\pm 2\sqrt{\frac{\bar\theta(1-\bar\theta)}{n}}$$
⚠ Using the recipe when no successes were seen

the formula still returns numbers

wrong$$0\text{ clicks in }30:\ \ 0\pm 2\sqrt{0\cdot 1/30}=[0,\,0]$$
right$$\text{flat prior: }\mathrm{Beta}(1,31),\ \ \theta_u=1-0.05^{1/31}\approx 0.092$$
From counts to a Bayesian answer

Any question that gives a Beta prior and yes/no data and asks for a posterior, an estimate or an interval.

  1. Count

    Write $N_1$ (successes) and $N_0$ (failures). Pool all batches first.

  2. Update

    Posterior $\mathrm{Beta}(N_1+a,\,N_0+b)$; successes feed the first parameter.

  3. Summarize

    MAP $\frac{a+N_1-1}{a+b+N-2}$ for the peak, $\frac{a+N_1}{a+b+N}$ for the mean.

  4. Interval

    Exact CDF when a parameter equals $1$; mean $\pm$ 2 sd when both counts are large.

  5. Sanity check

    Estimates lie in $[0,1]$ and between the prior's value and $N_1/N$.

Where it goes wrong
  • Adding $N$ instead of $N_0$ to the second parameter.

  • Using mean $\pm$ 2 sd with a handful of observations, which can run past $0$ or $1$.

A 95% confidence interval for a proportion

A frequentist interval for a rate from $n$ yes/no observations, with no prior involved.

  1. Estimate

    $\bar\theta=S_n/n$.

  2. Standard error

    $\mathrm{SE}=\sqrt{\bar\theta(1-\bar\theta)/n}$.

  3. Interval

    $\bar\theta\pm 2\,\mathrm{SE}$.

  4. Condition

    A common rule of thumb, not from the lecture: at least about $10$ successes and $10$ failures. Otherwise, say the approximation is rough.

  5. Sentence

    The recipe covers $\theta_{\mathrm{true}}$ about 95 times in 100; this one interval either does or does not.

Where it goes wrong
  • Dividing by $n$ instead of $\sqrt n$.

  • Reading the 95% as a probability about this one interval.

Which number to report

A question asks for the best estimate and you must decide which one it means.

  1. No prior wanted

    The MLE $N_1/N$, with a confidence interval.

  2. Only exact hits count

    The MAP, the most probable value under the posterior.

  3. Errors cost their square

    The posterior mean.

  4. Uncertainty matters

    A credible interval, not a single number.

Where it goes wrong
  • Reporting one number when the question asks how sure you are.

Credible interval: flat prior, 30 successes in 100

Flat prior, $30$ successes in $100$ trials. Give an approximate central 95% credible interval.

Find$\mathrm{Cred}_{0.05}(D)$ by mean $\pm$ 2 sd.
Givenposterior $\mathrm{Beta}(31,71)$
Solution

A hundred observations make the normal shortcut accurate enough by hand.

Mean and sd

$$m=\frac{31}{102}\approx 0.304,\qquad s=\sqrt{\frac{31\cdot 71}{102^2\cdot 103}}\approx 0.0453$$

A hundred observations make the posterior close to normal, so its mean and spread define the shortcut.

Interval

$$0.304\pm 2(0.0453)=[0.213,\;0.395]$$

About $0.95$ of a normal curve lies within two standard deviations of its mean.

Answer $$\boxed{[0.213,\;0.395]}$$
Check

The exact quantiles, computed numerically, give $[0.219,\,0.396]$.

Confidence interval: the same 30 successes in 100

The same data, no prior. Give the 95% confidence interval $\bar\theta\pm 2\,\mathrm{SE}$.

Find$\bar\theta\pm 2\,\mathrm{SE}$
Given$n=100$, $S_n=30$
Solution

Thirty successes and seventy failures are enough for the central-limit recipe.

Estimate and SE

$$\bar\theta=0.3,\qquad \mathrm{SE}=\sqrt{\frac{0.3\times 0.7}{100}}\approx 0.0458$$

The standard error of a sample fraction is the Bernoulli variance over $n$, square-rooted.

Interval

$$0.3\pm 2(0.0458)=[0.208,\;0.392]$$

The same two-standard-error recipe; no prior enters.

Answer $$\boxed{[0.208,\;0.392]}$$
Check

Its half-width $0.092$ is just under the worst case $2\sqrt{0.25/100}=0.1$.

The same $30$ successes in $100$ give nearly the same numbers, but only the credible interval says anything about where the rate is, given this data.

How to tell them apart

Ask what is random in the sentence: the rate given the data (credible), or the interval over repeated samples (confidence).

Peak of Beta(3, 12)

Posterior $\mathrm{Beta}(3,12)$. Find the MAP estimate.

Find$\hat\theta_{\mathrm{MAP}}$
Given$\alpha=3$, $\beta=12$
Solution

Subtract one from the top and two from the bottom.

Peak

$$\frac{3-1}{3+12-2}=\frac{2}{13}\approx 0.154$$

Mode of a Beta with both parameters above $1$.

Answer $$\boxed{\hat\theta_{\mathrm{MAP}}\approx 0.154}$$
Check

It is below the mean, as it must be for a posterior with a long right tail.

Mean of Beta(3, 12)

The same posterior $\mathrm{Beta}(3,12)$. Find the posterior mean.

Find$E[\Theta\mid D]$
Given$\alpha=3$, $\beta=12$
Solution

No subtractions: the mean uses the parameters as they are.

Mean

$$\frac{3}{3+12}=0.2$$

A Beta's mean uses the parameters as they are, with no subtraction.

Answer $$\boxed{E[\Theta\mid D]=0.2}$$
Check

The long right tail pulls the balance point above the peak.

One posterior, two summaries: the peak at $0.154$ and the balance point at $0.2$.

How to tell them apart

A $-1$ on top and a $-2$ below mean the peak (MAP); no subtractions mean the average (posterior mean).

Scaffolding comes off
The common skeleton
  1. Count: $N_1$ successes and $N_0$ failures, pooled over every batch.

  2. Update: the posterior is $\mathrm{Beta}(N_1+a,\,N_0+b)$.

  3. Summarize: the MAP $\frac{\alpha-1}{\alpha+\beta-2}$ or the mean $\frac{\alpha}{\alpha+\beta}$, whichever is asked.

  4. Check: the estimate lies between the prior's value and the MLE $N_1/N$.

1 · fully worked

Defect rate: a Beta(2, 8) prior, then 5 defects in 30 items

A factory's prior for a new line's defect rate is $\mathrm{Beta}(2,8)$. A sample of $30$ items has $5$ defects. Find the MAP estimate and the posterior mean, and check each against the prior and the MLE.

Find$\hat\theta_{\mathrm{MAP}}$ and $E[\Theta\mid D]$.
Given
  • prior $\mathrm{Beta}(2,8)$

  • $N_1=5$ defects, $N_0=25$ good items

Solution

All four skeleton steps in order; the check at the end is what catches a count fed into the wrong parameter.

Count

$$N_1=5,\qquad N_0=30-5=25$$

Defects are the successes here, since they are what we count.

Update

$$\Theta\mid D\sim\mathrm{Beta}(5+2,\;25+8)=\mathrm{Beta}(7,33)$$

Defects to the first parameter, good items to the second.

Summarize

$$\hat\theta_{\mathrm{MAP}}=\frac{7-1}{7+33-2}=\frac{6}{38}\approx 0.158$$

Both parameters exceed $1$, so the peak is interior and the mode formula applies.

$$E[\Theta\mid D]=\frac{7}{40}=0.175$$

The mean needs no subtraction: first parameter over the sum.

Check

$$\text{prior mode }\tfrac18=0.125<0.158<\tfrac16\approx 0.167=\text{MLE}$$

The MAP blends the prior's peak with the MLE.

$$\text{MLE }0.167<0.175<0.2=\text{prior mean}$$

The posterior mean blends the prior's mean with the MLE.

Answer $$\boxed{\hat\theta_{\mathrm{MAP}}\approx 0.158,\qquad E[\Theta\mid D]=0.175}$$
Check

Both estimates are weighted averages of a prior summary and the MLE $5/30$: $\tfrac{8}{38}\cdot\tfrac18+\tfrac{30}{38}\cdot\tfrac16=\tfrac{6}{38}\approx 0.158$ and $\tfrac{10}{40}(0.2)+\tfrac{30}{40}\cdot\tfrac16=0.175$, the same two numbers.

The check step is not decoration: an estimate outside its bracket means a count went into the wrong place.

2 · you write the reasoning

An easier one, and this time you supply the reasons. Flat prior $\mathrm{Beta}(1,1)$, then $3$ successes and $1$ failure. Find the posterior and the MAP, and for each line write why it is allowed.

  1. $N_1=3,\qquad N_0=1$

    reasoning

    Successes are the ones: three of them, and one failure.

  2. $\Theta\mid D\sim\mathrm{Beta}(1+3,\;1+1)=\mathrm{Beta}(4,2)$

    reasoning

    Conjugate update: successes go to the first parameter, failures to the second.

  3. $\hat\theta_{\mathrm{MAP}}=\dfrac{4-1}{4+2-2}=\dfrac34$

    reasoning

    Peak of $\mathrm{Beta}(\alpha,\beta)$ with $\alpha=4$ and $\beta=2$, both above $1$, so the peak is interior.

  4. $\hat\theta_{\mathrm{MLE}}=\dfrac34=\hat\theta_{\mathrm{MAP}}$

    reasoning

    The flat prior adds no pseudo-counts to the peak, so the MAP must equal the fraction of successes; this line is the check.

3 · find the buried error

Harder, and the work is done for you, with two errors buried in it. Prior $\mathrm{Beta}(4,2)$; data: $6$ successes and $14$ failures. Find the MAP estimate and the posterior mean.

  1. Step 1. $N_1=6,\ \allowbreak N_0=14,\ \allowbreak a=4,\ \allowbreak b=2$. Read off the counts and the prior.

  2. Step 2. $\Theta\mid D\sim\mathrm{Beta}(6+4,\;20+2)=\mathrm{Beta}(10,22)$. Conjugate update.

  3. Step 3. $\hat\theta_{\mathrm{MAP}}=\dfrac{10-1}{10+22-2}=\dfrac{9}{30}=0.3$. Peak of the posterior.

  4. Step 4. $E[\Theta\mid D]=\dfrac{10}{10+22-2}=\dfrac{10}{30}\approx 0.333$. Mean of the posterior.

  5. Step 5. $0.3$ and $0.333$ lie in $[0,1]$. Both estimates are valid rates.

the two buried errors (2)
⚠ step 2

The second parameter must get the failures, $N_0=14$, not the total $N=20$. The posterior is $\mathrm{Beta}(10,16)$.

The word total attaches itself to the second parameter easily, and $20$ is the number the problem shows most prominently.

right

$\Theta\mid D\sim\mathrm{Beta}(6+4,\;14+2)=\mathrm{Beta}(10,16)$, so the MAP is $\dfrac{9}{24}=0.375$.

⚠ step 4

The mean has no $-2$ in its denominator; that subtraction belongs to the peak. The mean of $\mathrm{Beta}(\alpha,\beta)$ is $\alpha/(\alpha+\beta)$.

The MAP was computed one line earlier with a $-2$, and the habit carries over to the next formula.

right

$E[\Theta\mid D]=\dfrac{10}{10+16}=\dfrac{10}{26}\approx 0.385$.

4 · the bare problem
§02.4 — MAP and mean on a conversion rate

An online store's conversion rate has the prior $\mathrm{Beta}(3,5)$ from last season. This week $18$ of $60$ visitors buy.

Find
  1. (a) Find the posterior.

  2. (b) Find the MAP estimate and the posterior mean.

  3. (c) Which of the two is closer to the MLE, and why?

Given
  • prior $\mathrm{Beta}(3,5)$

  • $18$ buyers, $42$ non-buyers

Hint 1/4

Follow the skeleton: counts, update, summaries, then compare with the MLE.

Hint 2/4

Posterior $\mathrm{Beta}(N_1+a,\,N_0+b)$; MAP $\frac{\alpha-1}{\alpha+\beta-2}$; mean $\frac{\alpha}{\alpha+\beta}$; MLE $N_1/N$.

Hint 3/4

$a=3$, $b=5$, $N_1=18$, $N_0=42$ give $\alpha=21$, $\beta=47$; the MLE is $18/60=0.3$.

Hint 4/4

Posterior $\mathrm{Beta}(21,47)$, MAP $20/66\approx 0.303$, mean $21/68\approx 0.309$; the MAP is closer to $0.3$.

Show solution

Write both summaries as weighted averages; that answers part (c) without extra algebra.

Update

$$\mathrm{Beta}(18+3,\;42+5)=\mathrm{Beta}(21,47)$$

Buyers are the counted outcome, so they join the first parameter; non-buyers join the second.

Summaries

$$\hat\theta_{\mathrm{MAP}}=\frac{20}{66}\approx 0.303,\qquad E[\Theta\mid D]=\frac{21}{68}\approx 0.309$$

One posterior gives both: subtract one and two for the peak, nothing for the mean.

Compare through the weights

$$\text{MAP}=\tfrac{6}{66}\cdot\tfrac13+\tfrac{60}{66}(0.3),\qquad \text{mean}=\tfrac{8}{68}(0.375)+\tfrac{60}{68}(0.3)$$

The peak's prior summary $1/3$ gets $6$ pseudo-counts; the mean's, $0.375$, gets $8$.

$$\vert 0.303-0.3\vert=0.003<0.009=\vert 0.309-0.3\vert$$

Less prior weight and a prior value nearer to $0.3$ keep the MAP closer.

Answer $$\boxed{\mathrm{Beta}(21,47);\ \ \hat\theta_{\mathrm{MAP}}\approx 0.303,\ \ E[\Theta\mid D]\approx 0.309}$$
Check

Both weighted averages reproduce the direct values: $\tfrac{6}{66}\cdot\tfrac13+\tfrac{60}{66}(0.3)=\tfrac{20}{66}$ and $\tfrac{8}{68}(0.375)+\tfrac{60}{68}(0.3)=\tfrac{21}{68}$.

To compare MAP and mean, compare their prior weights and their prior values; the data part is the same MLE in both.

Full exam-style question

Exam-style: packet loss on a sensor link, both waysexam format

A sensor node sends $n=200$ packets and $14$ are lost. Assume losses are i.i.d. $\mathrm{Ber}(\theta)$.

  • (a) Give the MLE of the loss rate and an approximate 95% confidence interval.
  • (b) The vendor's datasheet suggests the prior $\mathrm{Beta}(1,19)$. Find the posterior, the MAP estimate and the posterior mean.
  • (c) Give an approximate central 95% credible interval, and say in one sentence each what the two intervals mean.
FindMLE and confidence interval; posterior, MAP and mean; credible interval; the two readings.
Given
  • $n=200$, $S_n=14$ lost packets

  • prior $\mathrm{Beta}(1,19)$ for parts (b) and (c)

Solution

Part (a) needs no prior, so we do it first with the central-limit recipe; parts (b) and (c) reuse the same counts with the conjugate update, and the normal shortcut keeps the credible interval a hand calculation.

(a) MLE and confidence interval

$$\hat\theta_{\mathrm{MLE}}=\bar\theta=\frac{14}{200}=0.07$$

Without a prior, the best-fitting rate is the fraction of losses.

$$\mathrm{SE}=\sqrt{\frac{0.07\times 0.93}{200}}\approx 0.0180$$

Bernoulli variance over $n$, square-rooted.

$$0.07\pm 2(0.0180)=[0.034,\;0.106]$$

Fourteen losses and 186 deliveries are enough for the central-limit recipe.

(b) Posterior, MAP and mean

$$\Theta\mid D\sim\mathrm{Beta}(14+1,\;186+19)=\mathrm{Beta}(15,205)$$

Losses to the first parameter, deliveries to the second.

$$\hat\theta_{\mathrm{MAP}}=\frac{14}{218}\approx 0.064,\qquad E[\Theta\mid D]=\frac{15}{220}\approx 0.068$$

Peak $(\alpha-1)/(\alpha+\beta-2)$ and mean $\alpha/(\alpha+\beta)$.

(c) Credible interval and the two readings

$$\mathrm{sd}=\sqrt{\frac{15\cdot 205}{220^2\cdot 221}}\approx 0.0170$$

The posterior's spread sets the width of the normal shortcut.

$$0.068\pm 2(0.0170)=[0.034,\;0.102]$$

With $200$ observations the normal shortcut is usable, though rough here because the mean sits near $0$.

$$[0.034,\,0.106]\ \ \text{vs}\ \ [0.034,\,0.102]$$

The confidence interval's $0.95$ is about repeated samples of $200$ packets. The credible interval's $0.95$ is about the loss rate, given these $200$.

Answer $$\boxed{\begin{aligned}&\hat\theta_{\mathrm{MLE}}=0.07,\ \ \text{CI}\approx[0.034,\,0.106]\\&\mathrm{Beta}(15,205):\ \hat\theta_{\mathrm{MAP}}\approx 0.064,\ \ E[\Theta\mid D]\approx 0.068\\&\text{credible}\approx[0.034,\,0.102]\end{aligned}}$$
Check

The Bayesian estimates sit below $0.07$ because the prior mean is $1/20=0.05$. A numerical calculation of the exact $\mathrm{Beta}(15,205)$ quantiles gives about $[0.039,\,0.105]$, so the normal shortcut is rough at the left end, where this posterior is skewed.

Three tools on one set of counts: the MLE, the conjugate update and the central-limit recipe.

In an exam, name which interval you computed and give its reading in one sentence; the numbers alone do not show that you know the difference.

Practice

A · concept 4 questions
1§02.5 — a majority of heads and the posterior mean

A study partner offers a shortcut for Beta posteriors. You want to test it on a small case before trusting it.

Find(a) Is the claim true or false? Use the test case.
Given
  • Claim: if $N_1>N_0$, then $E[\Theta\mid D]>0.5$, whatever the Beta prior.

  • Test case: prior $\mathrm{Beta}(2,8)$, data $3$ heads and $2$ tails.

Hint 1/4

The claim ignores the prior, so check whether a strong prior can outvote a small majority of heads.

Hint 2/4

$E[\Theta\mid D]=\dfrac{a+N_1}{a+b+N}$.

Hint 3/4

With $a=2$, $b=8$, $N_1=3$ and $N=5$ this is $\dfrac{2+3}{2+8+5}$.

Hint 4/4

The posterior mean is $5/15\approx 0.333<0.5$, so the claim is false.

Show solution

A claim with 'whatever the prior' falls to one counterexample, so we compute the test case directly.

Update

$$\mathrm{Beta}(2+3,\;8+2)=\mathrm{Beta}(5,10)$$

The prior stays Beta: three heads join $2$, two tails join $8$.

Mean

$$E[\Theta\mid D]=\frac{5}{15}\approx 0.333$$

The posterior mean is the first parameter over the sum.

Verdict

$$0.333<0.5\ \Rightarrow\ \text{false}$$

Heads are the majority in the data, yet the estimate stays below one half.

Answer $$\boxed{\text{False: }E[\Theta\mid D]\approx 0.333}$$
Check

As a weighted average: $\tfrac{10}{15}(0.2)+\tfrac{5}{15}(0.6)=\tfrac{2+3}{15}\approx 0.333$, the same value.

Before trusting a data majority, count the prior's pseudo-flips; here they outnumber the real ones two to one.

2§02.2 — one more observation and the posterior variance

More data usually means less uncertainty. Check whether that holds after every single observation.

Find(a) True or false: adding one observation to a Beta posterior always makes its variance smaller.
Given
  • posterior before the next flip: $\mathrm{Beta}(1,10)$

  • the next flip lands heads

  • for $X\sim\mathrm{Beta}(x,y)$: $\mathrm{Var}(X)=\dfrac{xy}{(x+y)^2(x+y+1)}$

Hint 1/4

One counterexample settles a claim that says always; compute the variance before and after this head.

Hint 2/4

A head moves $\mathrm{Beta}(x,y)$ to $\mathrm{Beta}(x+1,y)$, and the variance is $\dfrac{xy}{(x+y)^2(x+y+1)}$.

Hint 3/4

Before: $\dfrac{1\cdot 10}{11^2\cdot 12}$. After: $\dfrac{2\cdot 10}{12^2\cdot 13}$.

Hint 4/4

The variance rises from about $0.0069$ to about $0.0107$, so the statement is false.

Show solution

Plugging into the Beta variance formula twice is quicker than any general argument.

Variance before

$$\frac{1\cdot 10}{121\cdot 12}=\frac{10}{1452}\approx 0.00689$$

The variance formula applied to the posterior before the flip, $x=1$ and $y=10$.

Variance after

$$\frac{2\cdot 10}{144\cdot 13}=\frac{20}{1872}\approx 0.01068$$

A head adds one to the first parameter only, so now $x=2$ and $y=10$.

Compare

$$0.01068>0.00689$$

The posterior got wider, so the claim fails.

Answer $$\boxed{\text{False}}$$
Check

The formula does shrink in the ordinary case: $\mathrm{Beta}(1,1)$ has variance $1/12\approx 0.083$ and $\mathrm{Beta}(2,1)$ has $1/18\approx 0.056$. Here the mean jumps from $1/11\approx 0.091$ to $1/6\approx 0.167$, and that surprise is what widens the posterior.

A surprising observation can widen the posterior; only on average does more data narrow it.

3§02.7 — what halves a confidence interval

A product team has a 95% confidence interval for a conversion rate and finds it too wide. They list four changes for the next study.

Find(a) Which change halves the width of the interval?
Given
  • interval: $\bar\theta\pm 2\sqrt{\bar\theta(1-\bar\theta)/n}$ from $n$ visitors

  • assume $\bar\theta$ comes out about the same next time

Hint 1/4

Look at how the width depends on $n$ when everything else stays fixed.

Hint 2/4

Width $=4\sqrt{\bar\theta(1-\bar\theta)/n}$, which is proportional to $1/\sqrt n$.

Hint 3/4

With $\bar\theta$ unchanged, halving $1/\sqrt n$ needs $\sqrt{n'}=2\sqrt n$, that is $n'=4n$.

Hint 4/4

Four times as many visitors halves the width.

Show solution

Only $n$ can shrink the width here; the other options change the method or the level.

Set up the ratio

$$\frac{\sqrt{n}}{\sqrt{n'}}=\frac12$$

New width over old width, with $\bar\theta$ unchanged.

Solve

$$n'=4n$$

Squaring removes the roots: the ratio of sample sizes is the square of the ratio of widths.

Answer $$\boxed{n'=4n}$$
Check

Doubling alone gives a ratio of $1/\sqrt2\approx 0.71$, not $0.5$.

Precision is expensive: each halving of the width costs four times the data.

4§02.2 — updating in pieces or all at once

Two analysts start from the same prior $\mathrm{Beta}(2,2)$ for a click rate. One updates after every single view during the day; the other waits and updates once on the day's totals.

Find(a) True or false: they end the day with the same posterior.
Given
  • the day's data: $6$ clicks and $9$ skips

  • both use the conjugate Beta update

Hint 1/4

Think of the update as bookkeeping: what gets added to which parameter, and does the order matter?

Hint 2/4

Each click adds $1$ to the first parameter and each skip adds $1$ to the second.

Hint 3/4

Over the day, $6$ ones are added to the first parameter and $9$ to the second, in whatever order the views came.

Hint 4/4

Both end at $\mathrm{Beta}(8,11)$, so the statement is true.

Show solution

The conjugate update is pure addition, and addition is order-free.

Batch

$$\mathrm{Beta}(2+6,\;2+9)=\mathrm{Beta}(8,11)$$

The batch analyst adds the day's totals in one step.

One at a time

$$2+\underbrace{1+\cdots+1}_{6}=8,\qquad 2+\underbrace{1+\cdots+1}_{9}=11$$

The same ones, added in the order the views arrived.

Answer $$\boxed{\text{True: }\mathrm{Beta}(8,11)}$$
Check

The parameter total $19=4+15$ matches the prior's $4$ plus the day's $15$ views in both cases.

This is why a Bayesian model can learn online: yesterday's posterior is today's prior, with no need to keep the raw data.

B · computation 8 questions
1§02.4 — posterior, MAP and mean for an adoption rate

A new feature's adoption rate gets the prior $\mathrm{Beta}(3,5)$ from a similar past feature. In a pilot, $10$ of $16$ users adopt it.

Find
  1. (a) Find the posterior.

  2. (b) Find the MAP estimate and the posterior mean.

  3. (c) Check both against the prior and the MLE.

Given
  • prior $\mathrm{Beta}(3,5)$

  • $10$ adopters, $6$ non-adopters

Hint 1/4

Everything follows from the posterior, so update first and summarize second.

Hint 2/4

Posterior $\mathrm{Beta}(a+N_1,\,b+N_0)$; MAP $\frac{\alpha-1}{\alpha+\beta-2}$; mean $\frac{\alpha}{\alpha+\beta}$.

Hint 3/4

$a=3$, $b=5$, $N_1=10$, $N_0=6$ give $\alpha=13$, $\beta=11$.

Hint 4/4

Posterior $\mathrm{Beta}(13,11)$, MAP $12/22\approx 0.545$, mean $13/24\approx 0.542$.

Show solution

The skeleton of the faded ladder: count, update, summarize, check.

Update

$$\mathrm{Beta}(10+3,\;6+5)=\mathrm{Beta}(13,11)$$

Adopters are the counted outcome, so they join the first parameter; non-adopters join the second.

Summaries

$$\hat\theta_{\mathrm{MAP}}=\frac{12}{22}\approx 0.545,\qquad E[\Theta\mid D]=\frac{13}{24}\approx 0.542$$

Both parameters exceed $1$, so the peak is interior; the mean needs no subtraction.

Brackets

$$\tfrac13<0.545<0.625,\qquad 0.375<0.542<0.625$$

Each estimate sits between its prior summary and the MLE.

Answer $$\boxed{\mathrm{Beta}(13,11);\ \ \hat\theta_{\mathrm{MAP}}\approx 0.545,\ \ E[\Theta\mid D]\approx 0.542}$$
Check

Parameter total $24=8+16$: the prior's $8$ pseudo-users plus the $16$ real ones.

The brackets in the last step are the fastest error detector you have.

2§02.3 — MLE from a raw record and a likelihood ratio

A spam filter's decisions on $12$ test emails are checked by hand, and each wrong decision is marked $1$.

Find
  1. (a) Find the MLE of the error rate.

  2. (b) How many times more probable is this record under the MLE than under $\theta=0.5$?

Given
  • record: $1, \allowbreak 0, \allowbreak 0, \allowbreak 1, \allowbreak 1, \allowbreak 0, \allowbreak 0, \allowbreak 0, \allowbreak 1, \allowbreak 0, \allowbreak 0, \allowbreak 0$

  • errors i.i.d. $\mathrm{Ber}(\theta)$

Hint 1/4

Only the counts of ones and zeros matter; count first.

Hint 2/4

$\hat\theta=N_1/N$ and $L(\theta)=\theta^{N_1}(1-\theta)^{N_0}$ for a specific record.

Hint 3/4

$N_1=4$, $N_0=8$; so compare $L(1/3)=(1/3)^4(2/3)^8$ with $L(1/2)=(1/2)^{12}$.

Hint 4/4

$\hat\theta=1/3$, and the ratio is $2^{20}/3^{12}\approx 1.97$.

Show solution

Writing both likelihoods with powers of $2$ and $3$ keeps the ratio exact.

Count and estimate

$$N_1=4,\ N_0=8\ \Rightarrow\ \hat\theta=\frac{4}{12}=\frac13$$

The likelihood of a 0/1 record depends only on the counts, and its peak is the fraction of ones.

Likelihood ratio

$$L(\tfrac13)=\frac{1}{3^4}\cdot\frac{2^8}{3^8}=\frac{2^8}{3^{12}},\qquad L(\tfrac12)=\frac{1}{2^{12}}$$

Probability of this exact record under each rate.

$$\frac{L(1/3)}{L(1/2)}=\frac{2^{20}}{3^{12}}=\frac{1048576}{531441}\approx 1.97$$

Dividing the two likelihoods puts $2^8\cdot 2^{12}=2^{20}$ on top.

Answer $$\boxed{\hat\theta=\tfrac13,\qquad \frac{L(1/3)}{L(1/2)}\approx 1.97}$$
Check

The ratio exceeds $1$, as it must: no rate can beat the MLE at explaining the record.

A likelihood ratio near $2$ is modest evidence; with $12$ observations the data cannot sharply separate $1/3$ from $1/2$.

3§02.6 — an exact tail probability

Three users try a new checkout flow: two finish and one abandons. With a flat prior, the posterior for the completion rate is $\mathrm{Beta}(3,2)$.

Find(a) Find $\Pr(\Theta>0.5\mid D)$ exactly.
Givenposterior density $p(\theta\mid D)=12\,\theta^{2}(1-\theta)$ on $[0,1]$
Hint 1/4

A posterior probability is an area under the posterior density.

Hint 2/4

$\Pr(\Theta>0.5\mid D)=\int_{0.5}^{1}12\,\theta^2(1-\theta)\,d\theta$.

Hint 3/4

With the density $12\theta^2(1-\theta)=12(\theta^2-\theta^3)$, the antiderivative is $12\big(\theta^3/3-\theta^4/4\big)$.

Hint 4/4

The area is $11/16=0.6875$.

Show solution

Small whole-number parameters make the density a polynomial, so the exact integral is short.

Antiderivative

$$\int 12(\theta^2-\theta^3)\,d\theta=4\theta^3-3\theta^4$$

Expanding $12\theta^2(1-\theta)$ first turns each term into a plain power.

Evaluate

$$\big[4\theta^3-3\theta^4\big]_{0.5}^{1}=(4-3)-\big(0.5-0.1875\big)$$

At $1$: $4-3$; at $0.5$: $4/8-3/16$.

$$=1-0.3125=0.6875=\tfrac{11}{16}$$

Upper value minus lower value is the area between them.

Answer $$\boxed{\Pr(\Theta>0.5\mid D)=\tfrac{11}{16}}$$
Check

Mirror the problem: flipping $\theta\to 1-\theta$ swaps the two counts, so $\Pr(\Theta>0.5\mid D)$ equals the area of $12\theta(1-\theta)^2$ below $0.5$: $\int_0^{0.5}12\theta(1-\theta)^2\,d\theta=12\big(\tfrac18-\tfrac1{12}+\tfrac1{64}\big)=\tfrac{11}{16}$.

Two finishes in three make a completion rate above one half likely, about $11$ chances in $16$, but not safe.

4§02.6 — credible bounds from a closed-form CDF

A new sensor is tested $9$ times and passes every test. With a flat prior on its pass rate, the posterior is $\mathrm{Beta}(10,1)$.

Find
  1. (a) Find the central 90% credible interval.

  2. (b) Find $\theta_l$ with $\Pr(\Theta\ge\theta_l\mid D)=0.95$.

Givenposterior CDF $F(\theta)=\theta^{10}$ on $[0,1]$
Hint 1/4

Every end point here is a point where the CDF takes a known value.

Hint 2/4

Central 90%: $F(\theta_l)=0.05$ and $F(\theta_u)=0.95$. One-sided: $F(\theta_l)=0.05$ as well.

Hint 3/4

With $F(\theta)=\theta^{10}$: $\theta_l=0.05^{1/10}$ and $\theta_u=0.95^{1/10}$.

Hint 4/4

(a) $[0.741,\,0.995]$; (b) $\theta_l\approx 0.741$.

Show solution

A single power as the CDF means each end point is a root, no approximation needed.

Central 90%

$$\theta_l=0.05^{1/10}\approx 0.741,\qquad \theta_u=0.95^{1/10}\approx 0.995$$

$0.05$ below $\theta_l$ and $0.05$ above $\theta_u$.

One-sided lower bound

$$\Pr(\Theta\ge\theta_l\mid D)=1-\theta_l^{10}=0.95\ \Rightarrow\ \theta_l=0.05^{1/10}\approx 0.741$$

Probability $0.05$ below the bound.

Answer $$\boxed{(a)\ [0.741,\;0.995]\qquad (b)\ \theta_l\approx 0.741}$$
Check

$0.741^{10}\approx 0.050$ and $0.995^{10}\approx 0.951$, confirming the tails.

The central 90% interval and the one-sided 95% bound share their lower end, because both leave $0.05$ below it.

5§02.7 — a confidence interval for a proportion

A survey of $600$ randomly chosen students finds that $150$ of them use the library's booking app.

Find(a) Give the 95% confidence interval $\bar\theta\pm 2\,\mathrm{SE}$ for the proportion of users.
Given$n=600$, $S_n=150$
Hint 1/4

Estimate first, then measure the spread of that estimate.

Hint 2/4

$\bar\theta=S_n/n$, $\mathrm{SE}=\sqrt{\bar\theta(1-\bar\theta)/n}$, interval $\bar\theta\pm 2\,\mathrm{SE}$.

Hint 3/4

With $n=600$ and $S_n=150$: $\bar\theta=0.25$ and $\mathrm{SE}=\sqrt{0.25\times 0.75/600}=\sqrt{0.0003125}$.

Hint 4/4

The interval is $0.25\pm 0.035=[0.215,\,0.285]$.

Show solution

$150$ users and $450$ non-users are plenty for the central-limit recipe.

Estimate

$$\bar\theta=\frac{150}{600}=0.25$$

The sample fraction is the natural estimate of the proportion of users.

Standard error

$$\sqrt{\frac{0.25\times 0.75}{600}}\approx 0.0177$$

Bernoulli variance over $n$, square-rooted.

Interval

$$0.25\pm 0.0354=[0.215,\;0.285]$$

With $n=600$ the sample fraction is close to normal, so two standard errors give about $0.95$ coverage.

Answer $$\boxed{[0.215,\;0.285]}$$
Check

The worst-case half-width at $n=600$ is $2\sqrt{0.25/600}\approx 0.041$, and ours, $0.035$, is below it.

State the result with its reading: the recipe covers the true proportion about 95 times in 100.

6§02.7 — planning a sample size

Before running a poll, you want the 95% interval $\bar\theta\pm 2\,\mathrm{SE}$ to have half-width at most $0.02$, whatever the result turns out to be.

Find(a) What is the smallest $n$ that guarantees this?
Given
  • half-width $2\sqrt{\bar\theta(1-\bar\theta)/n}\le 0.02$

  • $\bar\theta$ unknown in advance

Hint 1/4

Guard against the worst possible result, the one that makes the standard error largest.

Hint 2/4

$\bar\theta(1-\bar\theta)\le\tfrac14$, with equality at $\bar\theta=\tfrac12$.

Hint 3/4

The worst case needs $2\sqrt{1/(4n)}=1/\sqrt n\le 0.02$ to hold.

Hint 4/4

The smallest sample size that guarantees it is $n=2500$.

Show solution

The product $\bar\theta(1-\bar\theta)$ peaks at $1/2$, so designing for that value covers every outcome.

Worst-case half-width

$$2\sqrt{\frac{1/4}{n}}=\frac{1}{\sqrt n}$$

Plug the largest possible value of $\bar\theta(1-\bar\theta)$.

Solve

$$\frac{1}{\sqrt n}\le 0.02\ \Rightarrow\ \sqrt n\ge 50\ \Rightarrow\ n\ge 2500$$

Both sides are positive, so taking reciprocals flips the inequality and squaring keeps it.

Answer $$\boxed{n=2500}$$
Check

If the poll comes out at $\bar\theta=0.3$ instead, the half-width is $2\sqrt{0.21/2500}\approx 0.018$, inside the target, as a worst-case design should guarantee.

Design for $\bar\theta=1/2$ when nothing is known; any other outcome only makes the interval narrower.

7§02.5 — the posterior mean as a weighted average

A marketing team's prior for a campaign's response rate is $\mathrm{Beta}(6,14)$. The campaign then gets $35$ responses from $50$ contacts.

Find
  1. (a) Find the posterior mean.

  2. (b) Write it as $w\cdot(\text{prior mean})+(1-w)\cdot(\text{MLE})$ and find $w$.

Given
  • prior $\mathrm{Beta}(6,14)$

  • $N_1=35$, $N_0=15$

Hint 1/4

Part (a) is one formula; part (b) splits the same fraction into a prior piece and a data piece.

Hint 2/4

$E[\Theta\mid D]=\dfrac{a+N_1}{a+b+N}$ and $w=\dfrac{a+b}{a+b+N}$.

Hint 3/4

$a=6$, $b=14$, $N_1=35$, $N=50$; prior mean $6/20=0.3$, MLE $35/50=0.7$.

Hint 4/4

The mean is $41/70\approx 0.586$, with $w=20/70=2/7$.

Show solution

The direct formula gives the number; the weights explain it.

Direct

$$\frac{6+35}{6+14+50}=\frac{41}{70}\approx 0.586$$

The posterior is $\mathrm{Beta}(41,29)$; its mean is the first parameter over the sum.

Weights

$$w=\frac{20}{70}=\frac27,\qquad \tfrac27(0.3)+\tfrac57(0.7)=\frac{0.6+3.5}{7}\approx 0.586$$

The prior's share is its $20$ pseudo-counts out of $70$; the rest goes to the MLE $0.7$.

Answer $$\boxed{E[\Theta\mid D]\approx 0.586,\qquad w=\tfrac27}$$
Check

Both routes give $41/70$; and $0.586$ sits between $0.3$ and $0.7$, closer to the data, which carry $50$ of the $70$ counts.

The prior's share $w$ is its pseudo-count total over the grand total; it never depends on what the data say, only on how many there are.

8§02.6 — an approximate credible interval

With a flat prior, $60$ of $150$ trial users renew a subscription, so the posterior for the renewal rate is $\mathrm{Beta}(61,91)$.

Find
  1. (a) Find the posterior mean and standard deviation.

  2. (b) Give an approximate central 95% credible interval.

Given
  • posterior $\mathrm{Beta}(61,91)$

  • $\alpha=0.05$

Hint 1/4

For a posterior built on $150$ observations, a normal curve with the same mean and spread is a good stand-in.

Hint 2/4

Beta mean $\frac{x}{x+y}$ and variance $\frac{xy}{(x+y)^2(x+y+1)}$; interval $m\pm 2s$.

Hint 3/4

$x=61$, $y=91$: $m=61/152$ and $\mathrm{Var}=\dfrac{61\cdot 91}{152^2\cdot 153}$.

Hint 4/4

$m\approx 0.401$, $s\approx 0.0396$, interval $\approx[0.322,\,0.481]$.

Show solution

The exact quantiles need software; with this many observations the normal shortcut is accurate to about $0.003$.

Mean

$$m=\frac{61}{152}\approx 0.4013$$

The shortcut interval is centred at the posterior mean.

Standard deviation

$$\mathrm{Var}=\frac{61\cdot 91}{152^2\cdot 153}=\frac{5551}{3534912}\approx 0.00157,\qquad s\approx 0.0396$$

The spread of the posterior sets the half-width of the interval.

Interval

$$0.4013\pm 0.0793=[0.322,\;0.481]$$

Two posterior standard deviations each side of the mean give the approximate central 95%.

Answer $$\boxed{\mathrm{Cred}_{0.05}(D)\approx[0.322,\;0.481]}$$
Check

A numerical calculation of the exact quantiles gives $[0.325,\,0.480]$, within $0.003$ of the shortcut.

Check the shortcut's two conditions before using it: many observations, and a mean well away from $0$ and $1$.

C · exam level 4 questions
1§02.2 — two days of updates, then the mean

A click rate starts from the prior $\mathrm{Beta}(2,2)$. On day one, $3$ of $8$ views are clicks; on day two, $6$ of $8$.

Find(a) What is the posterior mean after both days?
Given
  • prior $\mathrm{Beta}(2,2)$

  • day 1: $3$ clicks, $5$ skips

  • day 2: $6$ clicks, $2$ skips

Hint 1/4

Both days update the same belief; the order and the split into days do not matter.

Hint 2/4

Posterior $\mathrm{Beta}(a+N_1,\,b+N_0)$ with the pooled counts, and mean $\alpha/(\alpha+\beta)$.

Hint 3/4

Pooled: $N_1=3+6=9$ and $N_0=5+2=7$, on top of $a=b=2$.

Hint 4/4

$\mathrm{Beta}(11,9)$, mean $11/20=0.550$.

Show solution

Pooling is legal because day one's posterior is day two's prior; it saves one update.

Pool

$$N_1=9,\qquad N_0=7$$

Clicks and skips over both days.

Update

$$\mathrm{Beta}(2+9,\;2+7)=\mathrm{Beta}(11,9)$$

Pooled clicks join the first parameter and pooled skips the second.

Mean

$$\frac{11}{20}=0.550$$

The question asks for the mean: first parameter over the sum.

Answer $$\boxed{E[\Theta\mid D]=0.550}$$
Check

Day by day: $\mathrm{Beta}(2,2)\to\mathrm{Beta}(5,7)\to\mathrm{Beta}(11,9)$, the same end point.

When data come in batches, add them all up before thinking about summaries.

2§02.4 — naming a rate when only an exact hit counts

A fraud team knows a merchant's chargeback rate is $0.2$, $0.4$ or $0.6$, with prior probabilities $0.5$, $0.35$ and $0.15$. In $5$ audited orders, $3$ are chargebacks. A bonus is paid only if the team names the rate exactly.

Find(a) Which rate should the team name, and how likely is it to be wrong?
Given
  • candidates $0.2,\ 0.4,\ 0.6$ with prior $0.5,\ 0.35,\ 0.15$

  • $3$ chargebacks in $5$ orders

Hint 1/4

Only an exact hit pays, so we want the value with the highest posterior probability, and one minus that probability.

Hint 2/4

$\Pr(\Theta=\theta\mid D)\propto\Pr(\Theta=\theta)\,\theta^{3}(1-\theta)^{2}$, normalized over the three candidates.

Hint 3/4

Likelihoods $0.00512,\ 0.02304,\ 0.03456$ for priors $0.5,\ 0.35,\ 0.15$ give weights $0.00256,\ 0.008064,\ 0.005184$.

Hint 4/4

Name $0.4$: its posterior probability is $0.510$, so it is wrong with probability about $0.49$.

Show solution

A three-row table gives the whole posterior; the binomial coefficient is the same in every row and drops out.

Likelihoods

$$0.2^3(0.8)^2=0.00512,\quad 0.4^3(0.6)^2=0.02304,\quad 0.6^3(0.4)^2=0.03456$$

Three chargebacks and two clean orders.

Posterior

$$0.00256,\ \ 0.008064,\ \ 0.005184;\qquad \text{sum }0.015808$$

Bayes rule weights each candidate by its prior times its likelihood.

$$\Pr(\Theta=\theta\mid D)\approx 0.162,\ \ 0.510,\ \ 0.328$$

Dividing by the total makes the three probabilities add to one.

Decision

$$\hat\theta_{\mathrm{MAP}}=0.4,\qquad \Pr(\text{wrong})=1-0.510=0.490$$

Exact-hit loss: one minus the posterior probability of the named value.

Answer $$\boxed{\text{name }0.4;\ \ \Pr(\text{wrong})\approx 0.49}$$
Check

Naming $0.6$, the likelihood's favourite, would be wrong with probability $1-0.328=0.672$; naming $0.2$, the prior's favourite, with $0.838$.

Neither the prior's favourite nor the data's favourite wins; the posterior's does.

3§02.7 — an A/B test interval

In an A/B test, version B converts $90$ of its $600$ visitors. The report needs a 95% confidence interval for version B's conversion rate.

Find(a) Which interval is $\bar\theta\pm 2\,\mathrm{SE}$?
Given$n=600$, $S_n=90$
Hint 1/4

Estimate the rate, then its standard error from all $600$ visitors.

Hint 2/4

$\bar\theta=S_n/n$, $\mathrm{SE}=\sqrt{\bar\theta(1-\bar\theta)/n}$, interval $\bar\theta\pm 2\,\mathrm{SE}$.

Hint 3/4

With $S_n=90$ and $n=600$: $\bar\theta=0.15$, $\mathrm{SE}=\sqrt{0.15\times 0.85/600}\approx 0.0146$.

Hint 4/4

$0.15\pm 0.029=[0.121,\,0.179]$.

Show solution

Ninety conversions and 510 non-conversions are enough for the central-limit recipe.

Estimate

$$\bar\theta=\frac{90}{600}=0.15$$

The sample fraction estimates version B's conversion rate.

Standard error

$$\sqrt{\frac{0.15\times 0.85}{600}}=\sqrt{0.0002125}\approx 0.0146$$

All $600$ visitors enter $n$.

Interval

$$0.15\pm 0.0292=[0.121,\;0.179]$$

Two standard errors, not one, for about $0.95$ coverage.

Answer $$\boxed{[0.121,\;0.179]}$$
Check

Half-width $0.029$ is below the worst case $2\sqrt{0.25/600}\approx 0.041$, as it should be for a rate of $0.15$.

In an A/B report, give each version's interval with the same recipe, then compare them.

4§02.6 — find the error in a credible interval

A student computes a central 90% credible interval for the posterior $\mathrm{Beta}(1,7)$, which came from $0$ successes in $6$ trials under a flat prior. The work is below.

Find(a) Which step contains the error?
Given
  • Step 1: $F(\theta)=1-(1-\theta)^{7}$.

  • Step 2: $F(\theta_l)=0.10$, so $\theta_l=1-0.9^{1/7}\approx 0.015$.

  • Step 3: $F(\theta_u)=0.95$, so $\theta_u=1-0.05^{1/7}\approx 0.348$.

Hint 1/4

Check each step against the definition of a central interval: how much probability belongs in each tail?

Hint 2/4

Central $1-\alpha$ interval: $F(\theta_l)=\alpha/2$ and $F(\theta_u)=1-\alpha/2$.

Hint 3/4

Here $\alpha=0.10$, so the steps should read $F(\theta_l)=0.05$ and $F(\theta_u)=0.95$, with $F(\theta)=1-(1-\theta)^7$.

Hint 4/4

Step 2 is wrong: $\theta_l=1-0.95^{1/7}\approx 0.0073$, and the interval is $[0.0073,\,0.348]$.

Show solution

Checking each step against $F(\theta_l)=\alpha/2$ and $F(\theta_u)=1-\alpha/2$ is quicker than redoing the whole problem.

Step 1

$$\int_0^{\theta}7(1-t)^6\,dt=1-(1-\theta)^7$$

Correct CDF of $\mathrm{Beta}(1,7)$.

Step 2

$$F(\theta_l)=0.05\ \Rightarrow\ \theta_l=1-0.95^{1/7}\approx 0.0073$$

The left tail of a central 90% interval holds $0.05$; the student used $0.10$.

Step 3

$$F(\theta_u)=0.95\ \Rightarrow\ \theta_u=1-0.05^{1/7}\approx 0.348$$

Correct: $0.05$ remains above $\theta_u$.

Answer $$\boxed{\text{Step 2};\ \ [0.0073,\;0.348]}$$
Check

With the fix, $F(0.348)-F(0.0073)=0.95-0.05=0.90$; with the student's end point it would be $0.95-0.10=0.85$.

Write $\alpha/2$ next to each tail before solving anything.

D · interleaved 3 questions
1§02.7 — a face classifier's test errors

A face classifier is evaluated on $n=500$ test images and gets $60$ of them wrong. You want an interval for its true error rate.

Find
  1. (a) Give an approximate 95% interval for the true error rate.

  2. (b) Using the inequality in the given list, give an interval whose 95% guarantee holds for every $n$.

  3. (c) Which interval is wider, and what does the wider one buy you?

Given
  • $n=500$ test images, $60$ errors, images independent

  • for independent $Z_i\in[0,1]$ with mean $\bar Z$: $\Pr(\vert\bar Z-E[\bar Z]\vert\ge\varepsilon)\le 2e^{-2n\varepsilon^{2}}$

Hint 1/4

Two different guarantees for the same average: one approximate for large $n$, one valid for every $n$. Compute both half-widths.

Hint 2/4

(a) $\mathrm{SE}=\sqrt{\bar\theta(1-\bar\theta)/n}$. (b) Solve $2e^{-2n\varepsilon^2}=0.05$ for $\varepsilon$.

Hint 3/4

With $n=500$ and $60$ errors, $\bar\theta=0.12$; $\mathrm{SE}=\sqrt{0.12\times 0.88/500}$, and $\varepsilon=\sqrt{\ln 40/(2\cdot 500)}$.

Hint 4/4

(a) $[0.091,\,0.149]$; (b) $\varepsilon\approx 0.061$, interval $[0.059,\,0.181]$; (c) the second is wider but holds for every $n$.

Show solution

The first interval uses the normal approximation, the second the inequality; putting them side by side shows the price of a guarantee.

Central-limit interval

$$\mathrm{SE}=\sqrt{\frac{0.12\times 0.88}{500}}\approx 0.0145\ \Rightarrow\ 0.12\pm 0.029=[0.091,\;0.149]$$

The central-limit recipe: estimate plus or minus two standard errors.

Inequality interval

$$2e^{-1000\,\varepsilon^2}=0.05\ \Rightarrow\ \varepsilon^2=\frac{\ln 40}{1000}\approx 0.00369$$

Take logs: $1000\varepsilon^2=\ln 40\approx 3.689$.

$$\varepsilon\approx 0.061\ \Rightarrow\ [0.059,\;0.181]$$

The inequality bounds the chance of missing by more than $\varepsilon$, so $\bar\theta\pm\varepsilon$ covers with probability at least $0.95$.

Compare

$$0.061\approx 2.1\times 0.029$$

The inequality needs no approximation and ignores the variance, so it pays in width.

Answer $$\boxed{[0.091,\,0.149]\ \ \text{vs}\ \ [0.059,\,0.181]}$$
Check

Put $\varepsilon$ back into the bound: $2e^{-1000\times 0.0607^2}=2e^{-3.68}\approx 0.050$, the level we asked for.

The central-limit interval is the everyday tool; the inequality is what you quote when $n$ is small or a guarantee must hold exactly.

2§02.4 — chips from two machines

Chips come from machine A, with defect rate $0.05$, or machine B, with defect rate $0.2$, and a box is equally likely to come from either. You test $10$ chips from one box and find $2$ defective.

Find
  1. (a) Find the probability that the box came from machine B.

  2. (b) You must name one machine and are paid only if you are right: which do you name, and how often will you be wrong?

Given
  • $\Pr(A)=\Pr(B)=0.5$

  • defect rates $0.05$ for A and $0.2$ for B

  • $2$ defective out of $10$, chips independent

Hint 1/4

The unknown is which machine, so this is Bayes rule with two hypotheses and a binomial likelihood.

Hint 2/4

$\Pr(B\mid D)=\dfrac{\Pr(D\mid B)\Pr(B)}{\Pr(D\mid A)\Pr(A)+\Pr(D\mid B)\Pr(B)}$ with $\Pr(D\mid\theta)=\binom{10}{2}\theta^2(1-\theta)^8$.

Hint 3/4

$\Pr(D\mid A)=45(0.05)^2(0.95)^8\approx 0.0746$ and $\Pr(D\mid B)=45(0.2)^2(0.8)^8\approx 0.3020$, each with prior $0.5$.

Hint 4/4

$\Pr(B\mid D)\approx 0.802$; the MAP guess is B, wrong with probability about $0.198$.

Show solution

Equal priors cancel, so the answer is the ratio of the two likelihoods, normalized.

Likelihoods

$$\Pr(D\mid A)=45(0.0025)(0.6634)\approx 0.0746$$

$\binom{10}{2}=45$, two defects and eight good chips.

$$\Pr(D\mid B)=45(0.04)(0.1678)\approx 0.3020$$

Same count, the other machine's rate.

Posterior

$$\Pr(B\mid D)=\frac{0.3020}{0.0746+0.3020}\approx 0.802$$

Equal priors cancel from top and bottom.

Decision

$$\text{MAP}=B,\qquad \Pr(\text{wrong})=1-0.802=0.198$$

Under the exact-hit loss, the best guess is the value with the higher posterior probability.

Answer $$\boxed{\Pr(B\mid D)\approx 0.802;\ \ \text{guess B, wrong w.p.}\approx 0.198}$$
Check

Two defects in ten is a fraction of $0.2$, exactly machine B's rate, so a strong lean toward B is expected; it is not certain because A produces this sample about $7$ times in $100$.

Choosing between two machines is the MAP rule on a two-point parameter: compute the posterior, name the larger.

3§02.7 — the spread of a sample mean

Let $X_1,\dots,X_n$ be i.i.d. $\mathrm{Ber}(\theta)$ and let $\bar\theta=(X_1+\cdots+X_n)/n$.

Find
  1. (a) Show that $\mathrm{Var}(\bar\theta)=\theta(1-\theta)/n$.

  2. (b) Compute the standard deviation of $\bar\theta$ for $\theta=0.2$ and $n=64$.

  3. (c) By what factor must $n$ grow to halve that standard deviation?

Given
  • $E[X_i]=\theta$ and $\mathrm{Var}(X_i)=\theta(1-\theta)$

  • for independent variables the variances add; $\mathrm{Var}(cY)=c^2\mathrm{Var}(Y)$

  • $\theta=0.2$ and $n=64$ for part (b)

Hint 1/4

Work on the sum first, then divide by $n$.

Hint 2/4

$\mathrm{Var}(X_1+\cdots+X_n)=\sum\mathrm{Var}(X_i)$ for independent terms, and $\mathrm{Var}(S/n)=\mathrm{Var}(S)/n^2$.

Hint 3/4

With $\mathrm{Var}(X_i)=\theta(1-\theta)$, $\theta=0.2$ and $n=64$: $\sqrt{0.2\times 0.8/64}$.

Hint 4/4

(a) $\theta(1-\theta)/n$; (b) $0.05$; (c) four times as large.

Show solution

Independence lets the variance of the sum split into $n$ equal pieces; the division by $n$ then comes out squared.

Variance of the sum

$$\mathrm{Var}\Big(\sum_{i=1}^n X_i\Big)=\sum_{i=1}^n\mathrm{Var}(X_i)=n\,\theta(1-\theta)$$

Independent terms: variances add.

Divide by n

$$\mathrm{Var}(\bar\theta)=\frac{1}{n^2}\,n\,\theta(1-\theta)=\frac{\theta(1-\theta)}{n}$$

A constant factor comes out squared.

Numbers

$$\sqrt{\frac{0.2\times 0.8}{64}}=\frac{0.4}{8}=0.05$$

Standard deviation for $\theta=0.2$, $n=64$.

$$\frac{0.4}{\sqrt{256}}=0.025$$

Quadrupling $n$ to $256$ halves it.

Answer $$\boxed{\mathrm{Var}(\bar\theta)=\frac{\theta(1-\theta)}{n};\ \ 0.05;\ \ \times 4}$$
Check

With $n=1$ the formula gives $\theta(1-\theta)$, the variance of a single flip, as it must.

This $1/\sqrt n$ is why the confidence interval narrows so slowly: precision grows with the square root of the data.

Mistake ledger (15 entries)
⚠ Treating the likelihood as a density over the rate

It is a nonnegative curve in $\theta$, so it looks like one.

wrong$$\int_0^1\theta^{N_1}(1-\theta)^{N_0}\,d\theta=1$$
right$$\int_0^1\theta^{4}\,d\theta=\tfrac15\neq 1\ \Rightarrow\ \text{normalize the posterior, not the likelihood}$$
⚠ Dropping a prior that is not flat

With the flat prior the posterior is just the normalized likelihood, and the habit sticks.

wrong$$p(\theta\mid D)\propto\theta^{2}(1-\theta)^{3}\ \ \text{under the prior }2\theta$$
right$$p(\theta\mid D)\propto\theta^{2}(1-\theta)^{3}\cdot 2\theta\propto\theta^{3}(1-\theta)^{3}$$
⚠ Adding all the flips to the second parameter

The second number is easy to read as the total.

wrong$$\mathrm{Beta}(N_1+a,\;N+b)$$
right$$\mathrm{Beta}(N_1+a,\;N_0+b)$$
⚠ Swapping the roles of heads and tails

The tails count is the second number in the data and feels like it belongs first somewhere.

wrong$$\mathrm{Beta}(N_0+a,\;N_1+b)$$
right$$\mathrm{Beta}(N_1+a,\;N_0+b)$$
⚠ Dividing heads by tails instead of by the total

The ratio of the counts is what the data look like.

wrong$$\hat\theta=\frac{N_1}{N_0}=\frac{5}{15}$$
right$$\hat\theta=\frac{N_1}{N_0+N_1}=\frac{5}{20}$$
⚠ Averaging batch fractions instead of pooling counts

Each batch already comes with a fraction, and averaging feels fair.

wrong$$\tfrac13\big(\tfrac{4}{20}+\tfrac{7}{30}+\tfrac{1}{10}\big)\approx 0.178$$
right$$\frac{4+7+1}{20+30+10}=0.2$$
⚠ Using the posterior-mean formula for the MAP

The two formulas differ only by the $-1$ and $-2$.

wrong$$\hat\theta_{\mathrm{MAP}}=\frac{a+N_1}{a+b+N}$$
right$$\hat\theta_{\mathrm{MAP}}=\frac{a+N_1-1}{a+b+N-2}$$
⚠ Reading the MAP as the whole answer

A single number feels like a conclusion.

wrong$$\hat\theta_{\mathrm{MAP}}=0\ \Rightarrow\ \text{the click rate is }0$$
right$$\hat\theta_{\mathrm{MAP}}=0\ \text{but}\ \Pr(\Theta>0.1\mid D)=0.9^{5}\approx 0.59$$
⚠ Letting the data speak alone once they arrive

The MLE is the familiar number and the prior feels like a starting guess only.

wrong$$E[\Theta\mid D]=\frac{N_1}{N}$$
right$$E[\Theta\mid D]=\frac{a+N_1}{a+b+N}$$
⚠ Borrowing the MAP's pseudo-counts for the mean

The $-2$ from the peak formula leaks into the weight.

wrong$$w=\frac{a+b-2}{a+b+N-2}\ \ \text{for the mean}$$
right$$w=\frac{a+b}{a+b+N}\ \ \text{for the mean}$$
⚠ Putting all of alpha into each tail

The level is remembered, the split is not.

wrong$$\Pr(\Theta<\theta_l\mid D)=0.05\ \ \text{for a central 95\% interval}$$
right$$\Pr(\Theta<\theta_l\mid D)=0.025\ \ \text{and}\ \ \Pr(\Theta>\theta_u\mid D)=0.025$$
⚠ Using mean plus or minus 2 sd on a tiny sample

It worked for two hundred users.

wrong$$\mathrm{Beta}(5,1):\ 0.833\pm 2(0.141)=[0.552,\,1.115]$$
right$$\text{exact: }[0.025^{1/5},\,0.975^{1/5}]=[0.478,\,0.995]$$
⚠ Giving one interval a probability it does not have

It is the sentence everyone wants to say.

wrong$$\Pr(0.18\le\theta_{\mathrm{true}}\le 0.26)=0.95$$
right$$\Pr_{\text{repeated samples}}\big(\bar\theta-2\,\mathrm{SE}\le\theta_{\mathrm{true}}\le\bar\theta+2\,\mathrm{SE}\big)\approx 0.95$$
⚠ Dividing by n instead of taking the square root

The variance of the sample mean has $n$ in its denominator, and the root gets lost.

wrong$$\bar\theta\pm 2\,\frac{\bar\theta(1-\bar\theta)}{n}$$
right$$\bar\theta\pm 2\sqrt{\frac{\bar\theta(1-\bar\theta)}{n}}$$
⚠ Using the recipe when no successes were seen

The formula still returns numbers.

wrong$$0\text{ clicks in }30:\ \ 0\pm 2\sqrt{0\cdot 1/30}=[0,\,0]$$
right$$\text{flat prior: }\mathrm{Beta}(1,31),\ \ \theta_u=1-0.05^{1/31}\approx 0.092$$
Formula card
Posterior from Bayes rule
$$p(\theta\mid D)\propto\theta^{N_1}(1-\theta)^{N_0}\,p_\Theta(\theta)$$

i.i.d. Bernoulli data; normalize so the area is $1$

Conjugate Beta update
$$\mathrm{Beta}(a,b)\ \to\ \mathrm{Beta}(N_1+a,\,N_0+b)$$

Beta prior, Bernoulli or binomial data

Maximum likelihood estimate
$$\hat\theta_{\mathrm{MLE}}=\frac{N_1}{N_0+N_1}$$

$N\ge 1$; holds at the endpoints too

MAP estimate
$$\hat\theta_{\mathrm{MAP}}=\frac{a+N_1-1}{a+b+N-2}$$

$a+N_1>1$ and $b+N_0>1$ for an interior peak

Posterior mean
$$E[\Theta\mid D]=\frac{a+N_1}{a+b+N}=w\frac{a}{a+b}+(1-w)\frac{N_1}{N},\ \ w=\frac{a+b}{a+b+N}$$

Beta prior

Credible interval
$$\Pr(\theta_l\le\Theta\le\theta_u\mid D)=1-\alpha,\ \ \text{central: }\alpha/2\ \text{per tail}$$

a posterior is available; exact CDF or, for large counts, mean $\pm$ 2 sd

95% confidence interval
$$\bar\theta\pm 2\sqrt{\bar\theta(1-\bar\theta)/n}$$

$n$ large, $\bar\theta$ away from $0$ and $1$

Check yourself

Close the page and write from memory: the posterior as likelihood times prior; the Beta update; the MLE, MAP and posterior-mean formulas and the loss each one wins under; how to get a central credible interval; the confidence-interval recipe and the one sentence that reads it correctly. Then compare with the formula card.

  • Write $p(\theta\mid D)$ up to a constant for any prior, and normalize it when the prior is flat?

    c-likelihood-posterior

  • Update $\mathrm{Beta}(a,b)$ on $N_1$ heads and $N_0$ tails without mixing up the parameters?

    c-beta-prior

  • Derive $N_1/N$ from the log-likelihood and handle the all-heads case at the endpoint?

    c-mle

  • Compute a MAP estimate and say which loss makes it the best guess?

    c-map

  • Compute a posterior mean and split it into prior and data weights?

    c-posterior-mean

  • Find a central or one-sided credible interval from a CDF, and know when mean $\pm$ 2 sd is safe?

    c-credible

  • Build $\bar\theta\pm 2\,\mathrm{SE}$ and say in one sentence what its 95% means?

    c-confidence

Glossary (15 terms)
likelihoodolabilirlik

The probability of the observed data viewed as a function of the parameter, with the data held fixed.

önsel dağılım

The distribution describing what is believed about a parameter before the data are seen.

sonsal dağılım

The distribution of a parameter after the data are seen, proportional to likelihood times prior.

eşlenik önsel

A prior for which the posterior stays in the same family; the Beta prior is conjugate to Bernoulli data.

Beta distributionBeta dağılımı

A distribution on $[0,1]$ with density proportional to $\theta^{a-1}(1-\theta)^{b-1}$.

maximum likelihood estimateen çok olabilirlik kestirimi

The parameter value that makes the observed data most probable.

MAP estimate

The maximum a posteriori estimate: the value at which the posterior density is highest.

posterior meansonsal ortalama

The expected value of the parameter under the posterior; the best single estimate under squared loss.

point estimatenokta kestirimi

A single number reported for a parameter, with no statement of uncertainty.

credible interval

An interval that contains the parameter with a stated posterior probability, given the observed data.

confidence intervalgüven aralığı

An interval from a recipe that covers the fixed true parameter with a stated probability over repeated samples.

standard errorstandart hata

The estimated standard deviation of an estimator; for a sample proportion it is $\sqrt{\bar\theta(1-\bar\theta)/n}$.

kayıp fonksiyonu

A rule that assigns a cost to reporting one value when the parameter takes another.

pseudo-count

A prior parameter read as a number of imaginary earlier observations.

coverage probabilitykapsama olasılığı

The probability, over repeated samples, that an interval recipe contains the true parameter.

What comes next
§03 · Linear regression and ordinary least squares

Here the unknown was a single rate and the data were yes/no counts. Next, the unknown becomes a set of coefficients that turn inputs into a real-valued prediction, and fitting them is the first full learning problem of the course.

Sources
  • textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The syllabus line for this week names no chapter of the book, so no section numbers are cited here.
  • course materialEEE 485 lecture slides and lecture notes, Chapter 2: Bayesian and frequentist machine learning Scope, order of topics and notation follow these materials; the explanations, examples and exercises here are original.
  • course materialEEE 485 syllabus, Fall 2026-27 Weekly topic line and assessment weights.
  • standard resultBeta and Bernoulli conjugacy; the two-standard-error interval for a proportion; the normal approximation to a Beta posterior Standard results. The exact coverage table, the exact Beta quantiles and the twenty simulated intervals were computed for this page.

Spotted something missing or wrong? tell us · share your own notes or an old exam.

Last updated .