A modest rate explains the data about one time in five, so neither 'the rate is 1' nor 'certainly above 0.9' follows; 'nothing at all' fails too, since $0.1$ is almost ruled out.
Answer $$\boxed{\theta=0.6\ \text{gives this result about 1 time in 5}}$$
Check
The ratio $0.729/0.216\approx 3.4$ says $0.9$ fits better than $0.6$, but by a factor of about three, not by certainty.
Three observations are weak evidence, not zero evidence and not proof; this section is about measuring exactly how weak.
A music app shows a new playlist suggestion to four listeners, and all four accept it. The dashboard divides $4$ by $4$, prints an acceptance rate of 100%, and the product team wants to announce that listeners always accept. How much should four out of four really convince you?
By the end you can turn those four acceptances into a full distribution over the true rate, read off that a rate above $0.9$ has probability only about $0.41$, and attach a credible or a to any rate estimated from counts.
In 60 seconds
Bayesian learning puts a distribution on the unknown rate and updates it by counting, Beta prior in and Beta posterior out; frequentist learning keeps the rate fixed and judges a recipe, the fraction of heads and its $\bar\theta\pm 2\,\mathrm{SE}$ interval, by how often it works.
a frequentist interval for a proportion, $n$ large
Three most common mistakes
Reading a 95% confidence interval as 'the true rate is in this interval with probability $0.95$'. The $0.95$ belongs to the recipe over repeated samples; that sentence needs a credible interval.
Using the posterior-mean formula for the MAP, or the reverse: the MAP has the $-1$ and $-2$, the mean does not.
Adding the total $N$ instead of the tails $N_0$ to the second Beta parameter.
Two course documents disagree on the weights. The STARS syllabus for Fall 2026-27 (printed 21 September) gives Midterm 30%, Final 30%, Problem sets + Quiz 20%, project 20%; the Chapter 1 slides (issued 15 and again 24 September) give 25%, 25%, 20% and 30% in the same order. Confirm with the course which split applies.
How much time do you have?
10 minutes
The update rule that answers most Bayesian questions, and the interval recipe with the one sentence that reads it correctly.
The 60-second card · The Beta prior · Confidence intervals · Formula card
45 minutes
Every estimate and interval of the chapter once, each with a worked example, then one ladder from a full solution to a bare problem.
The 60-second card · Likelihood and posterior · The Beta prior · Maximum likelihood · The MAP estimate · The posterior mean · Credible intervals · Confidence intervals · Scaffolding comes off · Formula card
full read
Where each formula comes from, when each estimate is the right one, and enough mixed practice to decide the method yourself.
The opening pages · Recall first · Likelihood and posterior · The Beta prior · Maximum likelihood · The MAP estimate · The posterior mean · Credible intervals · Confidence intervals · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
Write the likelihood of coin-flip data and turn it into a posterior density with Bayes rule, normalizing when needed.
Update a Beta prior on yes/no data to the posterior $\mathrm{Beta}(N_1+a,\,N_0+b)$, in one batch or one observation at a time.
Derive the MLE $N_1/N$, including the endpoint cases, and apply it to raw 0/1 records and to pooled batches.
Compute the MAP estimate under a Beta prior and explain why it wins when only an exact hit counts.
Compute the posterior mean, write it as a weighted average of the prior mean and the MLE, and justify it under squared loss.
Construct a central or one-sided credible interval from a posterior CDF, or from mean $\pm$ 2 sd when the counts are large.
Construct the 95% confidence interval $\bar\theta\pm 2\,\mathrm{SE}$ for a proportion and state correctly what its 95% means.
Syllabus coverage
covered
Bayesian and frequentist machine learning — covered
Likelihood and posterior for coin-flip data
the Beta prior and its conjugate update
the MLE
the MAP estimate and the posterior mean with the loss each one minimizes
why a hides uncertainty
the frequentist view, in which the parameter is fixed and the data are random
These are the lecture's sub-topics for the chapter, in the lecture's order; they fill the first five blocks and the start of the last one.
covered
credible intervals — covered
The definition, the central interval with $\alpha/2$ in each tail, exact intervals from a closed-form posterior CDF, and one-sided bounds when the posterior peaks at $0$ or $1$.
covered
confidence intervals — covered
The central limit theorem applied to the sample mean, the interval $\bar\theta\pm 2\sqrt{\bar\theta(1-\bar\theta)/n}$, and its reading as a statement about repeated experiments.
covered
their use in data science — covered
Click-through, conversion and open rates
A/B test intervals
a classifier's error rate on a test set
sensor packet loss
planning a sample size
The applications run through the examples of every block rather than sitting in one place.
off syllabus
normal approximation to a Beta posterior — off syllabus
Mean $\pm$ 2 sd as a hand-calculated credible interval.
Further reading, not on the lecture slides. It is used here only as a calculator-free shortcut, always next to the exact interval.
off syllabus
actual coverage of the two-standard-error recipe — off syllabus
How often the $\pm 2$ SE interval really covers the true rate for small samples and rates near $0$.
Further reading. The lecture states the recipe for large $n$; the table in that block shows by exact calculation how far the coverage falls when $n$ is small.
Recall first
Bayes rule for random variables
$p(y\mid x)=\dfrac{p(x\mid y)\,p(y)}{p(x)}$, with $p(x)=\int p(x\mid y')\,p(y')\,dy'$, or a sum when $y$ is discrete.
The posterior of this section is this formula with the unknown rate in the role of $y$.
Bernoulli and binomial distributions
$X\sim\mathrm{Ber}(\theta)$ takes $1$ with probability $\theta$. The number of ones in $n$ independent trials has $\Pr(S=k)=\binom{n}{k}\theta^k(1-\theta)^{n-k}$, mean $n\theta$ and variance $n\theta(1-\theta)$.
It is the likelihood of every data set on this page.
$X\sim\mathrm{Beta}(x,y)$ has density $\dfrac{1}{B(x,y)}t^{x-1}(1-t)^{y-1}$ on $[0,1]$, mean $\dfrac{x}{x+y}$ and variance $\dfrac{xy}{(x+y)^2(x+y+1)}$.
Priors and posteriors here are Betas; the mean and variance give the posterior mean and the approximate credible interval.
Expectation and variance rules
$E[aX+b]=aE[X]+b$, $\mathrm{Var}(aX)=a^2\mathrm{Var}(X)$, and variances of independent variables add. For an event $A$, $E[I(X\in A)]=\Pr(X\in A)$.
They give the of $\bar\theta$ and the expected losses that single out the MAP and the posterior mean.
Law of large numbers and central limit theorem
For i.i.d. $X_i$ with mean $\mu$ and variance $\sigma^2$, $\bar X_n\to\mu$, and $\dfrac{\bar X_n-\mu}{\sigma/\sqrt n}$ is approximately $\mathcal N(0,1)$ for large $n$; a standard normal lies in $[-2,2]$ with probability about $0.954$.
The confidence interval is built from exactly these two facts.
Hoeffding's inequality
For independent $Z_i\in[0,1]$ with average $\bar Z$: $\Pr(\vert\bar Z-E[\bar Z]\vert\ge\varepsilon)\le 2e^{-2n\varepsilon^2}$.
One practice question compares it with the central-limit interval.
Maximizing on a closed interval
A differentiable function on $[0,1]$ attains its maximum at a critical point or at an endpoint; since $\ln$ is increasing, $f$ and $\ln f$ peak at the same place.
The MLE and the MAP are maximizers, and the all-heads case is decided at an endpoint.
Try it yourself first (2 questions)
1§02.1 — Bayes rule with two coins
A drawer holds two coins: a fair one and a bent one that lands heads with probability $0.8$. You pick one at random and flip it twice: heads, heads.
Find(a) What is the probability that you picked the bent coin?
In odds form: prior odds $1$ times likelihood ratio $0.64/0.25=2.56$ gives posterior odds $2.56$, and $2.56/3.56\approx 0.719$. Two heads make the bent coin more likely than not, but far from certain.
Replace the two coins by a whole interval of possible rates and the sum becomes an integral: that is the posterior of this section.
2§02.7 — reading a confidence interval
A report on a survey says that the 95% confidence interval for the true rate is $[0.42,\,0.58]$. A classmate rewrites the sentence in her notes.
Find(a) Is her note a correct reading of the report? Answer true or false.
Given
95% confidence interval $[0.42,\,0.58]$, from one sample
her note: there is a $0.95$ probability that the true rate lies in $[0.42,\,0.58]$
Hint 1/4
Ask what is random in the frequentist model: the true rate, or the interval?
Hint 2/4
A confidence interval's $0.95$ is the probability, over repeated samples, that the recipe's interval covers the fixed true rate.
Hint 3/4
Here the recipe produced $[0.42,\,0.58]$ from one sample; the true rate is a fixed number either inside or outside it.
Hint 4/4
So the note is false: the $0.95$ belongs to the recipe, not to this interval.
Show solution
Naming the random quantity settles the question faster than any calculation.
The probability is about the recipe; once the numbers $0.42$ and $0.58$ are in, nothing random is left.
Answer $$\boxed{\text{False}}$$
Check
Try a fixed true rate: if it is $0.5$, the statement $0.42\le 0.5\le 0.58$ is simply true, probability $1$; if it is $0.6$, it is false, probability $0$. For a fixed number the answer is never $0.95$.
A probability statement about the parameter itself needs a posterior; that is what a credible interval gives.
Notation
symbol
reads as
means
watch out
$\Theta$
capital theta
the unknown rate treated as a random variable with values in $[0,1]$
Only in the Bayesian view; the frequentist writes the same unknown as a fixed $\theta_{\mathrm{true}}$.
$\theta$
theta
one particular value of the rate, for example $0.3$
Densities are functions of $\theta$; probabilities are statements about $\Theta$.
$N_1,\ N_0,\ N$
N one, N zero, N
numbers of heads (successes), tails (failures) and flips, with $N=N_0+N_1$
$N_1$ always feeds the first Beta parameter.
$D$
D
the observed data, $\{N_1\text{ heads},\,N_0\text{ tails}\}$
For coin-flip data only the counts matter, not the order.
$p_{D\mid\Theta}(D\mid\theta)$
the likelihood of D given theta
probability of the observed data if the rate were $\theta$
As a function of $\theta$ it does not integrate to $1$.
$p_{\Theta\mid D}(\theta\mid D)$
the posterior density
the density of the rate after seeing $D$
This one does integrate to $1$ over $\theta$.
$\mathrm{Beta}(a,b)$
Beta a b
the density $\theta^{a-1}(1-\theta)^{b-1}/B(a,b)$ on $[0,1]$
$B(a,b)$ is only the constant that makes the area $1$.
the maximizers of the likelihood and of the posterior
The hat marks a number computed from data.
$E[\Theta\mid D]$
the posterior mean
the average of $\Theta$ under the posterior
Equal to the MAP only when the posterior is symmetric.
$\mathrm{Cred}_\alpha(D)$
the credible interval at level one minus alpha
an interval $[\theta_l,\theta_u]$ holding posterior probability $1-\alpha$
$\alpha=0.05$ gives a 95% interval, with $0.025$ in each tail if it is central.
$\theta_{\mathrm{true}},\ S_n,\ \bar\theta$
theta true, S n, theta bar
the fixed unknown rate, the number of successes in $n$ trials, and $S_n/n$
$\bar\theta$ changes from sample to sample; $\theta_{\mathrm{true}}$ does not.
$\mathrm{SE}$
standard error
$\sqrt{\bar\theta(1-\bar\theta)/n}$, the estimated standard deviation of $\bar\theta$
It shrinks like $1/\sqrt n$, not like $1/n$.
Conventions used here
Capital and small theta.
$\Theta$ is the unknown rate as a random variable (the Bayesian view) and $\theta$ is a value in $[0,1]$. In the frequentist blocks the unknown rate is a fixed number, $\theta_{\mathrm{true}}$, and nothing about it is random.
Most interval mistakes come from forgetting which symbol carries the randomness.
Which count feeds which parameter.
$N_1$ counts the ones (heads, clicks, defects, whatever the question counts) and goes into the first Beta parameter; $N_0$ counts the zeros and goes into the second.
Defects or losses can be the ones; the rule follows what is counted, not whether it is good news.
The 95% multiplier.
Following the lecture, a 95% confidence interval steps $2$ standard errors each side. The exact normal value is $1.96$, which gives probability $0.950$ instead of $0.954$. If a question fixes the multiplier use it; otherwise say which one you used.
Answers built with $2$ and with $1.96$ differ in the third decimal, and both appear in textbooks.
Intervals are closed and live in [0, 1].
All intervals here are closed, $[\theta_l,\theta_u]$. If an approximate interval pokes outside $[0,1]$, the approximation is being used outside its safe range, and we say so instead of quietly clipping it.
A lower end of $-0.006$ is a warning sign about the method, not a rate.
What the probability refers to.
In a credible interval the probability is over $\Theta$, given the data you have. In a confidence interval it is over repeated data sets, with $\theta_{\mathrm{true}}$ held fixed.
The two can give almost the same numbers and still mean different things.
Rounding.
Intermediate steps keep at least four significant figures; final answers are rounded to three decimals unless an exact fraction is asked for.
Rounding a standard error early can move an interval end by one unit in the third decimal.
2.1Likelihood and posterior: one formula, read two ways
Treats the unknown rate as a random variable and uses Bayes rule to turn counts into a distribution over that rate.
The previous section gave us Bayes rule for random variables; here we point it at the unknown probability itself.
Solvable with what we have
Compute the chance of the data for a given rate: if $\theta=0.7$, four acceptances in four happen with probability $0.7^4\approx 0.24$.
Compare two named rates with Bayes rule, say $0.9$ against $0.5$.
Say where the fraction of successes is heading as the sample grows, by the law of large numbers.
Not solvable yet
Say how probable it is that the true rate exceeds $0.9$, given only four listeners.
Give an estimate after four out of four that is not the absurd value $1$.
Attach any error bar to a rate measured on four people.
Take the fraction as the truth: $\hat\theta=4/4=1$. Then the next $100$ listeners all accept with probability $1^{100}=1$, and the team can promise it.
Why it fails
A rate of $0.8$ produces four out of four with probability $0.8^4\approx 0.41$, which is hardly rare. The fraction makes a rate above $0.9$ certain; a distribution over $\theta$ will put that chance near $0.41$. One number throws the whole range of plausible rates away.
RuleBayes rule for a parameter
Conditions
$\Theta$ takes values in $[0,1]$ and has prior density $p_\Theta(\theta)$
given $\Theta=\theta$, the flips $X_1,\dots,X_N$ are i.i.d. $\mathrm{Ber}(\theta)$
the data $D$ record $N_1$ heads and $N_0$ tails, $N=N_1+N_0$
After the data, each candidate rate gets a weight equal to its prior weight times how well it explains the data. The denominator $p_D(D)$ does not involve $\theta$; it only rescales the curve so that its area is $1$.
Where the formula comes from
Take one particular sequence, say heads, heads, tails. Independence lets us multiply: $\theta\cdot\theta\cdot(1-\theta)$. Any sequence with $N_1$ heads and $N_0$ tails has probability $\theta^{N_1}(1-\theta)^{N_0}$, whatever the order.
The data record only the counts, and $\binom{N_1+N_0}{N_1}$ orders give the same counts, so $$p_{D\mid\Theta}(D\mid\theta)=\binom{N_1+N_0}{N_1}\theta^{N_1}(1-\theta)^{N_0}.$$
Bayes rule for random variables, with $\Theta$ as the unknown and $D$ as the observation, gives $$p_{\Theta\mid D}(\theta\mid D)=\frac{p_{D\mid\Theta}(D\mid\theta)\,p_\Theta(\theta)}{p_D(D)},\qquad p_D(D)=\int_0^1 p_{D\mid\Theta}(D\mid t)\,p_\Theta(t)\,dt.$$
Neither the binomial coefficient nor $p_D(D)$ contains $\theta$. Dropping them changes the height of the curve, not its shape, which is why we write $\propto$ and normalize at the end.
Flat prior (dashed) and the posterior $\textcolor{#d1690a}{5\theta^4}$ after four acceptances out of four. The shaded area, $1-0.9^5\approx 0.41$, is the posterior probability that the rate exceeds $0.9$.
Looks like this, but is not
The likelihood $\theta^{4}$ from four out of four is a nonnegative curve on $[0,1]$, so it looks like a density for $\theta$.
Its area is $\int_0^1\theta^4\,d\theta=\tfrac15$, not $1$. For each fixed $\theta$ the likelihood is a distribution over data; as a function of $\theta$ it is only a weight. Times the prior and divided by $p_D(D)$, it becomes the density $5\theta^4$.
Four acceptances out of four: the chance that the rate exceeds 0.9
A playlist suggestion is shown to $4$ listeners and all $4$ accept. Take a flat prior, $p_\Theta(\theta)=1$ on $[0,1]$, and find the posterior probability that the acceptance rate is above $0.9$.
Find$\Pr(\Theta>0.9\mid D)$
Given
$N_1=4$ acceptances, $N_0=0$ refusals
prior $p_\Theta(\theta)=1$ for $0\le\theta\le 1$
Solution
We write the posterior up to a constant and fix the constant with one easy integral; computing $p_D(D)$ separately would cost the same integral twice.
The posterior median solves $\theta^5=0.5$, so it is $0.5^{1/5}\approx 0.871<0.9$; the region above $0.9$ must therefore hold less than half the area, and $0.41$ does. Under the flat prior the same region held $0.1$, so four acceptances roughly quadruple it.
The dashboard's 100% was the wrong kind of answer: a rate above $0.9$ gets about 41 chances in 100. A normalized posterior turns vague plausibility into numbers you can report.
Two candidate rates: Bayes rule with a sum instead of an integral
A product lead believes a new suggestion is either average, $\theta=0.5$, or excellent, $\theta=0.9$, and puts probability $0.8$ on average. Then $4$ listeners out of $4$ accept. How probable is excellent now?
Find$\Pr(\Theta=0.9\mid D)$
Given
two candidate rates: $0.5$ with prior $0.8$, and $0.9$ with prior $0.2$
$4$ acceptances in $4$
Solution
With only two candidates the evidence $p_D(D)$ is a two-term sum, so we compute it directly instead of normalizing a curve.
In odds form: prior odds $0.2/0.8=0.25$ times likelihood ratio $0.6561/0.0625=10.4976$ gives posterior odds $2.624$, and $2.624/3.624\approx 0.724$. The data were enough to overturn a prior that was $4$ to $1$ against excellent.
Continuous or discrete, the recipe is the same: prior times likelihood, then divide by the total so the probabilities add to one.
Checkpoint
§02.1 — posterior shape from prior and likelihood
A colleague's prior for a coin's heads rate is tilted toward heads: $p_\Theta(\theta)=2\theta$ on $[0,1]$. The coin is then flipped $5$ times.
Find(a) Which expression is proportional to the posterior density $p_{\Theta\mid D}(\theta\mid D)$?
Given
$2$ heads and $3$ tails
prior $p_\Theta(\theta)=2\theta$ for $0\le\theta\le 1$
Hint 1/4
We need the shape of the posterior as a function of $\theta$, so any factor without $\theta$ can be dropped.
The result is symmetric about $0.5$: two heads plus a prior worth one extra head balance three tails, so a symmetric posterior is exactly what we should expect.
A prior of the form $\theta^{k}$ acts like $k$ extra heads; keep an eye out for that pattern in the next block.
⚠ Treating the likelihood as a density over the rate
it is a nonnegative curve in $\theta$, so it looks like one
Add the heads to the first parameter and the tails to the second. For the posterior mean, $\mathrm{Beta}(a,b)$ acts exactly like $a$ earlier heads and $b$ earlier tails, so a large $a+b$ means a stubborn prior.
Why the posterior stays a Beta
The prior density is $p_\Theta(\theta)=\dfrac{1}{B(a,b)}\theta^{a-1}(1-\theta)^{b-1}$ on $[0,1]$.
Multiply by the likelihood and keep only what depends on $\theta$: $$p(\theta\mid D)\propto\theta^{N_1}(1-\theta)^{N_0}\cdot\theta^{a-1}(1-\theta)^{b-1}=\theta^{N_1+a-1}(1-\theta)^{N_0+b-1}.$$
The right side has the $\mathrm{Beta}(N_1+a,\,N_0+b)$ shape. A density with that shape can have only one constant in front, $1/B(N_1+a,\,N_0+b)$, because its area must be $1$. So the posterior is exactly that Beta.
A prior whose posterior stays in the same family is called conjugate to the likelihood. The Beta family is conjugate to Bernoulli and binomial data.
Email open rate. Prior $\textcolor{#8250df}{\mathrm{Beta}(2,6)}$ from past campaigns (dashed), likelihood of $7$ opens in $10$ rescaled to area $1$ (dotted, blue), and posterior $\textcolor{#d1690a}{\mathrm{Beta}(9,9)}$. The posterior peak lands between the other two.
Looks like this, but is not
A normal prior on the rate, say $\mathcal N(0.5,\,0.1^2)$, looks as sensible as a Beta: smooth, one peak, centred where we want it.
Multiplied by $\theta^{N_1}(1-\theta)^{N_0}$ it gives neither a normal nor a Beta, so there is no closed form and every summary needs numerical integration. It also puts a little probability below $0$ and above $1$, where a rate cannot be.
Email open rate: a Beta(2, 6) prior, then 7 opens in 10
Past campaigns suggest open rates around $0.25$, encoded as the prior $\mathrm{Beta}(2,6)$. A new subject line goes to $10$ test users and $7$ open it. Find the posterior and the posterior probability that the new open rate exceeds $0.5$.
FindThe posterior and $\Pr(\Theta>0.5\mid D)$.
Given
prior $\mathrm{Beta}(2,6)$, mean $2/8=0.25$
$N_1=7$ opens, $N_0=3$ non-opens
Solution
The conjugate rule gives the posterior in one line, so no integral is needed for the first part; for the second, the posterior turns out symmetric, and symmetry answers without integrating.
The posterior mean $9/18=0.5$ lies between the prior mean $0.25$ and the data fraction $0.7$, as a compromise must. Ten real users against a prior worth $8$ earlier ones should land near the middle, and it does.
Seven opens in ten did not prove the new line beats one half; they pulled a pessimistic prior up to even odds.
Two days of clicks: updating twice gives the same posterior as updating once
Start from the flat prior $\mathrm{Beta}(1,1)$ for a click rate. Monday brings $3$ clicks in $8$ views; Tuesday brings $5$ clicks in $12$ views. Compare updating after each day with updating once on the pooled counts.
FindThe posterior after both days, computed both ways.
Given
prior $\mathrm{Beta}(1,1)$
Monday: $3$ clicks, $5$ skips
Tuesday: $5$ clicks, $7$ skips
Solution
Monday's posterior becomes Tuesday's prior; doing it both ways is the quickest proof that batching the data does not matter.
Both routes add the same numbers to the same parameters in a different order, so they must agree. The pooled fraction $8/20=0.4$ sits close to the posterior mean $0.409$, as it should with a prior worth only two views.
A running posterior is all a system has to store: two numbers, updated one observation at a time.
Checkpoint
§02.2 — conjugate Beta update
A coin collector believes her coins are close to fair and encodes that as the prior $\mathrm{Beta}(4,4)$. She flips one coin $20$ times.
Find(a) Which distribution is the posterior for the heads rate?
Given
prior $\mathrm{Beta}(4,4)$
$15$ heads in $20$ flips
Hint 1/4
Split the $20$ flips into heads and tails before touching the prior.
Hint 2/4
$\mathrm{Beta}(a,b)$ prior with $N_1$ heads and $N_0$ tails gives $\mathrm{Beta}(N_1+a,\,N_0+b)$.
Hint 3/4
Here $a=b=4$, $N_1=15$ and $N_0=20-15=5$.
Hint 4/4
The posterior is $\mathrm{Beta}(19,9)$.
Show solution
The conjugate rule needs the tails count, which the question hides inside the total; getting it first prevents the most common slip.
Get both counts
$$N_0=20-15=5$$
The second parameter wants tails, not flips.
Add them to the prior
$$\mathrm{Beta}(15+4,\;5+4)=\mathrm{Beta}(19,9)$$
A Beta prior with Bernoulli data stays Beta: heads join $a$, tails join $b$.
Answer $$\boxed{\mathrm{Beta}(19,9)}$$
Check
The parameters add to $28=8+20$: the prior's $8$ pseudo-flips plus the $20$ real ones. A total that does not match signals a slip.
Always check that the posterior parameters add up to $a+b+N$.
⚠ Adding all the flips to the second parameter
the second number is easy to read as the total
wrong$$\mathrm{Beta}(N_1+a,\;N+b)$$
right$$\mathrm{Beta}(N_1+a,\;N_0+b)$$
⚠ Swapping the roles of heads and tails
the tails count is the second number in the data and feels like it belongs first somewhere
wrong$$\mathrm{Beta}(N_0+a,\;N_1+b)$$
right$$\mathrm{Beta}(N_1+a,\;N_0+b)$$
Eight views arrive one at a time: click, skip, click, click, skip, click, skip, click. Step back to the flat prior and forward again. Each click adds $1$ to the first parameter, each skip adds $1$ to the second, and the curve narrows as the counts grow.
At the edges
0 views Beta(1, 1)
Nothing seen yet, so the posterior is the flat prior.
8 views Beta(6, 4)
Five clicks and three skips: the same curve you get by updating once on the totals.
2.3Maximum likelihood: the rate that makes the data most probable
Picks the value of $\theta$ that makes the observed data most probable; for coin flips, the fraction of heads.
A posterior needs a prior; maximum likelihood drops the prior and asks only which $\theta$ explains the data best.
Theorem for Bernoulli data
Conditions
i.i.d. $\mathrm{Ber}(\theta)$ flips with $N_1$ heads and $N_0$ tails, $N=N_1+N_0\ge 1$
Of all the rates, the one that makes the observed counts most probable is the plain fraction of heads. No prior enters; only the data and the form of the likelihood do.
Derivation, with the endpoint check
Maximize $L(\theta)=\theta^{N_1}(1-\theta)^{N_0}$ on $[0,1]$; the binomial coefficient does not move the maximizer. Since $\ln$ is increasing, $L$ and its logarithm peak at the same place: $$\ell(\theta)=N_1\ln\theta+N_0\ln(1-\theta),\qquad 0<\theta<1.$$
Set the derivative to zero: $$\ell'(\theta)=\frac{N_1}{\theta}-\frac{N_0}{1-\theta}=0\ \Longrightarrow\ N_1(1-\theta)=N_0\,\theta\ \Longrightarrow\ \theta=\frac{N_1}{N_0+N_1}.$$
It is a maximum: $\ell''(\theta)=-\dfrac{N_1}{\theta^2}-\dfrac{N_0}{(1-\theta)^2}<0$ on $(0,1)$ when both counts are positive.
Endpoints: with $N_1\ge 1$ and $N_0\ge 1$, $L(0)=L(1)=0$, so the interior point wins. With $N_0=0$, $L(\theta)=\theta^{N_1}$ rises all the way and the maximum is $\theta=1=N_1/N$; $N_1=0$ mirrors it at $0$. The formula holds in every case.
Likelihood of $3$ heads in $10$ flips, $\textcolor{#1f6feb}{L(\theta)=\theta^3(1-\theta)^7}$, in thousandths. The peak sits at $\textcolor{#d1690a}{\hat\theta_{\mathrm{MLE}}=0.3}$, and both ends are zero, so the endpoint check is quick here.
Looks like this, but is not
With all heads, say $N_1=3$ and $N_0=0$, it seems we can still find the MLE by solving $\ell'(\theta)=0$.
Now $\ell'(\theta)=3/\theta$ is positive on all of $(0,1)$ and never zero. The likelihood $\theta^3$ keeps rising, so the maximum is the endpoint $\theta=1$. The answer $N_1/N=1$ survives, but only the endpoint check finds it.
θ
L(θ) = θ³(1 − θ)⁷
$0.1$
$0.000478$
$0.2$
$0.001678$
$0.3$
$0.002224$
$0.4$
$0.001792$
$0.5$
$0.000977$
The column rises to $\theta=0.3$ and falls after it: $0.002224$ beats both neighbours, $0.001678$ and $0.001792$. The derivative found the same point without a grid.
Packet loss on a sensor link: the MLE from a 0/1 record
A sensor node sends $12$ packets and logs $1$ for a lost packet and $0$ for a delivered one: $$0,0,1,0,0,0,0,1,0,0,0,0.$$ Model the losses as i.i.d. $\mathrm{Ber}(\theta)$ and find the maximum likelihood estimate of the loss rate.
Find$\hat\theta_{\mathrm{MLE}}$
Given
the 12-packet record above
losses i.i.d. $\mathrm{Ber}(\theta)$
Solution
Counting first and applying the formula is faster than differentiating a 12-factor product; we still test the answer against its neighbours, because a miscount fails silently.
The log-likelihood $2\ln\theta+10\ln(1-\theta)$ has derivative $2/\theta-10/(1-\theta)$, and at $\theta=1/6$ this is $12-12=0$, so $1/6$ is the critical point.
With Bernoulli data the MLE never needs the full product: count, divide, and make sure the result is a fraction between $0$ and $1$.
Three days of one banner: pool the counts, do not average the rates
An online shop runs the same banner on three days: $4$ clicks in $20$ views, $7$ in $30$, and $1$ in $10$. Assuming one click rate $\theta$ for all views, find its MLE.
Find$\hat\theta_{\mathrm{MLE}}$ for the common rate.
Given
day 1: $4$ of $20$
day 2: $7$ of $30$
day 3: $1$ of $10$
all views i.i.d. $\mathrm{Ber}(\theta)$
Solution
Independent days multiply, so their counts add; averaging the three daily fractions would give a 10-view day the same say as a 30-view day.
Same maximization as before with $N_1=12$ and $N_0=48$.
Answer $$\boxed{\hat\theta_{\mathrm{MLE}}=0.2}$$
Check
At $\theta=0.2$ the expected number of clicks in $60$ views is $12$, exactly the total observed. The average of the daily fractions, $\tfrac13\big(\tfrac{4}{20}+\tfrac{7}{30}+\tfrac{1}{10}\big)\approx 0.178$, would predict only about $10.7$.
Pool counts, not fractions: the MLE weights every observation equally, whatever batch it came in.
Checkpoint
§02.3 — MLE from counts
A quality engineer tests $20$ solder joints and $5$ of them fail. She models failures as i.i.d. $\mathrm{Ber}(\theta)$.
Find(a) Which value of $\theta$ maximizes $\theta^{5}(1-\theta)^{15}$?
Given$N_1=5$ failures, $N_0=15$ good joints
Hint 1/4
We want the peak of the likelihood, and for Bernoulli data the peak has a closed form.
Hint 2/4
$\hat\theta_{\mathrm{MLE}}=N_1/(N_0+N_1)$.
Hint 3/4
Here $N_1=5$ and $N_0=15$, so $N=20$.
Hint 4/4
The maximizer is $5/20=0.25$.
Show solution
The closed form saves differentiating; a one-line check with the log-derivative confirms it.
Apply the formula
$$\hat\theta=\frac{5}{5+15}=0.25$$
The MLE of a Bernoulli rate is the count of ones over the total: failures over all joints.
Confirm with the derivative
$$\frac{5}{0.25}-\frac{15}{0.75}=20-20=0$$
The log-likelihood's slope is zero there.
Answer $$\boxed{\hat\theta_{\mathrm{MLE}}=0.25}$$
Check
The neighbours do worse: $L(0.2)\approx 1.13\times 10^{-5}$ and $L(0.3)\approx 1.15\times 10^{-5}$, against $L(0.25)\approx 1.31\times 10^{-5}$.
Divide by the total, never by the other count.
⚠ Dividing heads by tails instead of by the total
the ratio of the counts is what the data look like
The MAP estimate is the MLE computed on inflated counts: $a-1$ extra heads and $b-1$ extra tails. The flat prior $\mathrm{Beta}(1,1)$ adds no extras, so there MAP equals MLE.
Derivation, and why the peak is the right guess for exact hits
The posterior is $\mathrm{Beta}(\alpha,\beta)$ with $\alpha=a+N_1$ and $\beta=b+N_0$, so we maximize $\theta^{\alpha-1}(1-\theta)^{\beta-1}$. That is the MLE problem with $N_1$ replaced by $\alpha-1$ and $N_0$ by $\beta-1$: $$\hat\theta_{\mathrm{MAP}}=\frac{\alpha-1}{\alpha+\beta-2}=\frac{a+N_1-1}{a+b+N_0+N_1-2}.$$
Suppose $\Theta$ can take only a few values and you score only by naming it exactly. Under the loss $I(\hat\theta\neq\Theta)$ your expected loss is $$E[I(\hat\theta\neq\Theta)\mid D]=\Pr(\Theta\neq\hat\theta\mid D)=1-\Pr(\Theta=\hat\theta\mid D),$$ which is smallest for the most probable value.
For a continuous $\Theta$ every single point has probability $0$, so exact hits are impossible. The honest version counts a miss when $\vert\hat\theta-\Theta\vert>\varepsilon$; the best guess is the centre of the heaviest window of width $2\varepsilon$, and as $\varepsilon$ shrinks that centre approaches the peak.
Zero clicks in four views under a flat prior gives the posterior $\textcolor{#d1690a}{5(1-\theta)^4}$, that is $\mathrm{Beta}(1,5)$. Here $a+N_1=1$, so the peak sits on the boundary and the MAP is $0$, yet the shaded area above $0.1$ is $0.9^5\approx 0.59$. A point estimate says nothing about where the rest of the probability is.
Looks like this, but is not
The posterior mean $\dfrac{a+N_1}{a+b+N}$ looks like the MAP formula with the $-1$ and $-2$ tidied away, and for the posterior $\mathrm{Beta}(9,9)$ both give $0.5$.
They agree only when the posterior is symmetric. For $\mathrm{Beta}(2,5)$ the peak is $\tfrac{1}{5}=0.2$ but the mean is $\tfrac27\approx 0.286$. The subtractions are what move you from the average to the peak.
Click rate with a Beta(3, 7) prior after 12 clicks in 40 views
A recommender team encodes earlier experience as the prior $\mathrm{Beta}(3,7)$ for a new item's click rate. The item then gets $12$ clicks in $40$ views. Find the MAP estimate and compare it with the MLE and with the prior's own peak.
Find$\hat\theta_{\mathrm{MAP}}$, with $\hat\theta_{\mathrm{MLE}}$ and the prior mode for comparison.
Given
prior $\mathrm{Beta}(3,7)$
$N_1=12$ clicks, $N_0=28$ skips
Solution
Updating first and taking the peak second keeps each step checkable; one long substitution into the formula works too but hides where an error came from.
It sits between the prior mode $0.25$ and the MLE $0.3$, much closer to the MLE because $40$ real views outweigh $8$ pseudo-views; the weighted-average line reproduces it independently.
Read a MAP as data plus pseudo-counts: once the real counts dwarf $a+b-2$, MAP and MLE are nearly the same number.
Three possible defect rates: the MAP minimizes the chance of being wrong
A supplier's machine runs at one of three defect rates, $0.1$, $0.3$ or $0.5$, with prior probabilities $0.6$, $0.3$ and $0.1$. A sample of $5$ items has $2$ defectives. You must name the rate and are scored only on naming it exactly. Which do you name, and how often will you be wrong?
FindThe rate to name and the probability that it is wrong.
Given
candidates $0.1,\ 0.3,\ 0.5$ with prior $0.6,\ 0.3,\ 0.1$
$2$ defectives in $5$ items
Solution
With three candidates Bayes rule is a three-term sum, so we build the whole posterior table; the binomial coefficient is common to all rows and can be left out.
Naming $0.5$ would be wrong with probability $1-0.186=0.814$ and naming $0.1$ with $1-0.261=0.739$, both worse. By a hair, $0.5$ has the larger likelihood, $0.03125$ against $0.03087$, so the MLE among the three would be $0.5$: the prior is what moves the answer.
When only an exact hit counts, name the value with the highest posterior probability, even if another value explains the data slightly better.
Checkpoint
§02.4 — MAP with a Beta prior
A coach's prior for a player's free-throw rate is $\mathrm{Beta}(2,2)$, mild and centred at one half. In practice the player makes $9$ of $10$.
Find(a) What is the MAP estimate of the free-throw rate?
Given
prior $\mathrm{Beta}(2,2)$
$9$ made, $1$ missed
Hint 1/4
The MAP is the peak of the posterior, so get the posterior parameters first.
Average every candidate rate, weighted by its posterior density. For a Beta posterior the result is a weighted average of the prior mean and the fraction of heads, and the weight on the prior shrinks as $N$ grows.
Derivation, the weighted average, and why squared loss picks the mean
A $\mathrm{Beta}(x,y)$ variable has mean $x/(x+y)$. The posterior is $\mathrm{Beta}(a+N_1,\,b+N_0)$, which gives the boxed formula at once.
Split it to see the weights: $$\frac{a+N_1}{a+b+N}=\underbrace{\frac{a+b}{a+b+N}}_{w}\cdot\frac{a}{a+b}+\underbrace{\frac{N}{a+b+N}}_{1-w}\cdot\frac{N_1}{N}.$$
Under squared loss the expected loss of a guess $\hat\theta$ splits in two, with $m=E[\Theta\mid D]$: $$E[(\hat\theta-\Theta)^2\mid D]=(\hat\theta-m)^2+\mathrm{Var}(\Theta\mid D).$$ Only the first part depends on $\hat\theta$, and it vanishes exactly at $\hat\theta=m$.
The split comes from writing $\hat\theta-\Theta=(\hat\theta-m)+(m-\Theta)$ and expanding the square; the cross term has expectation $2(\hat\theta-m)\,E[m-\Theta\mid D]=0$.
Prior $\textcolor{#8250df}{\mathrm{Beta}(2,6)}$ with mean $0.25$, and data with $7$ heads in every $10$ flips, so the MLE is $\textcolor{#1f6feb}{0.7}$ throughout. The posterior mean $\textcolor{#d1690a}{(2+N_1)/(8+N)}$ slides toward the MLE as $N$ grows from $10$ to $160$.
Looks like this, but is not
The posterior mean averages everything the posterior knows, so it looks like the safest guess in any game.
With the three possible defect rates of the previous block, the posterior mean is $0.1(0.261)+0.3(0.553)+0.5(0.186)\approx 0.285$, which is not one of the three rates. If only an exact hit scores, it loses every time. The best estimate depends on the loss.
The posterior mean as a weighted average: a Beta(4, 6) prior and 30 heads in 40
The prior $\mathrm{Beta}(4,6)$ has mean $0.4$. The data are $30$ heads in $40$ flips. Find the posterior mean directly, then again as a weighted average of the prior mean and the MLE.
Find$E[\Theta\mid D]$, computed both ways.
Given
prior $\mathrm{Beta}(4,6)$
$N_1=30$, $N_0=10$, $N=40$
Solution
Computing it twice costs one extra line and catches the most common slip, feeding a count into the wrong parameter.
The prior counts as $10$ pseudo-flips against $40$ real ones.
$$0.2(0.4)+0.8(0.75)=0.08+0.6=0.68$$
The weight $w$ multiplies the prior mean $4/10=0.4$, and the rest goes to the MLE $30/40=0.75$.
Answer $$\boxed{E[\Theta\mid D]=0.68}$$
Check
The two routes agree, and $0.68$ lies between $0.4$ and $0.75$: $0.07$ from the MLE and $0.28$ from the prior mean, four times closer to the data because the data carry four times the weight.
Prior strength is measured in flips: $a+b$ pseudo-flips against $N$ real ones decides how far the estimate moves.
Zero clicks in four views: MAP or mean under squared loss?
A new ad gets $0$ clicks in $4$ views, and the prior is flat, so the posterior is $\mathrm{Beta}(1,5)$ with MAP $0$ and mean $1/6$. The team pays $(\hat\theta-\theta)^2$ for its estimate. Compare the expected cost of reporting the MAP with that of reporting the mean.
Find$E[(\hat\theta-\Theta)^2\mid D]$ for both estimates.
The MAP is $0$, so its cost is just $E[\Theta^2\mid D]$, which the Beta moment formula gives directly: $\frac{1\cdot 2}{6\cdot 7}=\frac{2}{42}\approx 0.0476$, the same number as the split.
Under squared loss, report the mean; for a skewed posterior it beats the peak by a wide margin.
Checkpoint
§02.5 — posterior mean
A coin is believed to be roughly fair, with prior $\mathrm{Beta}(3,3)$. It is flipped $8$ times and lands heads once.
Find(a) What is the posterior mean of the heads rate?
Given
prior $\mathrm{Beta}(3,3)$
$1$ head, $7$ tails
Hint 1/4
The mean of the posterior is asked, so find the posterior's parameters first.
Hint 2/4
$E[\Theta\mid D]=\dfrac{a+N_1}{a+b+N}$.
Hint 3/4
With $a=b=3$, $N_1=1$ and $N=8$: $\dfrac{3+1}{3+3+8}$.
Hint 4/4
$E[\Theta\mid D]=4/14\approx 0.286$.
Show solution
Update, then use the Beta mean; the weighted-average form doubles as a check.
Update
$$\mathrm{Beta}(3+1,\;3+7)=\mathrm{Beta}(4,10)$$
One head joins the first parameter, seven tails the second.
Mean
$$\frac{4}{4+10}=\frac{4}{14}\approx 0.286$$
The posterior is a Beta, and a Beta's mean is its first parameter over the sum.
Answer $$\boxed{E[\Theta\mid D]\approx 0.286}$$
Check
As a weighted average: $\tfrac{6}{14}(0.5)+\tfrac{8}{14}(0.125)=\tfrac{3+1}{14}\approx 0.286$, the same number.
A fair-leaning prior keeps one head in eight from dragging the estimate all the way down to $0.125$.
⚠ Letting the data speak alone once they arrive
the MLE is the familiar number and the prior feels like a starting guess only
wrong$$E[\Theta\mid D]=\frac{N_1}{N}$$
right$$E[\Theta\mid D]=\frac{a+N_1}{a+b+N}$$
⚠ Borrowing the MAP's pseudo-counts for the mean
the $-2$ from the peak formula leaks into the weight
wrong$$w=\frac{a+b-2}{a+b+N-2}\ \ \text{for the mean}$$
right$$w=\frac{a+b}{a+b+N}\ \ \text{for the mean}$$
2.6Credible intervals: where the posterior puts its probability
Reports an interval that holds the parameter with a stated posterior probability, usually cutting equal probability from each tail.
A point estimate hides the spread of the posterior; an interval reports it.
DefinitionCredible interval
Conditions
a posterior density $p(\theta\mid D)$ is available
Given the data you actually saw, $\Theta$ lies in $[\theta_l,\theta_u]$ with probability $1-\alpha$. Many intervals have that property; the central one cuts $\alpha/2$ of the posterior probability from each end.
Three ways to get the numbers
Closed-form CDF. When the posterior is $\mathrm{Beta}(a,1)$ or $\mathrm{Beta}(1,b)$, its CDF is $\theta^{a}$ or $1-(1-\theta)^{b}$, and each end point solves a one-line equation.
Integrate a polynomial. For small whole-number parameters the density is a polynomial, so a tail probability such as $\Pr(\Theta>0.5\mid D)$ is an exact integral.
Normal approximation (further reading, not on the lecture slides). For large counts a Beta posterior is close to a normal curve with the same mean and variance, so $$[\theta_l,\theta_u]\approx E[\Theta\mid D]\pm 2\sqrt{\mathrm{Var}(\Theta\mid D)}.$$ With few observations, or a mean near $0$ or $1$, compare it with the exact interval.
Four acceptances out of four with a flat prior give $\textcolor{#d1690a}{5\theta^4}$. Cutting $0.05$ of probability from each end leaves the central 90% credible interval $[0.549,\,0.990]$. Near $\theta=1$ the density is high, so the right tail is a thin strip.
Looks like this, but is not
It seems the central 95% interval must contain the most probable value, since that is where the density is highest.
For the posterior $\mathrm{Beta}(1,5)$ the density is highest at $\theta=0$, but the central interval is about $[0.005,\,0.522]$ and leaves $0$ out, because it must cut $0.025$ from the left end too. When the peak sits on the boundary, a one-sided interval $[0,\theta_u]$ is the better report.
posterior
exact (numerical)
mean ± 2 sd
$\mathrm{Beta}(5,1)$
$[0.478,\ 0.995]$
$[0.552,\ 1.115]$
$\mathrm{Beta}(3,60)$
$[0.010,\ 0.112]$
$[-0.006,\ 0.101]$
$\mathrm{Beta}(47,155)$
$[0.177,\ 0.293]$
$[0.173,\ 0.292]$
For $\mathrm{Beta}(5,1)$, four observations under a flat prior, the shortcut runs past $1$. With a mean near $0$ it dips below $0$. With about two hundred observations and a mean near $0.23$ it is off by less than $0.005$.
Four out of four: a 90% central credible interval from an exact CDF
Back to the four listeners who all accepted, with the flat prior. The posterior is $\mathrm{Beta}(5,1)$ with density $5\theta^4$. Find the central 90% credible interval for the acceptance rate.
Find$[\theta_l,\theta_u]$
Given
posterior density $5\theta^4$ on $[0,1]$
$\alpha=0.10$, so $0.05$ in each tail
Solution
The CDF of this posterior is a single power, so each end point is one root and no approximation is needed.
Write the CDF
$$F(\theta)=\int_0^{\theta}5t^4\,dt=\theta^{5}$$
The CDF is the area to the left of $\theta$, and a central interval is defined through it.
$F(0.990)-F(0.549)=0.95-0.05=0.90$. The interval also contains the posterior mean $5/6\approx 0.833$, a quick plausibility check.
Given these four listeners, the rate lies between $0.55$ and $0.99$ with probability $0.9$. That is the sentence the dashboard should have printed.
Two hundred users: an approximate 95% credible interval
With a flat prior, $46$ of $200$ users click a new button, so the posterior is $\mathrm{Beta}(47,155)$. Find an approximate central 95% credible interval from the posterior mean and standard deviation.
FindAn approximate $[\theta_l,\theta_u]$.
Given
posterior $\mathrm{Beta}(47,155)$
$\alpha=0.05$
Solution
With 200 observations the posterior is close to a normal curve, and mean $\pm$ 2 sd takes three lines; the exact Beta quantiles need a calculator or software.
A numerical calculation of the exact $\mathrm{Beta}(47,155)$ quantiles gives $[0.177,\,0.293]$, so the shortcut is off by about $0.004$ at the left end. The interval also brackets the data fraction $46/200=0.23$.
For a few hundred observations and a rate away from $0$ and $1$, mean $\pm$ 2 sd is a safe shortcut; with a handful of observations, go back to the exact CDF.
Checkpoint
§02.6 — one-sided credible bound
A new classifier is tested on $8$ images and makes no mistakes. With a flat prior on its error rate, the posterior is $\mathrm{Beta}(1,9)$.
Find(a) What is $\theta_u$ in the one-sided 95% credible interval $[0,\theta_u]$?
Given
posterior $\mathrm{Beta}(1,9)$
its CDF: $F(\theta)=1-(1-\theta)^{9}$
Hint 1/4
We need the point that cuts off the top tail of the posterior, leaving the stated probability to its left.
Hint 2/4
Solve $F(\theta_u)=0.95$.
Hint 3/4
$1-(1-\theta_u)^9=0.95$ with $F(\theta)=1-(1-\theta)^9$, so $(1-\theta_u)^9=0.05$.
Hint 4/4
$\theta_u=1-0.05^{1/9}\approx 0.283$.
Show solution
A one-sided interval starting at $0$ is the natural report here because the posterior peaks at $0$.
Take the fraction of successes and step two standard errors to each side. The 95% belongs to this recipe: over many repetitions of the experiment, about 95 of every 100 intervals it produces contain $\theta_{\mathrm{true}}$.
From the central limit theorem to the interval
Each $X_i$ has mean $\theta_{\mathrm{true}}$ and variance $\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})$, so $S_n$ has mean $n\theta_{\mathrm{true}}$ and variance $n\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})$.
By the central limit theorem, for large $n$ $$\frac{S_n-n\theta_{\mathrm{true}}}{\sqrt{n\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})}}\approx\mathcal N(0,1),\quad\text{so}\quad \bar\theta\approx\mathcal N\Big(\theta_{\mathrm{true}},\,\frac{\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})}{n}\Big).$$
A standard normal lies within $\pm 2$ with probability $0.954$. With $\sigma_n=\sqrt{\theta_{\mathrm{true}}(1-\theta_{\mathrm{true}})/n}$ this reads $$\Pr\big(\vert\bar\theta-\theta_{\mathrm{true}}\vert\le 2\sigma_n\big)\approx 0.95.$$
The event $\vert\bar\theta-\theta_{\mathrm{true}}\vert\le 2\sigma_n$ is the same as $\bar\theta-2\sigma_n\le\theta_{\mathrm{true}}\le\bar\theta+2\sigma_n$. Last, $\sigma_n$ contains the unknown $\theta_{\mathrm{true}}$; for large $n$ the law of large numbers puts $\bar\theta$ close to it, so we plug in $\bar\theta$.
Twenty simulated experiments, each $100$ flips of a coin with $\textcolor{#8250df}{\theta_{\mathrm{true}}=0.3}$. Each bar is one interval $\textcolor{#d1690a}{\bar\theta\pm 2\,\mathrm{SE}}$. In this run $19$ bars cross the dashed line and one, centred at $0.19$, misses. Nothing in that interval, taken alone, tells you it is the miss.
Looks like this, but is not
Our data gave the 95% confidence interval $[0.18,\,0.26]$, so it seems there is probability $0.95$ that $\theta_{\mathrm{true}}$ lies between $0.18$ and $0.26$.
In the frequentist model $\theta_{\mathrm{true}}$ is a fixed number: it is in $[0.18,\,0.26]$ or it is not, and no probability is left to assign. The $0.95$ describes the recipe over repeated samples. A probability about this one interval needs a posterior, which is what a credible interval supplies.
n
θtrue = 0.5
θtrue = 0.1
$10$
$89$
$65$
$30$
$96$
$81$
$100$
$94$
$93$
$400$
$95$
$95$
With $10$ samples and a rate of $0.1$, only about $65$ intervals in $100$ catch the truth, far from the promised $95$. The recipe earns its name once both the success count and the failure count are comfortably large.
Click-through rate: 88 clicks in 400 views
A recommendation widget is shown $400$ times and clicked $88$ times. Treat the views as i.i.d. $\mathrm{Ber}(\theta_{\mathrm{true}})$ and give a 95% confidence interval for the click-through rate.
FindThe interval $\bar\theta\pm 2\,\mathrm{SE}$.
Given$n=400$ views, $S_n=88$ clicks
Solution
With $88$ clicks and $312$ non-clicks both counts are large, so the central-limit recipe is safe and needs no prior.
Point estimate
$$\bar\theta=\frac{88}{400}=0.22$$
The mean of the 0/1 outcomes is the click fraction.
Plug the estimate into the Bernoulli variance $\theta(1-\theta)/n$.
Two standard errors each side
$$0.22\pm 2(0.0207)=0.22\pm 0.0414$$
The central limit theorem puts about $0.95$ of the sampling distribution within two standard errors.
$$[0.179,\;0.261]$$
The two ends, rounded to three decimals.
Answer $$\boxed{[0.179,\;0.261]}$$
Check
Scale check: the widest possible half-width at $n=400$ is $2\sqrt{0.25/400}=0.05$, and ours, $0.041$, is below it because $0.22$ is away from $0.5$.
Say it this way: the recipe that produced $[0.179,\,0.261]$ catches the true rate about 95 times in 100. Do not say the true rate is in this interval with probability $0.95$.
Twenty honest experiments: how many intervals miss?
A lab repeats the same $100$-flip experiment $20$ times and builds a 95% interval each time, as in the figure. Assume each interval covers $\theta_{\mathrm{true}}$ with probability $0.95$, independently of the others. How many misses should they expect, and how likely is at least one?
FindThe expected number of misses and $\Pr(\text{at least one miss})$.
Given
$K=20$ independent intervals
$0.95$ each
Solution
The number of misses counts independent yes/no events, so the binomial model from the previous section answers both parts.
Expected misses
$$E[\text{misses}]=20\times 0.05=1$$
A binomial count has mean $n$ times $p$.
At least one miss
$$\Pr(\text{no miss})=0.95^{20}\approx 0.358$$
All twenty independent intervals must cover.
$$\Pr(\text{at least one miss})=1-0.358=0.642$$
At least one miss is the opposite event of no miss.
Order of magnitude: with mean $1$ and small miss probability, the miss count is close to Poisson$(1)$, which gives $\Pr(\ge 1)\approx 1-e^{-1}\approx 0.632$, close to $0.642$. Exactly one miss has probability $20(0.05)(0.95)^{19}\approx 0.377$, the most likely count, so the figure's single miss is typical.
A 95% recipe guarantees misses in the long run; the guarantee is about how often they happen, never about which interval is one.
Checkpoint
§02.7 — 95% interval for a proportion
An exit poll asks $100$ voters a yes/no question, and exactly half say yes.
Find(a) Which is the 95% confidence interval $\bar\theta\pm 2\,\mathrm{SE}$?
Given
$n=100$
$\bar\theta=0.5$
Hint 1/4
Build the standard error first; the interval is the estimate plus or minus two of them.
Hint 2/4
$\mathrm{SE}=\sqrt{\bar\theta(1-\bar\theta)/n}$ and the interval is $\bar\theta\pm 2\,\mathrm{SE}$.
Hint 3/4
$\bar\theta=0.5$ and $n=100$ give $\mathrm{SE}=\sqrt{0.25/100}=0.05$.
Hint 4/4
The interval is $0.5\pm 0.1=[0.40,\,0.60]$.
Show solution
At $\bar\theta=0.5$ the standard error is as large as it gets, which makes the arithmetic clean.
Standard error
$$\sqrt{\frac{0.5\times 0.5}{100}}=0.05$$
Bernoulli variance over $n$, square-rooted.
Interval
$$0.5\pm 2(0.05)=[0.40,\;0.60]$$
Two standard errors, because a normal curve keeps about $0.95$ of its area within two standard deviations.
Answer $$\boxed{[0.40,\;0.60]}$$
Check
With $n=100$ the half-width $0.1$ equals $1/\sqrt{n}$, the known worst case for the two-standard-error recipe.
A poll of $100$ pins a proportion down to about $\pm 0.1$; for $\pm 0.05$ you need four times as many people.
⚠ Giving one interval a probability it does not have
It is below the mean, as it must be for a posterior with a long right tail.
Mean of Beta(3, 12)
The same posterior $\mathrm{Beta}(3,12)$. Find the posterior mean.
Find$E[\Theta\mid D]$
Given$\alpha=3$, $\beta=12$
Solution
No subtractions: the mean uses the parameters as they are.
Mean
$$\frac{3}{3+12}=0.2$$
A Beta's mean uses the parameters as they are, with no subtraction.
Answer $$\boxed{E[\Theta\mid D]=0.2}$$
Check
The long right tail pulls the balance point above the peak.
One posterior, two summaries: the peak at $0.154$ and the balance point at $0.2$.
How to tell them apart
A $-1$ on top and a $-2$ below mean the peak (MAP); no subtractions mean the average (posterior mean).
Scaffolding comes off
The common skeleton
Count: $N_1$ successes and $N_0$ failures, pooled over every batch.
Update: the posterior is $\mathrm{Beta}(N_1+a,\,N_0+b)$.
Summarize: the MAP $\frac{\alpha-1}{\alpha+\beta-2}$ or the mean $\frac{\alpha}{\alpha+\beta}$, whichever is asked.
Check: the estimate lies between the prior's value and the MLE $N_1/N$.
1 · fully worked
Defect rate: a Beta(2, 8) prior, then 5 defects in 30 items
A factory's prior for a new line's defect rate is $\mathrm{Beta}(2,8)$. A sample of $30$ items has $5$ defects. Find the MAP estimate and the posterior mean, and check each against the prior and the MLE.
Find$\hat\theta_{\mathrm{MAP}}$ and $E[\Theta\mid D]$.
Given
prior $\mathrm{Beta}(2,8)$
$N_1=5$ defects, $N_0=25$ good items
Solution
All four skeleton steps in order; the check at the end is what catches a count fed into the wrong parameter.
Count
$$N_1=5,\qquad N_0=30-5=25$$
Defects are the successes here, since they are what we count.
Both estimates are weighted averages of a prior summary and the MLE $5/30$: $\tfrac{8}{38}\cdot\tfrac18+\tfrac{30}{38}\cdot\tfrac16=\tfrac{6}{38}\approx 0.158$ and $\tfrac{10}{40}(0.2)+\tfrac{30}{40}\cdot\tfrac16=0.175$, the same two numbers.
The check step is not decoration: an estimate outside its bracket means a count went into the wrong place.
2 · you write the reasoning
An easier one, and this time you supply the reasons. Flat prior $\mathrm{Beta}(1,1)$, then $3$ successes and $1$ failure. Find the posterior and the MAP, and for each line write why it is allowed.
$N_1=3,\qquad N_0=1$
reasoning
Successes are the ones: three of them, and one failure.
The flat prior adds no pseudo-counts to the peak, so the MAP must equal the fraction of successes; this line is the check.
3 · find the buried error
Harder, and the work is done for you, with two errors buried in it. Prior $\mathrm{Beta}(4,2)$; data: $6$ successes and $14$ failures. Find the MAP estimate and the posterior mean.
Step 1.$N_1=6,\ \allowbreak N_0=14,\ \allowbreak a=4,\ \allowbreak b=2$. Read off the counts and the prior.
Both weighted averages reproduce the direct values: $\tfrac{6}{66}\cdot\tfrac13+\tfrac{60}{66}(0.3)=\tfrac{20}{66}$ and $\tfrac{8}{68}(0.375)+\tfrac{60}{68}(0.3)=\tfrac{21}{68}$.
To compare MAP and mean, compare their prior weights and their prior values; the data part is the same MLE in both.
Full exam-style question
Exam-style: packet loss on a sensor link, both waysexam format
A sensor node sends $n=200$ packets and $14$ are lost. Assume losses are i.i.d. $\mathrm{Ber}(\theta)$.
(a) Give the MLE of the loss rate and an approximate 95% confidence interval.
(b) The vendor's datasheet suggests the prior $\mathrm{Beta}(1,19)$. Find the posterior, the MAP estimate and the posterior mean.
(c) Give an approximate central 95% credible interval, and say in one sentence each what the two intervals mean.
FindMLE and confidence interval; posterior, MAP and mean; credible interval; the two readings.
Given
$n=200$, $S_n=14$ lost packets
prior $\mathrm{Beta}(1,19)$ for parts (b) and (c)
Solution
Part (a) needs no prior, so we do it first with the central-limit recipe; parts (b) and (c) reuse the same counts with the conjugate update, and the normal shortcut keeps the credible interval a hand calculation.
The Bayesian estimates sit below $0.07$ because the prior mean is $1/20=0.05$. A numerical calculation of the exact $\mathrm{Beta}(15,205)$ quantiles gives about $[0.039,\,0.105]$, so the normal shortcut is rough at the left end, where this posterior is skewed.
Three tools on one set of counts: the MLE, the conjugate update and the central-limit recipe.
In an exam, name which interval you computed and give its reading in one sentence; the numbers alone do not show that you know the difference.
Practice
A · concept 4 questions
1§02.5 — a majority of heads and the posterior mean
A study partner offers a shortcut for Beta posteriors. You want to test it on a small case before trusting it.
Find(a) Is the claim true or false? Use the test case.
Given
Claim: if $N_1>N_0$, then $E[\Theta\mid D]>0.5$, whatever the Beta prior.
Test case: prior $\mathrm{Beta}(2,8)$, data $3$ heads and $2$ tails.
Hint 1/4
The claim ignores the prior, so check whether a strong prior can outvote a small majority of heads.
Hint 2/4
$E[\Theta\mid D]=\dfrac{a+N_1}{a+b+N}$.
Hint 3/4
With $a=2$, $b=8$, $N_1=3$ and $N=5$ this is $\dfrac{2+3}{2+8+5}$.
Hint 4/4
The posterior mean is $5/15\approx 0.333<0.5$, so the claim is false.
Show solution
A claim with 'whatever the prior' falls to one counterexample, so we compute the test case directly.
Update
$$\mathrm{Beta}(2+3,\;8+2)=\mathrm{Beta}(5,10)$$
The prior stays Beta: three heads join $2$, two tails join $8$.
Mean
$$E[\Theta\mid D]=\frac{5}{15}\approx 0.333$$
The posterior mean is the first parameter over the sum.
Verdict
$$0.333<0.5\ \Rightarrow\ \text{false}$$
Heads are the majority in the data, yet the estimate stays below one half.
A head adds one to the first parameter only, so now $x=2$ and $y=10$.
Compare
$$0.01068>0.00689$$
The posterior got wider, so the claim fails.
Answer $$\boxed{\text{False}}$$
Check
The formula does shrink in the ordinary case: $\mathrm{Beta}(1,1)$ has variance $1/12\approx 0.083$ and $\mathrm{Beta}(2,1)$ has $1/18\approx 0.056$. Here the mean jumps from $1/11\approx 0.091$ to $1/6\approx 0.167$, and that surprise is what widens the posterior.
A surprising observation can widen the posterior; only on average does more data narrow it.
3§02.7 — what halves a confidence interval
A product team has a 95% confidence interval for a conversion rate and finds it too wide. They list four changes for the next study.
Find(a) Which change halves the width of the interval?
Given
interval: $\bar\theta\pm 2\sqrt{\bar\theta(1-\bar\theta)/n}$ from $n$ visitors
assume $\bar\theta$ comes out about the same next time
Hint 1/4
Look at how the width depends on $n$ when everything else stays fixed.
Hint 2/4
Width $=4\sqrt{\bar\theta(1-\bar\theta)/n}$, which is proportional to $1/\sqrt n$.
Hint 3/4
With $\bar\theta$ unchanged, halving $1/\sqrt n$ needs $\sqrt{n'}=2\sqrt n$, that is $n'=4n$.
Hint 4/4
Four times as many visitors halves the width.
Show solution
Only $n$ can shrink the width here; the other options change the method or the level.
Set up the ratio
$$\frac{\sqrt{n}}{\sqrt{n'}}=\frac12$$
New width over old width, with $\bar\theta$ unchanged.
Solve
$$n'=4n$$
Squaring removes the roots: the ratio of sample sizes is the square of the ratio of widths.
Answer $$\boxed{n'=4n}$$
Check
Doubling alone gives a ratio of $1/\sqrt2\approx 0.71$, not $0.5$.
Precision is expensive: each halving of the width costs four times the data.
4§02.2 — updating in pieces or all at once
Two analysts start from the same prior $\mathrm{Beta}(2,2)$ for a click rate. One updates after every single view during the day; the other waits and updates once on the day's totals.
Find(a) True or false: they end the day with the same posterior.
Given
the day's data: $6$ clicks and $9$ skips
both use the conjugate Beta update
Hint 1/4
Think of the update as bookkeeping: what gets added to which parameter, and does the order matter?
Hint 2/4
Each click adds $1$ to the first parameter and each skip adds $1$ to the second.
Hint 3/4
Over the day, $6$ ones are added to the first parameter and $9$ to the second, in whatever order the views came.
Hint 4/4
Both end at $\mathrm{Beta}(8,11)$, so the statement is true.
Show solution
The conjugate update is pure addition, and addition is order-free.
Batch
$$\mathrm{Beta}(2+6,\;2+9)=\mathrm{Beta}(8,11)$$
The batch analyst adds the day's totals in one step.
Mirror the problem: flipping $\theta\to 1-\theta$ swaps the two counts, so $\Pr(\Theta>0.5\mid D)$ equals the area of $12\theta(1-\theta)^2$ below $0.5$: $\int_0^{0.5}12\theta(1-\theta)^2\,d\theta=12\big(\tfrac18-\tfrac1{12}+\tfrac1{64}\big)=\tfrac{11}{16}$.
Two finishes in three make a completion rate above one half likely, about $11$ chances in $16$, but not safe.
4§02.6 — credible bounds from a closed-form CDF
A new sensor is tested $9$ times and passes every test. With a flat prior on its pass rate, the posterior is $\mathrm{Beta}(10,1)$.
Find
(a) Find the central 90% credible interval.
(b) Find $\theta_l$ with $\Pr(\Theta\ge\theta_l\mid D)=0.95$.
Givenposterior CDF $F(\theta)=\theta^{10}$ on $[0,1]$
Hint 1/4
Every end point here is a point where the CDF takes a known value.
Hint 2/4
Central 90%: $F(\theta_l)=0.05$ and $F(\theta_u)=0.95$. One-sided: $F(\theta_l)=0.05$ as well.
Hint 3/4
With $F(\theta)=\theta^{10}$: $\theta_l=0.05^{1/10}$ and $\theta_u=0.95^{1/10}$.
With $n=600$ the sample fraction is close to normal, so two standard errors give about $0.95$ coverage.
Answer $$\boxed{[0.215,\;0.285]}$$
Check
The worst-case half-width at $n=600$ is $2\sqrt{0.25/600}\approx 0.041$, and ours, $0.035$, is below it.
State the result with its reading: the recipe covers the true proportion about 95 times in 100.
6§02.7 — planning a sample size
Before running a poll, you want the 95% interval $\bar\theta\pm 2\,\mathrm{SE}$ to have half-width at most $0.02$, whatever the result turns out to be.
Find(a) What is the smallest $n$ that guarantees this?
Both sides are positive, so taking reciprocals flips the inequality and squaring keeps it.
Answer $$\boxed{n=2500}$$
Check
If the poll comes out at $\bar\theta=0.3$ instead, the half-width is $2\sqrt{0.21/2500}\approx 0.018$, inside the target, as a worst-case design should guarantee.
Design for $\bar\theta=1/2$ when nothing is known; any other outcome only makes the interval narrower.
7§02.5 — the posterior mean as a weighted average
A marketing team's prior for a campaign's response rate is $\mathrm{Beta}(6,14)$. The campaign then gets $35$ responses from $50$ contacts.
Find
(a) Find the posterior mean.
(b) Write it as $w\cdot(\text{prior mean})+(1-w)\cdot(\text{MLE})$ and find $w$.
Given
prior $\mathrm{Beta}(6,14)$
$N_1=35$, $N_0=15$
Hint 1/4
Part (a) is one formula; part (b) splits the same fraction into a prior piece and a data piece.
Hint 2/4
$E[\Theta\mid D]=\dfrac{a+N_1}{a+b+N}$ and $w=\dfrac{a+b}{a+b+N}$.
Hint 3/4
$a=6$, $b=14$, $N_1=35$, $N=50$; prior mean $6/20=0.3$, MLE $35/50=0.7$.
Hint 4/4
The mean is $41/70\approx 0.586$, with $w=20/70=2/7$.
Show solution
The direct formula gives the number; the weights explain it.
A numerical calculation of the exact quantiles gives $[0.325,\,0.480]$, within $0.003$ of the shortcut.
Check the shortcut's two conditions before using it: many observations, and a mean well away from $0$ and $1$.
C · exam level 4 questions
1§02.2 — two days of updates, then the mean
A click rate starts from the prior $\mathrm{Beta}(2,2)$. On day one, $3$ of $8$ views are clicks; on day two, $6$ of $8$.
Find(a) What is the posterior mean after both days?
Given
prior $\mathrm{Beta}(2,2)$
day 1: $3$ clicks, $5$ skips
day 2: $6$ clicks, $2$ skips
Hint 1/4
Both days update the same belief; the order and the split into days do not matter.
Hint 2/4
Posterior $\mathrm{Beta}(a+N_1,\,b+N_0)$ with the pooled counts, and mean $\alpha/(\alpha+\beta)$.
Hint 3/4
Pooled: $N_1=3+6=9$ and $N_0=5+2=7$, on top of $a=b=2$.
Hint 4/4
$\mathrm{Beta}(11,9)$, mean $11/20=0.550$.
Show solution
Pooling is legal because day one's posterior is day two's prior; it saves one update.
Pool
$$N_1=9,\qquad N_0=7$$
Clicks and skips over both days.
Update
$$\mathrm{Beta}(2+9,\;2+7)=\mathrm{Beta}(11,9)$$
Pooled clicks join the first parameter and pooled skips the second.
Mean
$$\frac{11}{20}=0.550$$
The question asks for the mean: first parameter over the sum.
Answer $$\boxed{E[\Theta\mid D]=0.550}$$
Check
Day by day: $\mathrm{Beta}(2,2)\to\mathrm{Beta}(5,7)\to\mathrm{Beta}(11,9)$, the same end point.
When data come in batches, add them all up before thinking about summaries.
2§02.4 — naming a rate when only an exact hit counts
A fraud team knows a merchant's chargeback rate is $0.2$, $0.4$ or $0.6$, with prior probabilities $0.5$, $0.35$ and $0.15$. In $5$ audited orders, $3$ are chargebacks. A bonus is paid only if the team names the rate exactly.
Find(a) Which rate should the team name, and how likely is it to be wrong?
Given
candidates $0.2,\ 0.4,\ 0.6$ with prior $0.5,\ 0.35,\ 0.15$
$3$ chargebacks in $5$ orders
Hint 1/4
Only an exact hit pays, so we want the value with the highest posterior probability, and one minus that probability.
Hint 2/4
$\Pr(\Theta=\theta\mid D)\propto\Pr(\Theta=\theta)\,\theta^{3}(1-\theta)^{2}$, normalized over the three candidates.
Hint 3/4
Likelihoods $0.00512,\ 0.02304,\ 0.03456$ for priors $0.5,\ 0.35,\ 0.15$ give weights $0.00256,\ 0.008064,\ 0.005184$.
Hint 4/4
Name $0.4$: its posterior probability is $0.510$, so it is wrong with probability about $0.49$.
Show solution
A three-row table gives the whole posterior; the binomial coefficient is the same in every row and drops out.
Two standard errors, not one, for about $0.95$ coverage.
Answer $$\boxed{[0.121,\;0.179]}$$
Check
Half-width $0.029$ is below the worst case $2\sqrt{0.25/600}\approx 0.041$, as it should be for a rate of $0.15$.
In an A/B report, give each version's interval with the same recipe, then compare them.
4§02.6 — find the error in a credible interval
A student computes a central 90% credible interval for the posterior $\mathrm{Beta}(1,7)$, which came from $0$ successes in $6$ trials under a flat prior. The work is below.
Find(a) Which step contains the error?
Given
Step 1: $F(\theta)=1-(1-\theta)^{7}$.
Step 2: $F(\theta_l)=0.10$, so $\theta_l=1-0.9^{1/7}\approx 0.015$.
Step 3: $F(\theta_u)=0.95$, so $\theta_u=1-0.05^{1/7}\approx 0.348$.
Hint 1/4
Check each step against the definition of a central interval: how much probability belongs in each tail?
Hint 2/4
Central $1-\alpha$ interval: $F(\theta_l)=\alpha/2$ and $F(\theta_u)=1-\alpha/2$.
Hint 3/4
Here $\alpha=0.10$, so the steps should read $F(\theta_l)=0.05$ and $F(\theta_u)=0.95$, with $F(\theta)=1-(1-\theta)^7$.
Hint 4/4
Step 2 is wrong: $\theta_l=1-0.95^{1/7}\approx 0.0073$, and the interval is $[0.0073,\,0.348]$.
Show solution
Checking each step against $F(\theta_l)=\alpha/2$ and $F(\theta_u)=1-\alpha/2$ is quicker than redoing the whole problem.
Put $\varepsilon$ back into the bound: $2e^{-1000\times 0.0607^2}=2e^{-3.68}\approx 0.050$, the level we asked for.
The central-limit interval is the everyday tool; the inequality is what you quote when $n$ is small or a guarantee must hold exactly.
2§02.4 — chips from two machines
Chips come from machine A, with defect rate $0.05$, or machine B, with defect rate $0.2$, and a box is equally likely to come from either. You test $10$ chips from one box and find $2$ defective.
Find
(a) Find the probability that the box came from machine B.
(b) You must name one machine and are paid only if you are right: which do you name, and how often will you be wrong?
Given
$\Pr(A)=\Pr(B)=0.5$
defect rates $0.05$ for A and $0.2$ for B
$2$ defective out of $10$, chips independent
Hint 1/4
The unknown is which machine, so this is Bayes rule with two hypotheses and a binomial likelihood.
Hint 2/4
$\Pr(B\mid D)=\dfrac{\Pr(D\mid B)\Pr(B)}{\Pr(D\mid A)\Pr(A)+\Pr(D\mid B)\Pr(B)}$ with $\Pr(D\mid\theta)=\binom{10}{2}\theta^2(1-\theta)^8$.
Hint 3/4
$\Pr(D\mid A)=45(0.05)^2(0.95)^8\approx 0.0746$ and $\Pr(D\mid B)=45(0.2)^2(0.8)^8\approx 0.3020$, each with prior $0.5$.
Hint 4/4
$\Pr(B\mid D)\approx 0.802$; the MAP guess is B, wrong with probability about $0.198$.
Show solution
Equal priors cancel, so the answer is the ratio of the two likelihoods, normalized.
Likelihoods
$$\Pr(D\mid A)=45(0.0025)(0.6634)\approx 0.0746$$
$\binom{10}{2}=45$, two defects and eight good chips.
Two defects in ten is a fraction of $0.2$, exactly machine B's rate, so a strong lean toward B is expected; it is not certain because A produces this sample about $7$ times in $100$.
Choosing between two machines is the MAP rule on a two-point parameter: compute the posterior, name the larger.
3§02.7 — the spread of a sample mean
Let $X_1,\dots,X_n$ be i.i.d. $\mathrm{Ber}(\theta)$ and let $\bar\theta=(X_1+\cdots+X_n)/n$.
Find
(a) Show that $\mathrm{Var}(\bar\theta)=\theta(1-\theta)/n$.
(b) Compute the standard deviation of $\bar\theta$ for $\theta=0.2$ and $n=64$.
(c) By what factor must $n$ grow to halve that standard deviation?
Given
$E[X_i]=\theta$ and $\mathrm{Var}(X_i)=\theta(1-\theta)$
for independent variables the variances add; $\mathrm{Var}(cY)=c^2\mathrm{Var}(Y)$
$\theta=0.2$ and $n=64$ for part (b)
Hint 1/4
Work on the sum first, then divide by $n$.
Hint 2/4
$\mathrm{Var}(X_1+\cdots+X_n)=\sum\mathrm{Var}(X_i)$ for independent terms, and $\mathrm{Var}(S/n)=\mathrm{Var}(S)/n^2$.
Hint 3/4
With $\mathrm{Var}(X_i)=\theta(1-\theta)$, $\theta=0.2$ and $n=64$: $\sqrt{0.2\times 0.8/64}$.
Hint 4/4
(a) $\theta(1-\theta)/n$; (b) $0.05$; (c) four times as large.
Show solution
Independence lets the variance of the sum split into $n$ equal pieces; the division by $n$ then comes out squared.
Close the page and write from memory: the posterior as likelihood times prior; the Beta update; the MLE, MAP and posterior-mean formulas and the loss each one wins under; how to get a central credible interval; the confidence-interval recipe and the one sentence that reads it correctly. Then compare with the formula card.
Write $p(\theta\mid D)$ up to a constant for any prior, and normalize it when the prior is flat?
c-likelihood-posterior
Update $\mathrm{Beta}(a,b)$ on $N_1$ heads and $N_0$ tails without mixing up the parameters?
c-beta-prior
Derive $N_1/N$ from the log-likelihood and handle the all-heads case at the endpoint?
c-mle
Compute a MAP estimate and say which loss makes it the best guess?
c-map
Compute a posterior mean and split it into prior and data weights?
c-posterior-mean
Find a central or one-sided credible interval from a CDF, and know when mean $\pm$ 2 sd is safe?
c-credible
Build $\bar\theta\pm 2\,\mathrm{SE}$ and say in one sentence what its 95% means?
c-confidence
Glossary (15 terms)
likelihoodolabilirlik
The probability of the observed data viewed as a function of the parameter, with the data held fixed.
önsel dağılım
The distribution describing what is believed about a parameter before the data are seen.
sonsal dağılım
The distribution of a parameter after the data are seen, proportional to likelihood times prior.
eşlenik önsel
A prior for which the posterior stays in the same family; the Beta prior is conjugate to Bernoulli data.
Beta distributionBeta dağılımı
A distribution on $[0,1]$ with density proportional to $\theta^{a-1}(1-\theta)^{b-1}$.
maximum likelihood estimateen çok olabilirlik kestirimi
The parameter value that makes the observed data most probable.
MAP estimate
The maximum a posteriori estimate: the value at which the posterior density is highest.
posterior meansonsal ortalama
The expected value of the parameter under the posterior; the best single estimate under squared loss.
point estimatenokta kestirimi
A single number reported for a parameter, with no statement of uncertainty.
credible interval
An interval that contains the parameter with a stated posterior probability, given the observed data.
confidence intervalgüven aralığı
An interval from a recipe that covers the fixed true parameter with a stated probability over repeated samples.
standard errorstandart hata
The estimated standard deviation of an estimator; for a sample proportion it is $\sqrt{\bar\theta(1-\bar\theta)/n}$.
kayıp fonksiyonu
A rule that assigns a cost to reporting one value when the parameter takes another.
pseudo-count
A prior parameter read as a number of imaginary earlier observations.
coverage probabilitykapsama olasılığı
The probability, over repeated samples, that an interval recipe contains the true parameter.
What comes next
§03 · Linear regression and ordinary least squares
Here the unknown was a single rate and the data were yes/no counts. Next, the unknown becomes a set of coefficients that turn inputs into a real-valued prediction, and fitting them is the first full learning problem of the course.
Sources
textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The syllabus line for this week names no chapter of the book, so no section numbers are cited here.
course materialEEE 485 lecture slides and lecture notes, Chapter 2: Bayesian and frequentist machine learning Scope, order of topics and notation follow these materials; the explanations, examples and exercises here are original.
course materialEEE 485 syllabus, Fall 2026-27 Weekly topic line and assessment weights.
standard resultBeta and Bernoulli conjugacy; the two-standard-error interval for a proportion; the normal approximation to a Beta posterior Standard results. The exact coverage table, the exact Beta quantiles and the twenty simulated intervals were computed for this page.