6 concepts19 worked examples31 exercises5 exam-level6 figures
What are you here for?
06 Linear regression from a Bayesian perspective: least squares, ridge and lasso as most probable coefficients, and the full posterior
Start with this
One question before you read anything. Getting it wrong is the point: it shows you what this section is for.
§06.1 — a noisier sensor and the most likely slope
A slope is fitted to $50$ readings by maximizing the likelihood under Gaussian noise with known standard deviation $1$. The same readings are then analysed again, assuming the noise standard deviation is $2$.
Find(a) What happens to the most likely slope?
Given
the same $50$ readings both times
noise model $\mathcal N(0,\sigma^2)$: first $\sigma=1$, then $\sigma=2$
Hint 1/4
Ask whether $\sigma$ changes where the peaks, or only its height and width.
Going from $\sigma=1$ to $\sigma=2$ changes the constant and multiplies the RSS term by $\frac14$; the $50$ readings, and so the RSS curve, stay the same.
Hint 4/4
The minimizer of the RSS does not change, so the most likely slope stays the same.
Show solution
Only where the two log-likelihoods peak answers the question, not their values, so we compare locations.
Answer $$\boxed{\text{the same slope, }\hat\beta^{\mathrm{LS}}}$$
Check
Edge case: as $\sigma\to\infty$ the log-likelihood flattens, but at every finite $\sigma$ its peak stays at the least squares slope.
Only a prior makes the noise level matter; that is what the ridge block adds.
Two groups fit a slope to the same three readings, with noise standard deviation 1. One minimizes squared misses plus 4 times the squared slope; the other has no penalty, only a belief that the slope is Gaussian around 0 with standard deviation 0.5, and reports the most probable slope. Both report 0.6; on any data they would agree.
By the end you can turn such a belief into a penalty and back, for ridge and for the lasso, and go past a single number: a full posterior for the coefficients and an error bar for every prediction.
In 60 seconds
Least squares, ridge and the lasso are the most probable coefficients under three models (Gaussian noise alone, plus a Gaussian prior, plus a Laplace prior); Bayesian linear regression keeps the whole Gaussian posterior, which gives every prediction an error bar.
a prediction with an honest error bar at a new input
Three most common mistakes
Writing $\lambda=1/(\text{prior variance})$: the noise variance belongs in the translation, $\lambda=\sigma^2/(\text{prior variance})$ for ridge and $b=2\sigma^2/\lambda$ for the Laplace scale.
Adding covariances instead of precisions: the posterior precision is $\Sigma_0^{-1}+X^T\Sigma_D^{-1}X$, and the posterior covariance is its inverse.
Dropping $\sigma^2$ from the predictive variance: a new reading carries its own noise, so the variance is $\sigma^2+x^T\Sigma_{\beta|D}x$ and never below $\sigma^2$.
Two course documents give different weights:
Chapter 1 slides, undergraduate line: midterm 25, final 25, four quizzes 20, two-phase project 30.
STARS syllabus page printed on 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. It is the later document; confirm which split applies.
The STARS weekly list puts ridge, lasso and Bayesian linear regression in one line; the lecture slides split them over this chapter and the one before.
How much time do you have?
10 minutes
The penalty-to-prior translation, then the two formulas you compute with: the Gaussian posterior and the predictive distribution.
The 60-second card · Ridge regression is the most probable β under a Gaussian prior · Bayesian linear regression · The posterior predictive distribution · Formula card
45 minutes
Every result once with a worked example, then the posterior-and-prediction calculation from a full solution down to a bare problem.
The 60-second card · Least squares is maximum likelihood when the noise is Gaussian · Ridge regression is the most probable β under a Gaussian prior · The lasso is the most probable β under a Laplace prior · Bayesian linear regression · The posterior sharpens reading by reading · The posterior predictive distribution · Scaffolding comes off · Formula card
full read
Adds the derivations, the look-alike pairs and enough mixed practice to decide the method yourself.
The opening pages · Recall first · Least squares is maximum likelihood when the noise is Gaussian · Ridge regression is the most probable β under a Gaussian prior · The lasso is the most probable β under a Laplace prior · Bayesian linear regression · The posterior sharpens reading by reading · The posterior predictive distribution · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
Derive the log-likelihood of a linear model with Gaussian noise of known variance, and show that its maximizer is the least squares estimate whatever $\sigma^2$ is.
Show that ridge regression is the MAP estimate under an i.i.d. Gaussian prior, and convert between $\lambda$ and the prior variance with $\lambda=\sigma^2/(\text{prior variance})$.
Show that the lasso is the MAP estimate under an i.i.d. Laplace prior with scale $2\sigma^2/\lambda$, and explain why the corner of that prior can set a coefficient exactly to $0$.
Compute the Gaussian posterior mean and covariance from the general formula, for the ridge prior, for a prior mean away from zero and for unequal noise levels.
Update a posterior one reading at a time and show that the result equals the posterior from all readings at once.
Compute the posterior predictive mean and variance at a new input, and split the variance into noise and coefficient uncertainty.
Syllabus coverage
covered
Linear regression from a Bayesian perspective — covered
Least squares as maximum likelihood under Gaussian noise of known variance
the likelihood and the log-likelihood
why the logarithm
the likelihood as the Gaussian vector density of y given β
The lecturer's chapter title. The chapter opens by rereading least squares as maximum likelihood; this block does that before any prior is added.
covered
Ridge regression — covered
Ridge as the MAP estimate under an i.i.d. Gaussian prior with variance σ²/λ; λ as noise variance over prior variance; the ridge posterior.
The ridge estimator itself comes from the previous section; here it is read as the peak of a posterior, and its full posterior follows in the Bayesian linear regression block.
covered
lasso — covered
The ; the lasso as the MAP estimate under an i.i.d. Laplace prior with scale 2σ²/λ; why the corner of that prior gives exact zeros; ridge and lasso as , least squares as unbiased.
The comparison of the three estimators from the lecture sits in this block's table.
covered
Bayesian linear regression — covered
Point estimates against a whole posterior
Gaussian prior and Gaussian likelihood give a Gaussian posterior
the posterior for ridge
a prior mean away from zero
unequal noise levels
the posterior predictive distribution and its error bars
Spread over three blocks: the posterior, how it sharpens reading by reading, and the predictive distribution.
off syllabus
One-coefficient lasso by cases — off syllabus
The closed form for a single coefficient, found by minimizing on each side of zero and testing the corner.
Further reading: the lecture says the lasso has no closed form in general. The one-coefficient case needs only calculus and shows why the corner of the Laplace prior produces exact zeros.
off syllabus
Proof of the posterior formula — off syllabus
Completing the square in the exponent to show that the posterior is Gaussian with the stated mean and covariance.
Further reading: the lecture states the formula. The proof is included for completeness, and no exercise depends on reproducing it.
Recall first
Least squares
For centered data $D=\{(x_i,y_i)\}_{i=1}^n$, $\mathrm{RSS}(\beta)=\sum_i(y_i-\beta^Tx_i)^2=(y-X\beta)^T(y-X\beta)$ and $\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$. For one input, $\hat\beta^{\mathrm{LS}}=\sum_ix_iy_i/\sum_ix_i^2$.
Every estimate here is compared with it, and it is the maximum likelihood answer.
Ridge and lasso
$\hat\beta^{\mathrm{Ridge}}=\arg\min_\beta\,\mathrm{RSS}(\beta)+\lambda\sum_j\beta_j^2$, equal to $(X^TX+\lambda I)^{-1}X^Ty$. $\hat\beta^{\mathrm{Lasso}}=\arg\min_\beta\,\mathrm{RSS}(\beta)+\lambda\sum_j\lvert\beta_j\rvert$ has no closed form in general.
This section explains where the two penalties come from.
MLE and MAP
The maximum likelihood estimate maximizes $p(D\mid\theta)$. The MAP estimate maximizes $p(\theta\mid D)=p(D\mid\theta)\,p(\theta)/p(D)$, and $p(D)$ does not depend on $\theta$.
Least squares, ridge and the lasso are each one of these two.
Gaussian densities
$\mathcal N(\mu,\sigma^2)$ has density $\frac{1}{\sqrt{2\pi}\sigma}e^{-(z-\mu)^2/(2\sigma^2)}$. A Gaussian vector $Z\sim\mathcal N(\mu,\Sigma)$ in $p$ dimensions has density $$f_Z(z)=\frac{1}{(2\pi)^{p/2}\lvert\Sigma\rvert^{1/2}}\exp\Big(-\tfrac12(z-\mu)^T\Sigma^{-1}(z-\mu)\Big).$$
The likelihood and the ridge prior both have this form.
Linear combinations of Gaussians
If $Z\sim\mathcal N(\mu,\Sigma)$ then $a^TZ+c\sim\mathcal N(a^T\mu+c,\ a^T\Sigma a)$. A sum of independent Gaussians is Gaussian, with the means and the variances added.
They give the predictive distribution in one line.
Logs and maximizers
For $f>0$, $\arg\max f=\arg\max\log f=\arg\min(-\log f)$. Adding a constant, or multiplying by a positive constant, does not move a maximizer.
Every derivation here maximizes a log and drops constants.
2 × 2 inverse
$\begin{pmatrix}a&b\\c&d\end{pmatrix}^{-1}=\frac{1}{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}$ when $ad-bc\neq0$.
Every two-coefficient posterior on this page is found with it.
Try it yourself first (2 questions)
1§06.2 — a belief about a coefficient, written as a penalty
An engineer believes, before any data, that a coefficient is Gaussian with mean $0$ and variance $0.25$. The noise variance of the readings is $1$.
Find(a) If this belief were written as a penalty $\lambda\beta^2$ added to the RSS, what would $\lambda$ be?
Given
prior $\mathcal N(0,\,0.25)$
noise variance $\sigma^2=1$
Hint 1/4
Write the log of the prior and compare it with the RSS term of the log-likelihood.
Hint 2/4
$-\log p(\beta)=\frac{\beta^2}{2\tau^2}+\text{const}$ and $-\log p(D\mid\beta)=\frac{\mathrm{RSS}}{2\sigma^2}+\text{const}$; multiply both by $2\sigma^2$.
Hint 3/4
Here $\tau^2=0.25$ and $\sigma^2=1$.
Hint 4/4
$\lambda=\sigma^2/\tau^2=1/0.25=4$.
Show solution
Scaling by 2σ² puts the expression in the RSS-plus-penalty form, where the penalty can be read off.
Multiplying by $2\sigma^2=2$ turns the likelihood part into the plain RSS.
$$\lambda=\frac{1}{0.25}=4$$
Read off the coefficient of $\beta^2$.
Answer $$\boxed{\lambda=4}$$
Check
A firmer belief, variance $0.01$, would give $\lambda=100$: a tighter belief means a larger penalty, as it should.
λ is the noise variance over the prior variance; the ridge block derives this.
2§06.6 — where a prediction is least certain
A line through the origin is fitted to $100$ readings whose inputs all lie between $-1$ and $1$. Predictions with error bars are wanted at $x=0.5$, $x=1$ and $x=10$.
Find(a) Where should the error bar of the prediction be widest?
Given
$100$ readings, inputs in $[-1,1]$
prediction inputs $0.5$, $1$ and $10$
Hint 1/4
A prediction is uncertain for two reasons: the new reading's noise, and our remaining doubt about the slope.
Hint 2/4
For a line through the origin the doubt about the slope enters as $x^2\,\mathrm{Var}(\beta\mid D)$, while the noise part is the same at every input.
Hint 3/4
Compare $x^2$ at the three inputs: $0.25$, $1$ and $100$.
Hint 4/4
The error bar is widest at $x=10$.
Show solution
Only the part of the variance that changes with x can rank the three inputs, so we isolate it.
$\mathrm{RSS}(\beta)+\lambda\lVert\beta\rVert_2^2$ and $\mathrm{RSS}(\beta)+\lambda\lVert\beta\rVert_1$
The regularization section wrote the estimates $\hat\beta^R$ and $\hat\beta^L$.
$I_p,\ I_n$
identity matrices
the $p\times p$ and $n\times n$ identities
Sizes follow the coefficients and the readings.
Conventions used here
No intercept.
As in the lecture, inputs and responses are centered, so the model $y_i=\beta^Tx_i+\varepsilon_i$ has no intercept. A few examples use a line that truly passes through the origin instead; the posterior formulas never need centering.
Without an intercept, the ridge and lasso priors treat every coefficient alike.
The noise variance is known.
Throughout, $\sigma^2$ is a given number and only $\beta$ is unknown. If $\sigma^2$ had to be estimated as well, the formulas here would change.
Every posterior and every prediction on this page relies on it.
Variances, not standard deviations.
$\mathcal N(\mu,v)$ always has the variance second. The ridge prior $\mathcal N(0,\sigma^2/\lambda)$ has variance $\sigma^2/\lambda$; when a question gives a prior standard deviation $\tau$, then $\lambda=\sigma^2/\tau^2$.
Mixing the two is the most common way to get λ wrong by a square.
Laplace scale.
$\mathrm{Lap}(\mu,b)$ uses the scale $b$: density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$, variance $2b^2$. Some books use the rate $1/b$; on this page $b$ is always the scale.
Scale and rate are reciprocals, and confusing them inverts the answer.
Precision.
The inverse of a variance or a covariance matrix is a precision. Posterior formulas add precisions and invert once at the end.
Adding covariances instead is the most common slip in this chapter.
Constants.
$p(\beta\mid D)\propto p(D\mid\beta)\,p(\beta)$ drops factors without $\beta$. After taking $-\log$ we write $+\text{const}$ where the lecture writes $\propto$; both mean that terms without $\beta$ were dropped.
A constant never moves a maximizer, so dropping it is safe.
Names of the estimates.
This chapter writes $\hat\beta^{\mathrm{LS}}$, $\hat\beta^{\mathrm{MLE}}$, $\hat\beta^{\mathrm{MAP}}$, $\hat\beta^{\mathrm{Ridge}}$ and $\hat\beta^{\mathrm{Lasso}}$; the regularization section wrote $\hat\beta^R$ and $\hat\beta^L$ for the last two.
Same objects, new names; the chapter's own notation is used here.
Rounding.
Intermediate steps keep at least four significant figures; final answers are rounded to three decimals unless an exact fraction is asked for.
Posterior means often differ from least squares only in the second decimal.
6.1Least squares is maximum likelihood when the noise is Gaussian
Turns the least squares recipe into a probability statement: the fitted coefficients are the ones that make the observed data most probable.
Least squares picked $\beta$ by a rule we chose, the smallest RSS; this block shows which noise model makes that rule the most probable answer.
Solvable with what we have
Fit least squares: $\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$, or $\sum_ix_iy_i/\sum_ix_i^2$ for one centered input.
Write the density of one Gaussian reading, $\frac{1}{\sqrt{2\pi}\sigma}e^{-r^2/(2\sigma^2)}$ for a residual $r$.
Maximize a likelihood for yes/no data, as in the Bayesian and frequentist section.
Not solvable yet
Say which slope makes a whole set of Gaussian readings most probable.
Say whether a noisier sensor should change that slope.
Compare two slopes on a thousand readings when the computer rounds the likelihood to 0.
Three centered readings, $(x,y)=(-1,-1),(0,1),(1,0)$, with Gaussian noise of standard deviation $1$. Each reading's density is largest when its own residual is $0$, so pick $\beta=1$: the first reading lands exactly on the line and its density reaches its top value, $0.399$.
Why it fails
One slope serves all three readings. At $\beta=1$ the product of the densities is $0.399\times0.242\times0.242=0.0234$; at $\beta=0.5$ it is $0.352\times0.242\times0.352=0.0300$, about $1.28$ times larger. Maximizing one factor is not maximizing the product.
TheoremLeast squares is the maximum likelihood estimate
Conditions
$y_i=(\beta^{\mathrm{true}})^Tx_i+\varepsilon_i$ for $i=1,\dots,n$, with the $x_i$ fixed
$\varepsilon_1,\dots,\varepsilon_n$ i.i.d. $\mathcal N(0,\sigma^2)$, and $\sigma^2$ is known
The probability of the data is a product of bell curves, one per reading. Its logarithm is a constant minus $\mathrm{RSS}/(2\sigma^2)$, so the slope that makes the data most probable is the least squares slope, whatever the known $\sigma$ is.
Why the log-likelihood peaks at least squares
Independence turns the joint density of $y_1,\dots,y_n$ into the product $L(\beta)$ in the box.
$\log$ is strictly increasing, so $L$ and $l=\log L$ peak at the same $\beta$; the log turns the product into a sum.
The first term of $l(\beta)$ has no $\beta$ in it, and $\frac{1}{2\sigma^2}>0$. So maximizing $l$ is minimizing $\mathrm{RSS}(\beta)$, and $\sigma^2$ only rescales the curve.
In vector form $\mathrm{RSS}(\beta)=(y-X\beta)^T(y-X\beta)$, so $L(\beta)$ is the $\mathcal N(X\beta,\sigma^2I_n)$ density evaluated at $y$: $$L(\beta)=\frac{1}{(2\pi)^{n/2}\sigma^n}\exp\Big(-\frac{1}{2\sigma^2}(y-X\beta)^T(y-X\beta)\Big).$$
The three readings $(-1,-1),(0,1),(1,0)$. The $\textcolor{#1f6feb}{\text{log-likelihood}}$ is a downward parabola in $\beta$, for $\sigma=1$ (solid) and $\sigma=0.8$ (dashed). Both peak at the $\textcolor{#d1690a}{\text{least squares slope }0.5}$. The slope $1$, which fits the first reading exactly, scores $-3.76$ against $-3.51$.
Looks like this, but is not
The noise level multiplies the RSS in $l(\beta)$, so a noisier sensor should pull the most likely slope toward $0$.
$\sigma^2$ scales the RSS term but does not move its minimum. For the three readings, $\sigma=1$ and $\sigma=0.8$ both peak at $\beta=0.5$; only the height and the curvature change. The noise level starts to matter once a prior enters, in the next block.
reading
residual at β = 0.5
density
residual at β = 1
density
$(-1,-1)$
$-0.5$
$0.352$
$0$
$0.399$
$(0,1)$
$1$
$0.242$
$1$
$0.242$
$(1,0)$
$-0.5$
$0.352$
$-1$
$0.242$
product
$0.0300$
$0.0234$
Moving from $0.5$ to $1$ lifts the first density from $0.352$ to $0.399$ but drops the third from $0.352$ to $0.242$; the product falls by a factor of $1.28$.
The slope that makes three Gaussian readings most probable
The readings $(x,y)=(-1,-1),(0,1),(1,0)$ follow $y_i=\beta x_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,1)$. Find the slope with the largest likelihood, and compare it with $\beta=1$, which puts the first reading exactly on the line.
Find$\hat\beta^{\mathrm{MLE}}$ and the likelihood ratio $L(\hat\beta^{\mathrm{MLE}})/L(1)$.
Given
readings $(-1,-1),\ (0,1),\ (1,0)$, centered
noise i.i.d. $\mathcal N(0,\sigma^2)$ with $\sigma=1$ known
Solution
We maximize $l(\beta)$, not $L(\beta)$: the log turns the product into a sum of squares we can differentiate, and it peaks at the same slope.
Direct products of the three densities: $0.352\times0.242\times0.352\approx0.0300$ at $\beta=0.5$ and $0.399\times0.242\times0.242\approx0.0234$ at $\beta=1$. Their ratio is $1.28$.
This settles the naive attempt above: one slope serves every reading, and the product of their densities is largest where the sum of squared residuals is smallest.
A thousand readings: comparing two slopes when the likelihood underflows
A sensor gives $n=1000$ centered readings with $\sum x_i^2=330$, $\sum x_iy_i=264$ and $\sum y_i^2=1161.2$. The noise is $\mathcal N(0,1)$. Compare the slopes $0.8$ and $0.9$ by their likelihood.
Find$l(0.8)$, the size of $L(0.8)$, and the ratio $L(0.8)/L(0.9)$.
$0.8$ is the least squares slope here, $264/330=0.8$, so it must win. Also $\mathrm{RSS}(0.9)-\mathrm{RSS}(0.8)$ must equal $330\,(0.9-0.8)^2=3.3$, and it does.
Compare likelihoods through differences of log-likelihoods: the constant $n\log\frac{1}{\sqrt{2\pi}\sigma}$ cancels and nothing underflows.
Checkpoint
§06.1 — the maximum likelihood slope and the noise level
A centered data set has $\sum_i x_i^2=8$ and $\sum_i x_iy_i=6$. The model is $y_i=\beta x_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,9)$, so $\sigma=3$ is known.
Find(a) Which slope maximizes the likelihood?
Given
$\sum_i x_i^2=8$, $\sum_i x_iy_i=6$
noise variance $\sigma^2=9$
Hint 1/4
The noise level appears in the log-likelihood; ask whether it changes where the maximum is, or only how high it is.
Hint 2/4
$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\sigma}-\frac{1}{2\sigma^2}\mathrm{RSS}(\beta)$, and for one centered input the RSS is smallest at $\sum_ix_iy_i/\sum_ix_i^2$.
Hint 3/4
Here $\sum_ix_i^2=8$, $\sum_ix_iy_i=6$ and $\sigma^2=9$, so the factor $\frac{1}{18}$ multiplies the RSS.
Hint 4/4
The maximizer is $6/8=0.75$, the same as for any other known $\sigma$.
Show solution
We differentiate $l(\beta)$ directly; the constant term drops out at once.
right$$\hat\beta^{\mathrm{MLE}}=\frac{\sum_i x_iy_i}{\sum_i x_i^2}\ \ \text{for every }\sigma^2>0$$
6.2Ridge regression is the most probable β under a Gaussian prior
Reads the ridge penalty as a belief: each coefficient is Gaussian around 0 with variance σ²/λ, and ridge returns the posterior's peak.
Maximum likelihood used the data alone; now we add a belief about $\beta$ before the data, and the most probable $\beta$ afterwards turns out to be the ridge estimate.
TheoremRidge regression is the MAP estimate under a Gaussian prior
Conditions
Gaussian likelihood as in the previous box, with $\sigma^2$ known
prior: $\beta_1,\dots,\beta_p$ i.i.d. $\mathcal N(0,\sigma^2/\lambda)$, that is $\beta\sim\mathcal N\big(0,\tfrac{\sigma^2}{\lambda}I_p\big)$, with $\lambda>0$
centered data, so there is no intercept left unpenalized
Minus the log-posterior is the ridge loss divided by $2\sigma^2$, plus terms without $\beta$. So the peak of the posterior is the ridge estimate, and the penalty weight is the noise variance over the prior variance: $\lambda=\sigma^2/(\text{prior variance})$.
From the log-posterior to the ridge loss
Bayes' rule gives $p(\beta\mid D)=p(D\mid\beta)\,p(\beta)/p(D)$, and $p(D)$ does not depend on $\beta$: $$-\log p(\beta\mid D)=-\log p(D\mid\beta)-\log p(\beta)+\log p(D).$$
The likelihood term is $\frac{1}{2\sigma^2}\sum_i(y_i-\beta^Tx_i)^2-n\log\frac{1}{\sqrt{2\pi}\sigma}$, from the previous box.
Each prior factor is $\frac{\sqrt\lambda}{\sqrt{2\pi}\sigma}\exp\big(-\frac{\lambda\beta_j^2}{2\sigma^2}\big)$, so $$-\log p(\beta)=\frac{\lambda}{2\sigma^2}\sum_{j=1}^p\beta_j^2-p\log\frac{\sqrt\lambda}{\sqrt{2\pi}\,\sigma}.$$
Adding, the $\beta$-terms are $\frac{1}{2\sigma^2}\mathrm{Loss}_{\mathrm{Ridge}}(\beta,\lambda)$ and the rest is constant in $\beta$. Minimizing it is minimizing the ridge loss, whose minimizer is $(X^TX+\lambda I_p)^{-1}X^Ty$.
The three readings $(-2,-2),(1,0),(1,2)$ with $\sigma=1$ and the $\textcolor{#8250df}{\text{prior }\mathcal N(0,0.25)}$. The $\textcolor{#1f6feb}{\text{likelihood}}$, rescaled to area $1$, peaks at the least squares slope $1$. Their product, the $\textcolor{#d1690a}{\text{posterior}}$, is narrower than both and peaks at $0.6$, the ridge slope for $\lambda=\sigma^2/0.25=4$.
Looks like this, but is not
A prior $\beta_j\sim\mathcal N(0,1/\lambda)$ gives ridge regression with penalty $\lambda$.
Only when $\sigma^2=1$. The log-prior brings $\beta_j^2/(2\cdot\text{variance})$ and the log-likelihood brings $\mathrm{RSS}/(2\sigma^2)$; matching them needs variance $\sigma^2/\lambda$. With variance $1/\lambda$ the MAP is ridge with penalty $\sigma^2\lambda$.
A penalty of 4 and a prior with standard deviation 0.5 give the same slope
Three centered readings $(x,y)=(-2,-2),(1,0),(1,2)$ have noise $\mathcal N(0,1)$. Group one minimizes $\mathrm{RSS}(\beta)+4\beta^2$. Group two puts the prior $\beta\sim\mathcal N(0,0.5^2)$ on the slope and reports the posterior's peak. Compute both.
Find$\hat\beta^{\mathrm{Ridge}}$ and $\hat\beta^{\mathrm{MAP}}$.
Given
readings $(-2,-2),\ (1,0),\ (1,2)$
$\sigma^2=1$ known
penalty $\lambda=4$; prior $\mathcal N(0,\tau^2)$ with $\tau=0.5$
Solution
With one coefficient every matrix is a number, so both answers come from setting a derivative to zero; no inverse is needed.
Weighted average: the data alone say $6/6=1$ with weight $6/\sigma^2=6$, the prior says $0$ with weight $1/\tau^2=4$, and $(6\cdot1+4\cdot0)/10=0.6$.
This is the opening puzzle: a penalty $\lambda$ is a prior variance $\sigma^2/\lambda$ in disguise, so the two groups solved the same problem.
From 'within ±0.4, almost surely' to a penalty weight
Before seeing data, an engineer expects about $19$ coefficients in $20$ to lie within $\pm0.4$ of $0$. The sensor noise has standard deviation $0.3$. Which ridge penalty encodes this belief, and what happens to it if the noise standard deviation doubles to $0.6$?
Find$\tau^2$, and $\lambda$ for $\sigma=0.3$ and for $\sigma=0.6$.
Given
belief: $\mathcal N(0,\tau^2)$ with about $19$ draws in $20$ inside $\pm0.4$
$\sigma=0.3$, then $\sigma=0.6$
a normal draw falls within $2$ standard deviations of its mean about $19$ times in $20$
Solution
We turn the interval into a standard deviation first, because the translation to λ needs the prior variance, not a range.
Set the box's prior variance equal to the analyst's.
Answer $$\boxed{\lambda=8}$$
Check
Directly: $-\log p(\beta)$ contributes $\beta_j^2/(2\cdot0.5)=\beta_j^2$, and multiplying $-\log p(\beta\mid D)$ by $2\sigma^2=8$ turns that into $8\beta_j^2$.
Always translate with the variance of the noise, not its standard deviation.
⚠ Forgetting the noise variance in the translation
the prior alone looks like the whole story, so λ gets read as its precision
wrong$$\lambda=\frac{1}{\tau^2}$$
right$$\lambda=\frac{\sigma^2}{\tau^2}$$
⚠ Reading σ²/λ as a standard deviation
$\mathcal N(\mu,\cdot)$ is written with a variance, while many texts quote spreads as standard deviations
With a Laplace prior the log-prior is a sum of absolute values, so minus the log-posterior is the lasso loss over $2\sigma^2$ plus a constant. The lasso is the peak of that posterior; there is still no closed form in general.
From the Laplace prior to the lasso loss
With $b=2\sigma^2/\lambda$ the Laplace density becomes $$\frac{1}{2b}e^{-\lvert\beta_j\rvert/b}=\frac{\lambda}{4\sigma^2}e^{-\lambda\lvert\beta_j\rvert/(2\sigma^2)},$$ so $$-\log p(\beta)=\frac{\lambda}{2\sigma^2}\sum_{j=1}^p\lvert\beta_j\rvert-p\log\frac{\lambda}{4\sigma^2}.$$
Add the likelihood term $\frac{1}{2\sigma^2}\sum_i(y_i-\beta^Tx_i)^2-n\log\frac{1}{\sqrt{2\pi}\sigma}$ and $\log p(D)$. The part with $\beta$ is $\frac{1}{2\sigma^2}\mathrm{Loss}_{\mathrm{Lasso}}(\beta,\lambda)$.
Minimizing $-\log p(\beta\mid D)$ is therefore minimizing the lasso loss.
Two $\textcolor{#8250df}{\text{priors}}$ with the same variance $0.5$: $\mathrm{Lap}(0,0.5)$, solid, and $\mathcal N(0,0.5)$, dashed. The Laplace density has a corner at $0$, where it reaches $1$, and heavier tails: at $2$ it is $0.018$ against the Gaussian's $0.010$.
Looks like this, but is not
The lasso sets $\hat\beta_j=0$, so the posterior says this coefficient is $0$ with high probability.
The posterior is a density, and any single value, $0$ included, has probability $0$. The lasso reports where the density peaks. Under the same posterior there is, for instance, a positive probability that $\beta_j>0.1$.
estimator
prior
estimate, Σxy = 6
estimate, Σxy = 1.5
on average
least squares
none
$1$
$0.25$
unbiased
ridge
$\mathcal N(0,\,0.25)$
$0.6$
$0.15$
biased toward 0
lasso
$\mathrm{Lap}(0,\,0.5)$
$0.667$
$0$
biased toward 0
Both priors pull toward $0$: ridge scales the least squares slope by $6/10$, the lasso subtracts $1/3$ and stops at $0$. Only least squares is unbiased; the other two trade some bias for lower variance.
Lasso with λ = 4 on the three readings, read as a Laplace prior
Same readings as the ridge example: $\sum x_i^2=6$, $\sum x_iy_i=6$, $\sum y_i^2=8$, $\sigma^2=1$. Find the Laplace prior that makes the lasso with $\lambda=4$ a MAP estimate, and compute that estimate.
FindThe prior $\mathrm{Lap}(0,b)$ and $\hat\beta^{\mathrm{Lasso}}$.
The absolute value has a corner at $0$, so we minimize separately on $\beta>0$ and on $\beta<0$, then check the corner. This rederives the one-predictor rule of the regularization section, further reading there and here; in general the lasso has no closed form.
Shortcut for one coefficient: $\big(\sum x_iy_i-\lambda/2\big)/\sum x_i^2=(6-2)/6=\tfrac23$ when $\sum x_iy_i>\lambda/2$. It sits between $0$ and the least squares slope $1$, as a shrunken estimate must.
The lasso subtracted a fixed $\lambda/(2\sum x_i^2)=\tfrac13$ from the least squares slope, while ridge with the same $\lambda$ multiplied it by $0.6$.
Weak evidence: the most probable slope is exactly zero
Now the readings are $(x,y)=(-2,-0.5),(1,0),(1,0.5)$, with $\sigma^2=1$ and $\lambda=4$ as before. Find the lasso MAP and the ridge MAP.
Find$\hat\beta^{\mathrm{Lasso}}$ and $\hat\beta^{\mathrm{Ridge}}$.
Given
$\sum_i x_i^2=6$, $\sum_i x_iy_i=1+0+0.5=1.5$
$\sigma^2=1$, $\lambda=4$
Solution
The data now pull only weakly, so we test the corner first: if the loss rises on both sides of $0$, nothing else needs checking.
Threshold check: the lasso slope is $0$ exactly when $\lvert\sum x_iy_i\rvert\le\lambda/2$, and here $1.5\le2$. The least squares slope, $1.5/6=0.25$, is small too.
The corner of the Laplace prior at $0$ absorbs weak evidence completely; that is how the lasso selects variables.
Checkpoint
§06.3 — the Laplace prior behind a lasso penalty
A lasso fit uses $\lambda=2$ on centered data with Gaussian noise of variance $\sigma^2=0.5$. We want the prior that makes this lasso fit a MAP estimate.
Find(a) Which i.i.d. prior on each coefficient does it?
Given
$\lambda=2$
$\sigma^2=0.5$, known
Hint 1/4
Match the exponent of the prior with the penalty term of $-\log p(\beta\mid D)$.
Hint 2/4
$\mathrm{Lap}(0,b)$ has $-\log p(\beta_j)=\lvert\beta_j\rvert/b+\text{const}$, and the lasso needs $\lambda\lvert\beta_j\rvert/(2\sigma^2)$, so $b=2\sigma^2/\lambda$.
Hint 3/4
Here $\sigma^2=0.5$ and $\lambda=2$.
Hint 4/4
$b=2(0.5)/2=0.5$, so the prior is $\mathrm{Lap}(0,0.5)$.
Show solution
The two expressions must penalize one unit of the absolute value by the same amount, so we equate their coefficients.
Plug back: $\frac{1}{2b}e^{-\lvert\beta_j\rvert/b}=e^{-2\lvert\beta_j\rvert}$, and $\lambda/(2\sigma^2)=2/1=2$ matches the coefficient in the exponent.
The Laplace scale uses $2\sigma^2/\lambda$; the ridge variance uses $\sigma^2/\lambda$. Keep the two apart.
⚠ Using the ridge pattern σ²/λ for the Laplace scale
the ridge translation is fresh in mind, and both formulas divide by λ
wrong$$b=\frac{\sigma^2}{\lambda}$$
right$$b=\frac{2\sigma^2}{\lambda}$$
⚠ Taking b² as the Laplace variance
for the Gaussian the second parameter is the variance, and the habit carries over
wrong$$\mathrm{Var}(Z)=b^2$$
right$$\mathrm{Var}(Z)=2b^2$$
⚠ Differentiating |β| at 0 as if β were positive
the absolute value is usually met on one side of zero only
right$$\text{slope }+1\text{ right of }0,\ \ -1\text{ left of }0:\ \text{check both sides}$$
The three readings of the ridge example: $\sum x_i^2=6$, $\sum x_iy_i=6$, $\sigma^2=1$. Each frame shows $\textcolor{#d1690a}{-\log p(\beta\mid D)}$, shifted to pass through $0$ at $\beta=0$, for a larger penalty. The minimum slides left as $\lambda$ grows, reaches the corner at $\lambda=12$ and stays there.
At the edges
λ = 0 MAP 1
No prior: the MAP estimate is the least squares slope.
λ ≥ 12 MAP 0
Once $\lambda$ reaches twice $\sum x_iy_i$, the corner holds the minimum at exactly $0$.
6.4Bayesian linear regression: the whole posterior, not just its peak
Keeps the whole posterior of β: with Gaussian prior and noise it is Gaussian, fixed by a mean and a covariance.
Least squares, ridge and the lasso each return one vector, the peak of a likelihood or a posterior; Bayesian linear regression keeps the whole posterior, and with it a measure of how sure we are.
TheoremGaussian prior and Gaussian likelihood give a Gaussian posterior
Conditions
prior $\beta\sim\mathcal N(\mu_0,\Sigma_0)$
likelihood $y\mid\beta\sim\mathcal N(X\beta,\Sigma_D)$ with $X$ fixed; $\Sigma_D=\sigma^2I_n$ for i.i.d. noise
$\Sigma_0$ and $\Sigma_D$ invertible covariance matrices
any prior is allowed in principle; the Gaussian pair is the one with this closed form
Precisions add: the posterior precision is the prior precision plus the data precision $X^T\Sigma_D^{-1}X$. The posterior mean is the posterior covariance times the sum of what the data say, $X^T\Sigma_D^{-1}y$, and what the prior says, $\Sigma_0^{-1}\mu_0$.
Why the posterior is Gaussian: completing the square (further reading)
Up to terms without $\beta$, $$-2\log p(\beta\mid D)=(y-X\beta)^T\Sigma_D^{-1}(y-X\beta)+(\beta-\mu_0)^T\Sigma_0^{-1}(\beta-\mu_0).$$
Expanding and keeping only the terms with $\beta$: $$\beta^T\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)\beta-2\beta^T\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big).$$
A Gaussian $\mathcal N(\mu,\Sigma)$ has $-2\log p=\beta^T\Sigma^{-1}\beta-2\beta^T\Sigma^{-1}\mu+\text{const}$. Matching the two lines gives $\Sigma^{-1}=\Sigma_0^{-1}+X^T\Sigma_D^{-1}X$ and $\Sigma^{-1}\mu=X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0$, the box.
A density proportional to $e^{-(\text{quadratic in }\beta)/2}$ with a positive definite matrix is that Gaussian; the normalizing constant takes care of itself.
Four centered readings with two inputs, $\sigma^2=1$, and the prior $\mathcal N(0,I_2)$, that is ridge with $\lambda=1$. Each curve joins the points one standard deviation from its centre. The $\textcolor{#d1690a}{\text{posterior}}$ is smaller than the $\textcolor{#8250df}{\text{prior}}$ circle and the $\textcolor{#1f6feb}{\text{likelihood}}$ ellipse, and its centre $(0.55,0.64)$ lies between theirs.
Looks like this, but is not
The posterior covariance depends on the readings $y$, so a surprising set of readings leaves us less sure about $\beta$.
$\Sigma_{\beta|D}=(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X)^{-1}$ has no $y$ in it. With $\sigma^2$ known, how sure we end up depends on where we measured and how noisy the sensor is, not on what the readings turned out to be; the readings move only the mean.
Three readings or twenty-three: one ridge estimate, two posteriors
Data set A is the three readings with $\sum x_i^2=6$, $\sum x_iy_i=6$. Data set B has $23$ readings with $\sum x_i^2=46$, $\sum x_iy_i=30$. Both use $\sigma^2=1$ and the prior $\mathcal N(0,0.25)$, which is ridge with $\lambda=4$. Compare the ridge estimates and the posteriors.
Find$\hat\beta^{\mathrm{Ridge}}$ and $\mathcal N(\mu_{\beta|D},\Sigma_{\beta|D})$ for A and for B.
Given
A: $\sum x_i^2=6$, $\sum x_iy_i=6$
B: $\sum x_i^2=46$, $\sum x_iy_i=30$
$\sigma^2=1$, prior $\mathcal N(0,\,0.25)$
Solution
With one coefficient the box's matrices are numbers: precisions add, and the mean is the posterior variance times the summed evidence.
Ridge formula directly: $6/(6+4)=0.6$ and $30/(46+4)=0.6$. The standard deviations differ by the factor $\sqrt{50/10}=\sqrt5\approx2.24$.
Two fits can report the same coefficient with very different certainty; only the posterior variance tells them apart.
The ridge posterior for two coefficients, by hand
Four centered readings with two inputs: rows $x_i=(1,1), \allowbreak (1,0), \allowbreak (-1,0), \allowbreak (-1,-1)$ and $y=(2,0,-1,-1)$. The noise is $\mathcal N(0,1)$ and the prior is $\mathcal N(0,I_2)$, the ridge prior with $\lambda=1$. Find the posterior and compare its mean with least squares.
Find$\Sigma_{\beta|D}$, $\mu_{\beta|D}$ and $\hat\beta^{\mathrm{LS}}$.
In the ridge case $\Sigma_0=(\sigma^2/\lambda)I$ and $\Sigma_D=\sigma^2I$, so the box becomes $\Sigma_{\beta|D}=\sigma^2(X^TX+\lambda I)^{-1}$ and $\mu_{\beta|D}=(X^TX+\lambda I)^{-1}X^Ty$. Only $X^TX$ and $X^Ty$ are needed.
$\mu_{\beta|D}$ must solve $(X^TX+I)\mu=X^Ty$: $5\cdot\tfrac{6}{11}+2\cdot\tfrac{7}{11}=4$ and $2\cdot\tfrac{6}{11}+3\cdot\tfrac{7}{11}=3$. The posterior variances, $0.27$ and $0.45$, are below the prior's $1$ and below the least squares variances $0.5$ and $1$.
One 2 × 2 inverse and two matrix-vector products; for many coefficients this is a job for a linear solver.
For a Gaussian posterior the mean is also the peak, so the posterior mean is the MAP estimate, here the ridge estimate.
A prior centred away from zero: an earlier study said about 2
An earlier study suggests the slope is near $2$, so the prior is $\beta\sim\mathcal N(2,0.5)$. The new readings are the three of the ridge example, $\sum x_i^2=6$, $\sum x_iy_i=6$, with $\sigma^2=1$. Find the posterior.
As a weighted average: the data alone say $6/6=1$ with weight $6$, the prior says $2$ with weight $2$, and $(6\cdot1+2\cdot2)/8=1.25$, between the two and closer to the heavier weight.
The posterior mean is a precision-weighted average of the prior mean and the least squares answer; a prior centred at $0$ is the special case ridge uses.
Checkpoint
§06.4 — a one-coefficient Gaussian posterior
One centered input gives $\sum_i x_i^2=8$ and $\sum_i x_iy_i=4$, with noise variance $\sigma^2=1$. The prior on the slope is $\mathcal N(0,0.5)$.
Find(a) Which is the posterior of the slope?
Given
$\sum_i x_i^2=8$, $\sum_i x_iy_i=4$
$\sigma^2=1$
prior $\mathcal N(0,\,0.5)$
Hint 1/4
Two numbers are needed, a variance and a mean; the variance comes first because the mean uses it.
Hint 2/4
$\Sigma_{\beta|D}^{-1}=\frac{1}{\text{prior variance}}+\frac{\sum x_i^2}{\sigma^2}$ and, with $\mu_0=0$, $\mu_{\beta|D}=\Sigma_{\beta|D}\,\frac{\sum x_iy_i}{\sigma^2}$.
Hint 3/4
Here the prior variance is $0.5$, $\sum x_i^2=8$, $\sum x_iy_i=4$ and $\sigma^2=1$, so the precision is $2+8$.
Hint 4/4
$\Sigma_{\beta|D}=0.1$ and $\mu_{\beta|D}=0.1\times4=0.4$: the posterior is $\mathcal N(0.4,0.1)$.
Show solution
The mean formula contains the posterior variance, so we find that first.
Precision
$$\Sigma_{\beta|D}^{-1}=\tfrac{1}{0.5}+8=10$$
Prior precision plus data precision.
Mean
$$\mu_{\beta|D}=\tfrac{1}{10}\cdot4=0.4$$
Posterior variance times $\sum x_iy_i/\sigma^2$.
Answer $$\boxed{\mathcal N(0.4,\ 0.1)}$$
Check
Ridge check: $\lambda=\sigma^2/0.5=2$ gives $4/(8+2)=0.4$, the same mean.
Compute the posterior variance first; the mean is that variance times the evidence.
⚠ Adding covariances instead of precisions
adding uncertainties sounds natural, but it is information that adds
Multiply yesterday's posterior by the new reading's likelihood and you have today's posterior. In numbers, the new reading adds $x_n^2/\sigma^2$ to the precision, and the mean moves to a precision-weighted blend of the old mean and the new evidence.
Why the sequential answer equals the batch answer
Bayes' rule with everything conditioned on the first $n-1$ readings: $p(\beta\mid D_{1:n})\propto p(y_n\mid\beta,D_{1:n-1})\,p(\beta\mid D_{1:n-1})$, and given $\beta$ the new reading does not depend on the old ones.
Unrolled from $n=1$: $p(\beta\mid D_{1:n})\propto p(\beta)\prod_{i=1}^n p(y_i\mid\beta)$, the batch posterior. The prior enters once.
The numeric step is the posterior box with $\Sigma_0=v_{n-1}$, $\mu_0=m_{n-1}$ and a single row $x_n$.
Twenty readings from a line with slope $0.8$ and noise standard deviation $0.5$, prior $\mathcal N(0,1)$. The $\textcolor{#d1690a}{\text{posterior standard deviation}}$ drops at every reading, by an amount set by $x_n^2$ and not by $y_n$: reading 4, at $x=0.97$, cuts it from $0.39$ to $0.31$; reading 18, at $x=0.02$, changes it by less than $0.0001$.
Looks like this, but is not
Updating one reading at a time uses the prior once per reading, so the sequential posterior leans harder toward the prior than the batch one.
The prior enters once, at the first step; every later step uses the previous posterior, not the original prior. Unrolled, the chain is the prior times all $n$ likelihood factors, exactly the batch posterior.
n
Σx²
Σxy
precision
mean
sd
0
0
0
1
0
1.000
1
1
0.5
5
0.400
0.447
2
1.25
0.85
6
0.567
0.408
5
2.723
2.25
11.89
0.757
0.290
10
4.182
3.56
17.73
0.803
0.237
20
6.859
5.785
28.44
0.814
0.188
The precision is $1+\sum x^2/0.25$ and only grows; the mean wanders early and settles near $0.8$. After $20$ readings the prior supplies $1$ of the precision $28.4$.
Two readings, one at a time and then both at once
Prior $\beta\sim\mathcal N(0,1)$, noise variance $\sigma^2=0.25$. Reading 1 is $(x,y)=(1,0.5)$ and reading 2 is $(-0.5,-0.7)$. Update after each reading, then compute the batch posterior from both.
Find$(m_1,v_1)$, $(m_2,v_2)$ and the batch posterior.
Given
prior $m_0=0$, $v_0=1$
$\sigma^2=0.25$
readings $(1,\,0.5)$ and $(-0.5,\,-0.7)$
Solution
We run the update rule twice and then the posterior box once; agreement between the two routes is the point of the example.
The posterior box with the prior $\mathcal N(0,1)$ and both rows at once.
Answer $$\boxed{\mathcal N(0.4,\ 0.2)\ \to\ \mathcal N(0.567,\ 0.167),\ \ \text{equal to the batch posterior}}$$
Check
Reverse the order: reading 2 first gives precision $2$ and mean $\tfrac{1.4}{2}=0.7$; then reading 1 gives precision $6$ and mean $\tfrac{2\cdot0.7+2}{6}=\tfrac{3.4}{6}$, the same.
Sequential and batch updating are the same computation in a different order; use whichever matches how the data arrive.
How many readings at x = 1 halve the posterior standard deviation?
Same prior $\mathcal N(0,1)$ and $\sigma^2=0.25$, and every reading is taken at $x=1$. After $2$ readings the posterior standard deviation is $1/3$. How many readings bring it to $1/6$ or below?
FindThe smallest $n$ with posterior standard deviation at most $1/6$.
Given
prior variance $1$, $\sigma^2=0.25$
all readings at $x=1$
Solution
Standard deviations do not add, precisions do, so we translate the target into a precision first.
Halving the standard deviation needs four times the precision, from $9$ to $36$.
Answer $$\boxed{n=9\ \text{readings}}$$
Check
$n=8$ gives $1/\sqrt{33}\approx0.174>0.167$ and $n=9$ gives $1/\sqrt{37}\approx0.164$, so $9$ is the first to pass.
Precision grows linearly in the number of readings, so the error bar on β shrinks like one over the square root.
Checkpoint
§06.5 — one more reading
After some readings the posterior of a slope is $\mathcal N(0.5,\,0.05)$. One more reading arrives, $(x,y)=(1,\,1.1)$, from a sensor with noise variance $\sigma^2=0.2$.
Find(a) What is the posterior after the new reading?
Given
current posterior $\mathcal N(0.5,\,0.05)$
new reading $(1,\,1.1)$
$\sigma^2=0.2$
Hint 1/4
Treat the current posterior as the prior for one new reading.
Hint 2/4
$\frac{1}{v_{\text{new}}}=\frac{1}{v_{\text{old}}}+\frac{x^2}{\sigma^2}$ and $m_{\text{new}}=v_{\text{new}}\big(\frac{m_{\text{old}}}{v_{\text{old}}}+\frac{xy}{\sigma^2}\big)$.
Hint 3/4
Here $m_{\text{old}}=0.5$, $v_{\text{old}}=0.05$, $x=1$, $y=1.1$ and $\sigma^2=0.2$: the precisions are $20$ and $5$.
Hint 4/4
$v_{\text{new}}=1/25=0.04$ and $m_{\text{new}}=0.04\,(10+5.5)=0.62$.
Show solution
One step of the update rule does it; no earlier reading needs to be revisited.
Precision
$$\tfrac{1}{0.05}+\tfrac{1^2}{0.2}=20+5=25$$
Old precision plus the new reading's $x^2/\sigma^2$.
A precision-weighted blend of the old mean and the new evidence.
Answer $$\boxed{\mathcal N(0.62,\ 0.04)}$$
Check
The new mean lies between $0.5$ and the reading's own slope $1.1$, four times closer to $0.5$ because the old posterior carries four times the precision: $0.62-0.5=0.12$ against $1.1-0.62=0.48$.
In a sequential update the old posterior is weighted by its precision, never by a plain average.
⚠ Restarting from the original prior at each step
the prior is written in the problem, the intermediate posterior is not
The same twenty readings. Step through the $\textcolor{#d1690a}{\text{posterior}}$ of the slope after $0$, $1$, $2$, $5$, $10$ and $20$ readings; the dashed curve is the $\textcolor{#8250df}{\text{prior}}$ $\mathcal N(0,1)$ and the dotted line the true slope $0.8$.
At the edges
n = 0 N(0, 1)
Nothing observed yet: the posterior is the prior.
n → ∞ sd → 0
With inputs spread over $[-1,1]$ each reading adds about $4/3$ to the precision on average, so the standard deviation shrinks like $1/\sqrt n$ and the mean settles at the true slope.
6.6The posterior predictive distribution: an error bar that knows where the data were
Gives each prediction an error bar: the new reading's noise $\sigma^2$ plus the coefficient uncertainty $x^T\Sigma_{\beta|D}x$ seen at input $x$.
A posterior for $\beta$ is not yet a statement about the next reading; averaging the model over the posterior gives one.
TheoremPosterior predictive distribution at a new input x
Predict with the posterior mean of the coefficients. The variance has two parts: $\sigma^2$, the noise any new reading carries however much data we have, and $x^T\Sigma_{\beta|D}x$, what we still do not know about $\beta$, seen through the new input.
Why the predictive distribution is this Gaussian
Given $D$, $x^T\beta$ is a linear function of the Gaussian vector $\beta$, so $x^T\beta\sim\mathcal N\big(x^T\mu_{\beta|D},\,x^T\Sigma_{\beta|D}x\big)$.
$\varepsilon$ is Gaussian and independent of $\beta$, and a sum of independent Gaussians is Gaussian with the means and the variances added. So $Y=x^T\beta+\varepsilon$ has mean $\mu_{\beta|D}^Tx$ and variance $x^T\Sigma_{\beta|D}x+\sigma^2$.
The posterior after the two readings of the sequential example, $\mathcal N(0.567,\,1/6)$, with $\sigma^2=0.25$. The $\textcolor{#d1690a}{\text{predictive band}}$, mean $\pm2$ standard deviations, stays close to the dashed noise band near the $\textcolor{#1f6feb}{\text{readings}}$ and widens away from them: at $x=3$ it is $1.70\pm2.65$.
Looks like this, but is not
With enough data the predictive variance goes to $0$, since the posterior of $\beta$ collapses onto one value.
Only $x^T\Sigma_{\beta|D}x$ shrinks with data. The $\sigma^2$ term belongs to the new reading itself: even with $\beta$ known exactly, a new reading scatters with variance $\sigma^2$.
Error bars at x = 0.5 and at x = 3 after two readings
Use the posterior after the two readings of the sequential example, $\beta\mid D\sim\mathcal N(17/30,\ 1/6)$, with $\sigma^2=0.25$. Find the predictive mean and standard deviation at $x=0.5$ and at $x=3$, and compare with a plug-in error bar that ignores the uncertainty in $\beta$.
FindThe predictive mean and standard deviation at both inputs, and the plug-in interval at $x=3$.
At $x=0$ the parameter term vanishes and the sd is exactly $\sigma=0.5$, the narrowest point of the band in the figure; both answers are above it, as they must be.
The predictive standard deviation is never below σ, and it grows with the distance from where the data were taken.
Two coefficients: which directions the data pinned down
Use the ridge posterior of the four two-input readings: $\mu_{\beta|D}=\tfrac{1}{11}(6,7)^T$ and $\Sigma_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}3&-2\\-2&5\end{pmatrix}$, with $\sigma^2=1$. Find the predictive mean and variance at $x=(1,1)$ and at $x=(1,-1)$.
Find$\mu_{\beta|D}^Tx$ and $\sigma^2+x^T\Sigma_{\beta|D}x$ at both inputs.
Why: along $(1,1)$ the four inputs give $x_i^T(1,1)^T=2, \allowbreak 1, \allowbreak -1, \allowbreak -2$, sum of squares $10$; along $(1,-1)$ they give $0,1,-1,0$, sum of squares $2$. The data spread five times as much along $(1,1)$, so that direction is pinned down better.
Predictive uncertainty depends on the direction of the new input: directions the data barely explored keep more of the prior's doubt.
Checkpoint
§06.6 — the distribution of the next reading
The posterior of a slope through the origin is $\mathcal N(0.6,\,0.1)$ and the noise variance is $\sigma^2=1$. A new input $x=2$ arrives.
Find(a) What is the distribution of the new reading $Y$?
Given
$\beta\mid D\sim\mathcal N(0.6,\,0.1)$
$\sigma^2=1$
new input $x=2$
Hint 1/4
The new reading is $2\beta$ plus fresh noise; find the mean and the variance of that sum.
Hint 2/4
For one coefficient, $Y\mid x,D\sim\mathcal N\big(\mu_{\beta|D}x,\ \sigma^2+x^2\Sigma_{\beta|D}\big)$.
Hint 3/4
Here $\mu_{\beta|D}=0.6$, $\Sigma_{\beta|D}=0.1$, $\sigma^2=1$ and $x=2$.
Hint 4/4
Mean $1.2$ and variance $1+4(0.1)=1.4$: $Y\sim\mathcal N(1.2,\,1.4)$.
Show solution
The new reading is a sum of two independent Gaussians, so we add means and add variances.
At $x$: mean $\mu_{\beta|D}^Tx$, variance $\sigma^2+x^T\Sigma_{\beta|D}x$.
Check
The mean lies between the prior mean and least squares, every posterior variance is below the prior's, and the predictive variance is at least $\sigma^2$.
Where it goes wrong
Adding covariances instead of precisions.
Forgetting $\Sigma_0^{-1}\mu_0$ when the prior mean is not $0$.
Leaving $\sigma^2$ out of the predictive variance.
One-coefficient lasso by cases (further reading)
A single coefficient, a lasso penalty λ and the two summaries Σx² and Σxy.
Summaries
$s=\sum_ix_i^2$ and $c=\sum_ix_iy_i$.
Corner test
If $\lvert c\rvert\le\lambda/2$, the answer is exactly $0$.
A Laplace prior absorbs evidence up to a threshold and sets the coefficient exactly to $0$.
Same readings, same $\lambda$, same noise; only the prior's shape changes, and the answer goes from $0.15$ to exactly $0$ because the absolute value has a corner where the square is flat.
How to tell them apart
A squared penalty (Gaussian prior) multiplies the least squares slope by $\sum x_i^2/(\sum x_i^2+\lambda)$; an absolute-value penalty (Laplace prior) subtracts $\lambda/(2\sum x_i^2)$ and stops at $0$.
A prediction near the data: the noise dominates
Posterior $\mathcal N(17/30,\ 1/6)$ for a slope through the origin, $\sigma^2=0.25$. Split the predictive variance at $x=0.5$ into its noise and parameter parts.
FindThe two parts and their shares.
Given
$\Sigma_{\beta|D}=1/6$, $\sigma^2=0.25$
$x=0.5$
Solution
The two parts of the predictive variance are separate terms, so we compute each and compare.
The shares have swapped: $6:1$ for the noise at $x=0.5$, $1:6$ at $x=3$.
Far from the data, more readings would narrow the error bar a lot; near the data, hardly at all.
Same posterior, same noise; at $x=0.5$ the noise is six sevenths of the predictive variance, at $x=3$ the slope's uncertainty is.
How to tell them apart
Compare $x^T\Sigma_{\beta|D}x$ with $\sigma^2$: if the first is small, more data will not help the prediction much; if it is large, the prediction is limited by what we know about $\beta$.
Scaffolding comes off
The common skeleton
Summaries: $s=\sum_ix_i^2$ and $c=\sum_ix_iy_i$.
Precisions: prior $1/\tau^2$ and data $s/\sigma^2$; add them and invert to get the posterior variance $v$.
Prediction at $x^\ast$: mean $m\,x^\ast$, variance $\sigma^2+(x^\ast)^2v$.
Check: $m$ lies between $\mu_0$ and $c/s$, and $v$ is below both $\tau^2$ and $\sigma^2/s$.
1 · fully worked
Posterior and prediction from four readings, prior N(0, 1)
Four centered readings $(x,y)=(-1,-0.8), \allowbreak (1,1.1), \allowbreak (2,1.9), \allowbreak (-2,-2.2)$, noise variance $\sigma^2=0.25$, prior $\mathcal N(0,1)$ on the slope. Find the posterior, then the predictive distribution and a 2-standard-deviation interval at $x^\ast=3$.
Find$m$, $v$, and the predictive mean, variance and interval at $x^\ast=3$.
Ridge check: $\lambda=\sigma^2/\tau^2=0.25$, and $c/(s+\lambda)=10.1/10.25\approx0.985$, the same mean.
With plenty of data the prior barely moves the mean, but a prediction far out still pays for the remaining slope uncertainty.
2 · you write the reasoning
An easier one, and this time you write the reasons. One reading $(x,y)=(2,3)$, $\sigma^2=1$, prior $\mathcal N(0,1)$; predict at $x^\ast=1$. For each line, write why it is allowed.
$s=4,\qquad c=6$
reasoning
One reading: $s=x^2=4$ and $c=xy=6$.
$\tfrac1v=1+4=5,\qquad v=0.2$
reasoning
Prior precision $1/\tau^2=1$ plus data precision $s/\sigma^2=4$; the posterior variance is the inverse of the sum.
$m=0.2\times6=1.2$
reasoning
The prior mean is $0$, so the mean is the variance times $c/\sigma^2=6$.
Predict with the posterior mean; the variance adds the noise $\sigma^2=1$ to $(x^\ast)^2v=0.2$. That mean and variance are both $1.2$ is a coincidence of these numbers.
The mean is shrunk from the data's $1.5$ toward the prior's $0$, and the variance $0.2$ is below both the prior's $1$ and the data's $\sigma^2/s=0.25$.
3 · find the buried error
Harder, with two errors buried in the solution. Four centered readings $(-2,-2.6), \allowbreak (-1,-0.9), \allowbreak (1,1.3), \allowbreak (2,2.2)$, $\sigma^2=0.5$, and an earlier study gives the prior $\mathcal N(1,\,0.5)$. Find the posterior mean and the predictive variance at $x^\ast=2$.
Step 1. $s=10$ and $c=5.2+0.9+1.3+4.4=11.8$: the summaries.
Step 2. $\tfrac1v=\tfrac{1}{0.5}+\tfrac{10}{0.5}=22$, so $v=\tfrac{1}{22}\approx0.0455$: precisions add.
Step 3.$m=v\cdot c/\sigma^2=23.6/22\approx1.073$: the posterior mean.
Step 4. The predictive variance at $x^\ast=2$ is $(x^\ast)^2v=4/22\approx0.182$.
Step 5. Check: $v\approx0.0455$ is below both $\tau^2=0.5$ and $\sigma^2/s=0.05$.
the two buried errors (2)
⚠ step 3
The prior mean is $1$, not $0$, so the term $\mu_0/\tau^2=2$ belongs in the bracket.
In the ridge case $\mu_0=0$ and the term vanishes, so the habit is to drop it.
A new reading carries its own noise: the predictive variance is $\sigma^2+(x^\ast)^2v$.
The posterior variance feels like the whole uncertainty.
right
$0.5+\tfrac{4}{22}\approx0.682$
4 · the bare problem
§06.6 — posterior and prediction interval from four readings
Four centered readings of a slope through the origin: $(-2,-1.3), \allowbreak (-1,-0.4), \allowbreak (1,0.7), \allowbreak (2,1.0)$. The noise variance is $\sigma^2=0.16$ and the prior on the slope is $\mathcal N(0,\,0.04)$.
Find
(a) Find the posterior mean and standard deviation of the slope.
(b) Find the predictive mean and a 2-standard-deviation interval at $x^\ast=1.5$.
Ridge check: $\lambda=0.16/0.04=4$ and $5.7/(10+4)\approx0.407$.
A tight prior can pull the estimate far from the data's own answer; the posterior variance shows how much the data were outweighed.
Full exam-style question
Exam-style: from likelihood to ridge posterior and two predictions, with two coefficientsexam format
Four centered readings with two inputs: rows $x_i=(1,0), \allowbreak (0,1), \allowbreak (1,1), \allowbreak (-2,-2)$ and responses $y=(1,-1,0.5,-0.5)$. The model is $y_i=\beta^Tx_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,1)$.
(a) Write the likelihood as a multivariate Gaussian and find $\hat\beta^{\mathrm{MLE}}$.
(b) With the prior $\beta\sim\mathcal N(0,I_2)$, find $\mu_{\beta|D}$ and $\Sigma_{\beta|D}$, and name the ridge estimate they correspond to.
(c) Find the predictive mean and variance at $x=(1,1)$ and at $x=(1,-1)$.
(d) Explain why the two predictive variances differ so much.
Find$\hat\beta^{\mathrm{MLE}}$; $\mu_{\beta|D}$ and $\Sigma_{\beta|D}$; two predictive distributions; an explanation.
Shrinkage along the two directions: least squares has $\hat\beta_1+\hat\beta_2\approx0.273$ and $\hat\beta_1-\hat\beta_2=2$; the posterior mean has $0.25=\tfrac{11}{12}(0.273)$ and $1=\tfrac12(2)$. The well-measured direction keeps eleven twelfths, the poorly measured one half.
In an exam, compute $X^TX$ and $X^Ty$ once; the MLE, the posterior and every prediction are built from them.
Practice
A · concept 4 questions
1§06.2 — is a MAP estimate unbiased?
A classmate argues that ridge regression is the most probable coefficient under a prior centred at $0$, so for $\lambda>0$ it must be an unbiased estimator of $\beta^{\mathrm{true}}$. You test the claim with one input.
Find(a) Is the claim true or false?
Given
model $y_i=\beta^{\mathrm{true}}x_i+\varepsilon_i$ with $E[\varepsilon_i]=0$ and the $x_i$ fixed
Edge case: with $\lambda=0$ the factor is $1$ and the estimate is unbiased; that is least squares.
MAP estimates with a prior centred at $0$ are biased toward $0$; the prior buys lower variance, not unbiasedness.
2§06.4 — what the posterior covariance depends on
Two labs measure the same slope at the same inputs, with the same known noise variance and the same Gaussian prior. Their readings $y$ differ a lot, and one lab's readings look surprising.
Find(a) True or false: the lab with the surprising readings ends up with the larger posterior covariance.
Given
same $X$, same $\sigma^2$, same prior $\mathcal N(\mu_0,\Sigma_0)$
different response vectors $y$
Hint 1/4
Separate what the posterior covariance is built from and what the posterior mean is built from.
Hint 2/4
$\Sigma_{\beta|D}=(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X)^{-1}$ and $\mu_{\beta|D}=\Sigma_{\beta|D}(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0)$.
Hint 3/4
The two labs share $X$, $\Sigma_D=\sigma^2I$ and $\Sigma_0$; only $y$ differs.
Hint 4/4
Their covariances are equal and only their means differ, so the statement is false.
Show solution
Reading off which symbols the formula contains is enough; no numbers are needed.
The difference between the labs shows up in the means only.
Answer $$\boxed{\text{False}}$$
Check
One-input check: $\Sigma_{\beta|D}=1/(1/\tau^2+\sum x_i^2/\sigma^2)$ has no $y_i$ in it.
When the noise variance is known, surprising readings move the estimate but do not change how sure we are.
3§06.2 — what a larger λ says about the prior
A team raises the ridge penalty $\lambda$ while the noise variance $\sigma^2$ of its sensor stays fixed and known. They want to explain the change to a colleague who thinks in priors.
Find(a) What does the larger $\lambda$ correspond to?
Answer $$\boxed{\text{smaller prior variance: a firmer belief near }0}$$
Check
Limits: $\lambda\to0$ gives an infinitely wide prior and least squares; $\lambda\to\infty$ pins the prior at $0$ and sends $\hat\beta\to0$.
Read $\lambda$ as confidence in zero: large $\lambda$, strong belief that the coefficients are small.
4§06.3 — what a zero from the lasso means
A lasso fit sets one coefficient exactly to $0$. A report then states that, under the matching Laplace prior, the posterior probability that this coefficient equals $0$ is high.
Find(a) Is the report's statement true or false?
Given
lasso = MAP estimate under i.i.d. $\mathrm{Lap}(0,2\sigma^2/\lambda)$ priors
the fitted coefficient is exactly $0$
Hint 1/4
Separate where a density is highest from how much probability sits at a single point.
Hint 2/4
For a continuous distribution $P(\beta_j=c)=\int_c^c p(\beta_j\mid D)\,d\beta_j=0$ for every $c$.
Hint 3/4
The Laplace prior and the Gaussian likelihood give a posterior with a density; its peak, the MAP estimate, has this coordinate at $0$.
Hint 4/4
The peak has $\beta_j=0$, but $P(\beta_j=0\mid D)=0$, so the statement is false.
Show solution
Two different questions hide in the report, where the density peaks and how much probability one point carries, so we answer them separately.
The zero is a coordinate of the most probable point, the mode of the posterior density.
Answer $$\boxed{\text{False}}$$
Check
Compare a Gaussian posterior: its mean is also its peak, and nobody reads the density there as the probability of that exact value.
Lasso zeros describe the posterior's peak, the MAP point estimate, not the probability that a coefficient is zero.
B · computation 8 questions
1§06.1 — maximum likelihood slope and a likelihood ratio
A centered data set with one input gives $\sum_i x_i^2=20$ and $\sum_i x_iy_i=14$. The noise is Gaussian with known variance $\sigma^2=2$.
Find
(a) Find $\hat\beta^{\mathrm{MLE}}$.
(b) By what factor is the data more probable under $\hat\beta^{\mathrm{MLE}}$ than under $\beta=0.5$?
Given
$\sum_i x_i^2=20$, $\sum_i x_iy_i=14$
$\sigma^2=2$
for one input, $\mathrm{RSS}(\beta)=\mathrm{RSS}(\hat\beta)+\big(\sum_ix_i^2\big)(\beta-\hat\beta)^2$ with $\hat\beta=\hat\beta^{\mathrm{LS}}$
Hint 1/4
Part (a) is the least squares slope; part (b) needs only the difference of two RSS values.
Hint 2/4
$$\frac{L(\hat\beta)}{L(\beta)}=\exp\Big(\frac{\mathrm{RSS}(\beta)-\mathrm{RSS}(\hat\beta)}{2\sigma^2}\Big),$$ and the RSS difference is $\sum x_i^2\,(\beta-\hat\beta)^2$.
Hint 3/4
Here $\hat\beta=14/20=0.7$, $\beta=0.5$, $\sum x_i^2=20$ and $\sigma^2=2$; the RSS difference is $20\,(0.2)^2=0.8$.
Hint 4/4
$\hat\beta^{\mathrm{MLE}}=0.7$ and the ratio is $e^{0.8/4}=e^{0.2}\approx1.221$.
Show solution
The constant term of $l(\beta)$ cancels in a ratio, so $\sum y_i^2$ is never needed.
With a smaller noise variance, say $\sigma^2=0.5$, the same RSS gap would give $e^{0.8}\approx2.23$: a sharper likelihood with the same peak.
Under Gaussian noise a likelihood ratio is the exponential of an RSS difference over $2\sigma^2$.
2§06.2 — from a prior's spread to a ridge estimate
Sensor noise has standard deviation $0.3$. Before seeing data, an engineer models each coefficient as $\mathcal N(0,0.15^2)$. A one-input centered data set gives $\sum_ix_i^2=12$ and $\sum_ix_iy_i=8$.
Find
(a) Find the ridge penalty $\lambda$ that matches the prior.
(b) Find the MAP slope and compare it with least squares.
Given
$\sigma=0.3$
prior $\mathcal N(0,\tau^2)$ with $\tau=0.15$
$\sum_ix_i^2=12$, $\sum_ix_iy_i=8$
Hint 1/4
Translate the prior into a penalty first; the estimate then follows from the ridge formula.
Hint 2/4
$\lambda=\sigma^2/\tau^2$, and for one input $\hat\beta^{\mathrm{Ridge}}=\sum x_iy_i/(\sum x_i^2+\lambda)$.
Hint 3/4
Here $\sigma^2=0.09$, $\tau^2=0.0225$, $\sum x_i^2=12$ and $\sum x_iy_i=8$.
Hint 4/4
$\lambda=4$ and $\hat\beta^{\mathrm{MAP}}=8/16=0.5$, against $8/12\approx0.667$ for least squares.
Show solution
The translation needs variances, so both spreads are squared before anything else.
Compare ridge at the same $\lambda$: its prior variance would be $\sigma^2/\lambda=0.25$, half the Laplace variance, yet the Laplace density at $0$ is higher, $1$ against $0.80$.
Same λ, different priors: the Laplace is wider on average but more peaked at zero.
4§06.3 — a one-coefficient lasso MAP by cases
One centered input gives $\sum_ix_i^2=8$ and $\sum_ix_iy_i=5$. The noise variance is $\sigma^2=1$ and the lasso penalty is $\lambda=6$.
Find
(a) Find the lasso MAP slope.
(b) Find the ridge MAP slope for the same $\lambda$.
(c) Find the smallest $\lambda$ that makes the lasso slope exactly $0$.
Shortcut: $(5-\lambda/2)/8=(5-3)/8=0.25$. The lasso moved the least squares slope $0.625$ down by the fixed amount $\lambda/(2\cdot8)=0.375$.
For one coefficient the lasso subtracts $\lambda/(2\sum x_i^2)$ from the least squares slope and stops at $0$.
5§06.4 — a posterior with a prior mean of 2
An earlier experiment suggests a slope near $2$, encoded as the prior $\mathcal N(2,0.25)$. New centered data with one input give $\sum_ix_i^2=4$ and $\sum_ix_iy_i=6$, with $\sigma^2=1$.
Find
(a) Find the posterior mean and variance.
(b) Where does the posterior mean sit between the least squares slope and the prior mean?
Limits: with a very wide prior the mean would approach $1.5$, with a very narrow one $2$; $1.75$ is between, as it must be.
When the prior and the data carry equal precision, the posterior mean is their midpoint.
6§06.4 — a two-coefficient posterior with an orthogonal design
Two centered inputs are uncorrelated in the data, so $X^TX$ is diagonal: $X^TX=\mathrm{diag}(2,8)$ and $X^Ty=(2,4)^T$. The noise variance is $\sigma^2=1$ and the prior is $\mathcal N(0,\tfrac12I_2)$, the ridge prior with $\lambda=2$.
Find
(a) Find $\Sigma_{\beta|D}$ and $\mu_{\beta|D}$.
(b) Compare $\mu_{\beta|D}$ with $\hat\beta^{\mathrm{LS}}$ coefficient by coefficient.
The better-measured coefficient, with $d_2=8$, shrinks less; as $\lambda\to0$ both factors go to $1$ and least squares returns.
Ridge shrinks weakly measured coefficients most, because there the prior's precision competes with little data precision.
7§06.6 — a two-coefficient prediction interval
A posterior from an orthogonal design is $\mu_{\beta|D}=(0.5,0.4)^T$, $\Sigma_{\beta|D}=\mathrm{diag}(0.25,0.1)$, and the noise variance is $\sigma^2=1$. A prediction is needed at the new input $x=(2,1)$.
Find
(a) Find the predictive mean and variance.
(b) Give the interval mean $\pm2$ standard deviations.
Split: noise $1$ and parameter part $1.1$, so slightly more than half the variance comes from not knowing $\beta$ at this input.
At inputs that load on weakly measured coefficients, the parameter part can rival the noise.
8§06.5 — two readings in sequence
A slope has the prior $\mathcal N(0,1)$ and the noise variance is $\sigma^2=0.25$. Reading 1 is $(x,y)=(1,\,0.9)$; reading 2 is $(2,\,1.5)$.
Find
(a) Update after reading 1.
(b) Update after reading 2.
(c) Confirm with the batch posterior.
Given
prior $m_0=0$, $v_0=1$
$\sigma^2=0.25$
readings $(1,\,0.9)$, then $(2,\,1.5)$
Hint 1/4
Run the update rule twice; the first posterior is the prior of the second step.
Hint 2/4
$\frac{1}{v_n}=\frac{1}{v_{n-1}}+\frac{x_n^2}{\sigma^2}$ and $m_n=v_n\big(\frac{m_{n-1}}{v_{n-1}}+\frac{x_ny_n}{\sigma^2}\big)$.
Hint 3/4
Reading 1, $(1,0.9)$, adds $1/0.25=4$ to the precision and $0.9/0.25=3.6$ to the evidence; reading 2, $(2,1.5)$, adds $4/0.25=16$ and $3/0.25=12$.
Hint 4/4
$\mathcal N(0.72,\,0.2)$ after reading 1, and $\mathcal N(15.6/21,\ 1/21)\approx\mathcal N(0.743,\,0.048)$ after reading 2, as in the batch computation.
Show solution
Each step needs only the previous posterior and one reading, so we never recompute from scratch.
Reading 2 sits at $x=2$, so it adds four times the precision of reading 1: $16$ against $4$.
Readings at larger $\lvert x\rvert$ carry more information about a slope, in proportion to $x^2$.
C · exam level 5 questions
1§06.3 — which prior produces a given penalty
A fitted model minimizes $\mathrm{RSS}(\beta)+3\sum_j\lvert\beta_j\rvert$. The noise is Gaussian with known variance $\sigma^2=1.5$, and the fit is to be read as a MAP estimate.
Find(a) Which i.i.d. prior on the coefficients makes this the MAP estimate?
Given
penalty $3\sum_j\lvert\beta_j\rvert$, so $\lambda=3$
$\sigma^2=1.5$
Hint 1/4
An absolute-value penalty points to one family of priors; then match the scale.
Hint 2/4
$\beta_j\sim\mathrm{Lap}(0,b)$ with $b=2\sigma^2/\lambda$ turns the MAP estimate into the lasso with penalty $\lambda$.
Hint 3/4
Here $\lambda=3$ and $\sigma^2=1.5$.
Hint 4/4
$b=2(1.5)/3=1$, so the prior is $\mathrm{Lap}(0,1)$.
Show solution
The penalty's shape fixes the family before any number is used.
Translate back: $2\sigma^2\cdot\lvert\beta_j\rvert/b=3\lvert\beta_j\rvert$, the given penalty.
Identify the family from the penalty's shape, then the scale from $2\sigma^2/\lambda$.
2§06.4 — more coefficients than readings
A model has two coefficients but only one reading so far: input $x_1=(1,2)$, response $y_1=3$, noise variance $\sigma^2=1$. The model has no intercept, and the prior is $\mathcal N(0,I_2)$.
Find
(a) Show that the maximum likelihood estimate is not unique.
(b) Find $\Sigma_{\beta|D}$ and $\mu_{\beta|D}$.
(c) Find the predictive variance at $x=(1,2)$ and at $x=(2,-1)$, and explain the difference.
Given
$X=(1\ \ 2)$, $y=3$
$\sigma^2=1$, prior $\mathcal N(0,I_2)$
Hint 1/4
With one equation and two unknowns the likelihood cannot single out a point; the prior can.
Hint 2/4
$L(\beta)$ depends on $\beta$ only through $\beta_1+2\beta_2$; the posterior uses $\Sigma_{\beta|D}^{-1}=I+X^TX/\sigma^2$.
Hint 3/4
Here $X^TX=\begin{pmatrix}1&2\\2&4\end{pmatrix}$, which is singular, and $X^Ty=(3,6)^T$; so $\Sigma_{\beta|D}^{-1}=\begin{pmatrix}2&2\\2&5\end{pmatrix}$.
Hint 4/4
$\mu_{\beta|D}=(0.5,1)^T$; the predictive variances are $1+\tfrac56$ at $(1,2)$ and $1+5=6$ at $(2,-1)$.
Show solution
The singular $X^TX$ blocks least squares, but the prior makes the posterior precision invertible, so we go straight to the posterior box.
$\mu_{\beta|D}$ lies on $\beta_1+2\beta_2=2.5$, not $3$: the prior pulls it toward $0$ along the measured direction, and along $(2,-1)$ it keeps the prior mean, since $2(0.5)-1=0$.
With more coefficients than readings the posterior still exists and is unique; directions the data never probed keep their prior uncertainty.
3§06.6 — find the error in a prediction interval
A student computes a 2-standard-deviation predictive interval for a line through the origin. The prior was $\mathcal N(0,1)$, $\sigma^2=0.25$, and after the data the posterior is $\mathcal N(0.8,\,0.04)$. The new input is $x=2$.
Find(a) Which step contains the error?
Given
Step 1: predictive mean $0.8\times2=1.6$.
Step 2: parameter part of the variance, $x^2\times1=4$, using the prior variance $1$.
Step 3: predictive variance $0.25+4=4.25$.
Step 4: interval $1.6\pm2\sqrt{4.25}=1.6\pm4.12$.
Hint 1/4
For each step, ask which distribution of β it should use now that the data are in.
Hint 2/4
$\mathrm{Var}(Y\mid x,D)=\sigma^2+x^2\,\Sigma_{\beta|D}$, with the posterior variance.
Hint 3/4
Here $\Sigma_{\beta|D}=0.04$, $\sigma^2=0.25$ and $x=2$; the student used $1$ in place of $0.04$.
Hint 4/4
Step 2 is wrong. Corrected: variance $0.25+0.16=0.41$ and interval $1.6\pm1.281$, that is $[0.319,\ 2.881]$.
Show solution
Checking each step against the formula is quicker than redoing the whole computation.
Sanity bound: the variance must lie between the noise alone, $0.25$, and the prior predictive, $4.25$; $0.41$ does.
Once data are in, every uncertainty about β comes from the posterior; the prior is used once, to build it.
4§06.4 — readings from sensors of different quality
A slope through the origin is measured with two sensors, and the prior is $\mathcal N(0,1)$. Reading A, $(x,y)=(1,0.5)$, comes from a precise sensor with $\sigma^2=0.25$; readings B, $(2,2)$, and C, $(1,3)$, come from a noisy one with $\sigma^2=1$.
Find(a) What is the posterior mean of the slope?
Given
prior $\mathcal N(0,1)$
A: $(1,\,0.5)$ with $\sigma_A^2=0.25$
B: $(2,\,2)$ and C: $(1,\,3)$, each with $\sigma^2=1$
$\Sigma_D=\mathrm{diag}(0.25,\,1,\,1)$
Hint 1/4
Each reading counts with its own noise variance: $\Sigma_D$ is diagonal but not a multiple of $I$.
Hint 2/4
$\Sigma_{\beta|D}^{-1}=\Sigma_0^{-1}+\sum_ix_i^2/\sigma_i^2$ and, with $\mu_0=0$, $\mu_{\beta|D}=\Sigma_{\beta|D}\sum_ix_iy_i/\sigma_i^2$.
Precision $10$, evidence $9$, posterior mean $0.9$.
Show solution
The general formula with $\Sigma_D^{-1}=\mathrm{diag}(4,1,1)$ weights each reading by one over its noise variance; for one coefficient it reduces to two sums.
Reading C alone suggests the slope $3$ with precision $1$; reading A suggests $0.5$ with precision $4$. The answer, $0.9$, leans toward the precise sensor, as it should.
With unequal noise, weight each reading by $1/\sigma_i^2$; the matrix $\Sigma_D$ does this automatically.
5§06.6 — the error bar after very many readings
Readings keep arriving at inputs spread evenly over $[-1,1]$, for a line through the origin with known noise variance $\sigma^2$ and a prior $\mathcal N(0,\tau^2)$ on the slope. Consider the predictive distribution at the fixed input $x^\ast=0.5$.
Find(a) What does the predictive variance at $x^\ast$ approach?
Given
$n\to\infty$ readings, inputs spread over $[-1,1]$
Each reading adds $x_i^2/\sigma^2\ge0$, and with inputs spread over $[-1,1]$ the sum grows about linearly in $n$.
Limit
$$\sigma^2+0.25\,\Sigma_{\beta|D}\to\sigma^2$$
The noise term does not depend on $n$.
Answer $$\boxed{\sigma^2}$$
Check
Cross-check with the sequential example: after $20$ readings $\Sigma_{\beta|D}\approx0.035$, so at $x^\ast=0.5$ the parameter part is already only $0.009$ against $\sigma^2=0.25$.
No amount of data removes the noise of the next reading; data only remove the doubt about β.
D · interleaved 4 questions
1§06.4 — a defect rate and a sensor slope
A factory's defect rate has the prior $\mathrm{Beta}(2,8)$, and $3$ of $10$ inspected items are defective. Separately, a sensor slope has the posterior $\mathcal N(0.6,\,0.1)$.
Find
(a) Find the MAP estimate and the posterior mean of the defect rate.
(b) Find the MAP estimate and the posterior mean of the slope.
(c) Why do the two estimates agree in one case and not in the other?
Given
defect rate: prior $\mathrm{Beta}(2,8)$; data: $3$ defective of $10$
$\mathrm{Beta}(a,b)$: mode $\frac{a-1}{a+b-2}$, mean $\frac{a}{a+b}$; Bernoulli data update it to $\mathrm{Beta}(a+N_1,\,b+N_0)$
slope: posterior $\mathcal N(0.6,\,0.1)$
Hint 1/4
Each part needs the posterior's peak and its average; find the posterior first where it is not given.
Hint 2/4
$\mathrm{Beta}(a,b)$ prior with $N_1$ ones and $N_0$ zeros gives $\mathrm{Beta}(a+N_1,\,b+N_0)$; a Gaussian's mode equals its mean.
Hint 3/4
Here $a=2$, $b=8$, $N_1=3$ and $N_0=7$, so the defect rate's posterior is $\mathrm{Beta}(5,15)$; the slope's posterior is $\mathcal N(0.6,0.1)$.
Hint 4/4
Defect rate: MAP $4/18\approx0.222$ and mean $5/20=0.25$. Slope: both $0.6$.
Show solution
The defect rate needs its posterior first; the slope's posterior is already given.
Defect rate
$$\mathrm{Beta}(2+3,\ 8+7)=\mathrm{Beta}(5,15)$$
Add the defective count to $a$ and the good count to $b$.
The defect rate's posterior is bounded at $0$ and has more mass to the right of its peak, which is why its mean exceeds its mode.
For Gaussian posteriors, as in Bayesian linear regression with a Gaussian prior, the MAP estimate and the posterior mean are the same vector.
2§06.2 — three penalties for one coefficient
One centered input with $\sum_ix_i^2=10$ is fixed; the labels follow $y_i=\beta^{\mathrm{true}}x_i+\varepsilon_i$ with $\beta^{\mathrm{true}}=0.5$ and noise variance $\sigma^2=1$. Ridge is run with $\lambda=0$, $2$ or $5$.
Find
(a) Find $E[\hat\beta^{\mathrm{Ridge}}]$ and $\mathrm{Var}(\hat\beta^{\mathrm{Ridge}})$ as functions of $\lambda$.
(b) Compute bias² plus variance for $\lambda=0,2,5$ and pick the best of the three.
Further reading: minimizing over every $\lambda$ gives $\lambda=\sigma^2/(\beta^{\mathrm{true}})^2=4$ with total $0.0714$, just below $0.0722$. That is the Bayesian $\lambda=\sigma^2/\tau^2$ with the prior variance equal to $(\beta^{\mathrm{true}})^2$.
The penalty that minimizes the expected error matches a prior whose spread is the size of the true coefficient.
3§06.4 — two reports on one slope
A centered one-input data set has $\sum_ix_i^2=8$, and the noise variance is $\sigma^2=2$. A frequentist reports the sampling variance of the least squares slope; a Bayesian with the prior $\mathcal N(0,1)$ reports a posterior variance.
Find
(a) Compute both variances.
(b) State what each one describes.
Given
$\sum_ix_i^2=8$, $\sigma^2=2$
$\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\sigma^2/\sum_ix_i^2$ over repeated data sets
prior $\mathcal N(0,1)$
Hint 1/4
One variance is about the estimator over repeated data sets, the other about β given this one data set.
Hint 2/4
$\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\sigma^2/\sum x_i^2$ and $\Sigma_{\beta|D}=1/(1/\tau^2+\sum x_i^2/\sigma^2)$.
Hint 3/4
Here $\sigma^2=2$, $\sum x_i^2=8$ and $\tau^2=1$.
Hint 4/4
$0.25$ and $1/(1+4)=0.2$.
Show solution
The two formulas answer different questions, so we compute each and then say what it describes.
Here the randomness is the noise in a new data set.
Bayesian
$$\Sigma_{\beta|D}=\tfrac{1}{1+8/2}=0.2$$
Here the randomness is our uncertainty about $\beta$; the prior adds precision $1$.
Answer $$\boxed{0.25\ \ \text{vs}\ \ 0.2}$$
Check
Flat-prior limit: with $\tau^2\to\infty$ the posterior variance becomes $1/(8/2)=0.25$, the frequentist number, although its meaning still differs.
A sampling variance describes an estimator; a posterior variance describes β.
4§06.6 — a combination of two correlated readings
A Gaussian random vector $Z=(Z_1,Z_2)$ has mean $(1,2)$ and covariance $\begin{pmatrix}2&1\\1&3\end{pmatrix}$. We need the distribution of $W=Z_1-Z_2+1$.
Find
(a) Find the distribution of $W$.
(b) Which quantities of the predictive formula play the roles of the vector $a$, the mean and the covariance here?
Given
$Z\sim\mathcal N\big((1,2)^T,\ \Sigma\big)$ with $\Sigma=\begin{pmatrix}2&1\\1&3\end{pmatrix}$
$W=a^TZ+c$ with $a=(1,-1)^T$ and $c=1$
Hint 1/4
$W$ is a linear function of a Gaussian vector plus a constant.
Hint 2/4
$a^TZ+c\sim\mathcal N(a^T\mu+c,\ a^T\Sigma a)$, with $a^T\Sigma a=\Sigma_{11}a_1^2+2\Sigma_{12}a_1a_2+\Sigma_{22}a_2^2$.
Hint 3/4
Here $a=(1,-1)$, $\mu=(1,2)$, $c=1$, $\Sigma_{11}=2$, $\Sigma_{12}=1$ and $\Sigma_{22}=3$.
Hint 4/4
Mean $1-2+1=0$ and variance $2-2+3=3$: $W\sim\mathcal N(0,3)$.
Show solution
The rule for linear functions of Gaussian vectors gives the answer in two lines.
the log-likelihood under Gaussian noise, and why its maximizer is least squares;
the two priors behind ridge and the lasso, with λ in each;
the posterior mean and covariance formulas, and the ridge special case;
the predictive mean and variance at a new input.
Then reopen the page and compare; whatever is missing is your reread list.
Derive $l(\beta)$ for Gaussian noise and show that its maximizer does not depend on $\sigma^2$?
c-ls-mle
Turn a prior $\mathcal N(0,\tau^2)$ into a ridge penalty, and a penalty into a prior variance?
c-ridge-map
Give the Laplace prior for a lasso penalty, and find a one-coefficient lasso estimate, including the zero case?
c-lasso-map
Compute $\mu_{\beta|D}$ and $\Sigma_{\beta|D}$ for two coefficients, with and without a nonzero prior mean?
c-bayes-posterior
Update a posterior one reading at a time and match the batch answer?
c-sequential
Compute a predictive mean, variance and interval, and say which part of the variance dominates at a given input?
c-predictive
Glossary (12 terms)
log-likelihoodlog olabilirlik
The natural logarithm of the likelihood, $l(\beta)=\log L(\beta)$. It has the same maximizer and turns a product over readings into a sum.
Minus the logarithm of the posterior density. Up to a constant it is the loss a MAP estimate minimizes, such as the ridge or lasso loss divided by $2\sigma^2$.
Gaussian prior
A normal distribution put on the coefficients before the data; with mean $0$ and variance $\sigma^2/\lambda$ its MAP estimate is ridge regression.
Laplace distributionLaplace dağılımı
$\mathrm{Lap}(\mu,b)$, with density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$, mean $\mu$ and variance $2b^2$; it has a sharp corner at its centre.
Laplace prior
An i.i.d. $\mathrm{Lap}(0,2\sigma^2/\lambda)$ distribution on the coefficients; its MAP estimate is the lasso.
modemod
The value at which a density is highest. The MAP estimate is the mode of the posterior; for a Gaussian the mode equals the mean.
biased estimatoryanlı kestirici
An estimator whose average over repeated data sets differs from the true value; ridge and the lasso are biased toward $0$ when $\lambda>0$.
Bayesian linear regressionBayesçi doğrusal regresyon
Linear regression that keeps the whole posterior distribution of the coefficients instead of one estimate; with a Gaussian prior and Gaussian noise the posterior is Gaussian.
precision
The inverse of a variance or of a covariance matrix. Large precision means small spread; in the Gaussian posterior the precisions of the prior and of the data add.
sequential updating
Using the posterior after some readings as the prior for the next ones; after all readings it equals the posterior computed from all of them at once.
posterior predictive distribution
The distribution of a new response given the data, $p(y\mid x,D)=\int p(y\mid x,\beta)\,p(\beta\mid D)\,d\beta$; here it is $\mathcal N\big(\mu_{\beta|D}^Tx,\ \sigma^2+x^T\Sigma_{\beta|D}x\big)$.
A prediction that treats an estimate of β as exact, so its error bar contains the noise only; it is too narrow far from the data.
What comes next
§07 · Generalized linear models
Here the noise was Gaussian, so minus the log-likelihood was a sum of squares. Next the response can be a yes/no label or a count, the likelihood changes shape, and the same recipe of writing it down and maximizing it leads to new models.
Sources
textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The weekly line names no chapter of the book, so no section numbers are cited here.
course materialEEE 485 lecture slides and lecture notes, chapter 6: Linear regression from a Bayesian perspective Scope, order of topics and notation (the LS, MLE, MAP, Ridge and Lasso estimates, the two losses, Lap(μ, b), the posterior mean and covariance, the predictive density) follow these materials; all explanations, data, figures and exercises here are original.
course materialEEE 485 syllabus page on STARS, Fall 2026-27, printed 21 September 2026 Assessment weights, the weekly topic list and the recommended books.
textbookK. P. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press, 2012 Recommended in the syllabus; the lecture's figures for Bayesian linear regression come from it.
textbookG. James, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning, Springer, 2013 Recommended in the syllabus; the lecture's picture of the ridge and lasso constraint regions comes from it.
standard resultCompleting the square for Gaussian densities; linear combinations of Gaussian vectors Standard results used in the derivations. Every number and figure on this page was computed for it, the simulated readings from a fixed random seed.