← back to EEE 485
Week 6110 min full read
6 concepts19 worked examples31 exercises5 exam-level6 figures
What are you here for?

06 Linear regression from a Bayesian perspective: least squares, ridge and lasso as most probable coefficients, and the full posterior

Start with this

One question before you read anything. Getting it wrong is the point: it shows you what this section is for.

§06.1 — a noisier sensor and the most likely slope

A slope is fitted to $50$ readings by maximizing the likelihood under Gaussian noise with known standard deviation $1$. The same readings are then analysed again, assuming the noise standard deviation is $2$.

Find(a) What happens to the most likely slope?
Given
  • the same $50$ readings both times

  • noise model $\mathcal N(0,\sigma^2)$: first $\sigma=1$, then $\sigma=2$

Hint 1/4

Ask whether $\sigma$ changes where the peaks, or only its height and width.

Hint 2/4

$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\sigma}-\frac{1}{2\sigma^2}\mathrm{RSS}(\beta)$.

Hint 3/4

Going from $\sigma=1$ to $\sigma=2$ changes the constant and multiplies the RSS term by $\frac14$; the $50$ readings, and so the RSS curve, stay the same.

Hint 4/4

The minimizer of the RSS does not change, so the most likely slope stays the same.

Show solution

Only where the two log-likelihoods peak answers the question, not their values, so we compare locations.

Write both log-likelihoods

$$l_1(\beta)=c_1-\tfrac12\,\mathrm{RSS}(\beta),\qquad l_2(\beta)=c_2-\tfrac18\,\mathrm{RSS}(\beta)$$

The constants $c_1$ and $c_2$ contain no $\beta$.

$$\arg\max l_1=\arg\max l_2=\arg\min\mathrm{RSS}(\beta)$$

Both are decreasing functions of the same RSS.

Answer $$\boxed{\text{the same slope, }\hat\beta^{\mathrm{LS}}}$$
Check

Edge case: as $\sigma\to\infty$ the log-likelihood flattens, but at every finite $\sigma$ its peak stays at the least squares slope.

Only a prior makes the noise level matter; that is what the ridge block adds.

Two groups fit a slope to the same three readings, with noise standard deviation 1. One minimizes squared misses plus 4 times the squared slope; the other has no penalty, only a belief that the slope is Gaussian around 0 with standard deviation 0.5, and reports the most probable slope. Both report 0.6; on any data they would agree.

By the end you can turn such a belief into a penalty and back, for ridge and for the lasso, and go past a single number: a full posterior for the coefficients and an error bar for every prediction.

In 60 seconds

Least squares, ridge and the lasso are the most probable coefficients under three models (Gaussian noise alone, plus a Gaussian prior, plus a Laplace prior); Bayesian linear regression keeps the whole Gaussian posterior, which gives every prediction an error bar.

Least squares = maximum likelihood
$$\varepsilon_i\sim\mathcal N(0,\sigma^2)\ \Rightarrow\ \arg\max_\beta l(\beta)=\arg\min_\beta\mathrm{RSS}(\beta)=(X^TX)^{-1}X^Ty$$

showing that least squares is the MLE; the known $\sigma^2$ never moves the answer

Ridge and lasso = MAP
$$\beta_j\sim\mathcal N\big(0,\tfrac{\sigma^2}{\lambda}\big)\Rightarrow\hat\beta^{\mathrm{Ridge}},\qquad \beta_j\sim\mathrm{Lap}\big(0,\tfrac{2\sigma^2}{\lambda}\big)\Rightarrow\hat\beta^{\mathrm{Lasso}}$$

translating a penalty into a prior and back

Gaussian posterior
$$\Sigma_{\beta|D}=\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)^{-1},\qquad \mu_{\beta|D}=\Sigma_{\beta|D}\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big)$$

a Gaussian prior and Gaussian noise with known variance: compute the whole posterior

Posterior predictive
$$Y\mid x,D\sim\mathcal N\big(\mu_{\beta|D}^Tx,\ \sigma^2+x^T\Sigma_{\beta|D}x\big)$$

a prediction with an honest error bar at a new input

Three most common mistakes
  1. Writing $\lambda=1/(\text{prior variance})$: the noise variance belongs in the translation, $\lambda=\sigma^2/(\text{prior variance})$ for ridge and $b=2\sigma^2/\lambda$ for the Laplace scale.

  2. Adding covariances instead of precisions: the posterior precision is $\Sigma_0^{-1}+X^T\Sigma_D^{-1}X$, and the posterior covariance is its inverse.

  3. Dropping $\sigma^2$ from the predictive variance: a new reading carries its own noise, so the variance is $\sigma^2+x^T\Sigma_{\beta|D}x$ and never below $\sigma^2$.

Two course documents give different weights:

  • Chapter 1 slides, undergraduate line: midterm 25, final 25, four quizzes 20, two-phase project 30.
  • STARS syllabus page printed on 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. It is the later document; confirm which split applies.
  • The STARS weekly list puts ridge, lasso and Bayesian linear regression in one line; the lecture slides split them over this chapter and the one before.
How much time do you have?
10 minutes

The penalty-to-prior translation, then the two formulas you compute with: the Gaussian posterior and the predictive distribution.

The 60-second card · Ridge regression is the most probable β under a Gaussian prior · Bayesian linear regression · The posterior predictive distribution · Formula card
45 minutes

Every result once with a worked example, then the posterior-and-prediction calculation from a full solution down to a bare problem.

The 60-second card · Least squares is maximum likelihood when the noise is Gaussian · Ridge regression is the most probable β under a Gaussian prior · The lasso is the most probable β under a Laplace prior · Bayesian linear regression · The posterior sharpens reading by reading · The posterior predictive distribution · Scaffolding comes off · Formula card
full read

Adds the derivations, the look-alike pairs and enough mixed practice to decide the method yourself.

The opening pages · Recall first · Least squares is maximum likelihood when the noise is Gaussian · Ridge regression is the most probable β under a Gaussian prior · The lasso is the most probable β under a Laplace prior · Bayesian linear regression · The posterior sharpens reading by reading · The posterior predictive distribution · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
  1. Derive the log-likelihood of a linear model with Gaussian noise of known variance, and show that its maximizer is the least squares estimate whatever $\sigma^2$ is.

  2. Show that ridge regression is the MAP estimate under an i.i.d. Gaussian prior, and convert between $\lambda$ and the prior variance with $\lambda=\sigma^2/(\text{prior variance})$.

  3. Show that the lasso is the MAP estimate under an i.i.d. Laplace prior with scale $2\sigma^2/\lambda$, and explain why the corner of that prior can set a coefficient exactly to $0$.

  4. Compute the Gaussian posterior mean and covariance from the general formula, for the ridge prior, for a prior mean away from zero and for unequal noise levels.

  5. Update a posterior one reading at a time and show that the result equals the posterior from all readings at once.

  6. Compute the posterior predictive mean and variance at a new input, and split the variance into noise and coefficient uncertainty.

Syllabus coverage

Linear regression from a Bayesian perspective — covered

  • Least squares as maximum likelihood under Gaussian noise of known variance
  • the likelihood and the log-likelihood
  • why the logarithm
  • the likelihood as the Gaussian vector density of y given β

The lecturer's chapter title. The chapter opens by rereading least squares as maximum likelihood; this block does that before any prior is added.

Ridge regression — covered

Ridge as the MAP estimate under an i.i.d. Gaussian prior with variance σ²/λ; λ as noise variance over prior variance; the ridge posterior.

The ridge estimator itself comes from the previous section; here it is read as the peak of a posterior, and its full posterior follows in the Bayesian linear regression block.

lasso — covered

The ; the lasso as the MAP estimate under an i.i.d. Laplace prior with scale 2σ²/λ; why the corner of that prior gives exact zeros; ridge and lasso as , least squares as unbiased.

The comparison of the three estimators from the lecture sits in this block's table.

Bayesian linear regression — covered

  • Point estimates against a whole posterior
  • Gaussian prior and Gaussian likelihood give a Gaussian posterior
  • the posterior for ridge
  • a prior mean away from zero
  • unequal noise levels
  • the posterior predictive distribution and its error bars

Spread over three blocks: the posterior, how it sharpens reading by reading, and the predictive distribution.

One-coefficient lasso by cases — off syllabus

The closed form for a single coefficient, found by minimizing on each side of zero and testing the corner.

Further reading: the lecture says the lasso has no closed form in general. The one-coefficient case needs only calculus and shows why the corner of the Laplace prior produces exact zeros.

Proof of the posterior formula — off syllabus

Completing the square in the exponent to show that the posterior is Gaussian with the stated mean and covariance.

Further reading: the lecture states the formula. The proof is included for completeness, and no exercise depends on reproducing it.

Recall first
Least squares

For centered data $D=\{(x_i,y_i)\}_{i=1}^n$, $\mathrm{RSS}(\beta)=\sum_i(y_i-\beta^Tx_i)^2=(y-X\beta)^T(y-X\beta)$ and $\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$. For one input, $\hat\beta^{\mathrm{LS}}=\sum_ix_iy_i/\sum_ix_i^2$.

Every estimate here is compared with it, and it is the maximum likelihood answer.

Ridge and lasso

$\hat\beta^{\mathrm{Ridge}}=\arg\min_\beta\,\mathrm{RSS}(\beta)+\lambda\sum_j\beta_j^2$, equal to $(X^TX+\lambda I)^{-1}X^Ty$. $\hat\beta^{\mathrm{Lasso}}=\arg\min_\beta\,\mathrm{RSS}(\beta)+\lambda\sum_j\lvert\beta_j\rvert$ has no closed form in general.

This section explains where the two penalties come from.

MLE and MAP

The maximum likelihood estimate maximizes $p(D\mid\theta)$. The MAP estimate maximizes $p(\theta\mid D)=p(D\mid\theta)\,p(\theta)/p(D)$, and $p(D)$ does not depend on $\theta$.

Least squares, ridge and the lasso are each one of these two.

Gaussian densities

$\mathcal N(\mu,\sigma^2)$ has density $\frac{1}{\sqrt{2\pi}\sigma}e^{-(z-\mu)^2/(2\sigma^2)}$. A Gaussian vector $Z\sim\mathcal N(\mu,\Sigma)$ in $p$ dimensions has density $$f_Z(z)=\frac{1}{(2\pi)^{p/2}\lvert\Sigma\rvert^{1/2}}\exp\Big(-\tfrac12(z-\mu)^T\Sigma^{-1}(z-\mu)\Big).$$

The likelihood and the ridge prior both have this form.

Linear combinations of Gaussians

If $Z\sim\mathcal N(\mu,\Sigma)$ then $a^TZ+c\sim\mathcal N(a^T\mu+c,\ a^T\Sigma a)$. A sum of independent Gaussians is Gaussian, with the means and the variances added.

They give the predictive distribution in one line.

Logs and maximizers

For $f>0$, $\arg\max f=\arg\max\log f=\arg\min(-\log f)$. Adding a constant, or multiplying by a positive constant, does not move a maximizer.

Every derivation here maximizes a log and drops constants.

2 × 2 inverse

$\begin{pmatrix}a&b\\c&d\end{pmatrix}^{-1}=\frac{1}{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}$ when $ad-bc\neq0$.

Every two-coefficient posterior on this page is found with it.

Try it yourself first (2 questions)
1§06.2 — a belief about a coefficient, written as a penalty

An engineer believes, before any data, that a coefficient is Gaussian with mean $0$ and variance $0.25$. The noise variance of the readings is $1$.

Find(a) If this belief were written as a penalty $\lambda\beta^2$ added to the RSS, what would $\lambda$ be?
Given
  • prior $\mathcal N(0,\,0.25)$

  • noise variance $\sigma^2=1$

Hint 1/4

Write the log of the prior and compare it with the RSS term of the log-likelihood.

Hint 2/4

$-\log p(\beta)=\frac{\beta^2}{2\tau^2}+\text{const}$ and $-\log p(D\mid\beta)=\frac{\mathrm{RSS}}{2\sigma^2}+\text{const}$; multiply both by $2\sigma^2$.

Hint 3/4

Here $\tau^2=0.25$ and $\sigma^2=1$.

Hint 4/4

$\lambda=\sigma^2/\tau^2=1/0.25=4$.

Show solution

Scaling by 2σ² puts the expression in the RSS-plus-penalty form, where the penalty can be read off.

Scale to the RSS

$$2\sigma^2\big(-\log p(\beta\mid D)\big)=\mathrm{RSS}(\beta)+\frac{\sigma^2}{\tau^2}\beta^2+\text{const}$$

Multiplying by $2\sigma^2=2$ turns the likelihood part into the plain RSS.

$$\lambda=\frac{1}{0.25}=4$$

Read off the coefficient of $\beta^2$.

Answer $$\boxed{\lambda=4}$$
Check

A firmer belief, variance $0.01$, would give $\lambda=100$: a tighter belief means a larger penalty, as it should.

λ is the noise variance over the prior variance; the ridge block derives this.

2§06.6 — where a prediction is least certain

A line through the origin is fitted to $100$ readings whose inputs all lie between $-1$ and $1$. Predictions with error bars are wanted at $x=0.5$, $x=1$ and $x=10$.

Find(a) Where should the error bar of the prediction be widest?
Given
  • $100$ readings, inputs in $[-1,1]$

  • prediction inputs $0.5$, $1$ and $10$

Hint 1/4

A prediction is uncertain for two reasons: the new reading's noise, and our remaining doubt about the slope.

Hint 2/4

For a line through the origin the doubt about the slope enters as $x^2\,\mathrm{Var}(\beta\mid D)$, while the noise part is the same at every input.

Hint 3/4

Compare $x^2$ at the three inputs: $0.25$, $1$ and $100$.

Hint 4/4

The error bar is widest at $x=10$.

Show solution

Only the part of the variance that changes with x can rank the three inputs, so we isolate it.

Split the variance

$$\mathrm{Var}(Y\mid x,D)=\sigma^2+x^2\,\mathrm{Var}(\beta\mid D)$$

The last block of this section derives this.

Compare the inputs

$$x^2=0.25,\ \ 1,\ \ 100$$

Only the second term changes with $x$, and it grows with $x^2$.

Answer $$\boxed{x=10}$$
Check

Edge case: at $x=0$ a line through the origin predicts $0$ with no slope uncertainty at all, so the error bar there is the noise alone.

Error bars should widen as a prediction moves away from the data; a single $\pm2\sigma$ everywhere hides that.

Notation
symbolreads asmeanswatch out
$D=\{(x_i,y_i)\}_{i=1}^n$

the data set D

$n$ centered pairs, $x_i\in\mathbb R^p$, $y_i\in\mathbb R$

Centered, so the model has no intercept.

$X,\ y$

X, y

the $n\times p$ design matrix with rows $x_i^T$, and the response vector

X is treated as fixed.

$\sigma^2$

sigma squared

the noise variance

Known throughout; never estimated here.

$L(\beta),\ l(\beta)$

the likelihood, the log-likelihood

$p(D\mid\beta)$ and its natural log

Functions of β, with the data held fixed.

$\hat\beta^{\mathrm{LS}},\ \hat\beta^{\mathrm{MLE}},\ \hat\beta^{\mathrm{MAP}}$

beta hat LS, MLE, MAP

least squares, maximum likelihood and maximum a posteriori estimates

LS and MLE coincide under Gaussian noise; MAP depends on the prior.

$\lambda,\ \tau^2$

lambda, tau squared

the penalty weight, and the prior variance of each coefficient

$\tau^2$ is our shorthand; the lecture writes the ridge prior variance as $\sigma^2/\lambda$.

$\mathrm{Lap}(\mu,b)$

Laplace with centre mu and scale b

density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$

Variance $2b^2$, not $b^2$.

$\mathcal N(\mu_0,\Sigma_0),\ \Sigma_D$

prior mean and covariance; likelihood covariance

the prior of β, and the covariance of $y$ given β

$\Sigma_D=\sigma^2I_n$ for i.i.d. noise.

$\mu_{\beta|D},\ \Sigma_{\beta|D}$

posterior mean, posterior covariance

the parameters of the Gaussian posterior

$\Sigma_{\beta|D}$ does not involve $y$.

$m_n,\ v_n$

m n, v n

posterior mean and variance of one slope after $n$ readings

$m_0=\mu_0$ and $v_0=\tau^2$.

$p(y\mid x,D)$

the posterior predictive density

the distribution of a new response at the input $x$

Its variance includes $\sigma^2$.

$\mathrm{Loss}_{\mathrm{Ridge}},\ \mathrm{Loss}_{\mathrm{Lasso}}$

ridge loss, lasso loss

$\mathrm{RSS}(\beta)+\lambda\lVert\beta\rVert_2^2$ and $\mathrm{RSS}(\beta)+\lambda\lVert\beta\rVert_1$

The regularization section wrote the estimates $\hat\beta^R$ and $\hat\beta^L$.

$I_p,\ I_n$

identity matrices

the $p\times p$ and $n\times n$ identities

Sizes follow the coefficients and the readings.

Conventions used here
No intercept.

As in the lecture, inputs and responses are centered, so the model $y_i=\beta^Tx_i+\varepsilon_i$ has no intercept. A few examples use a line that truly passes through the origin instead; the posterior formulas never need centering.

Without an intercept, the ridge and lasso priors treat every coefficient alike.

The noise variance is known.

Throughout, $\sigma^2$ is a given number and only $\beta$ is unknown. If $\sigma^2$ had to be estimated as well, the formulas here would change.

Every posterior and every prediction on this page relies on it.

Variances, not standard deviations.

$\mathcal N(\mu,v)$ always has the variance second. The ridge prior $\mathcal N(0,\sigma^2/\lambda)$ has variance $\sigma^2/\lambda$; when a question gives a prior standard deviation $\tau$, then $\lambda=\sigma^2/\tau^2$.

Mixing the two is the most common way to get λ wrong by a square.

Laplace scale.

$\mathrm{Lap}(\mu,b)$ uses the scale $b$: density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$, variance $2b^2$. Some books use the rate $1/b$; on this page $b$ is always the scale.

Scale and rate are reciprocals, and confusing them inverts the answer.

Precision.

The inverse of a variance or a covariance matrix is a precision. Posterior formulas add precisions and invert once at the end.

Adding covariances instead is the most common slip in this chapter.

Constants.

$p(\beta\mid D)\propto p(D\mid\beta)\,p(\beta)$ drops factors without $\beta$. After taking $-\log$ we write $+\text{const}$ where the lecture writes $\propto$; both mean that terms without $\beta$ were dropped.

A constant never moves a maximizer, so dropping it is safe.

Names of the estimates.

This chapter writes $\hat\beta^{\mathrm{LS}}$, $\hat\beta^{\mathrm{MLE}}$, $\hat\beta^{\mathrm{MAP}}$, $\hat\beta^{\mathrm{Ridge}}$ and $\hat\beta^{\mathrm{Lasso}}$; the regularization section wrote $\hat\beta^R$ and $\hat\beta^L$ for the last two.

Same objects, new names; the chapter's own notation is used here.

Rounding.

Intermediate steps keep at least four significant figures; final answers are rounded to three decimals unless an exact fraction is asked for.

Posterior means often differ from least squares only in the second decimal.

6.1Least squares is maximum likelihood when the noise is Gaussian

Turns the least squares recipe into a probability statement: the fitted coefficients are the ones that make the observed data most probable.

Least squares picked $\beta$ by a rule we chose, the smallest RSS; this block shows which noise model makes that rule the most probable answer.

Solvable with what we have
  • Fit least squares: $\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$, or $\sum_ix_iy_i/\sum_ix_i^2$ for one centered input.

  • Write the density of one Gaussian reading, $\frac{1}{\sqrt{2\pi}\sigma}e^{-r^2/(2\sigma^2)}$ for a residual $r$.

  • Maximize a likelihood for yes/no data, as in the Bayesian and frequentist section.

Not solvable yet
  • Say which slope makes a whole set of Gaussian readings most probable.

  • Say whether a noisier sensor should change that slope.

  • Compare two slopes on a thousand readings when the computer rounds the likelihood to 0.

Three centered readings, $(x,y)=(-1,-1),(0,1),(1,0)$, with Gaussian noise of standard deviation $1$. Each reading's density is largest when its own residual is $0$, so pick $\beta=1$: the first reading lands exactly on the line and its density reaches its top value, $0.399$.

Why it fails

One slope serves all three readings. At $\beta=1$ the product of the densities is $0.399\times0.242\times0.242=0.0234$; at $\beta=0.5$ it is $0.352\times0.242\times0.352=0.0300$, about $1.28$ times larger. Maximizing one factor is not maximizing the product.

TheoremLeast squares is the maximum likelihood estimate
Conditions
  • $y_i=(\beta^{\mathrm{true}})^Tx_i+\varepsilon_i$ for $i=1,\dots,n$, with the $x_i$ fixed

  • $\varepsilon_1,\dots,\varepsilon_n$ i.i.d. $\mathcal N(0,\sigma^2)$, and $\sigma^2$ is known

  • $X^TX$ invertible, for the closed form

$$\boxed{\begin{aligned}\textcolor{#1f6feb}{L(\beta)}&=p(D\mid\beta)=\prod_{i=1}^n\frac{1}{\sqrt{2\pi}\,\sigma}\exp\Big(-\frac{(y_i-\beta^Tx_i)^2}{2\sigma^2}\Big)\\ \textcolor{#1f6feb}{l(\beta)}&=n\log\frac{1}{\sqrt{2\pi}\,\sigma}-\frac{1}{2\sigma^2}\,\mathrm{RSS}(\beta)\\ \textcolor{#d1690a}{\hat\beta^{\mathrm{MLE}}}&=\arg\max_\beta\,l(\beta)=\arg\min_\beta\,\mathrm{RSS}(\beta)=\hat\beta^{\mathrm{LS}}\end{aligned}}$$

The probability of the data is a product of bell curves, one per reading. Its logarithm is a constant minus $\mathrm{RSS}/(2\sigma^2)$, so the slope that makes the data most probable is the least squares slope, whatever the known $\sigma$ is.

Why the log-likelihood peaks at least squares

Independence turns the joint density of $y_1,\dots,y_n$ into the product $L(\beta)$ in the box.

$\log$ is strictly increasing, so $L$ and $l=\log L$ peak at the same $\beta$; the log turns the product into a sum.

The first term of $l(\beta)$ has no $\beta$ in it, and $\frac{1}{2\sigma^2}>0$. So maximizing $l$ is minimizing $\mathrm{RSS}(\beta)$, and $\sigma^2$ only rescales the curve.

In vector form $\mathrm{RSS}(\beta)=(y-X\beta)^T(y-X\beta)$, so $L(\beta)$ is the $\mathcal N(X\beta,\sigma^2I_n)$ density evaluated at $y$: $$L(\beta)=\frac{1}{(2\pi)^{n/2}\sigma^n}\exp\Big(-\frac{1}{2\sigma^2}(y-X\beta)^T(y-X\beta)\Big).$$

Looks like this, but is not

The noise level multiplies the RSS in $l(\beta)$, so a noisier sensor should pull the most likely slope toward $0$.

$\sigma^2$ scales the RSS term but does not move its minimum. For the three readings, $\sigma=1$ and $\sigma=0.8$ both peak at $\beta=0.5$; only the height and the curvature change. The noise level starts to matter once a prior enters, in the next block.

readingresidual at β = 0.5densityresidual at β = 1density

$(-1,-1)$

$-0.5$

$0.352$

$0$

$0.399$

$(0,1)$

$1$

$0.242$

$1$

$0.242$

$(1,0)$

$-0.5$

$0.352$

$-1$

$0.242$

product

$0.0300$

$0.0234$

Moving from $0.5$ to $1$ lifts the first density from $0.352$ to $0.399$ but drops the third from $0.352$ to $0.242$; the product falls by a factor of $1.28$.

The slope that makes three Gaussian readings most probable

The readings $(x,y)=(-1,-1),(0,1),(1,0)$ follow $y_i=\beta x_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,1)$. Find the slope with the largest likelihood, and compare it with $\beta=1$, which puts the first reading exactly on the line.

Find$\hat\beta^{\mathrm{MLE}}$ and the likelihood ratio $L(\hat\beta^{\mathrm{MLE}})/L(1)$.
Given
  • readings $(-1,-1),\ (0,1),\ (1,0)$, centered

  • noise i.i.d. $\mathcal N(0,\sigma^2)$ with $\sigma=1$ known

Solution

We maximize $l(\beta)$, not $L(\beta)$: the log turns the product into a sum of squares we can differentiate, and it peaks at the same slope.

Write the log-likelihood

$$\mathrm{RSS}(\beta)=(-1+\beta)^2+1^2+(0-\beta)^2=2\beta^2-2\beta+2$$

Each residual is $y_i-\beta x_i$. The middle reading has $x=0$, so no slope can change its residual of $1$.

$$l(\beta)=3\log\tfrac{1}{\sqrt{2\pi}}-\tfrac12\,(2\beta^2-2\beta+2)$$

The box with $n=3$ and $\sigma=1$: a constant minus half the RSS.

Maximize

$$l'(\beta)=-\tfrac12\,(4\beta-2)=0\ \Rightarrow\ \beta=0.5$$

The parabola opens downward, so its one critical point is the maximum.

$$\frac{\sum_i x_iy_i}{\sum_i x_i^2}=\frac{(-1)(-1)+0+1\cdot0}{1+0+1}=\frac12$$

The least squares slope for one centered input gives the same number, as the box promises.

Compare with the naive slope

$$\frac{L(0.5)}{L(1)}=\exp\Big(\frac{\mathrm{RSS}(1)-\mathrm{RSS}(0.5)}{2}\Big)=e^{(2-1.5)/2}=e^{0.25}\approx1.284$$

The constants cancel in a ratio, so only the RSS difference matters.

Answer $$\boxed{\hat\beta^{\mathrm{MLE}}=\hat\beta^{\mathrm{LS}}=0.5,\qquad L(0.5)/L(1)\approx1.284}$$
Check

Direct products of the three densities: $0.352\times0.242\times0.352\approx0.0300$ at $\beta=0.5$ and $0.399\times0.242\times0.242\approx0.0234$ at $\beta=1$. Their ratio is $1.28$.

This settles the naive attempt above: one slope serves every reading, and the product of their densities is largest where the sum of squared residuals is smallest.

A thousand readings: comparing two slopes when the likelihood underflows

A sensor gives $n=1000$ centered readings with $\sum x_i^2=330$, $\sum x_iy_i=264$ and $\sum y_i^2=1161.2$. The noise is $\mathcal N(0,1)$. Compare the slopes $0.8$ and $0.9$ by their likelihood.

Find$l(0.8)$, the size of $L(0.8)$, and the ratio $L(0.8)/L(0.9)$.
Given
  • $n=1000$, $\sum_i x_i^2=330$, $\sum_i x_iy_i=264$, $\sum_i y_i^2=1161.2$

  • $\sigma=1$ known

Solution

We work only with $l(\beta)$ and the RSS summaries, because the product of $1000$ densities is too small for a computer to store.

RSS from the summaries

$$\mathrm{RSS}(\beta)=\textstyle\sum y_i^2-2\beta\sum x_iy_i+\beta^2\sum x_i^2$$

Expanding the square needs three sums, not the readings themselves.

$$\mathrm{RSS}(0.8)=1161.2-422.4+211.2=950$$

$2(0.8)(264)=422.4$ and $0.64\times330=211.2$.

$$\mathrm{RSS}(0.9)=1161.2-475.2+267.3=953.3$$

$2(0.9)(264)=475.2$ and $0.81\times330=267.3$.

Size of the likelihood

$$l(0.8)=-1000\log\sqrt{2\pi}-\tfrac{950}{2}\approx-918.94-475=-1393.94$$

Each reading contributes $-\log\sqrt{2\pi}\approx-0.9189$ to the constant.

$$L(0.8)=e^{-1393.94}=10^{-1393.94/2.3026}\approx4\times10^{-606}$$

The smallest positive double is about $5\times10^{-324}$, so a computer stores $L(0.8)$, and $L(0.9)$ too, as $0$.

Compare on the log scale

$$l(0.8)-l(0.9)=\tfrac{953.3-950}{2}=1.65\ \Rightarrow\ \tfrac{L(0.8)}{L(0.9)}=e^{1.65}\approx5.21$$

The shared constant cancels, so the difference of logs is exact even though both likelihoods underflow.

Answer $$\boxed{l(0.8)\approx-1393.94,\quad L(0.8)\approx4\times10^{-606},\quad L(0.8)/L(0.9)=e^{1.65}\approx5.21}$$
Check

$0.8$ is the least squares slope here, $264/330=0.8$, so it must win. Also $\mathrm{RSS}(0.9)-\mathrm{RSS}(0.8)$ must equal $330\,(0.9-0.8)^2=3.3$, and it does.

Compare likelihoods through differences of log-likelihoods: the constant $n\log\frac{1}{\sqrt{2\pi}\sigma}$ cancels and nothing underflows.

Checkpoint
§06.1 — the maximum likelihood slope and the noise level

A centered data set has $\sum_i x_i^2=8$ and $\sum_i x_iy_i=6$. The model is $y_i=\beta x_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,9)$, so $\sigma=3$ is known.

Find(a) Which slope maximizes the likelihood?
Given
  • $\sum_i x_i^2=8$, $\sum_i x_iy_i=6$

  • noise variance $\sigma^2=9$

Hint 1/4

The noise level appears in the log-likelihood; ask whether it changes where the maximum is, or only how high it is.

Hint 2/4

$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\sigma}-\frac{1}{2\sigma^2}\mathrm{RSS}(\beta)$, and for one centered input the RSS is smallest at $\sum_ix_iy_i/\sum_ix_i^2$.

Hint 3/4

Here $\sum_ix_i^2=8$, $\sum_ix_iy_i=6$ and $\sigma^2=9$, so the factor $\frac{1}{18}$ multiplies the RSS.

Hint 4/4

The maximizer is $6/8=0.75$, the same as for any other known $\sigma$.

Show solution

We differentiate $l(\beta)$ directly; the constant term drops out at once.

Locate the peak

$$l(\beta)=\text{const}-\tfrac{1}{18}\big(\textstyle\sum y_i^2-12\beta+8\beta^2\big)$$

The RSS expanded with the two given sums.

$$l'(\beta)=-\tfrac{1}{18}\,(16\beta-12)=0\ \Rightarrow\ \beta=0.75$$

The factor $\tfrac{1}{18}$ cannot make a nonzero bracket vanish, so it drops out.

Answer $$\boxed{\hat\beta^{\mathrm{MLE}}=0.75}$$
Check

With $\sigma^2=1$ the same equation reads $-\tfrac12(16\beta-12)=0$, again $\beta=0.75$.

A known noise variance never moves the maximum likelihood estimate; it only sets how sharply the likelihood peaks.

⚠ Maximizing one reading's density instead of the product

each factor is largest at a zero residual, and one reading can always be hit exactly

wrong$$\hat\beta=y_1/x_1=1\ \ \text{(reading 1 exact)}$$
right$$\hat\beta=\frac{\sum_i x_iy_i}{\sum_i x_i^2}=0.5$$
⚠ Letting the noise variance move the maximizer

$\sigma^2$ sits next to the RSS in $l(\beta)$, so it looks like part of the fit

wrong$$\hat\beta^{\mathrm{MLE}}=\frac{\sum_i x_iy_i}{\sum_i x_i^2+\sigma^2}$$
right$$\hat\beta^{\mathrm{MLE}}=\frac{\sum_i x_iy_i}{\sum_i x_i^2}\ \ \text{for every }\sigma^2>0$$

6.2Ridge regression is the most probable β under a Gaussian prior

Reads the ridge penalty as a belief: each coefficient is Gaussian around 0 with variance σ²/λ, and ridge returns the posterior's peak.

Maximum likelihood used the data alone; now we add a belief about $\beta$ before the data, and the most probable $\beta$ afterwards turns out to be the ridge estimate.

TheoremRidge regression is the MAP estimate under a Gaussian prior
Conditions
  • Gaussian likelihood as in the previous box, with $\sigma^2$ known

  • prior: $\beta_1,\dots,\beta_p$ i.i.d. $\mathcal N(0,\sigma^2/\lambda)$, that is $\beta\sim\mathcal N\big(0,\tfrac{\sigma^2}{\lambda}I_p\big)$, with $\lambda>0$

  • centered data, so there is no intercept left unpenalized

$$\boxed{\begin{aligned}-\log\textcolor{#d1690a}{p(\beta\mid D)}&=\frac{1}{2\sigma^2}\Big(\textcolor{#1f6feb}{\sum_{i=1}^n(y_i-\beta^Tx_i)^2}+\textcolor{#8250df}{\lambda\sum_{j=1}^p\beta_j^2}\Big)+\text{const}\\ \textcolor{#d1690a}{\hat\beta^{\mathrm{MAP}}}&=\arg\min_\beta\,\mathrm{Loss}_{\mathrm{Ridge}}(\beta,\lambda)=(X^TX+\lambda I_p)^{-1}X^Ty=\hat\beta^{\mathrm{Ridge}}\end{aligned}}$$

Minus the log-posterior is the ridge loss divided by $2\sigma^2$, plus terms without $\beta$. So the peak of the posterior is the ridge estimate, and the penalty weight is the noise variance over the prior variance: $\lambda=\sigma^2/(\text{prior variance})$.

From the log-posterior to the ridge loss

Bayes' rule gives $p(\beta\mid D)=p(D\mid\beta)\,p(\beta)/p(D)$, and $p(D)$ does not depend on $\beta$: $$-\log p(\beta\mid D)=-\log p(D\mid\beta)-\log p(\beta)+\log p(D).$$

The likelihood term is $\frac{1}{2\sigma^2}\sum_i(y_i-\beta^Tx_i)^2-n\log\frac{1}{\sqrt{2\pi}\sigma}$, from the previous box.

Each prior factor is $\frac{\sqrt\lambda}{\sqrt{2\pi}\sigma}\exp\big(-\frac{\lambda\beta_j^2}{2\sigma^2}\big)$, so $$-\log p(\beta)=\frac{\lambda}{2\sigma^2}\sum_{j=1}^p\beta_j^2-p\log\frac{\sqrt\lambda}{\sqrt{2\pi}\,\sigma}.$$

Adding, the $\beta$-terms are $\frac{1}{2\sigma^2}\mathrm{Loss}_{\mathrm{Ridge}}(\beta,\lambda)$ and the rest is constant in $\beta$. Minimizing it is minimizing the ridge loss, whose minimizer is $(X^TX+\lambda I_p)^{-1}X^Ty$.

Looks like this, but is not

A prior $\beta_j\sim\mathcal N(0,1/\lambda)$ gives ridge regression with penalty $\lambda$.

Only when $\sigma^2=1$. The log-prior brings $\beta_j^2/(2\cdot\text{variance})$ and the log-likelihood brings $\mathrm{RSS}/(2\sigma^2)$; matching them needs variance $\sigma^2/\lambda$. With variance $1/\lambda$ the MAP is ridge with penalty $\sigma^2\lambda$.

A penalty of 4 and a prior with standard deviation 0.5 give the same slope

Three centered readings $(x,y)=(-2,-2),(1,0),(1,2)$ have noise $\mathcal N(0,1)$. Group one minimizes $\mathrm{RSS}(\beta)+4\beta^2$. Group two puts the prior $\beta\sim\mathcal N(0,0.5^2)$ on the slope and reports the posterior's peak. Compute both.

Find$\hat\beta^{\mathrm{Ridge}}$ and $\hat\beta^{\mathrm{MAP}}$.
Given
  • readings $(-2,-2),\ (1,0),\ (1,2)$

  • $\sigma^2=1$ known

  • penalty $\lambda=4$; prior $\mathcal N(0,\tau^2)$ with $\tau=0.5$

Solution

With one coefficient every matrix is a number, so both answers come from setting a derivative to zero; no inverse is needed.

Data summaries

$$\textstyle\sum x_i^2=4+1+1=6,\qquad \sum x_iy_i=4+0+2=6$$

For one input these two sums carry everything the data say about the slope.

Group one: ridge

$$\tfrac{d}{d\beta}\big(\mathrm{RSS}(\beta)+4\beta^2\big)=-2\,(6-6\beta)+8\beta=0$$

The derivative of the RSS is $-2\big(\sum x_iy_i-\beta\sum x_i^2\big)$.

$$\hat\beta^{\mathrm{Ridge}}=\tfrac{6}{6+4}=0.6$$

The same as $(X^TX+\lambda)^{-1}X^Ty$ with $X^TX=6$ and $X^Ty=6$.

Group two: posterior peak

$$\log p(\beta\mid D)=\text{const}-\tfrac12\,\mathrm{RSS}(\beta)-\tfrac{\beta^2}{2(0.25)}$$

The log-likelihood with $\sigma^2=1$, plus the log of the $\mathcal N(0,0.25)$ prior.

$$\tfrac{d}{d\beta}:\ (6-6\beta)-4\beta=0\ \Rightarrow\ \hat\beta^{\mathrm{MAP}}=0.6$$

The prior contributes $-\beta/0.25=-4\beta$, exactly what the penalty contributes after dividing by $2\sigma^2$.

Why they agree

$$\lambda=\dfrac{\sigma^2}{\tau^2}=\dfrac{1}{0.25}=4$$

The penalty weight is the noise variance over the prior variance, as the box says.

Answer $$\boxed{\hat\beta^{\mathrm{Ridge}}=\hat\beta^{\mathrm{MAP}}=0.6}$$
Check

Weighted average: the data alone say $6/6=1$ with weight $6/\sigma^2=6$, the prior says $0$ with weight $1/\tau^2=4$, and $(6\cdot1+4\cdot0)/10=0.6$.

This is the opening puzzle: a penalty $\lambda$ is a prior variance $\sigma^2/\lambda$ in disguise, so the two groups solved the same problem.

From 'within ±0.4, almost surely' to a penalty weight

Before seeing data, an engineer expects about $19$ coefficients in $20$ to lie within $\pm0.4$ of $0$. The sensor noise has standard deviation $0.3$. Which ridge penalty encodes this belief, and what happens to it if the noise standard deviation doubles to $0.6$?

Find$\tau^2$, and $\lambda$ for $\sigma=0.3$ and for $\sigma=0.6$.
Given
  • belief: $\mathcal N(0,\tau^2)$ with about $19$ draws in $20$ inside $\pm0.4$

  • $\sigma=0.3$, then $\sigma=0.6$

  • a normal draw falls within $2$ standard deviations of its mean about $19$ times in $20$

Solution

We turn the interval into a standard deviation first, because the translation to λ needs the prior variance, not a range.

Belief to prior variance

$$2\tau=0.4\ \Rightarrow\ \tau=0.2,\qquad \tau^2=0.04$$

The range $\pm0.4$ is two standard deviations each side.

Prior variance to penalty

$$\lambda=\dfrac{\sigma^2}{\tau^2}=\dfrac{0.09}{0.04}=2.25$$

Match the box's prior $\mathcal N(0,\sigma^2/\lambda)$ to $\mathcal N(0,\tau^2)$.

Noisier sensor

$$\lambda=\dfrac{0.36}{0.04}=9$$

The belief is unchanged, but each reading now says less, so the same belief weighs four times as much against the data.

Answer $$\boxed{\tau^2=0.04;\qquad \lambda=2.25\ \text{for}\ \sigma=0.3,\qquad \lambda=9\ \text{for}\ \sigma=0.6}$$
Check

Translate back: $\lambda=9$ with $\sigma^2=0.36$ gives the prior variance $\sigma^2/\lambda=0.04$, the same belief, so $\pm2\tau=\pm0.4$ again.

$\lambda$ is a ratio of two variances: doubling the noise standard deviation multiplies it by $4$ for the same belief.

Checkpoint
§06.2 — the penalty that matches a Gaussian prior

A regression has Gaussian noise with $\sigma=2$, known. An analyst's prior puts each coefficient at $\mathcal N(0,0.5)$, independently.

Find(a) Which ridge penalty $\lambda$ makes $\hat\beta^{\mathrm{Ridge}}$ equal to $\hat\beta^{\mathrm{MAP}}$?
Given
  • $\sigma=2$, so $\sigma^2=4$

  • prior $\beta_j\sim\mathcal N(0,\,0.5)$, i.i.d.

Hint 1/4

The penalty has to reproduce the prior's term in $-\log p(\beta\mid D)$ once everything is multiplied by $2\sigma^2$.

Hint 2/4

The prior $\mathcal N(0,\sigma^2/\lambda)$ gives ridge with penalty $\lambda$, so $\lambda=\sigma^2/(\text{prior variance})$.

Hint 3/4

Here $\sigma^2=4$ and the prior variance is $0.5$.

Hint 4/4

$\lambda=4/0.5=8$.

Show solution

The box's prior has variance σ²/λ, so one equation fixes λ.

Solve

$$\frac{\sigma^2}{\lambda}=0.5\ \Rightarrow\ \lambda=\frac{4}{0.5}=8$$

Set the box's prior variance equal to the analyst's.

Answer $$\boxed{\lambda=8}$$
Check

Directly: $-\log p(\beta)$ contributes $\beta_j^2/(2\cdot0.5)=\beta_j^2$, and multiplying $-\log p(\beta\mid D)$ by $2\sigma^2=8$ turns that into $8\beta_j^2$.

Always translate with the variance of the noise, not its standard deviation.

⚠ Forgetting the noise variance in the translation

the prior alone looks like the whole story, so λ gets read as its precision

wrong$$\lambda=\frac{1}{\tau^2}$$
right$$\lambda=\frac{\sigma^2}{\tau^2}$$
⚠ Reading σ²/λ as a standard deviation

$\mathcal N(\mu,\cdot)$ is written with a variance, while many texts quote spreads as standard deviations

wrong$$\tau=\frac{\sigma^2}{\lambda}$$
right$$\tau^2=\frac{\sigma^2}{\lambda},\qquad \tau=\frac{\sigma}{\sqrt\lambda}$$

6.3The lasso is the most probable β under a Laplace prior

Swaps in a Laplace prior with scale 2σ²/λ; its corner at zero lets the lasso set coefficients exactly to 0.

Ridge came from a Gaussian prior; keep the same likelihood, change the prior's shape, and the MAP estimate becomes the lasso.

TheoremThe lasso is the MAP estimate under a Laplace prior
Conditions
  • Gaussian likelihood with $\sigma^2$ known, as before

  • prior: $\beta_1,\dots,\beta_p$ i.i.d. $\mathrm{Lap}(0,b)$ with $b=2\sigma^2/\lambda$

  • $Z\sim\mathrm{Lap}(\mu,b)$ has density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$, mean $\mu$ and variance $2b^2$

$$\boxed{\begin{aligned}\textcolor{#8250df}{p(\beta_j)}&=\frac{\lambda}{4\sigma^2}\exp\Big(-\frac{\lambda\lvert\beta_j\rvert}{2\sigma^2}\Big)\\ -\log\textcolor{#d1690a}{p(\beta\mid D)}&=\frac{1}{2\sigma^2}\Big(\textcolor{#1f6feb}{\sum_{i=1}^n(y_i-\beta^Tx_i)^2}+\textcolor{#8250df}{\lambda\sum_{j=1}^p\lvert\beta_j\rvert}\Big)+\text{const}\\ \textcolor{#d1690a}{\hat\beta^{\mathrm{MAP}}}&=\arg\min_\beta\,\mathrm{Loss}_{\mathrm{Lasso}}(\beta,\lambda)=\hat\beta^{\mathrm{Lasso}}\end{aligned}}$$

With a Laplace prior the log-prior is a sum of absolute values, so minus the log-posterior is the lasso loss over $2\sigma^2$ plus a constant. The lasso is the peak of that posterior; there is still no closed form in general.

From the Laplace prior to the lasso loss

With $b=2\sigma^2/\lambda$ the Laplace density becomes $$\frac{1}{2b}e^{-\lvert\beta_j\rvert/b}=\frac{\lambda}{4\sigma^2}e^{-\lambda\lvert\beta_j\rvert/(2\sigma^2)},$$ so $$-\log p(\beta)=\frac{\lambda}{2\sigma^2}\sum_{j=1}^p\lvert\beta_j\rvert-p\log\frac{\lambda}{4\sigma^2}.$$

Add the likelihood term $\frac{1}{2\sigma^2}\sum_i(y_i-\beta^Tx_i)^2-n\log\frac{1}{\sqrt{2\pi}\sigma}$ and $\log p(D)$. The part with $\beta$ is $\frac{1}{2\sigma^2}\mathrm{Loss}_{\mathrm{Lasso}}(\beta,\lambda)$.

Minimizing $-\log p(\beta\mid D)$ is therefore minimizing the lasso loss.

Looks like this, but is not

The lasso sets $\hat\beta_j=0$, so the posterior says this coefficient is $0$ with high probability.

The posterior is a density, and any single value, $0$ included, has probability $0$. The lasso reports where the density peaks. Under the same posterior there is, for instance, a positive probability that $\beta_j>0.1$.

estimatorpriorestimate, Σxy = 6estimate, Σxy = 1.5on average

least squares

none

$1$

$0.25$

unbiased

ridge

$\mathcal N(0,\,0.25)$

$0.6$

$0.15$

biased toward 0

lasso

$\mathrm{Lap}(0,\,0.5)$

$0.667$

$0$

biased toward 0

Both priors pull toward $0$: ridge scales the least squares slope by $6/10$, the lasso subtracts $1/3$ and stops at $0$. Only least squares is unbiased; the other two trade some bias for lower variance.

Lasso with λ = 4 on the three readings, read as a Laplace prior

Same readings as the ridge example: $\sum x_i^2=6$, $\sum x_iy_i=6$, $\sum y_i^2=8$, $\sigma^2=1$. Find the Laplace prior that makes the lasso with $\lambda=4$ a MAP estimate, and compute that estimate.

FindThe prior $\mathrm{Lap}(0,b)$ and $\hat\beta^{\mathrm{Lasso}}$.
Given
  • $\sum_i x_i^2=6$, $\sum_i x_iy_i=6$, $\sum_i y_i^2=8$

  • $\sigma^2=1$, $\lambda=4$

Solution

The absolute value has a corner at $0$, so we minimize separately on $\beta>0$ and on $\beta<0$, then check the corner. This rederives the one-predictor rule of the regularization section, further reading there and here; in general the lasso has no closed form.

The prior

$$b=\frac{2\sigma^2}{\lambda}=\frac{2}{4}=0.5,\qquad \mathrm{Var}=2b^2=0.5$$

The box's scale; the prior is $\mathrm{Lap}(0,0.5)$, with twice the variance of the ridge prior at the same $\lambda$.

Minimize on each side

$$\beta>0:\ \ \tfrac{d}{d\beta}\big(6\beta^2-12\beta+8+4\beta\big)=12\beta-8=0\ \Rightarrow\ \beta=\tfrac23$$

On $\beta>0$, $\lvert\beta\rvert=\beta$. The answer is positive, so it is a genuine candidate.

$$\beta<0:\ \ 12\beta-16=0\ \Rightarrow\ \beta=\tfrac43>0$$

On $\beta<0$, $\lvert\beta\rvert=-\beta$. The solution is not negative, so this side has no minimum inside it.

Check the corner

$$\text{right slope at }0:\ -12+4=-8<0$$

The loss still falls as $\beta$ leaves $0$ to the right, so $0$ is not the minimum.

Answer $$\boxed{\mathrm{Lap}(0,\,0.5);\qquad \hat\beta^{\mathrm{MAP}}=\hat\beta^{\mathrm{Lasso}}=\tfrac23\approx0.667}$$
Check

Shortcut for one coefficient: $\big(\sum x_iy_i-\lambda/2\big)/\sum x_i^2=(6-2)/6=\tfrac23$ when $\sum x_iy_i>\lambda/2$. It sits between $0$ and the least squares slope $1$, as a shrunken estimate must.

The lasso subtracted a fixed $\lambda/(2\sum x_i^2)=\tfrac13$ from the least squares slope, while ridge with the same $\lambda$ multiplied it by $0.6$.

Weak evidence: the most probable slope is exactly zero

Now the readings are $(x,y)=(-2,-0.5),(1,0),(1,0.5)$, with $\sigma^2=1$ and $\lambda=4$ as before. Find the lasso MAP and the ridge MAP.

Find$\hat\beta^{\mathrm{Lasso}}$ and $\hat\beta^{\mathrm{Ridge}}$.
Given
  • $\sum_i x_i^2=6$, $\sum_i x_iy_i=1+0+0.5=1.5$

  • $\sigma^2=1$, $\lambda=4$

Solution

The data now pull only weakly, so we test the corner first: if the loss rises on both sides of $0$, nothing else needs checking.

Slopes of the lasso loss at 0

$$\mathrm{Loss}_{\mathrm{Lasso}}(\beta)=6\beta^2-3\beta+\textstyle\sum y_i^2+4\lvert\beta\rvert$$

$\mathrm{RSS}(\beta)=\sum y_i^2-2(1.5)\beta+6\beta^2$.

$$\text{right slope: }-3+4=1>0,\qquad \text{left slope: }-3-4=-7<0$$

On the right $\lvert\beta\rvert$ adds $+4$, on the left $-4$; the loss rises when leaving $0$ either way.

Compare with ridge

$$\hat\beta^{\mathrm{Ridge}}=\tfrac{1.5}{6+4}=0.15$$

The squared penalty has slope $0$ at $\beta=0$, so any evidence moves the estimate a little.

Answer $$\boxed{\hat\beta^{\mathrm{Lasso}}=0,\qquad \hat\beta^{\mathrm{Ridge}}=0.15}$$
Check

Threshold check: the lasso slope is $0$ exactly when $\lvert\sum x_iy_i\rvert\le\lambda/2$, and here $1.5\le2$. The least squares slope, $1.5/6=0.25$, is small too.

The corner of the Laplace prior at $0$ absorbs weak evidence completely; that is how the lasso selects variables.

Checkpoint
§06.3 — the Laplace prior behind a lasso penalty

A lasso fit uses $\lambda=2$ on centered data with Gaussian noise of variance $\sigma^2=0.5$. We want the prior that makes this lasso fit a MAP estimate.

Find(a) Which i.i.d. prior on each coefficient does it?
Given
  • $\lambda=2$

  • $\sigma^2=0.5$, known

Hint 1/4

Match the exponent of the prior with the penalty term of $-\log p(\beta\mid D)$.

Hint 2/4

$\mathrm{Lap}(0,b)$ has $-\log p(\beta_j)=\lvert\beta_j\rvert/b+\text{const}$, and the lasso needs $\lambda\lvert\beta_j\rvert/(2\sigma^2)$, so $b=2\sigma^2/\lambda$.

Hint 3/4

Here $\sigma^2=0.5$ and $\lambda=2$.

Hint 4/4

$b=2(0.5)/2=0.5$, so the prior is $\mathrm{Lap}(0,0.5)$.

Show solution

The two expressions must penalize one unit of the absolute value by the same amount, so we equate their coefficients.

Match

$$\frac{\lvert\beta_j\rvert}{b}=\frac{\lambda\lvert\beta_j\rvert}{2\sigma^2}\ \Rightarrow\ b=\frac{2\sigma^2}{\lambda}=\frac{2(0.5)}{2}=0.5$$

Both sides must charge the same amount per unit of $\lvert\beta_j\rvert$.

Answer $$\boxed{\beta_j\ \text{i.i.d.}\ \mathrm{Lap}(0,\,0.5)}$$
Check

Plug back: $\frac{1}{2b}e^{-\lvert\beta_j\rvert/b}=e^{-2\lvert\beta_j\rvert}$, and $\lambda/(2\sigma^2)=2/1=2$ matches the coefficient in the exponent.

The Laplace scale uses $2\sigma^2/\lambda$; the ridge variance uses $\sigma^2/\lambda$. Keep the two apart.

⚠ Using the ridge pattern σ²/λ for the Laplace scale

the ridge translation is fresh in mind, and both formulas divide by λ

wrong$$b=\frac{\sigma^2}{\lambda}$$
right$$b=\frac{2\sigma^2}{\lambda}$$
⚠ Taking b² as the Laplace variance

for the Gaussian the second parameter is the variance, and the habit carries over

wrong$$\mathrm{Var}(Z)=b^2$$
right$$\mathrm{Var}(Z)=2b^2$$
⚠ Differentiating |β| at 0 as if β were positive

the absolute value is usually met on one side of zero only

wrong$$\tfrac{d}{d\beta}\lvert\beta\rvert\,\Big|_{\beta=0}=1$$
right$$\text{slope }+1\text{ right of }0,\ \ -1\text{ left of }0:\ \text{check both sides}$$
−0.500.511.5−4−20246810slope β−log posterior, shiftedλ = 16: MAP stays exactly 0, ridge gives 0.27MAP = 0

The three readings of the ridge example: $\sum x_i^2=6$, $\sum x_iy_i=6$, $\sigma^2=1$. Each frame shows $\textcolor{#d1690a}{-\log p(\beta\mid D)}$, shifted to pass through $0$ at $\beta=0$, for a larger penalty. The minimum slides left as $\lambda$ grows, reaches the corner at $\lambda=12$ and stays there.

At the edges
λ = 0 MAP 1

No prior: the MAP estimate is the least squares slope.

λ ≥ 12 MAP 0

Once $\lambda$ reaches twice $\sum x_iy_i$, the corner holds the minimum at exactly $0$.

6.4Bayesian linear regression: the whole posterior, not just its peak

Keeps the whole posterior of β: with Gaussian prior and noise it is Gaussian, fixed by a mean and a covariance.

Least squares, ridge and the lasso each return one vector, the peak of a likelihood or a posterior; Bayesian linear regression keeps the whole posterior, and with it a measure of how sure we are.

TheoremGaussian prior and Gaussian likelihood give a Gaussian posterior
Conditions
  • prior $\beta\sim\mathcal N(\mu_0,\Sigma_0)$

  • likelihood $y\mid\beta\sim\mathcal N(X\beta,\Sigma_D)$ with $X$ fixed; $\Sigma_D=\sigma^2I_n$ for i.i.d. noise

  • $\Sigma_0$ and $\Sigma_D$ invertible covariance matrices

  • any prior is allowed in principle; the Gaussian pair is the one with this closed form

$$\boxed{\begin{aligned}\beta\mid D&\sim\mathcal N\big(\textcolor{#d1690a}{\mu_{\beta|D}},\,\textcolor{#d1690a}{\Sigma_{\beta|D}}\big)\\ \textcolor{#d1690a}{\Sigma_{\beta|D}}&=\big(\textcolor{#8250df}{\Sigma_0^{-1}}+\textcolor{#1f6feb}{X^T\Sigma_D^{-1}X}\big)^{-1}\\ \textcolor{#d1690a}{\mu_{\beta|D}}&=\textcolor{#d1690a}{\Sigma_{\beta|D}}\big(\textcolor{#1f6feb}{X^T\Sigma_D^{-1}y}+\textcolor{#8250df}{\Sigma_0^{-1}\mu_0}\big)\end{aligned}}$$

Precisions add: the posterior precision is the prior precision plus the data precision $X^T\Sigma_D^{-1}X$. The posterior mean is the posterior covariance times the sum of what the data say, $X^T\Sigma_D^{-1}y$, and what the prior says, $\Sigma_0^{-1}\mu_0$.

Why the posterior is Gaussian: completing the square (further reading)

Up to terms without $\beta$, $$-2\log p(\beta\mid D)=(y-X\beta)^T\Sigma_D^{-1}(y-X\beta)+(\beta-\mu_0)^T\Sigma_0^{-1}(\beta-\mu_0).$$

Expanding and keeping only the terms with $\beta$: $$\beta^T\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)\beta-2\beta^T\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big).$$

A Gaussian $\mathcal N(\mu,\Sigma)$ has $-2\log p=\beta^T\Sigma^{-1}\beta-2\beta^T\Sigma^{-1}\mu+\text{const}$. Matching the two lines gives $\Sigma^{-1}=\Sigma_0^{-1}+X^T\Sigma_D^{-1}X$ and $\Sigma^{-1}\mu=X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0$, the box.

A density proportional to $e^{-(\text{quadratic in }\beta)/2}$ with a positive definite matrix is that Gaussian; the normalizing constant takes care of itself.

Looks like this, but is not

The posterior covariance depends on the readings $y$, so a surprising set of readings leaves us less sure about $\beta$.

$\Sigma_{\beta|D}=(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X)^{-1}$ has no $y$ in it. With $\sigma^2$ known, how sure we end up depends on where we measured and how noisy the sensor is, not on what the readings turned out to be; the readings move only the mean.

Three readings or twenty-three: one ridge estimate, two posteriors

Data set A is the three readings with $\sum x_i^2=6$, $\sum x_iy_i=6$. Data set B has $23$ readings with $\sum x_i^2=46$, $\sum x_iy_i=30$. Both use $\sigma^2=1$ and the prior $\mathcal N(0,0.25)$, which is ridge with $\lambda=4$. Compare the ridge estimates and the posteriors.

Find$\hat\beta^{\mathrm{Ridge}}$ and $\mathcal N(\mu_{\beta|D},\Sigma_{\beta|D})$ for A and for B.
Given
  • A: $\sum x_i^2=6$, $\sum x_iy_i=6$

  • B: $\sum x_i^2=46$, $\sum x_iy_i=30$

  • $\sigma^2=1$, prior $\mathcal N(0,\,0.25)$

Solution

With one coefficient the box's matrices are numbers: precisions add, and the mean is the posterior variance times the summed evidence.

Posterior precisions

$$A:\ \tfrac{1}{0.25}+\tfrac{6}{1}=10,\qquad B:\ \tfrac{1}{0.25}+\tfrac{46}{1}=50$$

Prior precision $\Sigma_0^{-1}=4$ plus data precision $X^T\Sigma_D^{-1}X=\sum x_i^2/\sigma^2$.

Means and variances

$$A:\ \mu=\tfrac{6}{10}=0.6,\ \ \Sigma=0.1;\qquad B:\ \mu=\tfrac{30}{50}=0.6,\ \ \Sigma=0.02$$

$\mu_0=0$, so only $X^T\Sigma_D^{-1}y=\sum x_iy_i/\sigma^2$ is left in the bracket.

$$\text{sd: }\sqrt{0.1}\approx0.316\quad\text{vs}\quad\sqrt{0.02}\approx0.141$$

The posterior standard deviation is what a point estimate cannot report.

Answer $$\boxed{\hat\beta^{\mathrm{Ridge}}=0.6\ \text{for both};\quad A:\ \mathcal N(0.6,\,0.1),\quad B:\ \mathcal N(0.6,\,0.02)}$$
Check

Ridge formula directly: $6/(6+4)=0.6$ and $30/(46+4)=0.6$. The standard deviations differ by the factor $\sqrt{50/10}=\sqrt5\approx2.24$.

Two fits can report the same coefficient with very different certainty; only the posterior variance tells them apart.

The ridge posterior for two coefficients, by hand

Four centered readings with two inputs: rows $x_i=(1,1), \allowbreak (1,0), \allowbreak (-1,0), \allowbreak (-1,-1)$ and $y=(2,0,-1,-1)$. The noise is $\mathcal N(0,1)$ and the prior is $\mathcal N(0,I_2)$, the ridge prior with $\lambda=1$. Find the posterior and compare its mean with least squares.

Find$\Sigma_{\beta|D}$, $\mu_{\beta|D}$ and $\hat\beta^{\mathrm{LS}}$.
Given
  • $X=\begin{pmatrix}1&1\\1&0\\-1&0\\-1&-1\end{pmatrix}$, $y=(2,\,0,\,-1,\,-1)^T$

  • $\sigma^2=1$; prior $\mu_0=0$, $\Sigma_0=I_2$

Solution

In the ridge case $\Sigma_0=(\sigma^2/\lambda)I$ and $\Sigma_D=\sigma^2I$, so the box becomes $\Sigma_{\beta|D}=\sigma^2(X^TX+\lambda I)^{-1}$ and $\mu_{\beta|D}=(X^TX+\lambda I)^{-1}X^Ty$. Only $X^TX$ and $X^Ty$ are needed.

Data summaries

$$X^TX=\begin{pmatrix}4&2\\2&2\end{pmatrix},\qquad X^Ty=\begin{pmatrix}4\\3\end{pmatrix}$$

Sums of squares and products of the columns; for example the $(1,2)$ entry is $1+0+0+1=2$.

Posterior covariance

$$\Sigma_{\beta|D}^{-1}=I_2+X^TX=\begin{pmatrix}5&2\\2&3\end{pmatrix},\qquad \det=11$$

Prior precision $I_2$ plus data precision $X^TX/\sigma^2$.

$$\Sigma_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}3&-2\\-2&5\end{pmatrix}$$

Swap the diagonal, negate the off-diagonal, divide by the determinant.

Posterior mean

$$\mu_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}3\cdot4-2\cdot3\\-2\cdot4+5\cdot3\end{pmatrix}=\tfrac{1}{11}\begin{pmatrix}6\\7\end{pmatrix}\approx\begin{pmatrix}0.545\\0.636\end{pmatrix}$$

$\mu_0=0$, so the bracket is $X^Ty/\sigma^2$.

Least squares for comparison

$$\hat\beta^{\mathrm{LS}}=\tfrac14\begin{pmatrix}2&-2\\-2&4\end{pmatrix}\begin{pmatrix}4\\3\end{pmatrix}=\begin{pmatrix}0.5\\1\end{pmatrix}$$

$(X^TX)^{-1}$ with $\det X^TX=4$.

Answer $$\boxed{\Sigma_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}3&-2\\-2&5\end{pmatrix},\qquad \mu_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}6\\7\end{pmatrix}=\hat\beta^{\mathrm{Ridge}}\ (\lambda=1)}$$
Check

$\mu_{\beta|D}$ must solve $(X^TX+I)\mu=X^Ty$: $5\cdot\tfrac{6}{11}+2\cdot\tfrac{7}{11}=4$ and $2\cdot\tfrac{6}{11}+3\cdot\tfrac{7}{11}=3$. The posterior variances, $0.27$ and $0.45$, are below the prior's $1$ and below the least squares variances $0.5$ and $1$.

One 2 × 2 inverse and two matrix-vector products; for many coefficients this is a job for a linear solver.

For a Gaussian posterior the mean is also the peak, so the posterior mean is the MAP estimate, here the ridge estimate.

A prior centred away from zero: an earlier study said about 2

An earlier study suggests the slope is near $2$, so the prior is $\beta\sim\mathcal N(2,0.5)$. The new readings are the three of the ridge example, $\sum x_i^2=6$, $\sum x_iy_i=6$, with $\sigma^2=1$. Find the posterior.

Find$\mu_{\beta|D}$ and $\Sigma_{\beta|D}$.
Given
  • prior $\mathcal N(\mu_0,\Sigma_0)=\mathcal N(2,\,0.5)$

  • $\sum x_i^2=6$, $\sum x_iy_i=6$, $\sigma^2=1$

Solution

Now $\mu_0\neq0$, so the term $\Sigma_0^{-1}\mu_0$ in the box no longer vanishes; ridge cannot express this case.

Precision

$$\Sigma_{\beta|D}^{-1}=\tfrac{1}{0.5}+\tfrac{6}{1}=8,\qquad \Sigma_{\beta|D}=0.125$$

Prior precision $2$ plus data precision $6$.

Mean

$$\mu_{\beta|D}=0.125\,\big(6+2\cdot2\big)=0.125\times10=1.25$$

Data term $\sum x_iy_i/\sigma^2=6$ plus prior term $\Sigma_0^{-1}\mu_0=2\times2=4$.

Answer $$\boxed{\beta\mid D\sim\mathcal N(1.25,\ 0.125)}$$
Check

As a weighted average: the data alone say $6/6=1$ with weight $6$, the prior says $2$ with weight $2$, and $(6\cdot1+2\cdot2)/8=1.25$, between the two and closer to the heavier weight.

The posterior mean is a precision-weighted average of the prior mean and the least squares answer; a prior centred at $0$ is the special case ridge uses.

Checkpoint
§06.4 — a one-coefficient Gaussian posterior

One centered input gives $\sum_i x_i^2=8$ and $\sum_i x_iy_i=4$, with noise variance $\sigma^2=1$. The prior on the slope is $\mathcal N(0,0.5)$.

Find(a) Which is the posterior of the slope?
Given
  • $\sum_i x_i^2=8$, $\sum_i x_iy_i=4$

  • $\sigma^2=1$

  • prior $\mathcal N(0,\,0.5)$

Hint 1/4

Two numbers are needed, a variance and a mean; the variance comes first because the mean uses it.

Hint 2/4

$\Sigma_{\beta|D}^{-1}=\frac{1}{\text{prior variance}}+\frac{\sum x_i^2}{\sigma^2}$ and, with $\mu_0=0$, $\mu_{\beta|D}=\Sigma_{\beta|D}\,\frac{\sum x_iy_i}{\sigma^2}$.

Hint 3/4

Here the prior variance is $0.5$, $\sum x_i^2=8$, $\sum x_iy_i=4$ and $\sigma^2=1$, so the precision is $2+8$.

Hint 4/4

$\Sigma_{\beta|D}=0.1$ and $\mu_{\beta|D}=0.1\times4=0.4$: the posterior is $\mathcal N(0.4,0.1)$.

Show solution

The mean formula contains the posterior variance, so we find that first.

Precision

$$\Sigma_{\beta|D}^{-1}=\tfrac{1}{0.5}+8=10$$

Prior precision plus data precision.

Mean

$$\mu_{\beta|D}=\tfrac{1}{10}\cdot4=0.4$$

Posterior variance times $\sum x_iy_i/\sigma^2$.

Answer $$\boxed{\mathcal N(0.4,\ 0.1)}$$
Check

Ridge check: $\lambda=\sigma^2/0.5=2$ gives $4/(8+2)=0.4$, the same mean.

Compute the posterior variance first; the mean is that variance times the evidence.

⚠ Adding covariances instead of precisions

adding uncertainties sounds natural, but it is information that adds

wrong$$\Sigma_{\beta|D}=\Sigma_0+\big(X^T\Sigma_D^{-1}X\big)^{-1}$$
right$$\Sigma_{\beta|D}=\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)^{-1}$$
⚠ Dropping the prior-mean term

in the ridge case $\mu_0=0$ and the term vanishes, so it is easy to forget it exists

wrong$$\mu_{\beta|D}=\Sigma_{\beta|D}X^T\Sigma_D^{-1}y\ \ \text{although}\ \mu_0\neq0$$
right$$\mu_{\beta|D}=\Sigma_{\beta|D}\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big)$$

6.5The posterior sharpens reading by reading: yesterday's posterior is today's prior

Updates the posterior one reading at a time: each reading adds x²/σ² to the precision, and the batch answer comes out.

The box computed the posterior from all readings at once; here the same posterior is built one reading at a time, the way data often arrive.

RuleSequential updating for one coefficient
Conditions
  • readings independent given $\beta$, Gaussian noise with known $\sigma^2$

  • $m_0=\mu_0$ and $v_0=\tau^2$, the prior mean and variance

  • several coefficients: the same step with matrices, as in the posterior box

$$\boxed{\begin{aligned}p(\beta\mid D_{1:n})&\propto p(y_n\mid\beta)\,\textcolor{#8250df}{p(\beta\mid D_{1:n-1})}\\ \textcolor{#d1690a}{\frac{1}{v_n}}&=\textcolor{#8250df}{\frac{1}{v_{n-1}}}+\frac{x_n^2}{\sigma^2},\qquad \textcolor{#d1690a}{m_n}=v_n\Big(\textcolor{#8250df}{\frac{m_{n-1}}{v_{n-1}}}+\frac{x_ny_n}{\sigma^2}\Big)\end{aligned}}$$

Multiply yesterday's posterior by the new reading's likelihood and you have today's posterior. In numbers, the new reading adds $x_n^2/\sigma^2$ to the precision, and the mean moves to a precision-weighted blend of the old mean and the new evidence.

Why the sequential answer equals the batch answer

Bayes' rule with everything conditioned on the first $n-1$ readings: $p(\beta\mid D_{1:n})\propto p(y_n\mid\beta,D_{1:n-1})\,p(\beta\mid D_{1:n-1})$, and given $\beta$ the new reading does not depend on the old ones.

Unrolled from $n=1$: $p(\beta\mid D_{1:n})\propto p(\beta)\prod_{i=1}^n p(y_i\mid\beta)$, the batch posterior. The prior enters once.

The numeric step is the posterior box with $\Sigma_0=v_{n-1}$, $\mu_0=m_{n-1}$ and a single row $x_n$.

Looks like this, but is not

Updating one reading at a time uses the prior once per reading, so the sequential posterior leans harder toward the prior than the batch one.

The prior enters once, at the first step; every later step uses the previous posterior, not the original prior. Unrolled, the chain is the prior times all $n$ likelihood factors, exactly the batch posterior.

nΣx²Σxyprecisionmeansd

0

0

0

1

0

1.000

1

1

0.5

5

0.400

0.447

2

1.25

0.85

6

0.567

0.408

5

2.723

2.25

11.89

0.757

0.290

10

4.182

3.56

17.73

0.803

0.237

20

6.859

5.785

28.44

0.814

0.188

The precision is $1+\sum x^2/0.25$ and only grows; the mean wanders early and settles near $0.8$. After $20$ readings the prior supplies $1$ of the precision $28.4$.

Two readings, one at a time and then both at once

Prior $\beta\sim\mathcal N(0,1)$, noise variance $\sigma^2=0.25$. Reading 1 is $(x,y)=(1,0.5)$ and reading 2 is $(-0.5,-0.7)$. Update after each reading, then compute the batch posterior from both.

Find$(m_1,v_1)$, $(m_2,v_2)$ and the batch posterior.
Given
  • prior $m_0=0$, $v_0=1$

  • $\sigma^2=0.25$

  • readings $(1,\,0.5)$ and $(-0.5,\,-0.7)$

Solution

We run the update rule twice and then the posterior box once; agreement between the two routes is the point of the example.

After reading 1

$$\tfrac{1}{v_1}=1+\tfrac{1^2}{0.25}=5,\qquad v_1=0.2$$

The reading at $x=1$ adds $4$ to the precision.

$$m_1=0.2\,\big(0+\tfrac{1\cdot0.5}{0.25}\big)=0.2\times2=0.4$$

The old mean is 0, so only the new evidence counts.

After reading 2

$$\tfrac{1}{v_2}=5+\tfrac{0.25}{0.25}=6,\qquad v_2=\tfrac16$$

At $x=-0.5$ the reading adds only $1$: inputs near $0$ say little about a slope.

$$m_2=\tfrac16\,\big(5\cdot0.4+\tfrac{(-0.5)(-0.7)}{0.25}\big)=\tfrac{2+1.4}{6}\approx0.567$$

Yesterday's posterior plays the prior: its precision $5$ times its mean $0.4$.

Batch, from both readings

$$\textstyle\sum x^2=1.25,\ \ \sum xy=0.85:\ \ \tfrac1v=1+\tfrac{1.25}{0.25}=6,\ \ m=\tfrac{0.85/0.25}{6}=\tfrac{3.4}{6}$$

The posterior box with the prior $\mathcal N(0,1)$ and both rows at once.

Answer $$\boxed{\mathcal N(0.4,\ 0.2)\ \to\ \mathcal N(0.567,\ 0.167),\ \ \text{equal to the batch posterior}}$$
Check

Reverse the order: reading 2 first gives precision $2$ and mean $\tfrac{1.4}{2}=0.7$; then reading 1 gives precision $6$ and mean $\tfrac{2\cdot0.7+2}{6}=\tfrac{3.4}{6}$, the same.

Sequential and batch updating are the same computation in a different order; use whichever matches how the data arrive.

How many readings at x = 1 halve the posterior standard deviation?

Same prior $\mathcal N(0,1)$ and $\sigma^2=0.25$, and every reading is taken at $x=1$. After $2$ readings the posterior standard deviation is $1/3$. How many readings bring it to $1/6$ or below?

FindThe smallest $n$ with posterior standard deviation at most $1/6$.
Given
  • prior variance $1$, $\sigma^2=0.25$

  • all readings at $x=1$

Solution

Standard deviations do not add, precisions do, so we translate the target into a precision first.

Precision after n readings

$$\tfrac{1}{v_n}=1+n\cdot\tfrac{1}{0.25}=1+4n$$

Each reading at $x=1$ adds $1/\sigma^2=4$.

$$n=2:\ \ v_2=\tfrac19,\ \ \text{sd}=\tfrac13$$

Matches the statement.

Target

$$\text{sd}\le\tfrac16\iff 1+4n\ge36\iff n\ge8.75$$

Halving the standard deviation needs four times the precision, from $9$ to $36$.

Answer $$\boxed{n=9\ \text{readings}}$$
Check

$n=8$ gives $1/\sqrt{33}\approx0.174>0.167$ and $n=9$ gives $1/\sqrt{37}\approx0.164$, so $9$ is the first to pass.

Precision grows linearly in the number of readings, so the error bar on β shrinks like one over the square root.

Checkpoint
§06.5 — one more reading

After some readings the posterior of a slope is $\mathcal N(0.5,\,0.05)$. One more reading arrives, $(x,y)=(1,\,1.1)$, from a sensor with noise variance $\sigma^2=0.2$.

Find(a) What is the posterior after the new reading?
Given
  • current posterior $\mathcal N(0.5,\,0.05)$

  • new reading $(1,\,1.1)$

  • $\sigma^2=0.2$

Hint 1/4

Treat the current posterior as the prior for one new reading.

Hint 2/4

$\frac{1}{v_{\text{new}}}=\frac{1}{v_{\text{old}}}+\frac{x^2}{\sigma^2}$ and $m_{\text{new}}=v_{\text{new}}\big(\frac{m_{\text{old}}}{v_{\text{old}}}+\frac{xy}{\sigma^2}\big)$.

Hint 3/4

Here $m_{\text{old}}=0.5$, $v_{\text{old}}=0.05$, $x=1$, $y=1.1$ and $\sigma^2=0.2$: the precisions are $20$ and $5$.

Hint 4/4

$v_{\text{new}}=1/25=0.04$ and $m_{\text{new}}=0.04\,(10+5.5)=0.62$.

Show solution

One step of the update rule does it; no earlier reading needs to be revisited.

Precision

$$\tfrac{1}{0.05}+\tfrac{1^2}{0.2}=20+5=25$$

Old precision plus the new reading's $x^2/\sigma^2$.

Mean

$$0.04\,\big(20\cdot0.5+\tfrac{1\cdot1.1}{0.2}\big)=0.04\times15.5=0.62$$

A precision-weighted blend of the old mean and the new evidence.

Answer $$\boxed{\mathcal N(0.62,\ 0.04)}$$
Check

The new mean lies between $0.5$ and the reading's own slope $1.1$, four times closer to $0.5$ because the old posterior carries four times the precision: $0.62-0.5=0.12$ against $1.1-0.62=0.48$.

In a sequential update the old posterior is weighted by its precision, never by a plain average.

⚠ Restarting from the original prior at each step

the prior is written in the problem, the intermediate posterior is not

wrong$$\tfrac{1}{v_2}=\tfrac{1}{v_0}+\tfrac{x_2^2}{\sigma^2}$$
right$$\tfrac{1}{v_2}=\tfrac{1}{v_1}+\tfrac{x_2^2}{\sigma^2}=\tfrac{1}{v_0}+\tfrac{x_1^2+x_2^2}{\sigma^2}$$
⚠ Averaging the old mean and the new reading's slope

an average is the familiar way to combine two numbers

wrong$$m_{\text{new}}=\tfrac12\big(m_{\text{old}}+\tfrac{y}{x}\big)$$
right$$m_{\text{new}}=v_{\text{new}}\Big(\tfrac{m_{\text{old}}}{v_{\text{old}}}+\tfrac{xy}{\sigma^2}\Big)$$
−101200.511.52slope βdensitydotted line: true slope 0.8n = 20: mean 0.81, sd 0.19prior

The same twenty readings. Step through the $\textcolor{#d1690a}{\text{posterior}}$ of the slope after $0$, $1$, $2$, $5$, $10$ and $20$ readings; the dashed curve is the $\textcolor{#8250df}{\text{prior}}$ $\mathcal N(0,1)$ and the dotted line the true slope $0.8$.

At the edges
n = 0 N(0, 1)

Nothing observed yet: the posterior is the prior.

n → ∞ sd → 0

With inputs spread over $[-1,1]$ each reading adds about $4/3$ to the precision on average, so the standard deviation shrinks like $1/\sqrt n$ and the mean settles at the true slope.

6.6The posterior predictive distribution: an error bar that knows where the data were

Gives each prediction an error bar: the new reading's noise $\sigma^2$ plus the coefficient uncertainty $x^T\Sigma_{\beta|D}x$ seen at input $x$.

A posterior for $\beta$ is not yet a statement about the next reading; averaging the model over the posterior gives one.

TheoremPosterior predictive distribution at a new input x
Conditions
  • posterior $\beta\mid D\sim\mathcal N(\mu_{\beta|D},\Sigma_{\beta|D})$

  • new reading $Y=\beta^Tx+\varepsilon$ with $\varepsilon\sim\mathcal N(0,\sigma^2)$ independent of $\beta$ and of $D$

  • $x$ may be a vector of features, such as $(x,x^2)$ for a degree-2 polynomial; the formula does not change

$$\boxed{\begin{aligned}p(y\mid x,D)&=\int p(y\mid x,\beta)\,\textcolor{#d1690a}{p(\beta\mid D)}\,d\beta\\ Y\mid x,D&\sim\mathcal N\big(\textcolor{#d1690a}{\mu_{\beta|D}^T}x,\ \sigma^2+x^T\textcolor{#d1690a}{\Sigma_{\beta|D}}\,x\big)\end{aligned}}$$

Predict with the posterior mean of the coefficients. The variance has two parts: $\sigma^2$, the noise any new reading carries however much data we have, and $x^T\Sigma_{\beta|D}x$, what we still do not know about $\beta$, seen through the new input.

Why the predictive distribution is this Gaussian

Given $D$, $x^T\beta$ is a linear function of the Gaussian vector $\beta$, so $x^T\beta\sim\mathcal N\big(x^T\mu_{\beta|D},\,x^T\Sigma_{\beta|D}x\big)$.

$\varepsilon$ is Gaussian and independent of $\beta$, and a sum of independent Gaussians is Gaussian with the means and the variances added. So $Y=x^T\beta+\varepsilon$ has mean $\mu_{\beta|D}^Tx$ and variance $x^T\Sigma_{\beta|D}x+\sigma^2$.

Looks like this, but is not

With enough data the predictive variance goes to $0$, since the posterior of $\beta$ collapses onto one value.

Only $x^T\Sigma_{\beta|D}x$ shrinks with data. The $\sigma^2$ term belongs to the new reading itself: even with $\beta$ known exactly, a new reading scatters with variance $\sigma^2$.

Error bars at x = 0.5 and at x = 3 after two readings

Use the posterior after the two readings of the sequential example, $\beta\mid D\sim\mathcal N(17/30,\ 1/6)$, with $\sigma^2=0.25$. Find the predictive mean and standard deviation at $x=0.5$ and at $x=3$, and compare with a plug-in error bar that ignores the uncertainty in $\beta$.

FindThe predictive mean and standard deviation at both inputs, and the plug-in interval at $x=3$.
Given
  • $\mu_{\beta|D}=17/30\approx0.567$, $\Sigma_{\beta|D}=1/6$

  • $\sigma^2=0.25$

  • new inputs $x=0.5$ and $x=3$

Solution

With one coefficient $x^T\Sigma x=x^2\Sigma$, so the predictive variance is a one-line computation at each input.

At x = 0.5

$$\text{mean}=0.567\times0.5\approx0.283,\qquad \text{var}=0.25+0.25\cdot\tfrac16\approx0.292$$

Noise $0.25$ plus $x^2\Sigma_{\beta|D}=0.25/6$.

$$\text{sd}\approx0.540,\qquad \text{mean}\pm2\,\text{sd}\approx[-0.797,\ 1.364]$$

Close to the noise-only sd $0.5$: near the data the slope is known well enough.

At x = 3

$$\text{mean}=0.567\times3=1.70,\qquad \text{var}=0.25+9\cdot\tfrac16=1.75$$

Now $x^2\Sigma_{\beta|D}=1.5$ is six times the noise variance.

$$\text{sd}\approx1.323,\qquad \text{mean}\pm2\,\text{sd}\approx[-0.946,\ 4.346]$$

Far from the data the slope's uncertainty dominates.

Plug-in comparison

$$\text{plug-in: }1.70\pm2(0.5)=[0.70,\ 2.70]$$

Treating $\beta=0.567$ as exact keeps only the noise, a band about $2.6$ times too narrow at $x=3$.

Answer $$\boxed{x=0.5:\ 0.283\pm2(0.540);\qquad x=3:\ 1.70\pm2(1.323)}$$
Check

At $x=0$ the parameter term vanishes and the sd is exactly $\sigma=0.5$, the narrowest point of the band in the figure; both answers are above it, as they must be.

The predictive standard deviation is never below σ, and it grows with the distance from where the data were taken.

Two coefficients: which directions the data pinned down

Use the ridge posterior of the four two-input readings: $\mu_{\beta|D}=\tfrac{1}{11}(6,7)^T$ and $\Sigma_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}3&-2\\-2&5\end{pmatrix}$, with $\sigma^2=1$. Find the predictive mean and variance at $x=(1,1)$ and at $x=(1,-1)$.

Find$\mu_{\beta|D}^Tx$ and $\sigma^2+x^T\Sigma_{\beta|D}x$ at both inputs.
Given
  • $\mu_{\beta|D}=\tfrac{1}{11}(6,\,7)^T$

  • $\Sigma_{\beta|D}=\tfrac{1}{11}\begin{pmatrix}3&-2\\-2&5\end{pmatrix}$

  • $\sigma^2=1$; new inputs $(1,1)$ and $(1,-1)$

Solution

The quadratic form $x^T\Sigma x=\Sigma_{11}x_1^2+2\Sigma_{12}x_1x_2+\Sigma_{22}x_2^2$ is quicker by hand than two matrix products.

At (1, 1)

$$\mu^Tx=\tfrac{6+7}{11}\approx1.182,\qquad x^T\Sigma x=\tfrac{3-4+5}{11}=\tfrac{4}{11}$$

The cross term $2\Sigma_{12}x_1x_2=2\big(-\tfrac{2}{11}\big)(1)(1)$ is negative here.

$$\text{var}=1+\tfrac{4}{11}\approx1.364,\qquad \text{sd}\approx1.168$$

Noise plus a small parameter term.

At (1, −1)

$$\mu^Tx=\tfrac{6-7}{11}\approx-0.091,\qquad x^T\Sigma x=\tfrac{3+4+5}{11}=\tfrac{12}{11}$$

The same cross term now enters with a plus sign, since $x_1x_2=-1$.

$$\text{var}=1+\tfrac{12}{11}\approx2.091,\qquad \text{sd}\approx1.446$$

Three times the parameter variance of the other direction.

Answer $$\boxed{(1,1):\ \mathcal N(1.182,\ 1.364);\qquad (1,-1):\ \mathcal N(-0.091,\ 2.091)}$$
Check

Why: along $(1,1)$ the four inputs give $x_i^T(1,1)^T=2, \allowbreak 1, \allowbreak -1, \allowbreak -2$, sum of squares $10$; along $(1,-1)$ they give $0,1,-1,0$, sum of squares $2$. The data spread five times as much along $(1,1)$, so that direction is pinned down better.

Predictive uncertainty depends on the direction of the new input: directions the data barely explored keep more of the prior's doubt.

Checkpoint
§06.6 — the distribution of the next reading

The posterior of a slope through the origin is $\mathcal N(0.6,\,0.1)$ and the noise variance is $\sigma^2=1$. A new input $x=2$ arrives.

Find(a) What is the distribution of the new reading $Y$?
Given
  • $\beta\mid D\sim\mathcal N(0.6,\,0.1)$

  • $\sigma^2=1$

  • new input $x=2$

Hint 1/4

The new reading is $2\beta$ plus fresh noise; find the mean and the variance of that sum.

Hint 2/4

For one coefficient, $Y\mid x,D\sim\mathcal N\big(\mu_{\beta|D}x,\ \sigma^2+x^2\Sigma_{\beta|D}\big)$.

Hint 3/4

Here $\mu_{\beta|D}=0.6$, $\Sigma_{\beta|D}=0.1$, $\sigma^2=1$ and $x=2$.

Hint 4/4

Mean $1.2$ and variance $1+4(0.1)=1.4$: $Y\sim\mathcal N(1.2,\,1.4)$.

Show solution

The new reading is a sum of two independent Gaussians, so we add means and add variances.

Mean

$$E[Y]=0.6\times2=1.2$$

Predict with the posterior mean of the slope.

Variance

$$\mathrm{Var}(Y)=1+2^2(0.1)=1.4$$

Noise plus $x^2$ times the posterior variance.

Answer $$\boxed{Y\mid x=2,\,D\sim\mathcal N(1.2,\ 1.4)}$$
Check

Lower bound: the variance must be at least $\sigma^2=1$, and $1.4$ is; at $x=0$ it would be exactly $1$.

The predictive variance adds the noise to the parameter uncertainty; neither part may be dropped.

⚠ Dropping the noise term

the posterior of β feels like the whole uncertainty

wrong$$\mathrm{Var}(Y\mid x,D)=x^T\Sigma_{\beta|D}x$$
right$$\mathrm{Var}(Y\mid x,D)=\sigma^2+x^T\Sigma_{\beta|D}x$$
⚠ Using the prior covariance instead of the posterior

$\Sigma_0$ is the covariance written in the problem

wrong$$\sigma^2+x^T\Sigma_0\,x$$
right$$\sigma^2+x^T\Sigma_{\beta|D}\,x$$
⚠ Putting a variance where the error bar needs a standard deviation

the formula delivers a variance, and the interval is written right after it

wrong$$\mu_{\beta|D}^Tx\pm2\big(\sigma^2+x^T\Sigma_{\beta|D}x\big)$$
right$$\mu_{\beta|D}^Tx\pm2\sqrt{\sigma^2+x^T\Sigma_{\beta|D}x}$$
From a prior to a penalty, and back

A question gives a prior on the coefficients and Gaussian noise and asks for the MAP estimate, or gives a penalty and asks for the prior.

  1. Log-likelihood

    Write $-\log p(D\mid\beta)=\frac{1}{2\sigma^2}\mathrm{RSS}(\beta)+\text{const}$.

  2. Log-prior

    Write $-\log p(\beta)$: $\sum_j\beta_j^2/(2\tau^2)$ for $\mathcal N(0,\tau^2)$, $\sum_j\lvert\beta_j\rvert/b$ for $\mathrm{Lap}(0,b)$.

  3. Scale by 2σ²

    Multiply the sum by $2\sigma^2$ so that the likelihood part becomes the plain RSS.

  4. Read off λ

    Gaussian: $\lambda=\sigma^2/\tau^2$. Laplace: $\lambda=2\sigma^2/b$, so $b=2\sigma^2/\lambda$.

  5. Solve

    Ridge: $(X^TX+\lambda I)^{-1}X^Ty$. Lasso: numerically, or by cases for one coefficient.

Where it goes wrong
  • Dropping $\sigma^2$ and writing $\lambda=1/\tau^2$.

  • Using the ridge factor $\sigma^2/\lambda$ for the Laplace scale, which needs $2\sigma^2/\lambda$.

  • Reading the second argument of $\mathcal N(0,\cdot)$ as a standard deviation.

Gaussian posterior and prediction by hand

A Gaussian prior, Gaussian noise of known variance and a small X: find the posterior mean and covariance, then a predictive distribution.

  1. Set up

    Write $\mu_0$, $\Sigma_0$ and $\Sigma_D$; for ridge, $\mu_0=0$, $\Sigma_0=(\sigma^2/\lambda)I$ and $\Sigma_D=\sigma^2I$.

  2. Data terms

    Compute $X^T\Sigma_D^{-1}X$ and $X^T\Sigma_D^{-1}y$; with $\Sigma_D=\sigma^2I$ they are $X^TX/\sigma^2$ and $X^Ty/\sigma^2$.

  3. Add precisions

    $\Sigma_{\beta|D}^{-1}=\Sigma_0^{-1}+X^T\Sigma_D^{-1}X$, then invert.

  4. Mean

    $\mu_{\beta|D}=\Sigma_{\beta|D}\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big)$.

  5. Predict

    At $x$: mean $\mu_{\beta|D}^Tx$, variance $\sigma^2+x^T\Sigma_{\beta|D}x$.

  6. Check

    The mean lies between the prior mean and least squares, every posterior variance is below the prior's, and the predictive variance is at least $\sigma^2$.

Where it goes wrong
  • Adding covariances instead of precisions.

  • Forgetting $\Sigma_0^{-1}\mu_0$ when the prior mean is not $0$.

  • Leaving $\sigma^2$ out of the predictive variance.

One-coefficient lasso by cases (further reading)

A single coefficient, a lasso penalty λ and the two summaries Σx² and Σxy.

  1. Summaries

    $s=\sum_ix_i^2$ and $c=\sum_ix_iy_i$.

  2. Corner test

    If $\lvert c\rvert\le\lambda/2$, the answer is exactly $0$.

  3. Otherwise

    $\hat\beta=\big(c-\operatorname{sign}(c)\,\lambda/2\big)/s$.

  4. Check the sign

    The answer must have the sign of $c$; compare it with ridge, $c/(s+\lambda)$.

Where it goes wrong
  • Differentiating $\lvert\beta\rvert$ at $0$ as if its slope were $+1$.

  • Subtracting $\lambda$ instead of $\lambda/2$ from $c$.

  • Keeping a negative answer from the positive branch.

Weak evidence under a Gaussian prior: ridge moves a little

Readings with $\sum x_i^2=6$ and $\sum x_iy_i=1.5$, $\sigma^2=1$, prior $\mathcal N(0,0.25)$, which is ridge with $\lambda=4$. Find the MAP slope.

Find$\hat\beta^{\mathrm{Ridge}}$.
Given
  • $\sum x_i^2=6$, $\sum x_iy_i=1.5$

  • $\sigma^2=1$, $\lambda=4$

Solution

The ridge loss is smooth, so setting its derivative to zero finds the minimum directly.

Minimize

$$-2\,(1.5-6\beta)+8\beta=0\ \Rightarrow\ \beta=\tfrac{1.5}{10}=0.15$$

The squared penalty is flat at $\beta=0$, so it cannot hold the estimate there.

Answer $$\boxed{\hat\beta^{\mathrm{Ridge}}=0.15}$$
Check

Shrink factor: $0.15=\tfrac{6}{10}\times0.25$, the least squares slope $0.25$ scaled by $0.6$.

A Gaussian prior scales the estimate toward $0$ but never reaches it.

The same evidence under a Laplace prior: the lasso stops at zero

Same readings and $\lambda=4$, now with the prior $\mathrm{Lap}(0,0.5)$, which is the lasso. Find the MAP slope.

Find$\hat\beta^{\mathrm{Lasso}}$.
Given
  • $\sum x_i^2=6$, $\sum x_iy_i=1.5$

  • $\sigma^2=1$, $\lambda=4$

Solution

The lasso loss has a corner at $0$, so we test the corner first.

Test the corner

$$\text{right slope }-3+4=1>0,\qquad \text{left slope }-3-4=-7<0$$

The loss rises on both sides of $0$.

Answer $$\boxed{\hat\beta^{\mathrm{Lasso}}=0}$$
Check

Threshold: $\lvert\sum x_iy_i\rvert=1.5\le\lambda/2=2$.

A Laplace prior absorbs evidence up to a threshold and sets the coefficient exactly to $0$.

Same readings, same $\lambda$, same noise; only the prior's shape changes, and the answer goes from $0.15$ to exactly $0$ because the absolute value has a corner where the square is flat.

How to tell them apart

A squared penalty (Gaussian prior) multiplies the least squares slope by $\sum x_i^2/(\sum x_i^2+\lambda)$; an absolute-value penalty (Laplace prior) subtracts $\lambda/(2\sum x_i^2)$ and stops at $0$.

A prediction near the data: the noise dominates

Posterior $\mathcal N(17/30,\ 1/6)$ for a slope through the origin, $\sigma^2=0.25$. Split the predictive variance at $x=0.5$ into its noise and parameter parts.

FindThe two parts and their shares.
Given
  • $\Sigma_{\beta|D}=1/6$, $\sigma^2=0.25$

  • $x=0.5$

Solution

The two parts of the predictive variance are separate terms, so we compute each and compare.

Split

$$\sigma^2=0.25,\qquad x^2\Sigma_{\beta|D}=\tfrac{0.25}{6}\approx0.042$$

Noise part, then parameter part.

$$\tfrac{0.25}{0.292}\approx0.86$$

The noise is about six sevenths of the variance.

Answer $$\boxed{0.25+0.042=0.292}$$
Check

The two parts are in the exact ratio $6:1$, since $x^2\Sigma_{\beta|D}=\sigma^2/6$ here.

Near the data, knowing $\beta$ better would barely narrow the error bar.

A prediction far from the data: the slope's uncertainty dominates

Same posterior and noise, now at $x=3$.

FindThe two parts and their shares.
Given
  • $\Sigma_{\beta|D}=1/6$, $\sigma^2=0.25$

  • $x=3$

Solution

The same split as near the data; only $x^2$ changes.

Split

$$\sigma^2=0.25,\qquad x^2\Sigma_{\beta|D}=\tfrac{9}{6}=1.5$$

The parameter part grows with $x^2$.

$$\tfrac{1.5}{1.75}\approx0.86$$

Now the parameter part is six sevenths.

Answer $$\boxed{0.25+1.5=1.75}$$
Check

The shares have swapped: $6:1$ for the noise at $x=0.5$, $1:6$ at $x=3$.

Far from the data, more readings would narrow the error bar a lot; near the data, hardly at all.

Same posterior, same noise; at $x=0.5$ the noise is six sevenths of the predictive variance, at $x=3$ the slope's uncertainty is.

How to tell them apart

Compare $x^T\Sigma_{\beta|D}x$ with $\sigma^2$: if the first is small, more data will not help the prediction much; if it is large, the prediction is limited by what we know about $\beta$.

Scaffolding comes off
The common skeleton
  1. Summaries: $s=\sum_ix_i^2$ and $c=\sum_ix_iy_i$.

  2. Precisions: prior $1/\tau^2$ and data $s/\sigma^2$; add them and invert to get the posterior variance $v$.

  3. Posterior mean: $m=v\big(c/\sigma^2+\mu_0/\tau^2\big)$.

  4. Prediction at $x^\ast$: mean $m\,x^\ast$, variance $\sigma^2+(x^\ast)^2v$.

  5. Check: $m$ lies between $\mu_0$ and $c/s$, and $v$ is below both $\tau^2$ and $\sigma^2/s$.

1 · fully worked

Posterior and prediction from four readings, prior N(0, 1)

Four centered readings $(x,y)=(-1,-0.8), \allowbreak (1,1.1), \allowbreak (2,1.9), \allowbreak (-2,-2.2)$, noise variance $\sigma^2=0.25$, prior $\mathcal N(0,1)$ on the slope. Find the posterior, then the predictive distribution and a 2-standard-deviation interval at $x^\ast=3$.

Find$m$, $v$, and the predictive mean, variance and interval at $x^\ast=3$.
Given
  • readings $(-1,-0.8),\ \allowbreak (1,1.1),\ \allowbreak (2,1.9),\ \allowbreak (-2,-2.2)$

  • $\sigma^2=0.25$, prior $\mathcal N(0,1)$

  • new input $x^\ast=3$

Solution

One coefficient, so the five lines of the skeleton are the whole computation; no matrix is needed.

1 · Summaries

$$s=1+1+4+4=10,\qquad c=0.8+1.1+3.8+4.4=10.1$$

Each $x_iy_i$ is positive: all four readings agree on a positive slope.

2 · Precisions

$$\tfrac1v=\tfrac11+\tfrac{10}{0.25}=41,\qquad v=\tfrac{1}{41}\approx0.0244$$

The data precision $40$ dwarfs the prior's $1$.

3 · Mean

$$m=\tfrac{1}{41}\cdot\tfrac{10.1}{0.25}=\tfrac{40.4}{41}\approx0.985$$

$\mu_0=0$, so only the data term is left.

4 · Prediction

$$m\,x^\ast\approx2.956,\qquad \sigma^2+(x^\ast)^2v=0.25+\tfrac{9}{41}\approx0.470$$

Noise plus nine times the posterior variance.

$$2.956\pm2\sqrt{0.470}\approx2.956\pm1.370=[1.586,\ 4.327]$$

Two standard deviations each side.

5 · Check

$$0<0.985<\tfrac{c}{s}=1.01,\qquad v\approx0.0244<\min(1,\ 0.025)$$

The mean is shrunk slightly toward $0$, and the variance is below both the prior's and the data's.

Answer $$\boxed{\beta\mid D\sim\mathcal N(0.985,\ 0.0244);\qquad Y\mid x^\ast=3\sim\mathcal N(2.956,\ 0.470),\ \ [1.586,\ 4.327]}$$
Check

Ridge check: $\lambda=\sigma^2/\tau^2=0.25$, and $c/(s+\lambda)=10.1/10.25\approx0.985$, the same mean.

With plenty of data the prior barely moves the mean, but a prediction far out still pays for the remaining slope uncertainty.

2 · you write the reasoning

An easier one, and this time you write the reasons. One reading $(x,y)=(2,3)$, $\sigma^2=1$, prior $\mathcal N(0,1)$; predict at $x^\ast=1$. For each line, write why it is allowed.

  1. $s=4,\qquad c=6$

    reasoning

    One reading: $s=x^2=4$ and $c=xy=6$.

  2. $\tfrac1v=1+4=5,\qquad v=0.2$

    reasoning

    Prior precision $1/\tau^2=1$ plus data precision $s/\sigma^2=4$; the posterior variance is the inverse of the sum.

  3. $m=0.2\times6=1.2$

    reasoning

    The prior mean is $0$, so the mean is the variance times $c/\sigma^2=6$.

  4. $\text{mean }1.2\times1=1.2,\qquad \text{variance }1+1^2(0.2)=1.2$

    reasoning

    Predict with the posterior mean; the variance adds the noise $\sigma^2=1$ to $(x^\ast)^2v=0.2$. That mean and variance are both $1.2$ is a coincidence of these numbers.

  5. $0<1.2<\tfrac{c}{s}=1.5,\qquad 0.2<\min(1,\ 0.25)$

    reasoning

    The mean is shrunk from the data's $1.5$ toward the prior's $0$, and the variance $0.2$ is below both the prior's $1$ and the data's $\sigma^2/s=0.25$.

3 · find the buried error

Harder, with two errors buried in the solution. Four centered readings $(-2,-2.6), \allowbreak (-1,-0.9), \allowbreak (1,1.3), \allowbreak (2,2.2)$, $\sigma^2=0.5$, and an earlier study gives the prior $\mathcal N(1,\,0.5)$. Find the posterior mean and the predictive variance at $x^\ast=2$.

  1. Step 1. $s=10$ and $c=5.2+0.9+1.3+4.4=11.8$: the summaries.

  2. Step 2. $\tfrac1v=\tfrac{1}{0.5}+\tfrac{10}{0.5}=22$, so $v=\tfrac{1}{22}\approx0.0455$: precisions add.

  3. Step 3. $m=v\cdot c/\sigma^2=23.6/22\approx1.073$: the posterior mean.

  4. Step 4. The predictive variance at $x^\ast=2$ is $(x^\ast)^2v=4/22\approx0.182$.

  5. Step 5. Check: $v\approx0.0455$ is below both $\tau^2=0.5$ and $\sigma^2/s=0.05$.

the two buried errors (2)
⚠ step 3

The prior mean is $1$, not $0$, so the term $\mu_0/\tau^2=2$ belongs in the bracket.

In the ridge case $\mu_0=0$ and the term vanishes, so the habit is to drop it.

right

$m=\tfrac{1}{22}\,(23.6+2)=\tfrac{25.6}{22}\approx1.164$

⚠ step 4

A new reading carries its own noise: the predictive variance is $\sigma^2+(x^\ast)^2v$.

The posterior variance feels like the whole uncertainty.

right

$0.5+\tfrac{4}{22}\approx0.682$

4 · the bare problem
§06.6 — posterior and prediction interval from four readings

Four centered readings of a slope through the origin: $(-2,-1.3), \allowbreak (-1,-0.4), \allowbreak (1,0.7), \allowbreak (2,1.0)$. The noise variance is $\sigma^2=0.16$ and the prior on the slope is $\mathcal N(0,\,0.04)$.

Find
  1. (a) Find the posterior mean and standard deviation of the slope.

  2. (b) Find the predictive mean and a 2-standard-deviation interval at $x^\ast=1.5$.

Given
  • readings $(-2,-1.3),\ \allowbreak (-1,-0.4),\ \allowbreak (1,0.7),\ \allowbreak (2,1.0)$

  • $\sigma^2=0.16$, prior $\mathcal N(0,\,0.04)$

  • new input $x^\ast=1.5$

Hint 1/4

Run the five-line skeleton: summaries, precisions, mean, prediction, check.

Hint 2/4

$\frac1v=\frac{1}{\tau^2}+\frac{s}{\sigma^2}$, $m=v\,\frac{c}{\sigma^2}$, and the predictive variance is $\sigma^2+(x^\ast)^2v$.

Hint 3/4

Here $s=10$, $c=2.6+0.4+0.7+2.0=5.7$, $\tau^2=0.04$, $\sigma^2=0.16$ and $x^\ast=1.5$.

Hint 4/4

$m\approx0.407$ and sd $\approx0.107$; the prediction is $0.611\pm2(0.431)$, about $[-0.251,\ 1.473]$.

Show solution

The same five lines as the fully worked rung; the only new feature is a prior tight enough to matter.

Summaries

$$s=4+1+1+4=10,\qquad c=2.6+0.4+0.7+2.0=5.7$$

All four products are positive.

Posterior

$$\tfrac1v=\tfrac{1}{0.04}+\tfrac{10}{0.16}=25+62.5=87.5$$

A tight prior: its precision $25$ is two fifths of the data's $62.5$.

$$m=\tfrac{5.7/0.16}{87.5}=\tfrac{35.625}{87.5}\approx0.407,\qquad \text{sd}=\sqrt v\approx0.107$$

The prior mean is 0.

Prediction

$$0.407\times1.5\approx0.611,\qquad 0.16+2.25\,v\approx0.186$$

Noise plus $(x^\ast)^2v$.

$$0.611\pm2\sqrt{0.186}\approx[-0.251,\ 1.473]$$

Two standard deviations each side.

Check

$$0<0.407<\tfrac{5.7}{10}=0.57$$

The tight prior pulls the mean well below the data's $0.57$.

Answer $$\boxed{m\approx0.407,\ \ \text{sd}\approx0.107;\qquad [-0.251,\ 1.473]\ \text{at}\ x^\ast=1.5}$$
Check

Ridge check: $\lambda=0.16/0.04=4$ and $5.7/(10+4)\approx0.407$.

A tight prior can pull the estimate far from the data's own answer; the posterior variance shows how much the data were outweighed.

Full exam-style question

Exam-style: from likelihood to ridge posterior and two predictions, with two coefficientsexam format

Four centered readings with two inputs: rows $x_i=(1,0), \allowbreak (0,1), \allowbreak (1,1), \allowbreak (-2,-2)$ and responses $y=(1,-1,0.5,-0.5)$. The model is $y_i=\beta^Tx_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,1)$.

  • (a) Write the likelihood as a multivariate Gaussian and find $\hat\beta^{\mathrm{MLE}}$.
  • (b) With the prior $\beta\sim\mathcal N(0,I_2)$, find $\mu_{\beta|D}$ and $\Sigma_{\beta|D}$, and name the ridge estimate they correspond to.
  • (c) Find the predictive mean and variance at $x=(1,1)$ and at $x=(1,-1)$.
  • (d) Explain why the two predictive variances differ so much.
Find$\hat\beta^{\mathrm{MLE}}$; $\mu_{\beta|D}$ and $\Sigma_{\beta|D}$; two predictive distributions; an explanation.
Given
  • $X=\begin{pmatrix}1&0\\0&1\\1&1\\-2&-2\end{pmatrix}$, $y=(1,\,-1,\,0.5,\,-0.5)^T$

  • $\sigma^2=1$; prior $\mathcal N(0,I_2)$

Solution

Everything runs through $X^TX$ and $X^Ty$, so we compute them once and reuse them in every part.

(a) Likelihood and MLE

$$p(y\mid X,\beta)=\mathcal N\big(y;\,X\beta,\ I_4\big)\ \Rightarrow\ \hat\beta^{\mathrm{MLE}}=\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$$

Independent Gaussian noise with $\sigma^2=1$ makes $y$ a Gaussian vector with covariance $\sigma^2I_4$.

$$X^TX=\begin{pmatrix}6&5\\5&6\end{pmatrix},\qquad X^Ty=\begin{pmatrix}2.5\\0.5\end{pmatrix}$$

For example $(X^TX)_{11}=1+0+1+4=6$ and $(X^Ty)_1=1+0+0.5+1=2.5$.

$$\hat\beta^{\mathrm{LS}}=\tfrac{1}{11}\begin{pmatrix}6&-5\\-5&6\end{pmatrix}\begin{pmatrix}2.5\\0.5\end{pmatrix}=\tfrac{1}{22}\begin{pmatrix}25\\-19\end{pmatrix}\approx\begin{pmatrix}1.136\\-0.864\end{pmatrix}$$

$\det X^TX=36-25=11$.

(b) Posterior

$$\Sigma_{\beta|D}=\big(I_2+X^TX\big)^{-1}=\begin{pmatrix}7&5\\5&7\end{pmatrix}^{-1}=\tfrac{1}{24}\begin{pmatrix}7&-5\\-5&7\end{pmatrix}$$

Ridge case with $\sigma^2=1$ and $\lambda=1$; $\det=49-25=24$.

$$\mu_{\beta|D}=\tfrac{1}{24}\begin{pmatrix}7(2.5)-5(0.5)\\-5(2.5)+7(0.5)\end{pmatrix}=\begin{pmatrix}0.625\\-0.375\end{pmatrix}=\hat\beta^{\mathrm{Ridge}}\ (\lambda=1)$$

For a Gaussian the mean is the , so this is also the MAP estimate.

(c) Predictions

$$x=(1,1):\ \ \mu^Tx=0.25,\ \ x^T\Sigma x=\tfrac{7-10+7}{24}=\tfrac16,\ \ \mathrm{Var}=\tfrac76\approx1.167$$

Noise $1$ plus a small parameter term.

$$x=(1,-1):\ \ \mu^Tx=1,\ \ x^T\Sigma x=\tfrac{7+10+7}{24}=1,\ \ \mathrm{Var}=2$$

The parameter term is six times larger in this direction.

(d) Why

$$X^TX\begin{pmatrix}1\\1\end{pmatrix}=11\begin{pmatrix}1\\1\end{pmatrix},\qquad X^TX\begin{pmatrix}1\\-1\end{pmatrix}=1\begin{pmatrix}1\\-1\end{pmatrix}$$

The data carry eleven times as much information along $(1,1)$ as along $(1,-1)$.

$$\text{posterior variance along a unit direction: }\tfrac{1}{11+1}\ \ \text{vs}\ \ \tfrac{1}{1+1}$$

Data precision $11$ against $1$, plus the prior's $1$ in both: the direction $(1,-1)$ was barely measured.

Answer $$\boxed{\begin{aligned}&\hat\beta^{\mathrm{MLE}}\approx(1.136,\,-0.864),\qquad \mu_{\beta|D}=(0.625,\,-0.375)=\hat\beta^{\mathrm{Ridge}}\\ &\mathrm{Var}(Y)=\tfrac76\ \text{at}\ (1,1),\qquad \mathrm{Var}(Y)=2\ \text{at}\ (1,-1)\end{aligned}}$$
Check

Shrinkage along the two directions: least squares has $\hat\beta_1+\hat\beta_2\approx0.273$ and $\hat\beta_1-\hat\beta_2=2$; the posterior mean has $0.25=\tfrac{11}{12}(0.273)$ and $1=\tfrac12(2)$. The well-measured direction keeps eleven twelfths, the poorly measured one half.

In an exam, compute $X^TX$ and $X^Ty$ once; the MLE, the posterior and every prediction are built from them.

Practice

A · concept 4 questions
1§06.2 — is a MAP estimate unbiased?

A classmate argues that ridge regression is the most probable coefficient under a prior centred at $0$, so for $\lambda>0$ it must be an unbiased estimator of $\beta^{\mathrm{true}}$. You test the claim with one input.

Find(a) Is the claim true or false?
Given
  • model $y_i=\beta^{\mathrm{true}}x_i+\varepsilon_i$ with $E[\varepsilon_i]=0$ and the $x_i$ fixed

  • $\sum_i x_i^2=10$, $\lambda=5$, $\beta^{\mathrm{true}}=2$

  • $\hat\beta^{\mathrm{Ridge}}=\sum_ix_iy_i/(\sum_ix_i^2+\lambda)$

Hint 1/4

Unbiased means that the average of the estimate over repeated data sets equals $\beta^{\mathrm{true}}$; compute that average.

Hint 2/4

$E\big[\sum_ix_iy_i\big]=\sum_ix_i\,E[y_i]=\beta^{\mathrm{true}}\sum_ix_i^2$.

Hint 3/4

Here $\sum_ix_i^2=10$, $\lambda=5$ and $\beta^{\mathrm{true}}=2$.

Hint 4/4

$E[\hat\beta^{\mathrm{Ridge}}]=\frac{10}{15}\cdot2\approx1.33$, so the claim is false.

Show solution

A single computed average settles a claim about unbiasedness, so we compute it instead of arguing in general.

Average the estimate

$$E[\hat\beta^{\mathrm{Ridge}}]=\frac{E[\sum_ix_iy_i]}{\sum_ix_i^2+\lambda}=\frac{\beta^{\mathrm{true}}\sum_ix_i^2}{\sum_ix_i^2+\lambda}$$

The denominator is fixed, so the expectation passes through to the sum.

$$=\frac{10}{15}\cdot2\approx1.33$$

Substitute the given numbers.

Answer $$\boxed{\text{False: }E[\hat\beta^{\mathrm{Ridge}}]\approx1.33\neq2}$$
Check

Edge case: with $\lambda=0$ the factor is $1$ and the estimate is unbiased; that is least squares.

MAP estimates with a prior centred at $0$ are biased toward $0$; the prior buys lower variance, not unbiasedness.

2§06.4 — what the posterior covariance depends on

Two labs measure the same slope at the same inputs, with the same known noise variance and the same Gaussian prior. Their readings $y$ differ a lot, and one lab's readings look surprising.

Find(a) True or false: the lab with the surprising readings ends up with the larger posterior covariance.
Given
  • same $X$, same $\sigma^2$, same prior $\mathcal N(\mu_0,\Sigma_0)$

  • different response vectors $y$

Hint 1/4

Separate what the posterior covariance is built from and what the posterior mean is built from.

Hint 2/4

$\Sigma_{\beta|D}=(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X)^{-1}$ and $\mu_{\beta|D}=\Sigma_{\beta|D}(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0)$.

Hint 3/4

The two labs share $X$, $\Sigma_D=\sigma^2I$ and $\Sigma_0$; only $y$ differs.

Hint 4/4

Their covariances are equal and only their means differ, so the statement is false.

Show solution

Reading off which symbols the formula contains is enough; no numbers are needed.

Inspect the formula

$$\Sigma_{\beta|D}=\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)^{-1}$$

No $y$ appears, so equal $X$, $\Sigma_D$ and $\Sigma_0$ give equal covariances.

$$\mu_{\beta|D}^{(1)}-\mu_{\beta|D}^{(2)}=\Sigma_{\beta|D}X^T\Sigma_D^{-1}\big(y^{(1)}-y^{(2)}\big)$$

The difference between the labs shows up in the means only.

Answer $$\boxed{\text{False}}$$
Check

One-input check: $\Sigma_{\beta|D}=1/(1/\tau^2+\sum x_i^2/\sigma^2)$ has no $y_i$ in it.

When the noise variance is known, surprising readings move the estimate but do not change how sure we are.

3§06.2 — what a larger λ says about the prior

A team raises the ridge penalty $\lambda$ while the noise variance $\sigma^2$ of its sensor stays fixed and known. They want to explain the change to a colleague who thinks in priors.

Find(a) What does the larger $\lambda$ correspond to?
Given
  • ridge prior $\beta_j\sim\mathcal N(0,\sigma^2/\lambda)$, i.i.d.

  • $\sigma^2$ fixed

Hint 1/4

Translate $\lambda$ into the prior's parameters; the mean and the shape may or may not depend on it.

Hint 2/4

Ridge with penalty $\lambda$ is the MAP estimate under $\beta_j\sim\mathcal N(0,\sigma^2/\lambda)$.

Hint 3/4

With $\sigma^2$ fixed, raising $\lambda$ changes only the variance $\sigma^2/\lambda$.

Hint 4/4

The prior variance shrinks, so the belief that each coefficient is near $0$ becomes firmer.

Show solution

The prior has three features (centre, shape, spread), so we check which of them contains λ.

Translate

$$\text{prior variance}=\frac{\sigma^2}{\lambda}\ \text{falls as}\ \lambda\ \text{rises}$$

$\sigma^2$ is fixed, so the variance is inversely proportional to $\lambda$.

$$\text{centre }0,\ \text{Gaussian shape: unchanged}$$

Neither depends on $\lambda$.

Answer $$\boxed{\text{smaller prior variance: a firmer belief near }0}$$
Check

Limits: $\lambda\to0$ gives an infinitely wide prior and least squares; $\lambda\to\infty$ pins the prior at $0$ and sends $\hat\beta\to0$.

Read $\lambda$ as confidence in zero: large $\lambda$, strong belief that the coefficients are small.

4§06.3 — what a zero from the lasso means

A lasso fit sets one coefficient exactly to $0$. A report then states that, under the matching Laplace prior, the posterior probability that this coefficient equals $0$ is high.

Find(a) Is the report's statement true or false?
Given
  • lasso = MAP estimate under i.i.d. $\mathrm{Lap}(0,2\sigma^2/\lambda)$ priors

  • the fitted coefficient is exactly $0$

Hint 1/4

Separate where a density is highest from how much probability sits at a single point.

Hint 2/4

For a continuous distribution $P(\beta_j=c)=\int_c^c p(\beta_j\mid D)\,d\beta_j=0$ for every $c$.

Hint 3/4

The Laplace prior and the Gaussian likelihood give a posterior with a density; its peak, the MAP estimate, has this coordinate at $0$.

Hint 4/4

The peak has $\beta_j=0$, but $P(\beta_j=0\mid D)=0$, so the statement is false.

Show solution

Two different questions hide in the report, where the density peaks and how much probability one point carries, so we answer them separately.

Point probability

$$P(\beta_j=0\mid D)=\int_0^0 p(\beta_j\mid D)\,d\beta_j=0$$

Any single value of a continuous variable has probability 0.

What the zero does mean

$$\hat\beta^{\mathrm{MAP}}=\arg\max_\beta p(\beta\mid D)\ \ \text{has}\ \ \hat\beta_j=0$$

The zero is a coordinate of the most probable point, the mode of the posterior density.

Answer $$\boxed{\text{False}}$$
Check

Compare a Gaussian posterior: its mean is also its peak, and nobody reads the density there as the probability of that exact value.

Lasso zeros describe the posterior's peak, the MAP point estimate, not the probability that a coefficient is zero.

B · computation 8 questions
1§06.1 — maximum likelihood slope and a likelihood ratio

A centered data set with one input gives $\sum_i x_i^2=20$ and $\sum_i x_iy_i=14$. The noise is Gaussian with known variance $\sigma^2=2$.

Find
  1. (a) Find $\hat\beta^{\mathrm{MLE}}$.

  2. (b) By what factor is the data more probable under $\hat\beta^{\mathrm{MLE}}$ than under $\beta=0.5$?

Given
  • $\sum_i x_i^2=20$, $\sum_i x_iy_i=14$

  • $\sigma^2=2$

  • for one input, $\mathrm{RSS}(\beta)=\mathrm{RSS}(\hat\beta)+\big(\sum_ix_i^2\big)(\beta-\hat\beta)^2$ with $\hat\beta=\hat\beta^{\mathrm{LS}}$

Hint 1/4

Part (a) is the least squares slope; part (b) needs only the difference of two RSS values.

Hint 2/4

$$\frac{L(\hat\beta)}{L(\beta)}=\exp\Big(\frac{\mathrm{RSS}(\beta)-\mathrm{RSS}(\hat\beta)}{2\sigma^2}\Big),$$ and the RSS difference is $\sum x_i^2\,(\beta-\hat\beta)^2$.

Hint 3/4

Here $\hat\beta=14/20=0.7$, $\beta=0.5$, $\sum x_i^2=20$ and $\sigma^2=2$; the RSS difference is $20\,(0.2)^2=0.8$.

Hint 4/4

$\hat\beta^{\mathrm{MLE}}=0.7$ and the ratio is $e^{0.8/4}=e^{0.2}\approx1.221$.

Show solution

The constant term of $l(\beta)$ cancels in a ratio, so $\sum y_i^2$ is never needed.

Maximum likelihood slope

$$\hat\beta^{\mathrm{MLE}}=\hat\beta^{\mathrm{LS}}=\tfrac{14}{20}=0.7$$

Gaussian noise with known σ: maximum likelihood is least squares.

RSS difference

$$\mathrm{RSS}(0.5)-\mathrm{RSS}(0.7)=20\,(0.5-0.7)^2=0.8$$

The given identity; the cross term vanishes at the least squares slope.

Ratio

$$\tfrac{L(0.7)}{L(0.5)}=\exp\big(\tfrac{0.8}{2\cdot2}\big)=e^{0.2}\approx1.221$$

Divide the RSS difference by $2\sigma^2=4$.

Answer $$\boxed{\hat\beta^{\mathrm{MLE}}=0.7,\qquad L(0.7)/L(0.5)=e^{0.2}\approx1.221}$$
Check

With a smaller noise variance, say $\sigma^2=0.5$, the same RSS gap would give $e^{0.8}\approx2.23$: a sharper likelihood with the same peak.

Under Gaussian noise a likelihood ratio is the exponential of an RSS difference over $2\sigma^2$.

2§06.2 — from a prior's spread to a ridge estimate

Sensor noise has standard deviation $0.3$. Before seeing data, an engineer models each coefficient as $\mathcal N(0,0.15^2)$. A one-input centered data set gives $\sum_ix_i^2=12$ and $\sum_ix_iy_i=8$.

Find
  1. (a) Find the ridge penalty $\lambda$ that matches the prior.

  2. (b) Find the MAP slope and compare it with least squares.

Given
  • $\sigma=0.3$

  • prior $\mathcal N(0,\tau^2)$ with $\tau=0.15$

  • $\sum_ix_i^2=12$, $\sum_ix_iy_i=8$

Hint 1/4

Translate the prior into a penalty first; the estimate then follows from the ridge formula.

Hint 2/4

$\lambda=\sigma^2/\tau^2$, and for one input $\hat\beta^{\mathrm{Ridge}}=\sum x_iy_i/(\sum x_i^2+\lambda)$.

Hint 3/4

Here $\sigma^2=0.09$, $\tau^2=0.0225$, $\sum x_i^2=12$ and $\sum x_iy_i=8$.

Hint 4/4

$\lambda=4$ and $\hat\beta^{\mathrm{MAP}}=8/16=0.5$, against $8/12\approx0.667$ for least squares.

Show solution

The translation needs variances, so both spreads are squared before anything else.

Penalty

$$\lambda=\frac{0.3^2}{0.15^2}=\frac{0.09}{0.0225}=4$$

Square both standard deviations before dividing.

MAP slope

$$\hat\beta^{\mathrm{MAP}}=\frac{8}{12+4}=0.5$$

Ridge with the matching penalty.

$$\hat\beta^{\mathrm{LS}}=\frac{8}{12}\approx0.667$$

For comparison: the data alone.

Answer $$\boxed{\lambda=4,\qquad \hat\beta^{\mathrm{MAP}}=0.5}$$
Check

Weighted average: data weight $12/0.09\approx133.3$ at $0.667$, prior weight $1/0.0225\approx44.4$ at $0$, giving $133.3\times0.667/177.8\approx0.5$.

The shrink factor is $\sum x_i^2/(\sum x_i^2+\lambda)$; here a quarter of the least squares slope is given up.

3§06.3 — the Laplace prior behind a lasso fit

A lasso fit uses $\lambda=0.8$ on centered data whose Gaussian noise has variance $\sigma^2=0.2$. We want to describe the prior this fit assumes.

Find
  1. (a) Find the scale $b$ of the matching Laplace prior.

  2. (b) Find its variance, its standard deviation and its density at $0$.

Given
  • $\lambda=0.8$

  • $\sigma^2=0.2$

  • $\mathrm{Lap}(\mu,b)$: density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$, variance $2b^2$

Hint 1/4

The scale comes from matching exponents; the rest are properties of the Laplace density.

Hint 2/4

$b=2\sigma^2/\lambda$, the variance is $2b^2$, and the density at the centre is $1/(2b)$.

Hint 3/4

Here $\sigma^2=0.2$ and $\lambda=0.8$.

Hint 4/4

$b=0.5$, variance $0.5$, standard deviation $\approx0.707$, density at $0$ equal to $1$.

Show solution

Every property of the Laplace density is a simple function of its scale, so the scale comes first.

Scale

$$b=\frac{2\sigma^2}{\lambda}=\frac{0.4}{0.8}=0.5$$

Match $\lvert\beta_j\rvert/b$ with $\lambda\lvert\beta_j\rvert/(2\sigma^2)$.

Properties

$$\mathrm{Var}=2b^2=0.5,\qquad \text{sd}=\sqrt{0.5}\approx0.707$$

The Laplace variance is twice the squared scale.

$$p(0)=\frac{1}{2b}=1$$

The height of the corner.

Answer $$\boxed{b=0.5,\quad \mathrm{Var}=0.5,\quad \text{sd}\approx0.707,\quad p(0)=1}$$
Check

Compare ridge at the same $\lambda$: its prior variance would be $\sigma^2/\lambda=0.25$, half the Laplace variance, yet the Laplace density at $0$ is higher, $1$ against $0.80$.

Same λ, different priors: the Laplace is wider on average but more peaked at zero.

4§06.3 — a one-coefficient lasso MAP by cases

One centered input gives $\sum_ix_i^2=8$ and $\sum_ix_iy_i=5$. The noise variance is $\sigma^2=1$ and the lasso penalty is $\lambda=6$.

Find
  1. (a) Find the lasso MAP slope.

  2. (b) Find the ridge MAP slope for the same $\lambda$.

  3. (c) Find the smallest $\lambda$ that makes the lasso slope exactly $0$.

Given
  • $\sum_ix_i^2=8$, $\sum_ix_iy_i=5$

  • $\sigma^2=1$, $\lambda=6$

  • $\mathrm{Loss}(\beta)=\sum y_i^2-2\beta\sum x_iy_i+\beta^2\sum x_i^2+\lambda\lvert\beta\rvert$

Hint 1/4

Minimize on $\beta>0$ and on $\beta<0$ separately, then test the corner at $0$.

Hint 2/4

On $\beta>0$ the derivative is $2\beta\sum x_i^2-2\sum x_iy_i+\lambda$; the corner holds the minimum when $\lvert\sum x_iy_i\rvert\le\lambda/2$.

Hint 3/4

Here the positive side gives $2\cdot8\,\beta-2\cdot5+6=0$; for ridge, $5/(8+6)$.

Hint 4/4

Lasso $0.25$, ridge $\approx0.357$, and the lasso slope is $0$ from $\lambda=10$ on.

Show solution

Because $\sum x_iy_i>0$ the minimum can only be at $\beta\ge0$, so we try the positive side first and confirm the sign.

Positive side

$$16\beta-10+6=0\ \Rightarrow\ \beta=0.25>0$$

A positive solution on the positive side is a valid candidate.

$$\text{right slope at }0:\ -10+6=-4<0$$

The loss falls as $\beta$ leaves $0$ to the right, so the corner is not the minimum.

Ridge

$$\hat\beta^{\mathrm{Ridge}}=\tfrac{5}{8+6}\approx0.357$$

The squared penalty scales instead of subtracting.

Threshold

$$\lvert\textstyle\sum x_iy_i\rvert\le\lambda/2\iff\lambda\ge10$$

At $\lambda=10$ the right slope at $0$ is $-10+10=0$ and the minimum reaches the corner.

Answer $$\boxed{\hat\beta^{\mathrm{Lasso}}=0.25,\quad \hat\beta^{\mathrm{Ridge}}\approx0.357,\quad \lambda\ge10\Rightarrow\hat\beta^{\mathrm{Lasso}}=0}$$
Check

Shortcut: $(5-\lambda/2)/8=(5-3)/8=0.25$. The lasso moved the least squares slope $0.625$ down by the fixed amount $\lambda/(2\cdot8)=0.375$.

For one coefficient the lasso subtracts $\lambda/(2\sum x_i^2)$ from the least squares slope and stops at $0$.

5§06.4 — a posterior with a prior mean of 2

An earlier experiment suggests a slope near $2$, encoded as the prior $\mathcal N(2,0.25)$. New centered data with one input give $\sum_ix_i^2=4$ and $\sum_ix_iy_i=6$, with $\sigma^2=1$.

Find
  1. (a) Find the posterior mean and variance.

  2. (b) Where does the posterior mean sit between the least squares slope and the prior mean?

Given
  • prior $\mathcal N(\mu_0,\Sigma_0)=\mathcal N(2,\,0.25)$

  • $\sum_ix_i^2=4$, $\sum_ix_iy_i=6$, $\sigma^2=1$

Hint 1/4

The prior mean is not $0$, so both terms in the bracket of the mean formula count.

Hint 2/4

$\Sigma_{\beta|D}^{-1}=\Sigma_0^{-1}+\sum x_i^2/\sigma^2$ and $\mu_{\beta|D}=\Sigma_{\beta|D}\big(\sum x_iy_i/\sigma^2+\Sigma_0^{-1}\mu_0\big)$.

Hint 3/4

Here $\Sigma_0^{-1}=4$, $\mu_0=2$, $\sum x_i^2=4$, $\sum x_iy_i=6$ and $\sigma^2=1$.

Hint 4/4

Precision $8$, variance $0.125$, mean $(6+8)/8=1.75$.

Show solution

The general formula with $\mu_0\neq0$; with one coefficient it is two lines of arithmetic.

Precision and variance

$$\Sigma_{\beta|D}^{-1}=4+4=8,\qquad \Sigma_{\beta|D}=0.125$$

Prior precision $1/0.25=4$ plus data precision $4/1=4$.

Mean

$$\mu_{\beta|D}=0.125\,(6+4\cdot2)=0.125\times14=1.75$$

Data term $6$ plus prior term $\Sigma_0^{-1}\mu_0=8$.

Interpret

$$\hat\beta^{\mathrm{LS}}=\tfrac64=1.5,\qquad \tfrac{4\cdot1.5+4\cdot2}{8}=1.75$$

Equal precisions give the midpoint.

Answer $$\boxed{\beta\mid D\sim\mathcal N(1.75,\ 0.125)}$$
Check

Limits: with a very wide prior the mean would approach $1.5$, with a very narrow one $2$; $1.75$ is between, as it must be.

When the prior and the data carry equal precision, the posterior mean is their midpoint.

6§06.4 — a two-coefficient posterior with an orthogonal design

Two centered inputs are uncorrelated in the data, so $X^TX$ is diagonal: $X^TX=\mathrm{diag}(2,8)$ and $X^Ty=(2,4)^T$. The noise variance is $\sigma^2=1$ and the prior is $\mathcal N(0,\tfrac12I_2)$, the ridge prior with $\lambda=2$.

Find
  1. (a) Find $\Sigma_{\beta|D}$ and $\mu_{\beta|D}$.

  2. (b) Compare $\mu_{\beta|D}$ with $\hat\beta^{\mathrm{LS}}$ coefficient by coefficient.

Given
  • $X^TX=\begin{pmatrix}2&0\\0&8\end{pmatrix}$, $X^Ty=\begin{pmatrix}2\\4\end{pmatrix}$

  • $\sigma^2=1$, prior $\mathcal N(0,\tfrac12I_2)$

Hint 1/4

A diagonal design keeps the two coefficients apart; each is a one-input problem.

Hint 2/4

$\Sigma_{\beta|D}=\sigma^2(X^TX+\lambda I)^{-1}$ and $\mu_{\beta|D}=(X^TX+\lambda I)^{-1}X^Ty$.

Hint 3/4

Here $X^TX+\lambda I=\mathrm{diag}(4,10)$ and $X^Ty=(2,4)^T$.

Hint 4/4

$\Sigma_{\beta|D}=\mathrm{diag}(0.25,0.1)$ and $\mu_{\beta|D}=(0.5,0.4)^T$; least squares is $(1,0.5)^T$, so the shrink factors are $0.5$ and $0.8$.

Show solution

Inverting a diagonal matrix is dividing entry by entry, so the matrix formula costs two divisions.

Covariance

$$X^TX+2I=\begin{pmatrix}4&0\\0&10\end{pmatrix}\ \Rightarrow\ \Sigma_{\beta|D}=\begin{pmatrix}0.25&0\\0&0.1\end{pmatrix}$$

The ridge case with $\sigma^2=1$.

Mean

$$\mu_{\beta|D}=\big(\tfrac24,\ \tfrac{4}{10}\big)^T=(0.5,\ 0.4)^T$$

Divide each entry of $X^Ty$ by the matching diagonal entry.

Compare

$$\hat\beta^{\mathrm{LS}}=\big(\tfrac22,\ \tfrac48\big)^T=(1,\ 0.5)^T;\qquad \tfrac24=0.5,\ \ \tfrac{8}{10}=0.8$$

The shrink factor of coefficient $j$ is $d_j/(d_j+\lambda)$ for the diagonal entry $d_j$.

Answer $$\boxed{\Sigma_{\beta|D}=\mathrm{diag}(0.25,\,0.1),\qquad \mu_{\beta|D}=(0.5,\,0.4)^T}$$
Check

The better-measured coefficient, with $d_2=8$, shrinks less; as $\lambda\to0$ both factors go to $1$ and least squares returns.

Ridge shrinks weakly measured coefficients most, because there the prior's precision competes with little data precision.

7§06.6 — a two-coefficient prediction interval

A posterior from an orthogonal design is $\mu_{\beta|D}=(0.5,0.4)^T$, $\Sigma_{\beta|D}=\mathrm{diag}(0.25,0.1)$, and the noise variance is $\sigma^2=1$. A prediction is needed at the new input $x=(2,1)$.

Find
  1. (a) Find the predictive mean and variance.

  2. (b) Give the interval mean $\pm2$ standard deviations.

Given
  • $\mu_{\beta|D}=(0.5,\,0.4)^T$, $\Sigma_{\beta|D}=\mathrm{diag}(0.25,\,0.1)$

  • $\sigma^2=1$

  • new input $x=(2,\,1)^T$

Hint 1/4

The predictive distribution needs a mean and a variance with two parts.

Hint 2/4

The mean is $\mu_{\beta|D}^Tx$ and the variance $\sigma^2+x^T\Sigma_{\beta|D}x$; with a diagonal $\Sigma$, $x^T\Sigma x=\sum_j\Sigma_{jj}x_j^2$.

Hint 3/4

Here $x=(2,1)$, so $x_1^2=4$ and $x_2^2=1$; $\Sigma_{11}=0.25$, $\Sigma_{22}=0.1$ and $\sigma^2=1$.

Hint 4/4

Mean $1.4$, variance $2.1$, sd $\approx1.449$, interval $\approx[-1.498,\ 4.298]$.

Show solution

A diagonal covariance has no cross term, so the quadratic form is a weighted sum of squares.

Mean

$$\mu_{\beta|D}^Tx=2(0.5)+1(0.4)=1.4$$

Predict with the posterior mean.

Variance

$$\sigma^2+x^T\Sigma_{\beta|D}x=1+4(0.25)+1(0.1)=2.1$$

Noise plus the parameter term.

Interval

$$1.4\pm2\sqrt{2.1}=1.4\pm2.898$$

Two standard deviations, from the square root of the variance.

Answer $$\boxed{\mathcal N(1.4,\ 2.1);\qquad [-1.498,\ 4.298]}$$
Check

Split: noise $1$ and parameter part $1.1$, so slightly more than half the variance comes from not knowing $\beta$ at this input.

At inputs that load on weakly measured coefficients, the parameter part can rival the noise.

8§06.5 — two readings in sequence

A slope has the prior $\mathcal N(0,1)$ and the noise variance is $\sigma^2=0.25$. Reading 1 is $(x,y)=(1,\,0.9)$; reading 2 is $(2,\,1.5)$.

Find
  1. (a) Update after reading 1.

  2. (b) Update after reading 2.

  3. (c) Confirm with the batch posterior.

Given
  • prior $m_0=0$, $v_0=1$

  • $\sigma^2=0.25$

  • readings $(1,\,0.9)$, then $(2,\,1.5)$

Hint 1/4

Run the update rule twice; the first posterior is the prior of the second step.

Hint 2/4

$\frac{1}{v_n}=\frac{1}{v_{n-1}}+\frac{x_n^2}{\sigma^2}$ and $m_n=v_n\big(\frac{m_{n-1}}{v_{n-1}}+\frac{x_ny_n}{\sigma^2}\big)$.

Hint 3/4

Reading 1, $(1,0.9)$, adds $1/0.25=4$ to the precision and $0.9/0.25=3.6$ to the evidence; reading 2, $(2,1.5)$, adds $4/0.25=16$ and $3/0.25=12$.

Hint 4/4

$\mathcal N(0.72,\,0.2)$ after reading 1, and $\mathcal N(15.6/21,\ 1/21)\approx\mathcal N(0.743,\,0.048)$ after reading 2, as in the batch computation.

Show solution

Each step needs only the previous posterior and one reading, so we never recompute from scratch.

Reading 1

$$\tfrac{1}{v_1}=1+4=5,\qquad m_1=\tfrac15\,(0+3.6)=0.72$$

Evidence $x_1y_1/\sigma^2=0.9/0.25=3.6$.

Reading 2

$$\tfrac{1}{v_2}=5+16=21,\qquad m_2=\tfrac{1}{21}\,(5\cdot0.72+12)=\tfrac{15.6}{21}\approx0.743$$

Old precision times old mean is $3.6$; the new evidence is $2\cdot1.5/0.25=12$.

Batch

$$1+\tfrac{5}{0.25}=21,\qquad \tfrac{3.9/0.25}{21}=\tfrac{15.6}{21}$$

$\sum x^2=1+4$ and $\sum xy=0.9+3$.

Answer $$\boxed{\mathcal N(0.72,\ 0.2)\ \to\ \mathcal N(0.743,\ 0.048)}$$
Check

Reading 2 sits at $x=2$, so it adds four times the precision of reading 1: $16$ against $4$.

Readings at larger $\lvert x\rvert$ carry more information about a slope, in proportion to $x^2$.

C · exam level 5 questions
1§06.3 — which prior produces a given penalty

A fitted model minimizes $\mathrm{RSS}(\beta)+3\sum_j\lvert\beta_j\rvert$. The noise is Gaussian with known variance $\sigma^2=1.5$, and the fit is to be read as a MAP estimate.

Find(a) Which i.i.d. prior on the coefficients makes this the MAP estimate?
Given
  • penalty $3\sum_j\lvert\beta_j\rvert$, so $\lambda=3$

  • $\sigma^2=1.5$

Hint 1/4

An absolute-value penalty points to one family of priors; then match the scale.

Hint 2/4

$\beta_j\sim\mathrm{Lap}(0,b)$ with $b=2\sigma^2/\lambda$ turns the MAP estimate into the lasso with penalty $\lambda$.

Hint 3/4

Here $\lambda=3$ and $\sigma^2=1.5$.

Hint 4/4

$b=2(1.5)/3=1$, so the prior is $\mathrm{Lap}(0,1)$.

Show solution

The penalty's shape fixes the family before any number is used.

Family

$$\lambda\textstyle\sum_j\lvert\beta_j\rvert\ \leftrightarrow\ -\log p(\beta)=\sum_j\tfrac{\lvert\beta_j\rvert}{b}+\text{const}$$

Only a Laplace prior has an absolute value in its log.

Scale

$$b=\frac{2\sigma^2}{\lambda}=\frac{2(1.5)}{3}=1$$

Match the coefficient of $\lvert\beta_j\rvert$ after multiplying by $2\sigma^2$.

Answer $$\boxed{\beta_j\ \text{i.i.d.}\ \mathrm{Lap}(0,\,1)}$$
Check

Translate back: $2\sigma^2\cdot\lvert\beta_j\rvert/b=3\lvert\beta_j\rvert$, the given penalty.

Identify the family from the penalty's shape, then the scale from $2\sigma^2/\lambda$.

2§06.4 — more coefficients than readings

A model has two coefficients but only one reading so far: input $x_1=(1,2)$, response $y_1=3$, noise variance $\sigma^2=1$. The model has no intercept, and the prior is $\mathcal N(0,I_2)$.

Find
  1. (a) Show that the maximum likelihood estimate is not unique.

  2. (b) Find $\Sigma_{\beta|D}$ and $\mu_{\beta|D}$.

  3. (c) Find the predictive variance at $x=(1,2)$ and at $x=(2,-1)$, and explain the difference.

Given
  • $X=(1\ \ 2)$, $y=3$

  • $\sigma^2=1$, prior $\mathcal N(0,I_2)$

Hint 1/4

With one equation and two unknowns the likelihood cannot single out a point; the prior can.

Hint 2/4

$L(\beta)$ depends on $\beta$ only through $\beta_1+2\beta_2$; the posterior uses $\Sigma_{\beta|D}^{-1}=I+X^TX/\sigma^2$.

Hint 3/4

Here $X^TX=\begin{pmatrix}1&2\\2&4\end{pmatrix}$, which is singular, and $X^Ty=(3,6)^T$; so $\Sigma_{\beta|D}^{-1}=\begin{pmatrix}2&2\\2&5\end{pmatrix}$.

Hint 4/4

$\mu_{\beta|D}=(0.5,1)^T$; the predictive variances are $1+\tfrac56$ at $(1,2)$ and $1+5=6$ at $(2,-1)$.

Show solution

The singular $X^TX$ blocks least squares, but the prior makes the posterior precision invertible, so we go straight to the posterior box.

(a) The likelihood

$$L(\beta)\propto\exp\big(-\tfrac12\,(3-\beta_1-2\beta_2)^2\big)$$

It takes its largest value on the whole line $\beta_1+2\beta_2=3$.

$$\det X^TX=1\cdot4-2\cdot2=0$$

So $(X^TX)^{-1}$ does not exist and least squares has no unique answer.

(b) The posterior

$$\Sigma_{\beta|D}^{-1}=\begin{pmatrix}2&2\\2&5\end{pmatrix},\ \ \det=6,\ \ \Sigma_{\beta|D}=\tfrac16\begin{pmatrix}5&-2\\-2&2\end{pmatrix}$$

Prior precision $I_2$ plus data precision $X^TX$.

$$\mu_{\beta|D}=\tfrac16\begin{pmatrix}5\cdot3-2\cdot6\\-2\cdot3+2\cdot6\end{pmatrix}=\begin{pmatrix}0.5\\1\end{pmatrix}$$

The posterior covariance times $X^Ty=(3,6)^T$.

(c) Two predictions

$$x=(1,2):\ \tfrac16\,(5-8+8)=\tfrac56;\qquad x=(2,-1):\ \tfrac16\,(20+8+2)=5$$

The quadratic form $\Sigma_{11}x_1^2+2\Sigma_{12}x_1x_2+\Sigma_{22}x_2^2$.

$$1+\tfrac56\approx1.833\qquad\text{vs}\qquad 1+5=6$$

At $(2,-1)$, perpendicular to the measured direction, the variance equals the prior's $1+\lVert x\rVert^2=6$.

Answer $$\boxed{\mu_{\beta|D}=(0.5,\,1)^T;\qquad \mathrm{Var}(Y)=1.833\ \text{at}\ (1,2),\quad 6\ \text{at}\ (2,-1)}$$
Check

$\mu_{\beta|D}$ lies on $\beta_1+2\beta_2=2.5$, not $3$: the prior pulls it toward $0$ along the measured direction, and along $(2,-1)$ it keeps the prior mean, since $2(0.5)-1=0$.

With more coefficients than readings the posterior still exists and is unique; directions the data never probed keep their prior uncertainty.

3§06.6 — find the error in a prediction interval

A student computes a 2-standard-deviation predictive interval for a line through the origin. The prior was $\mathcal N(0,1)$, $\sigma^2=0.25$, and after the data the posterior is $\mathcal N(0.8,\,0.04)$. The new input is $x=2$.

Find(a) Which step contains the error?
Given
  • Step 1: predictive mean $0.8\times2=1.6$.

  • Step 2: parameter part of the variance, $x^2\times1=4$, using the prior variance $1$.

  • Step 3: predictive variance $0.25+4=4.25$.

  • Step 4: interval $1.6\pm2\sqrt{4.25}=1.6\pm4.12$.

Hint 1/4

For each step, ask which distribution of β it should use now that the data are in.

Hint 2/4

$\mathrm{Var}(Y\mid x,D)=\sigma^2+x^2\,\Sigma_{\beta|D}$, with the posterior variance.

Hint 3/4

Here $\Sigma_{\beta|D}=0.04$, $\sigma^2=0.25$ and $x=2$; the student used $1$ in place of $0.04$.

Hint 4/4

Step 2 is wrong. Corrected: variance $0.25+0.16=0.41$ and interval $1.6\pm1.281$, that is $[0.319,\ 2.881]$.

Show solution

Checking each step against the formula is quicker than redoing the whole computation.

Locate

$$x^2\,\Sigma_0=4\times1=4\qquad\text{vs}\qquad x^2\,\Sigma_{\beta|D}=4\times0.04=0.16$$

Once the data are in, the uncertainty about $\beta$ is the posterior's.

Fix

$$0.25+0.16=0.41,\qquad 1.6\pm2\sqrt{0.41}\approx1.6\pm1.281$$

Steps 1, 3 and 4 were right and stay as they are.

Answer $$\boxed{\text{Step 2};\qquad [0.319,\ 2.881]}$$
Check

Sanity bound: the variance must lie between the noise alone, $0.25$, and the prior predictive, $4.25$; $0.41$ does.

Once data are in, every uncertainty about β comes from the posterior; the prior is used once, to build it.

4§06.4 — readings from sensors of different quality

A slope through the origin is measured with two sensors, and the prior is $\mathcal N(0,1)$. Reading A, $(x,y)=(1,0.5)$, comes from a precise sensor with $\sigma^2=0.25$; readings B, $(2,2)$, and C, $(1,3)$, come from a noisy one with $\sigma^2=1$.

Find(a) What is the posterior mean of the slope?
Given
  • prior $\mathcal N(0,1)$

  • A: $(1,\,0.5)$ with $\sigma_A^2=0.25$

  • B: $(2,\,2)$ and C: $(1,\,3)$, each with $\sigma^2=1$

  • $\Sigma_D=\mathrm{diag}(0.25,\,1,\,1)$

Hint 1/4

Each reading counts with its own noise variance: $\Sigma_D$ is diagonal but not a multiple of $I$.

Hint 2/4

$\Sigma_{\beta|D}^{-1}=\Sigma_0^{-1}+\sum_ix_i^2/\sigma_i^2$ and, with $\mu_0=0$, $\mu_{\beta|D}=\Sigma_{\beta|D}\sum_ix_iy_i/\sigma_i^2$.

Hint 3/4

Precision: $1+\frac{1}{0.25}+\frac{4}{1}+\frac{1}{1}$. Evidence: $\frac{0.5}{0.25}+\frac{4}{1}+\frac{3}{1}$.

Hint 4/4

Precision $10$, evidence $9$, posterior mean $0.9$.

Show solution

The general formula with $\Sigma_D^{-1}=\mathrm{diag}(4,1,1)$ weights each reading by one over its noise variance; for one coefficient it reduces to two sums.

Weighted precision

$$\Sigma_{\beta|D}^{-1}=1+\tfrac{1^2}{0.25}+\tfrac{2^2}{1}+\tfrac{1^2}{1}=10$$

The prior's $1$, then $x_i^2/\sigma_i^2$ for each reading.

Weighted evidence and mean

$$\textstyle\sum_i\tfrac{x_iy_i}{\sigma_i^2}=\tfrac{0.5}{0.25}+\tfrac{4}{1}+\tfrac{3}{1}=9,\qquad \mu_{\beta|D}=\tfrac{9}{10}=0.9$$

The precise reading counts four times per unit of $xy$.

Answer $$\boxed{\mu_{\beta|D}=0.9,\qquad \Sigma_{\beta|D}=0.1}$$
Check

Reading C alone suggests the slope $3$ with precision $1$; reading A suggests $0.5$ with precision $4$. The answer, $0.9$, leans toward the precise sensor, as it should.

With unequal noise, weight each reading by $1/\sigma_i^2$; the matrix $\Sigma_D$ does this automatically.

5§06.6 — the error bar after very many readings

Readings keep arriving at inputs spread evenly over $[-1,1]$, for a line through the origin with known noise variance $\sigma^2$ and a prior $\mathcal N(0,\tau^2)$ on the slope. Consider the predictive distribution at the fixed input $x^\ast=0.5$.

Find(a) What does the predictive variance at $x^\ast$ approach?
Given
  • $n\to\infty$ readings, inputs spread over $[-1,1]$

  • noise variance $\sigma^2$, known; prior $\mathcal N(0,\tau^2)$

  • fixed new input $x^\ast=0.5$

Hint 1/4

Split the predictive variance into the part data can shrink and the part they cannot.

Hint 2/4

$\mathrm{Var}(Y\mid x^\ast,D)=\sigma^2+(x^\ast)^2\Sigma_{\beta|D}$, with $\Sigma_{\beta|D}=1/\big(1/\tau^2+\sum_ix_i^2/\sigma^2\big)$.

Hint 3/4

Here $\sum_ix_i^2$ grows without bound as readings accumulate, while $x^\ast=0.5$ stays fixed.

Hint 4/4

$\Sigma_{\beta|D}\to0$, so the predictive variance approaches $\sigma^2$.

Show solution

The two parts behave differently as n grows, so we take the limit of each.

Parameter part

$$\Sigma_{\beta|D}=\frac{1}{1/\tau^2+\sum_ix_i^2/\sigma^2}\to0$$

Each reading adds $x_i^2/\sigma^2\ge0$, and with inputs spread over $[-1,1]$ the sum grows about linearly in $n$.

Limit

$$\sigma^2+0.25\,\Sigma_{\beta|D}\to\sigma^2$$

The noise term does not depend on $n$.

Answer $$\boxed{\sigma^2}$$
Check

Cross-check with the sequential example: after $20$ readings $\Sigma_{\beta|D}\approx0.035$, so at $x^\ast=0.5$ the parameter part is already only $0.009$ against $\sigma^2=0.25$.

No amount of data removes the noise of the next reading; data only remove the doubt about β.

D · interleaved 4 questions
1§06.4 — a defect rate and a sensor slope

A factory's defect rate has the prior $\mathrm{Beta}(2,8)$, and $3$ of $10$ inspected items are defective. Separately, a sensor slope has the posterior $\mathcal N(0.6,\,0.1)$.

Find
  1. (a) Find the MAP estimate and the posterior mean of the defect rate.

  2. (b) Find the MAP estimate and the posterior mean of the slope.

  3. (c) Why do the two estimates agree in one case and not in the other?

Given
  • defect rate: prior $\mathrm{Beta}(2,8)$; data: $3$ defective of $10$

  • $\mathrm{Beta}(a,b)$: mode $\frac{a-1}{a+b-2}$, mean $\frac{a}{a+b}$; Bernoulli data update it to $\mathrm{Beta}(a+N_1,\,b+N_0)$

  • slope: posterior $\mathcal N(0.6,\,0.1)$

Hint 1/4

Each part needs the posterior's peak and its average; find the posterior first where it is not given.

Hint 2/4

$\mathrm{Beta}(a,b)$ prior with $N_1$ ones and $N_0$ zeros gives $\mathrm{Beta}(a+N_1,\,b+N_0)$; a Gaussian's mode equals its mean.

Hint 3/4

Here $a=2$, $b=8$, $N_1=3$ and $N_0=7$, so the defect rate's posterior is $\mathrm{Beta}(5,15)$; the slope's posterior is $\mathcal N(0.6,0.1)$.

Hint 4/4

Defect rate: MAP $4/18\approx0.222$ and mean $5/20=0.25$. Slope: both $0.6$.

Show solution

The defect rate needs its posterior first; the slope's posterior is already given.

Defect rate

$$\mathrm{Beta}(2+3,\ 8+7)=\mathrm{Beta}(5,15)$$

Add the defective count to $a$ and the good count to $b$.

$$\text{MAP}=\tfrac{5-1}{5+15-2}=\tfrac{4}{18}\approx0.222,\qquad \text{mean}=\tfrac{5}{20}=0.25$$

The two Beta formulas.

Slope

$$\text{MAP}=\text{mean}=0.6$$

A Gaussian density peaks at its mean.

Answer $$\boxed{\text{rate: }0.222\ \text{and}\ 0.25;\qquad \text{slope: }0.6\ \text{and}\ 0.6}$$
Check

The defect rate's posterior is bounded at $0$ and has more mass to the right of its peak, which is why its mean exceeds its mode.

For Gaussian posteriors, as in Bayesian linear regression with a Gaussian prior, the MAP estimate and the posterior mean are the same vector.

2§06.2 — three penalties for one coefficient

One centered input with $\sum_ix_i^2=10$ is fixed; the labels follow $y_i=\beta^{\mathrm{true}}x_i+\varepsilon_i$ with $\beta^{\mathrm{true}}=0.5$ and noise variance $\sigma^2=1$. Ridge is run with $\lambda=0$, $2$ or $5$.

Find
  1. (a) Find $E[\hat\beta^{\mathrm{Ridge}}]$ and $\mathrm{Var}(\hat\beta^{\mathrm{Ridge}})$ as functions of $\lambda$.

  2. (b) Compute bias² plus variance for $\lambda=0,2,5$ and pick the best of the three.

Given
  • $\sum_ix_i^2=10$, $\beta^{\mathrm{true}}=0.5$, $\sigma^2=1$

  • $\hat\beta^{\mathrm{Ridge}}=\sum_ix_iy_i/(10+\lambda)$

  • error $=\text{bias}^2+\text{variance}$ of $\hat\beta^{\mathrm{Ridge}}$ around $\beta^{\mathrm{true}}$

Hint 1/4

$\hat\beta^{\mathrm{Ridge}}$ is a fixed multiple of $\sum_ix_iy_i$, so its mean and variance follow from those of that sum.

Hint 2/4

$E[\sum_ix_iy_i]=\beta^{\mathrm{true}}\sum_ix_i^2$ and $\mathrm{Var}(\sum_ix_iy_i)=\sigma^2\sum_ix_i^2$.

Hint 3/4

With $\sum x_i^2=10$, $\beta^{\mathrm{true}}=0.5$ and $\sigma^2=1$: $E=\frac{10}{10+\lambda}\cdot0.5$ and $\mathrm{Var}=\frac{10}{(10+\lambda)^2}$.

Hint 4/4

The totals are $0.1$, $0.0764$ and $0.0722$, so $\lambda=5$ is the best of the three.

Show solution

The estimate is linear in the labels, so its mean and variance come from the rules for sums.

Mean and variance

$$E[\hat\beta]=\tfrac{0.5\cdot10}{10+\lambda}=\tfrac{5}{10+\lambda},\qquad \mathrm{Var}(\hat\beta)=\tfrac{1\cdot10}{(10+\lambda)^2}$$

Constants come out of expectations unchanged and out of variances squared.

Three totals

$$\lambda=0:\ \ 0+0.1=0.1$$

Least squares: unbiased, all variance.

$$\lambda=2:\ \ (0.4167-0.5)^2+\tfrac{10}{144}\approx0.0069+0.0694=0.0764$$

A little bias buys a smaller variance.

$$\lambda=5:\ \ (0.3333-0.5)^2+\tfrac{10}{225}\approx0.0278+0.0444=0.0722$$

More bias, even less variance, a smaller total.

Answer $$\boxed{0.1,\ \ 0.0764,\ \ 0.0722:\qquad \lambda=5\ \text{wins}}$$
Check

Further reading: minimizing over every $\lambda$ gives $\lambda=\sigma^2/(\beta^{\mathrm{true}})^2=4$ with total $0.0714$, just below $0.0722$. That is the Bayesian $\lambda=\sigma^2/\tau^2$ with the prior variance equal to $(\beta^{\mathrm{true}})^2$.

The penalty that minimizes the expected error matches a prior whose spread is the size of the true coefficient.

3§06.4 — two reports on one slope

A centered one-input data set has $\sum_ix_i^2=8$, and the noise variance is $\sigma^2=2$. A frequentist reports the sampling variance of the least squares slope; a Bayesian with the prior $\mathcal N(0,1)$ reports a posterior variance.

Find
  1. (a) Compute both variances.

  2. (b) State what each one describes.

Given
  • $\sum_ix_i^2=8$, $\sigma^2=2$

  • $\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\sigma^2/\sum_ix_i^2$ over repeated data sets

  • prior $\mathcal N(0,1)$

Hint 1/4

One variance is about the estimator over repeated data sets, the other about β given this one data set.

Hint 2/4

$\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\sigma^2/\sum x_i^2$ and $\Sigma_{\beta|D}=1/(1/\tau^2+\sum x_i^2/\sigma^2)$.

Hint 3/4

Here $\sigma^2=2$, $\sum x_i^2=8$ and $\tau^2=1$.

Hint 4/4

$0.25$ and $1/(1+4)=0.2$.

Show solution

The two formulas answer different questions, so we compute each and then say what it describes.

Frequentist

$$\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\tfrac{2}{8}=0.25$$

Here the randomness is the noise in a new data set.

Bayesian

$$\Sigma_{\beta|D}=\tfrac{1}{1+8/2}=0.2$$

Here the randomness is our uncertainty about $\beta$; the prior adds precision $1$.

Answer $$\boxed{0.25\ \ \text{vs}\ \ 0.2}$$
Check

Flat-prior limit: with $\tau^2\to\infty$ the posterior variance becomes $1/(8/2)=0.25$, the frequentist number, although its meaning still differs.

A sampling variance describes an estimator; a posterior variance describes β.

4§06.6 — a combination of two correlated readings

A Gaussian random vector $Z=(Z_1,Z_2)$ has mean $(1,2)$ and covariance $\begin{pmatrix}2&1\\1&3\end{pmatrix}$. We need the distribution of $W=Z_1-Z_2+1$.

Find
  1. (a) Find the distribution of $W$.

  2. (b) Which quantities of the predictive formula play the roles of the vector $a$, the mean and the covariance here?

Given
  • $Z\sim\mathcal N\big((1,2)^T,\ \Sigma\big)$ with $\Sigma=\begin{pmatrix}2&1\\1&3\end{pmatrix}$

  • $W=a^TZ+c$ with $a=(1,-1)^T$ and $c=1$

Hint 1/4

$W$ is a linear function of a Gaussian vector plus a constant.

Hint 2/4

$a^TZ+c\sim\mathcal N(a^T\mu+c,\ a^T\Sigma a)$, with $a^T\Sigma a=\Sigma_{11}a_1^2+2\Sigma_{12}a_1a_2+\Sigma_{22}a_2^2$.

Hint 3/4

Here $a=(1,-1)$, $\mu=(1,2)$, $c=1$, $\Sigma_{11}=2$, $\Sigma_{12}=1$ and $\Sigma_{22}=3$.

Hint 4/4

Mean $1-2+1=0$ and variance $2-2+3=3$: $W\sim\mathcal N(0,3)$.

Show solution

The rule for linear functions of Gaussian vectors gives the answer in two lines.

Mean

$$a^T\mu+c=1-2+1=0$$

Expectation is linear.

Variance

$$a^T\Sigma a=2(1)^2+2(1)(1)(-1)+3(-1)^2=3$$

The cross term enters with the sign of $a_1a_2$.

Analogy

$$x^T\beta\ \ \text{with}\ \ \beta\mid D\sim\mathcal N(\mu_{\beta|D},\Sigma_{\beta|D})$$

The predictive variance is $x^T\Sigma_{\beta|D}x$ plus the independent noise $\sigma^2$.

Answer $$\boxed{W\sim\mathcal N(0,\ 3)}$$
Check

Directly: $\mathrm{Var}(Z_1-Z_2)=\mathrm{Var}Z_1+\mathrm{Var}Z_2-2\,\mathrm{Cov}(Z_1,Z_2)=2+3-2=3$.

Every predictive variance on this page is one of these quadratic forms, plus $\sigma^2$.

Mistake ledger (14 entries)
⚠ Maximizing one reading's density instead of the product

Each factor is largest at a zero residual, and one reading can always be hit exactly.

wrong$$\hat\beta=y_1/x_1=1\ \ \text{(reading 1 exact)}$$
right$$\hat\beta=\frac{\sum_i x_iy_i}{\sum_i x_i^2}=0.5$$
⚠ Letting the noise variance move the maximizer

$\sigma^2$ sits next to the RSS in $l(\beta)$, so it looks like part of the fit.

wrong$$\hat\beta^{\mathrm{MLE}}=\frac{\sum_i x_iy_i}{\sum_i x_i^2+\sigma^2}$$
right$$\hat\beta^{\mathrm{MLE}}=\frac{\sum_i x_iy_i}{\sum_i x_i^2}\ \ \text{for every }\sigma^2>0$$
⚠ Forgetting the noise variance in the translation

The prior alone looks like the whole story, so λ gets read as its precision.

wrong$$\lambda=\frac{1}{\tau^2}$$
right$$\lambda=\frac{\sigma^2}{\tau^2}$$
⚠ Reading σ²/λ as a standard deviation

$\mathcal N(\mu,\cdot)$ is written with a variance, while many texts quote spreads as standard deviations.

wrong$$\tau=\frac{\sigma^2}{\lambda}$$
right$$\tau^2=\frac{\sigma^2}{\lambda},\qquad \tau=\frac{\sigma}{\sqrt\lambda}$$
⚠ Using the ridge pattern σ²/λ for the Laplace scale

The ridge translation is fresh in mind, and both formulas divide by λ.

wrong$$b=\frac{\sigma^2}{\lambda}$$
right$$b=\frac{2\sigma^2}{\lambda}$$
⚠ Taking b² as the Laplace variance

For the Gaussian the second parameter is the variance, and the habit carries over.

wrong$$\mathrm{Var}(Z)=b^2$$
right$$\mathrm{Var}(Z)=2b^2$$
⚠ Differentiating |β| at 0 as if β were positive

The absolute value is usually met on one side of zero only.

wrong$$\tfrac{d}{d\beta}\lvert\beta\rvert\,\Big|_{\beta=0}=1$$
right$$\text{slope }+1\text{ right of }0,\ \ -1\text{ left of }0:\ \text{check both sides}$$
⚠ Adding covariances instead of precisions

Adding uncertainties sounds natural, but it is information that adds.

wrong$$\Sigma_{\beta|D}=\Sigma_0+\big(X^T\Sigma_D^{-1}X\big)^{-1}$$
right$$\Sigma_{\beta|D}=\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)^{-1}$$
⚠ Dropping the prior-mean term

In the ridge case $\mu_0=0$ and the term vanishes, so it is easy to forget it exists.

wrong$$\mu_{\beta|D}=\Sigma_{\beta|D}X^T\Sigma_D^{-1}y\ \ \text{although}\ \mu_0\neq0$$
right$$\mu_{\beta|D}=\Sigma_{\beta|D}\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big)$$
⚠ Restarting from the original prior at each step

The prior is written in the problem, the intermediate posterior is not.

wrong$$\tfrac{1}{v_2}=\tfrac{1}{v_0}+\tfrac{x_2^2}{\sigma^2}$$
right$$\tfrac{1}{v_2}=\tfrac{1}{v_1}+\tfrac{x_2^2}{\sigma^2}=\tfrac{1}{v_0}+\tfrac{x_1^2+x_2^2}{\sigma^2}$$
⚠ Averaging the old mean and the new reading's slope

An average is the familiar way to combine two numbers.

wrong$$m_{\text{new}}=\tfrac12\big(m_{\text{old}}+\tfrac{y}{x}\big)$$
right$$m_{\text{new}}=v_{\text{new}}\Big(\tfrac{m_{\text{old}}}{v_{\text{old}}}+\tfrac{xy}{\sigma^2}\Big)$$
⚠ Dropping the noise term

The posterior of β feels like the whole uncertainty.

wrong$$\mathrm{Var}(Y\mid x,D)=x^T\Sigma_{\beta|D}x$$
right$$\mathrm{Var}(Y\mid x,D)=\sigma^2+x^T\Sigma_{\beta|D}x$$
⚠ Using the prior covariance instead of the posterior

$\Sigma_0$ is the covariance written in the problem.

wrong$$\sigma^2+x^T\Sigma_0\,x$$
right$$\sigma^2+x^T\Sigma_{\beta|D}\,x$$
⚠ Putting a variance where the error bar needs a standard deviation

The formula delivers a variance, and the interval is written right after it.

wrong$$\mu_{\beta|D}^Tx\pm2\big(\sigma^2+x^T\Sigma_{\beta|D}x\big)$$
right$$\mu_{\beta|D}^Tx\pm2\sqrt{\sigma^2+x^T\Sigma_{\beta|D}x}$$
Formula card
Least squares = maximum likelihood
$$\hat\beta^{\mathrm{MLE}}=\arg\max_\beta l(\beta)=\arg\min_\beta\mathrm{RSS}(\beta)=(X^TX)^{-1}X^Ty$$

i.i.d. Gaussian noise with known $\sigma^2$; $X^TX$ invertible

Ridge = MAP under a Gaussian prior
$$\beta_j\sim\mathcal N\big(0,\tfrac{\sigma^2}{\lambda}\big)\ \Rightarrow\ \hat\beta^{\mathrm{MAP}}=(X^TX+\lambda I)^{-1}X^Ty$$

$\lambda=\sigma^2/\tau^2$ for a prior variance $\tau^2$

Lasso = MAP under a Laplace prior
$$\beta_j\sim\mathrm{Lap}\big(0,\tfrac{2\sigma^2}{\lambda}\big)\ \Rightarrow\ \hat\beta^{\mathrm{MAP}}=\hat\beta^{\mathrm{Lasso}}$$

$\mathrm{Lap}(\mu,b)$ has density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$ and variance $2b^2$

One-coefficient lasso (further reading)
$$\hat\beta=\operatorname{sign}(c)\,\max\Big(\frac{\lvert c\rvert-\lambda/2}{s},\,0\Big),\qquad s=\textstyle\sum x_i^2,\ c=\sum x_iy_i$$

one coefficient, centered data

Gaussian posterior
$$\begin{aligned}\Sigma_{\beta|D}&=\big(\Sigma_0^{-1}+X^T\Sigma_D^{-1}X\big)^{-1}\\ \mu_{\beta|D}&=\Sigma_{\beta|D}\big(X^T\Sigma_D^{-1}y+\Sigma_0^{-1}\mu_0\big)\end{aligned}$$

prior $\mathcal N(\mu_0,\Sigma_0)$, likelihood $\mathcal N(X\beta,\Sigma_D)$; ridge: $\Sigma_{\beta|D}=\sigma^2(X^TX+\lambda I)^{-1}$

Sequential update
$$\frac{1}{v_n}=\frac{1}{v_{n-1}}+\frac{x_n^2}{\sigma^2},\qquad m_n=v_n\Big(\frac{m_{n-1}}{v_{n-1}}+\frac{x_ny_n}{\sigma^2}\Big)$$

one coefficient; $m_0=\mu_0$, $v_0=\tau^2$

Posterior predictive
$$Y\mid x,D\sim\mathcal N\big(\mu_{\beta|D}^Tx,\ \sigma^2+x^T\Sigma_{\beta|D}x\big)$$

Gaussian posterior; new noise independent of β

Check yourself

Close the page and write down from memory:

  • the log-likelihood under Gaussian noise, and why its maximizer is least squares;
  • the two priors behind ridge and the lasso, with λ in each;
  • the posterior mean and covariance formulas, and the ridge special case;
  • the predictive mean and variance at a new input.

Then reopen the page and compare; whatever is missing is your reread list.

  • Derive $l(\beta)$ for Gaussian noise and show that its maximizer does not depend on $\sigma^2$?

    c-ls-mle

  • Turn a prior $\mathcal N(0,\tau^2)$ into a ridge penalty, and a penalty into a prior variance?

    c-ridge-map

  • Give the Laplace prior for a lasso penalty, and find a one-coefficient lasso estimate, including the zero case?

    c-lasso-map

  • Compute $\mu_{\beta|D}$ and $\Sigma_{\beta|D}$ for two coefficients, with and without a nonzero prior mean?

    c-bayes-posterior

  • Update a posterior one reading at a time and match the batch answer?

    c-sequential

  • Compute a predictive mean, variance and interval, and say which part of the variance dominates at a given input?

    c-predictive

Glossary (12 terms)
log-likelihoodlog olabilirlik

The natural logarithm of the likelihood, $l(\beta)=\log L(\beta)$. It has the same maximizer and turns a product over readings into a sum.

Minus the logarithm of the posterior density. Up to a constant it is the loss a MAP estimate minimizes, such as the ridge or lasso loss divided by $2\sigma^2$.

Gaussian prior

A normal distribution put on the coefficients before the data; with mean $0$ and variance $\sigma^2/\lambda$ its MAP estimate is ridge regression.

Laplace distributionLaplace dağılımı

$\mathrm{Lap}(\mu,b)$, with density $\frac{1}{2b}e^{-\lvert z-\mu\rvert/b}$, mean $\mu$ and variance $2b^2$; it has a sharp corner at its centre.

Laplace prior

An i.i.d. $\mathrm{Lap}(0,2\sigma^2/\lambda)$ distribution on the coefficients; its MAP estimate is the lasso.

modemod

The value at which a density is highest. The MAP estimate is the mode of the posterior; for a Gaussian the mode equals the mean.

biased estimatoryanlı kestirici

An estimator whose average over repeated data sets differs from the true value; ridge and the lasso are biased toward $0$ when $\lambda>0$.

Bayesian linear regressionBayesçi doğrusal regresyon

Linear regression that keeps the whole posterior distribution of the coefficients instead of one estimate; with a Gaussian prior and Gaussian noise the posterior is Gaussian.

precision

The inverse of a variance or of a covariance matrix. Large precision means small spread; in the Gaussian posterior the precisions of the prior and of the data add.

sequential updating

Using the posterior after some readings as the prior for the next ones; after all readings it equals the posterior computed from all of them at once.

posterior predictive distribution

The distribution of a new response given the data, $p(y\mid x,D)=\int p(y\mid x,\beta)\,p(\beta\mid D)\,d\beta$; here it is $\mathcal N\big(\mu_{\beta|D}^Tx,\ \sigma^2+x^T\Sigma_{\beta|D}x\big)$.

A prediction that treats an estimate of β as exact, so its error bar contains the noise only; it is too narrow far from the data.

What comes next
§07 · Generalized linear models

Here the noise was Gaussian, so minus the log-likelihood was a sum of squares. Next the response can be a yes/no label or a count, the likelihood changes shape, and the same recipe of writing it down and maximizing it leads to new models.

Sources
  • textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The weekly line names no chapter of the book, so no section numbers are cited here.
  • course materialEEE 485 lecture slides and lecture notes, chapter 6: Linear regression from a Bayesian perspective Scope, order of topics and notation (the LS, MLE, MAP, Ridge and Lasso estimates, the two losses, Lap(μ, b), the posterior mean and covariance, the predictive density) follow these materials; all explanations, data, figures and exercises here are original.
  • course materialEEE 485 syllabus page on STARS, Fall 2026-27, printed 21 September 2026 Assessment weights, the weekly topic list and the recommended books.
  • textbookK. P. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press, 2012 Recommended in the syllabus; the lecture's figures for Bayesian linear regression come from it.
  • textbookG. James, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning, Springer, 2013 Recommended in the syllabus; the lecture's picture of the ridge and lasso constraint regions comes from it.
  • standard resultCompleting the square for Gaussian densities; linear combinations of Gaussian vectors Standard results used in the derivations. Every number and figure on this page was computed for it, the simulated readings from a fixed random seed.

Spotted something missing or wrong? tell us · share your own notes or an old exam.

Last updated .