← back to EEE 485
Week 5126 min full read
7 concepts18 worked examples31 exercises5 exam-level7 figures
What are you here for?

05 Regularized regression: ridge and lasso

Start with this

One question before you read anything. Getting it wrong is the point: it shows you what this section is for.

§05.2 — what a does to the training RSS

A least squares fit to a training set has $\mathrm{RSS}=20$. On the same training set you now minimize $\mathrm{RSS}(\beta)+3\sum_j\beta_j^2$ instead.

Find(a) Before computing anything, what can you say about the training RSS of the new minimizer?
Given
  • least squares on the training set: $\mathrm{RSS}=20$

  • new objective: $\mathrm{RSS}(\beta)+3\sum_{j}\beta_j^2$, same data

Hint 1/4

Compare the new minimizer with the one fit whose training RSS is known to be as small as possible.

Hint 2/4

Least squares minimizes $\mathrm{RSS}(\beta)$ over all $\beta$: $\mathrm{RSS}(\beta)\ge\mathrm{RSS}(\hat\beta^{\mathrm{LS}})$ for every $\beta$.

Hint 3/4

Here $\mathrm{RSS}(\hat\beta^{\mathrm{LS}})=20$, and the penalized minimizer is one more $\beta$ on the same training set.

Hint 4/4

The new training RSS is at least $20$.

Show solution

One inequality settles it, so we use the definition of least squares instead of computing a fit.

Use the definition of least squares

$$\mathrm{RSS}(\hat\beta^R)\ \ge\ \min_\beta\mathrm{RSS}(\beta)=\mathrm{RSS}(\hat\beta^{\mathrm{LS}})=20$$

The penalized minimizer is one particular $\beta$, and no $\beta$ goes below the minimum.

See what the penalty buys

$$\mathrm{RSS}(\hat\beta^R)+3P(\hat\beta^R)\ \le\ 20+3P(\hat\beta^{\mathrm{LS}})$$

The penalized fit wins on its own objective; here $P=\sum_j\beta_j^2$.

$$\Rightarrow\ \ P(\hat\beta^R)\ \le\ P(\hat\beta^{\mathrm{LS}})$$

A higher RSS can only be paid for with a smaller penalty.

Answer $$\boxed{\mathrm{RSS}(\hat\beta^R)\ \ge\ 20}$$
Check

On the kiosk data of this section the RSS rises from $20$ to $25.625$ at $\lambda=3$, while $\sum_j\beta_j^2$ falls from $10$ to $1.625$.

A penalty always costs training fit; whether it pays off is decided on data the fit has not seen.

Two thermometers hang side by side outside a kiosk. On five days they read $18, 20, 21, 22, 24$ and $20, 18, 21, 22, 24$ degrees, and a least squares fit of daily sales on both readings gives the second thermometer a negative coefficient: by the fit, its warmth costs sales. Had the two thermometers swapped their first two readings, the fit would have blamed the first thermometer instead.

By the end you can compute, by hand, a fit that gives both thermometers a small positive weight, choose from the data how strongly to shrink, and say why a second kind of penalty keeps only one thermometer.

In 60 seconds

Ridge and the lasso are least squares plus a price $\lambda$ on the size of the coefficients: ridge's $\lambda\sum_j\beta_j^2$ shrinks every coefficient smoothly, with the closed form $(X^TX+\lambda I)^{-1}X^Ty$, and the lasso's $\lambda\sum_j\lvert\beta_j\rvert$ sets some coefficients exactly to zero, with $\lambda$ chosen by cross-validation.

Ridge estimate
$$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$$

centered response, standardized predictors and any $\lambda>0$, even when $X^TX$ is singular

One standardized predictor
$$\hat\beta^R=\frac{n}{n+\lambda}\,\hat\beta^{\mathrm{LS}},\qquad \hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\Big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\frac{\lambda}{2n}\Big)_+$$

one predictor with $\sum_ix_i^2=n$, or uncorrelated standardized predictors one at a time

Prediction at a new input
$$\tilde z_j=\frac{z_j-\bar x_j}{\sigma_j},\qquad \hat y=\bar y+\tilde z^T\hat\beta^R$$

any new raw input; the means and standard deviations come from the training data

Constrained forms
$$\min_\beta\,\mathrm{RSS}(\beta)\ \ \text{s.t.}\ \ \sum_j\beta_j^2\le s\ \ \text{or}\ \ \sum_j\lvert\beta_j\rvert\le s$$

geometry questions; each budget $s$ matches one $\lambda$

Three most common mistakes
  1. Fitting ridge or the lasso on raw predictors: the penalty then depends on the units, so recording income in thousands of TL instead of TL changes the predictions. Standardize first, with the training means and standard deviations.

  2. Choosing $\lambda$ by the training RSS: it never falls as $\lambda$ grows, so it always picks $\lambda=0$, plain least squares. Use cross-validation.

  3. Expecting zeros from ridge: $(X^TX+\lambda I)^{-1}X^Ty$ shrinks the coefficients but, apart from coincidences, leaves every one of them nonzero. Only the lasso holds coefficients at exactly zero.

Two course documents give different weights:

  • Chapter 1 slides, undergraduate line: midterm 25, final 25, four quizzes 20, two-phase project 30.
  • STARS syllabus page printed on 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. It is the later document; confirm which split applies.
  • The syllabus assesses the outcome 'apply probability and linear algebra knowledge to analyze the performance of statistical learning algorithms' in the midterm, the final, the problem sets and quizzes, and the project.
How much time do you have?
10 minutes

The ridge formula with the hand computation that settles the opening puzzle, and the lasso's one-predictor rule with the reason it produces zeros.

The 60-second card · Ridge regression · The lasso · Formula card
45 minutes

Every block once with its first worked example and checkpoint, then one ladder from a full ridge computation down to a bare problem.

The 60-second card · Ridge regression · The tuning parameter λ · Scale matters for ridge · Why shrinking helps · Ridge always has one answer · The lasso · Budgets and shapes · Scaffolding comes off · Formula card
full read

The derivations, the look-alike pairs and enough mixed practice to decide on your own which tool a question needs.

The opening pages · Recall first · Ridge regression · The tuning parameter λ · Scale matters for ridge · Why shrinking helps · Ridge always has one answer · The lasso · Budgets and shapes · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
  1. Compute the ridge estimate $(X^TX+\lambda I)^{-1}X^Ty$ for two predictors by hand, and derive it by setting the gradient of the ridge loss to zero.

  2. Describe how the ridge coefficients move as $\lambda$ runs from $0$ to $\infty$, and choose $\lambda$ by cross-validation.

  3. Standardize the predictors, show why ridge needs this while least squares does not, and turn a new raw input into a prediction.

  4. Quantify the bias and the variance of a ridge coefficient, and find the $\lambda$ with the smallest expected test error for one predictor.

  5. Prove that $X^TX+\lambda I$ is invertible for every $\lambda>0$, and use ridge when least squares has no unique solution.

  6. Compute lasso estimates for one predictor and for uncorrelated standardized predictors, and explain why the lasso selects variables and ridge does not.

  7. Rewrite ridge and the lasso as constrained problems, match a budget $s$ to a $\lambda$, and explain the lasso's zeros with the diamond picture.

Syllabus coverage

Regularized regression — covered

Why shrink at all: least squares has low bias but can have high variance when predictors are nearly collinear or $p$ is close to $n$, and a penalty on the size of the coefficients trades a little bias for less variance.

The lecturer's chapter title, 'Regularized regression: ridge and lasso'. The motivation opens the first block and the tradeoff is worked out in the bias-variance block.

Ridge regression — covered

  • The ridge loss and its closed form
  • the choice of $\lambda$ by cross-validation
  • bias and variance against $\lambda$
  • invertibility of $X^TX+\lambda I$
  • ridge keeps every predictor

Official weekly line. Ridge fills the first five blocks, in the lecture's order.

lasso — covered

The lasso loss, its limits as $\lambda\to0$ and $\lambda\to\infty$, and sparse models, the constrained forms and the diamond picture.

The lasso has no closed form in general; the one-predictor formula is the exception.

Bayesian linear regression — deferred

Regression with a prior distribution on the coefficients.

The official line groups it with ridge and the lasso; it is the fourth line of the STARS weekly list. The lecturer teaches it as the next chapter, 'Linear regression from a Bayesian perspective', and the next section covers it. Check the course schedule for the week.

The one-predictor lasso formula — off syllabus

: with $b=\hat\beta^{\mathrm{LS}}$, $\hat\beta^L=\operatorname{sign}(b)(\lvert b\rvert-\lambda/2n)_+$ for one standardized predictor, and coefficient by coefficient for uncorrelated ones.

Further reading. The slides say the lasso has no closed form in general; in one dimension it has one, and the formula makes the zeros visible.

Checking a zero coefficient — off syllabus

A lasso coefficient belongs at $0$ exactly when $\lvert\partial\mathrm{RSS}/\partial\beta_j\rvert\le\lambda$ there; all coefficients are $0$ once $\lambda\ge2\max_j\lvert x_j^Ty\rvert$.

Further reading: an elementary check used in two worked examples, the exam example and one practice question.

The best $\lambda$ for one predictor — off syllabus

$\lambda^*=\sigma^2/\beta^2$ from the exact bias and variance of a ridge coefficient.

Further reading: the lecture shows the tradeoff as curves; this block derives them for one predictor so that every number can be checked.

Ridge as $\lambda\to0$ with a singular $X^TX$ — off syllabus

The limit is the least squares fit with the smallest sum of squared coefficients.

Further reading, used in one figure, one worked example and one practice question.

Recall first
Least squares in matrix form

$\mathrm{RSS}(\beta)=(y-X\beta)^T(y-X\beta)$, and the normal equations $X^TX\hat\beta=X^Ty$ give $\hat\beta=(X^TX)^{-1}X^Ty$ when the columns of $X$ are linearly independent.

Ridge changes one term of these equations, and least squares is the case $\lambda=0$ throughout.

Matrix derivatives, row convention

$\frac{\partial}{\partial\beta}(a^T\beta)=a^T$ and, for symmetric $A$, $\frac{\partial}{\partial\beta}(\beta^TA\beta)=2\beta^TA$. Setting a row derivative to zero and transposing gives the same equations as a column gradient.

The ridge formula comes from one such derivative.

The 2 × 2 inverse and invertibility

$\begin{pmatrix}a&b\\c&d\end{pmatrix}^{-1}=\frac{1}{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}$ when $ad-bc\neq0$. A square matrix $M$ is invertible exactly when $Mv=0$ has only the solution $v=0$.

Every hand computation here has two predictors, and ridge's main advantage is an invertibility argument.

Directions a matrix only stretches

If $Av=d\,v$ for some $v\neq0$, then $(A+\lambda I)v=(d+\lambda)v$ and $(A+\lambda I)^{-1}v=\frac{v}{d+\lambda}$. For $A=\begin{pmatrix}a&c\\c&a\end{pmatrix}$ the directions $(1,1)$ and $(1,-1)$ are stretched by $a+c$ and $a-c$.

It shows why ridge shrinks some combinations of the coefficients much harder than others.

Bias-variance decomposition

Expected test error $=\text{bias}^2+\text{variance}+\text{noise}$: bias² compares the average model with the expected label, and the variance measures how far one fit strays from the average model.

Ridge buys a little bias for a large cut in variance.

Cross-validation

$\mathrm{CV}(k)=\frac1k\sum_{i=1}^k\mathrm{MSE}_i$, where fold $i$ is scored by a model fitted without it; with $k=n$ this is leave-one-out, $\mathrm{CV}(n)$.

$\lambda$ is chosen this way.

Scaled estimates

$E[cZ]=c\,E[Z]$ and $\mathrm{Var}(cZ)=c^2\mathrm{Var}(Z)$. For one centered predictor with fixed inputs and noise variance $\sigma^2$, $\mathrm{Var}(\hat\beta_1)=\sigma^2/\sum_ix_i^2$.

With one predictor a ridge coefficient is a scaled least squares coefficient.

Confidence intervals

For an unbiased, normally distributed estimate, $\hat\theta\pm2\,\mathrm{SE}$ covers the true value about $95$ times in $100$ repetitions.

One interleaved question asks what happens to this recipe when the estimate is biased.

Try it yourself first (2 questions)
1§05.5 — adding to the diagonal of a singular matrix

Two predictors give $X^TX=\begin{pmatrix}4&2\\2&1\end{pmatrix}$. Before any regression, check two matrices for an inverse.

Find(a) Which of $X^TX$ and $X^TX+I$ has an inverse?
Given
  • $X^TX=\begin{pmatrix}4&2\\2&1\end{pmatrix}$

  • $I=\begin{pmatrix}1&0\\0&1\end{pmatrix}$

Hint 1/4

An inverse exists exactly when a determinant is nonzero; decide which two determinants to compute.

Hint 2/4

$\det\begin{pmatrix}a&b\\c&d\end{pmatrix}=ad-bc$, and adding $I$ adds $1$ to each diagonal entry.

Hint 3/4

The two matrices are $X^TX=\begin{pmatrix}4&2\\2&1\end{pmatrix}$ and $X^TX+I=\begin{pmatrix}5&2\\2&2\end{pmatrix}$.

Hint 4/4

$\det(X^TX)=0$ and $\det(X^TX+I)=6$, so only $X^TX+I$ has an inverse.

Show solution

A $2\times2$ matrix is invertible exactly when its determinant is nonzero, which is quicker to check than attempting the inverse.

The original matrix

$$\det(X^TX)=4\cdot1-2\cdot2=0$$

The second column is half the first, so the columns are dependent.

After adding the identity

$$\det(X^TX+I)=5\cdot2-2\cdot2=6\neq0$$

Only the diagonal grew, and that alone broke the dependence.

Answer $$\boxed{\text{only }X^TX+I\text{ is invertible}}$$
Check

Direct check: $X^TX\,(1,\,-2)^T=(0,\,0)^T$, so $X^TX$ sends a nonzero vector to zero and has no inverse, while $(X^TX+I)(1,\,-2)^T=(1,\,-2)^T\neq0$.

A singular $X^TX$ blocks least squares, and adding $\lambda I$ with $\lambda>0$ removes the block.

2§05.3 — least squares after a change of units

A least squares fit predicts monthly spending from income recorded in TL, and the income coefficient is $0.0004$. The same data are refitted with income recorded in thousands of TL.

Find(a) What is the income coefficient after the refit?
Given
  • income coefficient with income in TL: $0.0004$

  • new unit: $1$ thousand TL

Hint 1/4

Least squares keeps its predictions under a change of units; ask what the coefficient must do to keep them.

Hint 2/4

If every value of a predictor is multiplied by $c$, the least squares coefficient is divided by $c$.

Hint 3/4

Income in thousands of TL is income in TL times $c=\frac{1}{1000}$, and the old coefficient is $0.0004$.

Hint 4/4

The new coefficient is $0.0004\times1000=0.4$.

Show solution

of least squares turns this into one multiplication, so we use it instead of refitting.

Keep the predictions

$$\hat\beta_{\text{new}}\cdot\frac{x}{1000}=\hat\beta_{\text{old}}\cdot x$$

Least squares sees the same data in new units and returns the same fitted values.

Solve for the new coefficient

$$\hat\beta_{\text{new}}=1000\times0.0004=0.4$$

The equation must hold for every income $x$, so the factors in front of $x$ must match.

Answer $$\boxed{0.4\ \text{per thousand TL}}$$
Check

Units check: $0.0004$ per TL times $1000$ TL per unit is $0.4$ per unit, and an income of $25\,000$ TL gives $0.0004\times25\,000=10=0.4\times25$ either way.

Least squares absorbs any change of units into its coefficients; ridge will not, which is why this section standardizes.

Notation
symbolreads asmeanswatch out
$\lambda$

lambda, the tuning parameter

the price per unit of penalty, fixed before fitting, $\lambda\ge0$

Chosen by cross-validation, never by the training RSS.

$\mathrm{Loss}_R(\beta,\lambda),\ \mathrm{Loss}_L(\beta,\lambda)$

ridge loss, lasso loss

$\mathrm{RSS}(\beta)$ plus $\lambda\sum_j\beta_j^2$, or plus $\lambda\sum_j\lvert\beta_j\rvert$

The sums start at $j=1$: the intercept $\beta_0$ is never penalized.

$\hat\beta^{\mathrm{LS}},\ \hat\beta^R,\ \hat\beta^L$

beta hat LS, R, L

the least squares, ridge and lasso estimates

The least squares section wrote $\hat\beta_{\mathrm{RSS}}$; the superscript leaves room for an index, as in $\hat\beta^R_j$.

$\hat\beta^R_\lambda$

ridge estimate at lambda

the ridge estimate for one particular $\lambda$

Used when several values of $\lambda$ are compared.

$p,\ n$

p, n

the number of predictors and the number of observations

Least squares gets into trouble when $p$ is close to $n$ or larger.

$\bar x_j,\ \sigma_j,\ \bar y$

x bar j, sigma j, y bar

mean and standard deviation of predictor $j$, and the mean response, all on the training data

$\sigma_j$ divides by $n$, as in the lecture, not by $n-1$.

$\tilde x_{ij},\ \tilde z_j$

x tilde i j, z tilde j

a standardized training value and a standardized new input

The new input uses the training $\bar x_j$ and $\sigma_j$.

$\lVert\beta\rVert_2^2,\ \lVert\beta\rVert_1$

squared 2-norm, 1-norm

$\sum_{j=1}^p\beta_j^2$ and $\sum_{j=1}^p\lvert\beta_j\rvert$

Both leave out $\beta_0$.

$s$

the budget

the bound in the constrained forms, $\lVert\beta\rVert_2^2\le s$ or $\lVert\beta\rVert_1\le s$

A larger budget matches a smaller $\lambda$.

$(u)_+,\ \operatorname{sign}(u)$

positive part, sign

$\max(u,0)$, and $+1$, $-1$ or $0$ by the sign of $u$

Used only in the lasso formula for one predictor.

$I$

the

$p\times p$, ones on the diagonal and zeros elsewhere

$\lambda I$ adds $\lambda$ to the diagonal only.

Conventions used here
The intercept is never penalized.

Every penalty sums over $j=1,\dots,p$. After centering the response and every predictor the fitted intercept is $0$, so $X$ has no column of ones in this section, and on standardized inputs the model reads $\hat y=\bar y+\sum_j\hat\beta_j\tilde x_j$.

Penalizing the intercept would make predictions depend on where zero sits on the response scale.

Standardization divides by n.

$\sigma_j=\sqrt{\frac1n\sum_i(x_{ij}-\bar x_j)^2}$, as in the lecture, so every standardized column has $\sum_i\tilde x_{ij}=0$ and $\sum_i\tilde x_{ij}^2=n$.

Software that divides by $n-1$ produces the same family of fits at slightly different values of $\lambda$.

Our λ is the lecture's λ.

The losses are $\mathrm{RSS}+\lambda\cdot\text{penalty}$, with no $\frac12$ and no $\frac1n$ in front of the RSS. Books and packages that minimize $\frac{1}{2n}\mathrm{RSS}+\lambda'\cdot\text{penalty}$ use $\lambda'=\lambda/(2n)$.

The same fit can carry very different $\lambda$ labels in different sources.

Zero means exactly zero.

A coefficient 'at zero' is exactly $0$, which the lasso produces over whole ranges of $\lambda$. A ridge coefficient can pass through $0$ at one isolated $\lambda$; that does not count as selecting a variable.

Selection is about ranges of λ, not about a curve crossing an axis.

Log scales for λ.

Figures and grids of $\lambda$ use powers of ten or of three, because the interesting range spans several orders of magnitude.

A linear grid from 0 to 1000 would spend almost every value on the flat end.

Rounding.

Intermediate steps keep at least four significant figures; final answers are rounded to three decimals unless an exact fraction is asked for.

Ridge and least squares answers often differ only in the second decimal.

5.1Ridge regression: least squares plus a price on the size of the coefficients

Adds $\lambda\sum_j\beta_j^2$ to the RSS so that unstable coefficients shrink; on centered, standardized data the fit is $(X^TX+\lambda I)^{-1}X^Ty$.

Least squares gave us the smallest RSS; on the kiosk data the smallest RSS is exactly what goes wrong.

Solvable with what we have
  • Fit $\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$ whenever the columns of $X$ are independent.

  • Split a model's expected test error into bias², variance and noise.

  • Choose among a few candidate models by cross-validation.

Not solvable yet
  • Get stable coefficients when two predictors almost copy each other.

  • Fit at all when $X^TX$ is singular, for instance when $p>n$.

  • Turn the complexity of a model up or down smoothly instead of dropping whole predictors.

Fit least squares anyway. The five kiosk days, standardized, give $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$ and $X^Ty=(11,\,7)^T$, so $\hat\beta^{\mathrm{LS}}=(3,\,-1)$: each standardized degree on the second thermometer costs one unit of sales.

Why it fails

The data fix the sum $\beta_1+\beta_2$ from how sales follow the average reading, but the difference $\beta_1-\beta_2$ only from days 1 and 2, the two days the thermometers disagree. Least squares sets $\hat\beta_1-\hat\beta_2$ to the sales gap between those two days, $31-27=4$, noise and all; swapping two readings turns it into $-4$.

DefinitionThe ridge loss and its minimizer
Conditions
  • $\lambda\ge0$ is fixed before fitting, and the penalty starts at $j=1$: the intercept is not penalized

  • for the closed form, the response is centered and the predictors are centered and standardized, so $X$ has no column of ones

$$\boxed{\begin{aligned}\mathrm{Loss}_R(\beta,\lambda)&=\mathrm{RSS}(\beta)+\lambda\sum_{j=1}^p\beta_j^2\\ \mathrm{RSS}(\beta)&=\sum_{i=1}^n\Big(y_i-\beta_0-\sum_{j=1}^p\beta_jx_{ij}\Big)^2\\ \textcolor{#1f6feb}{\hat\beta^R}&=\arg\min_\beta\,\mathrm{Loss}_R(\beta,\lambda)=(X^TX+\lambda I)^{-1}X^Ty\end{aligned}}$$

The ridge loss is the least squares score plus a charge of $\lambda$ for every unit of $\sum_j\beta_j^2$. Its minimizer solves $(X^TX+\lambda I)\beta=X^Ty$, the normal equations with $\lambda$ added to each diagonal entry; that is the only change the penalty makes.

Deriving the closed form

The intercept first. It is not penalized, so setting $\partial\mathrm{Loss}_R/\partial\beta_0=-2\sum_i\big(y_i-\beta_0-\sum_j\beta_jx_{ij}\big)$ to zero gives $\hat\beta_0=\bar y-\sum_j\beta_j\bar x_j$. With centered columns every $\bar x_j=0$, and with a centered response $\hat\beta_0=0$: the column of ones drops out.

Write the rest with vectors; it is the kiosk computation with letters instead of numbers. Starting from $(y-X\beta)^T(y-X\beta)+\lambda\beta^T\beta$:

$$\mathrm{Loss}_R=y^Ty-2y^TX\beta+\beta^T(X^TX+\lambda I)\beta.$$

By the two derivative rules recalled above, the row derivative is $-2y^TX+2\beta^T(X^TX+\lambda I)$. Setting it to zero and transposing gives $(X^TX+\lambda I)\beta=X^Ty$, because $X^TX+\lambda I$ is symmetric.

For $\lambda>0$ the matrix $X^TX+\lambda I$ is invertible, and the loss has the second derivative $2(X^TX+\lambda I)$, which is ; the invertibility block below proves both. The loss is therefore a bowl with one lowest point, $\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$.

Looks like this, but is not

Fitting least squares and then multiplying $\hat\beta^{\mathrm{LS}}$ by one shrink factor, say $0.5$, looks like ridge: every coefficient gets smaller.

Ridge does not shrink all directions alike. On the kiosk data it multiplies the well measured sum $\beta_1+\beta_2$ by $\frac{9}{9+\lambda}$ and the badly measured difference by $\frac{1}{1+\lambda}$; at $\lambda=3$ that is $0.75$ against $0.25$. One common factor equals ridge only when $X^TX$ is a multiple of $I$.

Two thermometers, five days: least squares against ridge at λ = 3

The kiosk's readings are $T_1=18, \allowbreak 20, \allowbreak 21, \allowbreak 22, \allowbreak 24$ and $T_2=20, \allowbreak 18, \allowbreak 21, \allowbreak 22, \allowbreak 24$ degrees, with sales $27,31,26,32,34$. Standardize the readings, center the sales, and fit least squares and ridge with $\lambda=3$. Then predict sales on a day with $T_1=23$ and $T_2=19$.

Find$\hat\beta^{\mathrm{LS}}$, $\hat\beta^R$ and both predictions for $T_1=23$, $T_2=19$.
Given
  • $T_1$: $18,20,21,22,24$; $T_2$: $20,18,21,22,24$ (degrees)

  • sales: $27,31,26,32,34$

  • $\lambda=3$

Solution

With two predictors the $2\times2$ inverse is quicker than elimination, and the same pattern serves $\lambda=0$ and $\lambda=3$.

Standardize the readings and center the sales

$$\bar T_1=\bar T_2=21,\qquad \sigma_1=\sigma_2=\sqrt{\tfrac{9+1+0+1+9}{5}}=2,\qquad \bar y=30$$

The lecture's spread divides by $n=5$; the two thermometers happen to share mean and spread.

$$x_1=(-1.5,\,-0.5,\,0,\,0.5,\,1.5),\quad x_2=(-0.5,\,-1.5,\,0,\,0.5,\,1.5),\quad y=(-3,\,1,\,-4,\,2,\,4)$$

Subtract the mean and divide by $2$; centering the sales removes the intercept.

Form the two products

$$X^TX=\begin{pmatrix}\sum x_1^2&\sum x_1x_2\\ \sum x_1x_2&\sum x_2^2\end{pmatrix}=\begin{pmatrix}5&4\\4&5\end{pmatrix}$$

$\sum x_1^2=5=n$, as standardization promises; the cross term is $0.75+0.75+0+0.25+2.25=4$.

$$X^Ty=\begin{pmatrix}4.5-0.5+0+1+6\\ 1.5-1.5+0+1+6\end{pmatrix}=\begin{pmatrix}11\\7\end{pmatrix}$$

Each entry pairs one standardized column with the centered sales.

Least squares, λ = 0

$$\hat\beta^{\mathrm{LS}}=\frac{1}{25-16}\begin{pmatrix}5&-4\\-4&5\end{pmatrix}\begin{pmatrix}11\\7\end{pmatrix}=\frac19\begin{pmatrix}27\\-9\end{pmatrix}=\begin{pmatrix}3\\-1\end{pmatrix}$$

Determinant $9$: swap the diagonal entries and negate the off-diagonal ones.

Ridge, λ = 3

$$X^TX+3I=\begin{pmatrix}8&4\\4&8\end{pmatrix},\qquad \det=64-16=48$$

The penalty adds $\lambda$ to the diagonal and touches nothing else.

$$\textcolor{#1f6feb}{\hat\beta^R}=\frac1{48}\begin{pmatrix}8&-4\\-4&8\end{pmatrix}\begin{pmatrix}11\\7\end{pmatrix}=\frac1{48}\begin{pmatrix}60\\12\end{pmatrix}=\begin{pmatrix}1.25\\0.25\end{pmatrix}$$

Same inverse pattern; the larger determinant does the shrinking.

Predict the day the thermometers disagree

$$\tilde z=\Big(\tfrac{23-21}{2},\ \tfrac{19-21}{2}\Big)=(1,\,-1)$$

A new day is standardized with the training means and spreads.

$$\text{LS: }30+3(1)-1(-1)=34,\qquad \text{ridge: }30+1.25(1)+0.25(-1)=31$$

Add $\bar y=30$ back, since both fits were made on centered sales.

Answer $$\boxed{\hat\beta^{\mathrm{LS}}=(3,\,-1),\quad \hat\beta^R=(1.25,\,0.25),\quad \hat y:\ 34\ \text{vs}\ 31}$$
Check

Multiply back: $(X^TX+3I)\hat\beta^R=(8\cdot1.25+4\cdot0.25,\ 4\cdot1.25+8\cdot0.25)=(11,\,7)=X^Ty$. The sum $1.25+0.25=1.5$ is $\frac{9}{12}$ of the least squares sum $2$, and the difference $1$ is $\frac14$ of $4$: the two shrink factors of the counterexample above.

One $2\times2$ inverse per value of $\lambda$; the products $X^TX$ and $X^Ty$ are computed once.

This settles the opening puzzle: ridge gives both thermometers a positive weight, and on a day when they disagree by four degrees it predicts $31$, not $34$, because the difference least squares read off two noisy days is cut to a quarter.

One standardized predictor: ridge is least squares times n/(n + λ)

A single standardized predictor has $\sum_ix_i^2=n=20$ and $\sum_ix_iy_i=30$ on centered data. Compute $\hat\beta^{\mathrm{LS}}$ and $\hat\beta^R$ for $\lambda=5$ and $\lambda=20$, and find the factor that turns one into the other.

FindBoth ridge estimates and the factor from $\hat\beta^{\mathrm{LS}}$ to $\hat\beta^R$.
Given
  • $n=20$, $\sum_ix_i^2=20$, $\sum_ix_iy_i=30$

  • $\lambda\in\{5,\,20\}$

Solution

With one column, $X^TX$ is the number $\sum_ix_i^2$, so the matrix formula collapses to one division.

Least squares

$$\hat\beta^{\mathrm{LS}}=\frac{\sum_ix_iy_i}{\sum_ix_i^2}=\frac{30}{20}=1.5$$

The $1\times1$ normal equation.

Ridge

$$\hat\beta^R=\frac{\sum_ix_iy_i}{\sum_ix_i^2+\lambda}=\frac{30}{20+\lambda}$$

$X^TX+\lambda I$ is the number $20+\lambda$.

$$\lambda=5:\ \tfrac{30}{25}=1.2,\qquad \lambda=20:\ \tfrac{30}{40}=0.75$$

Only the denominator depends on $\lambda$, so one numerator serves both values.

The shrink factor

$$\hat\beta^R=\frac{n}{n+\lambda}\cdot\frac{\sum_ix_iy_i}{n}=\frac{n}{n+\lambda}\,\hat\beta^{\mathrm{LS}}$$

Standardization makes $\sum_ix_i^2=n$, so the factor depends only on $n$ and $\lambda$.

Answer $$\boxed{\hat\beta^R=\tfrac{20}{20+\lambda}\cdot1.5:\quad 1.2\ \ (\lambda=5),\qquad 0.75\ \ (\lambda=20)}$$
Check

Factor check: $\frac{20}{25}=0.8$ and $0.8\times1.5=1.2$; at $\lambda=n$ the factor is exactly $\frac12$, and $0.75$ is half of $1.5$.

With one standardized predictor, ridge never changes the sign of the fit and never reaches $0$: it multiplies by $\frac{n}{n+\lambda}$, which lies strictly between $0$ and $1$.

Checkpoint
§05.1 — one ridge solve with two predictors

Two standardized predictors on centered data give the products below. You fit ridge with $\lambda=2$.

Find(a) Which vector is $\hat\beta^R$?
Given
  • $X^TX=\begin{pmatrix}6&2\\2&6\end{pmatrix}$

  • $X^Ty=(9,\,6)^T$

  • $\lambda=2$

Hint 1/4

The ridge estimate solves one linear system; decide which matrix and which right-hand side it uses before computing.

Hint 2/4

$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$, and $\begin{pmatrix}a&b\\b&a\end{pmatrix}^{-1}=\frac{1}{a^2-b^2}\begin{pmatrix}a&-b\\-b&a\end{pmatrix}$.

Hint 3/4

Here $X^TX+2I=\begin{pmatrix}8&2\\2&8\end{pmatrix}$ with determinant $60$, and $X^Ty=(9,\,6)^T$.

Hint 4/4

$\hat\beta^R=\frac1{60}(72-12,\ -18+48)=(1,\ 0.5)$.

Show solution

The $2\times2$ inverse is one line, so we invert rather than eliminate.

Add λ to the diagonal

$$X^TX+2I=\begin{pmatrix}8&2\\2&8\end{pmatrix},\qquad \det=64-4=60$$

Only the diagonal changes; the right-hand side stays $X^Ty$.

Invert and multiply

$$\hat\beta^R=\frac1{60}\begin{pmatrix}8&-2\\-2&8\end{pmatrix}\begin{pmatrix}9\\6\end{pmatrix}=\frac1{60}\begin{pmatrix}60\\30\end{pmatrix}=\begin{pmatrix}1\\0.5\end{pmatrix}$$

Swap the diagonal, negate the off-diagonal, divide by the determinant.

Answer $$\boxed{\hat\beta^R=(1,\ 0.5)}$$
Check

Multiply back: $8(1)+2(0.5)=9$ and $2(1)+8(0.5)=6$, which is $X^Ty$.

Whatever the data, the penalty touches only the diagonal of $X^TX$; the right-hand side $X^Ty$ is the same as for least squares.

⚠ Adding λ to every entry, not only the diagonal

the phrase 'add λ' is remembered without the identity matrix that says where

wrong$$X^TX+\lambda\mathbf 1\mathbf 1^T=\begin{pmatrix}5+\lambda&4+\lambda\\4+\lambda&5+\lambda\end{pmatrix}$$
right$$X^TX+\lambda I=\begin{pmatrix}5+\lambda&4\\4&5+\lambda\end{pmatrix}$$
⚠ Penalizing the intercept

the penalty sum is written from j = 0 out of habit

wrong$$\lambda\sum_{j=0}^p\beta_j^2$$
right$$\lambda\sum_{j=1}^p\beta_j^2\qquad(\hat\beta_0=\bar y\ \text{is left alone})$$
⚠ Inverting first and adding λ afterwards

both formulas contain the same three pieces, and the order of inverting and adding is easy to swap

wrong$$\hat\beta^R=\big((X^TX)^{-1}+\lambda I\big)X^Ty$$
right$$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$$

5.2The tuning parameter λ: from least squares to the mean, chosen by cross-validation

Shows what each $\lambda$ does, from $\hat\beta^{\mathrm{LS}}$ at $\lambda=0$ to the mean prediction as $\lambda\to\infty$, and picks $\lambda$ by cross-validation.

The kiosk fit used $\lambda=3$ without a reason; this block shows what the other values do and how the data choose one.

RuleThe two ends of λ, and how λ is chosen
Conditions
  • centered response and standardized predictors

  • for the limit $\lambda\to0$: $X^TX$ invertible

$$\boxed{\begin{aligned}\lambda\to0:&\quad \hat\beta^R_\lambda\to\textcolor{#8250df}{\hat\beta^{\mathrm{LS}}}\\ \lambda\to\infty:&\quad \hat\beta^R_\lambda\to0,\qquad \hat y\to\bar y\\ \text{choose }\lambda:&\quad \hat\lambda=\arg\min_{\lambda\in\text{grid}}\mathrm{CV}(k)\end{aligned}}$$

At $\lambda=0$ the penalty is free and ridge is least squares; as $\lambda$ grows the coefficients are pulled toward zero, and for a huge $\lambda$ every prediction is the mean response. The training RSS cannot choose between these fits, because it is smallest at $\lambda=0$; cross-validation over a grid of values can.

Why the two ends and the ordering hold

Small $\lambda$: when $X^TX$ is invertible, $(X^TX+\lambda I)^{-1}$ depends continuously on $\lambda$, so it tends to $(X^TX)^{-1}$ and $\hat\beta^R_\lambda$ tends to $\hat\beta^{\mathrm{LS}}$.

Large $\lambda$: $(X^TX+\lambda I)^{-1}=\frac1\lambda\big(I+\frac1\lambda X^TX\big)^{-1}$ and the bracket tends to $I$, so $\hat\beta^R_\lambda\approx\frac1\lambda X^Ty\to0$.

The training RSS never falls as $\lambda$ grows. Take $\lambda_1<\lambda_2$ with fits $\beta_1,\beta_2$, RSS values $R_1,R_2$ and penalties $P_1,P_2$. Each fit is optimal for its own loss: $R_1+\lambda_1P_1\le R_2+\lambda_1P_2$ and $R_2+\lambda_2P_2\le R_1+\lambda_2P_1$.

Adding the two gives $(\lambda_2-\lambda_1)(P_1-P_2)\ge0$, so $P_1\ge P_2$: the penalty never grows. The first inequality then gives $R_1\le R_2+\lambda_1(P_2-P_1)\le R_2$.

Looks like this, but is not

Since a larger $\lambda$ shrinks the fit, every coefficient looks as if it must move toward $0$ as $\lambda$ grows.

Only the whole vector is guaranteed to shrink: $\sum_j\beta_j^2$ never grows with $\lambda$. Single coefficients can grow. In the kiosk fit $\hat\beta_2$ climbs from $-1$ through $0$ to about $0.31$ before it heads back to $0$.

λthermometer 1thermometer 2sum of squares

$0$

$3$

$-1$

$10$

$1$

$1.9$

$-0.1$

$3.62$

$3$

$1.25$

$0.25$

$1.625$

$9$

$0.7$

$0.3$

$0.58$

$27$

$0.321$

$0.179$

$0.135$

$81$

$0.124$

$0.076$

$0.021$

The sum of squares falls at every step, as the proof in the box says it must; thermometer 2's coefficient does not, since it climbs from $-1$ to $0.3$.

The kiosk at λ = 1000: every prediction lands near the mean

Fit ridge to the kiosk data with $\lambda=1000$, compare the coefficients with $X^Ty/\lambda$, and compute the training RSS.

Find$\hat\beta^R$, its approximation and the training RSS.
Given
  • $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$, $X^Ty=(11,\,7)^T$

  • centered sales $y=(-3,\, \allowbreak 1,\, \allowbreak -4,\, \allowbreak 2,\, \allowbreak 4)$

  • $\lambda=1000$

Solution

For a huge $\lambda$ the matrix is almost $\lambda I$, so the approximation $X^Ty/\lambda$ says what to expect before the exact inverse confirms it.

Approximate

$$X^TX+1000I\approx1000\,I\ \Rightarrow\ \hat\beta^R\approx\frac{X^Ty}{1000}=(0.011,\ 0.007)$$

The entries $5$ and $4$ are tiny next to $1000$.

Exact

$$\det=1005^2-16=1\,010\,009,\qquad \hat\beta^R=\frac{(1005\cdot11-4\cdot7,\ \ -4\cdot11+1005\cdot7)}{1\,010\,009}$$

The exact inverse confirms the approximation; the determinant is dominated by $1005^2$.

$$\hat\beta^R=(0.01092,\ 0.00692)$$

Close to the approximation: $0.0109$ against $0.011$, and $0.0069$ against $0.007$.

Training RSS

$$\lvert\hat y_i-30\rvert\le0.03\ \Rightarrow\ \mathrm{RSS}=45.66\approx\textstyle\sum_iy_i^2=46$$

Every fitted value is within $0.03$ of the mean, so the residuals are almost the centered sales.

Answer $$\boxed{\hat\beta^R\approx(0.0109,\ 0.0069),\qquad \hat y\approx30\ \text{on every day},\qquad \mathrm{RSS}\approx45.66}$$
Check

The two ends bracket every fit: $\mathrm{RSS}=20$ at $\lambda=0$ (least squares) and $46=\sum_iy_i^2$ as $\lambda\to\infty$ (predict $\bar y$); $45.66$ lies between them, near the top.

A very large $\lambda$ does not predict $0$; it predicts the mean response, because the intercept is never penalized.

Choosing λ for the kiosk by leave-one-out cross-validation

Each kiosk day was left out in turn, ridge was refitted on the other four days (standardizing with their means and spreads), and the left-out day's squared error was recorded. Average the errors for each $\lambda$, choose $\lambda$, and refit on all five days.

Find$\mathrm{CV}(5)$ for each $\lambda$, the chosen $\lambda$ and the refitted coefficients.
Given
  • $\lambda=0$: $165.31,\ \allowbreak 165.31,\ \allowbreak 25.00,\ \allowbreak 1.80,\ \allowbreak 11.11$

  • $\lambda=1$: $0.23,\ \allowbreak 39.85,\ \allowbreak 25.00,\ \allowbreak 2.21,\ \allowbreak 12.71$

  • $\lambda=3$: $1.54,\ \allowbreak 24.10,\ \allowbreak 25.00,\ \allowbreak 2.84,\ \allowbreak 15.04$

  • $\lambda=9$: $5.61,\ \allowbreak 12.17,\ \allowbreak 25.00,\ \allowbreak 3.95,\ \allowbreak 18.68$

  • $\lambda=27$: $9.75,\ \allowbreak 5.34,\ \allowbreak 25.00,\ \allowbreak 5.10,\ \allowbreak 22.00$

  • all five days: $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$, $X^Ty=(11,\,7)^T$, mean sales $30$

Solution

With five days, leave-one-out is five-fold cross-validation, the most any split can hold out; the errors are given, so the work is five averages and one refit.

Average each row

$$\mathrm{CV}(5)=\tfrac15\textstyle\sum_{i=1}^5\mathrm{MSE}_i:\quad 73.71,\ \ 16.00,\ \ 13.70,\ \ 13.08,\ \ 13.44$$

Each row holds one squared error per left-out day.

Choose

$$\hat\lambda=9,\qquad \mathrm{CV}(5)=13.08$$

The smallest average wins. Least squares is worst by far: without day 1 or day 2, a single day is left to separate the thermometers.

Refit on all five days

$$X^TX+9I=\begin{pmatrix}14&4\\4&14\end{pmatrix},\qquad \det=196-16=180$$

The chosen $\lambda$ is used once more, now with every day.

$$\hat\beta^R=\frac1{180}\begin{pmatrix}14&-4\\-4&14\end{pmatrix}\begin{pmatrix}11\\7\end{pmatrix}=\frac1{180}\begin{pmatrix}126\\54\end{pmatrix}=\begin{pmatrix}0.7\\0.3\end{pmatrix}$$

Same $X^Ty$ as before; only the diagonal changed.

Answer $$\boxed{\hat\lambda=9,\qquad \hat\beta^R=(0.7,\ 0.3),\qquad \hat y=30+0.7\,\tilde x_1+0.3\,\tilde x_2}$$
Check

Day 3 contributes $25.00$ to every row: its readings equal the other four days' means, so every fit predicts their mean sales, $31$, against the actual $26$. The ranking comes from the other days, above all days 1 and 2, the two that separate the thermometers.

Five values of $\lambda$ times five folds is $25$ ridge fits, plus one refit.

Cross-validation picks $\lambda$ from errors on held-out days, which the training RSS cannot do; the winner is then refitted on all the data.

Checkpoint
§05.2 — choosing λ by the training RSS

A friend fits ridge with $\lambda=0,1,10,100$ on a training set and keeps the $\lambda$ whose fit has the smallest RSS on that same training set.

Find(a) Which λ will the friend choose?
Given
  • candidates: $\lambda\in\{0,\,1,\,10,\,100\}$

  • score: RSS on the training data used for fitting

Hint 1/4

The friend's score is measured on the data the fits were made from; ask which fit is best at exactly that.

Hint 2/4

Least squares minimizes RSS over all $\beta$, and the training RSS of $\hat\beta^R_\lambda$ never falls as $\lambda$ grows.

Hint 3/4

The candidates are $\lambda=0,1,10,100$, and $\lambda=0$ gives $\hat\beta^{\mathrm{LS}}$.

Hint 4/4

The friend always picks $\lambda=0$, plain least squares.

Show solution

The ordering proved in the box answers this for every data set at once, so no fit needs computing.

Least squares is the minimum

$$\mathrm{RSS}(\hat\beta^R_0)=\mathrm{RSS}(\hat\beta^{\mathrm{LS}})=\min_\beta\mathrm{RSS}(\beta)$$

At $\lambda=0$ the ridge loss is the RSS itself.

The rest can only be worse

$$\mathrm{RSS}(\hat\beta^R_0)\le\mathrm{RSS}(\hat\beta^R_1)\le\mathrm{RSS}(\hat\beta^R_{10})\le\mathrm{RSS}(\hat\beta^R_{100})$$

The training RSS never falls as λ grows.

Answer $$\boxed{\hat\lambda_{\text{friend}}=0}$$
Check

The kiosk numbers show it: training RSS $20$ at $\lambda=0$, $25.625$ at $\lambda=3$ and $45.66$ at $\lambda=1000$, while leave-one-out preferred $\lambda=9$.

Never choose a tuning parameter by the error on the data the fit was trained on; hold data out.

⚠ Choosing λ by the training RSS

it is the error at hand, and for most fitting questions it is the right score

wrong$$\hat\lambda=\arg\min_\lambda\mathrm{RSS}(\hat\beta^R_\lambda)=0$$
right$$\hat\lambda=\arg\min_\lambda\mathrm{CV}(k)\ \ \text{over a grid}$$
⚠ Thinking a huge λ predicts 0

on centered data the fitted values do go to 0, and the step back to raw units is forgotten

wrong$$\lambda\to\infty:\ \hat y\to0$$
right$$\lambda\to\infty:\ \hat y\to\bar y\quad(\text{the intercept is not penalized})$$
⚠ Keeping a fold's fit instead of refitting

the fold fits are already computed, and refitting looks like extra work

wrong$$\hat\beta=\text{one of the }k\text{ fold fits at }\hat\lambda$$
right$$\hat\beta=\hat\beta^R_{\hat\lambda}\ \text{refitted on all }n\text{ points}$$

5.3Scale matters for ridge: standardize with the training statistics

Explains why ridge depends on units, how standardizing removes that, and how to predict at a new raw input.

Every example so far came with standardized readings; here is why the lecture insists on them, and how to get back to raw units.

MethodStandardize, fit, predict
Conditions
  • $\bar x_j$, $\sigma_j$ and $\bar y$ are computed on the training data only

  • $\sigma_j>0$; a predictor that never changes carries no information and is dropped

$$\boxed{\begin{aligned}\tilde x_{ij}&=\frac{x_{ij}-\bar x_j}{\sigma_j},\qquad \sigma_j=\sqrt{\tfrac1n\textstyle\sum_{i=1}^n(x_{ij}-\bar x_j)^2}\\ \textcolor{#1f6feb}{\hat\beta^R}&=(\tilde X^T\tilde X+\lambda I)^{-1}\tilde X^T(y-\bar y\mathbf 1)\\ \hat y(z)&=\bar y+\sum_{j=1}^p\hat\beta^R_j\,\frac{z_j-\bar x_j}{\sigma_j}\end{aligned}}$$

Measure every predictor in standard deviations from its own mean, center the response, and fit ridge there. A new input goes through the same two operations, with the training means and spreads, before $\bar y$ is added back. Least squares predicts the same with or without this step; ridge does not.

Why least squares ignores units and ridge does not

Record predictor $j$ in a new unit, $x_{ij}\to c\,x_{ij}$ with $c>0$. The RSS depends on $\beta_j$ only through the products $\beta_jx_{ij}$, so least squares answers with $\hat\beta_j\to\hat\beta_j/c$ and every prediction stays the same.

Ridge also charges $\lambda\beta_j^2$. Keeping the old fit would need the coefficient $\beta_j/c$, whose charge is $\lambda\beta_j^2/c^2$: a change of unit acts like the penalty $\lambda/c^2$ on that predictor. Going from TL to thousands of TL, $c=\frac1{1000}$, makes that penalty a million times stronger.

After standardizing, $\tilde x_{ij}=(x_{ij}-\bar x_j)/\sigma_j$ is the same number in every unit, because $c$ cancels between the numerator and $\sigma_j$. The standardized fit and its predictions no longer depend on units.

Looks like this, but is not

Rescaling every predictor by the same factor looks harmless for ridge, since the predictors keep their relative sizes.

A common factor $c$ still turns $\lambda$ into $\lambda/c^2$ for every coefficient, so the amount of shrinkage changes. It looks harmless only because least squares, which has no $\lambda$, really is unaffected.

Income in TL or in thousands of TL: least squares shrugs, ridge does not

Four people earn $13, 19, 21, 27$ thousand TL a month and spend $5, 11, 9, 15$ hundred TL. Fit least squares and ridge with $\lambda=100$ on centered data, once with income in thousands of TL and once in TL, and predict spending at $30$ thousand TL.

FindBoth slopes in both units, and the four predictions.
Given
  • incomes $13,19,21,27$ (thousand TL); spending $5,11,9,15$ (hundred TL)

  • $\lambda=100$ on centered, unstandardized data

Solution

One predictor makes each fit a single division, so we redo it in the second unit and compare instead of arguing in general.

Center

$$\bar x=20,\ \ x^c=(-7,\,-1,\,1,\,7);\qquad \bar y=10,\ \ y^c=(-5,\,1,\,-1,\,5)$$

Centering removes the intercept from both fits.

Thousands of TL

$$\textstyle\sum x^cy^c=35-1-1+35=68,\qquad \sum(x^c)^2=49+1+1+49=100$$

With centered data these two sums are all a one-predictor fit needs.

$$\hat\beta^{\mathrm{LS}}=\tfrac{68}{100}=0.68,\qquad \hat\beta^R=\tfrac{68}{100+100}=0.34$$

One column: divide by $\sum(x^c)^2$, plus $\lambda$ for ridge.

TL

$$x^c\to1000\,x^c:\quad \textstyle\sum x^cy^c=68\,000,\qquad \sum(x^c)^2=10^8$$

Every centered income is now a thousand times larger.

$$\hat\beta^{\mathrm{LS}}=0.00068,\qquad \hat\beta^R=\tfrac{68\,000}{10^8+100}\approx0.00068$$

The $100$ is negligible next to $10^8$: in TL, ridge hardly shrinks at all.

Predict at 30 thousand TL

$$\text{thousands: }\ 10+0.68(10)=16.8,\qquad 10+0.34(10)=13.4$$

The new income is $10$ thousand above the mean.

$$\text{TL: }\ 10+0.00068(10\,000)=16.8\ \ \text{for both fits}$$

The same person is $10\,000$ TL above the mean.

Answer $$\boxed{\text{LS: }16.8\ \text{in both units};\qquad \text{ridge: }13.4\ \text{(thousands)}\ \ \text{vs}\ \ 16.8\ \text{(TL)}}$$
Check

The least squares slopes differ by exactly the unit factor, $0.68=1000\times0.00068$, as scale invariance requires. The ridge slopes do not: $0.34$ is half of $0.68$, while $0.00068$ is essentially the least squares value.

Same people, same $\lambda$, two predictions: on raw data the unit decides how much ridge shrinks, which is why the lecture standardizes first.

Standardize, fit, predict: the same four people

Standardize the four incomes, fit ridge with $\lambda=4$ on standardized income and centered spending, and predict spending at $30$ thousand TL. Then check that recording income in TL changes nothing.

Find$\hat\beta^R$ and the prediction at $30$ thousand TL.
Given
  • incomes $13,19,21,27$ thousand TL; spending $5,11,9,15$ hundred TL

  • $\lambda=4$ on standardized data

Solution

We compute the training mean and spread once and reuse them for the new person, the step most often done wrong.

Training statistics

$$\bar x=20,\qquad \sigma=\sqrt{\tfrac{49+1+1+49}{4}}=5,\qquad \bar y=10$$

The lecture's spread divides by $n=4$.

Standardize and fit

$$\tilde x=(-1.4,\,-0.2,\,0.2,\,1.4),\qquad \textstyle\sum\tilde x^2=4=n$$

Each centered income divided by $5$.

$$\hat\beta^R=\frac{\sum\tilde x\,y^c}{n+\lambda}=\frac{7-0.2-0.2+7}{4+4}=\frac{13.6}{8}=1.7$$

One column, so the ridge formula is a single division.

Predict

$$\tilde z=\frac{30-20}{5}=2,\qquad \hat y=10+1.7\times2=13.4$$

The new income uses the training mean and spread, never its own.

Change the unit

$$\text{TL: }\ \sigma=5000,\qquad \tilde z=\frac{30\,000-20\,000}{5000}=2$$

The unit cancels in the ratio, so $\tilde x$, $\hat\beta^R$ and $\hat y$ do not move.

Answer $$\boxed{\hat\beta^R=1.7\ \text{per standard deviation},\qquad \hat y=13.4\ \ (1340\ \text{TL})}$$
Check

The previous example's ridge in thousands of TL used $\lambda=100=4\sigma^2$ and also predicted $13.4$: dividing a column by $\sigma$ has the same effect as multiplying its penalty by $\sigma^2$.

Standardize with the training data, fit, send every new input through the same means and spreads, then add $\bar y$ back.

Checkpoint
§05.3 — predicting with the training statistics

A ridge model was trained on standardized ages and centered responses. It now has to score a test batch whose ages have a different mean and spread from the training ages.

Find(a) What does the model predict for the 50-year-old?
Given
  • training ages: mean $40$, standard deviation $10$; age coefficient $\hat\beta^R=2$; mean response $\bar y=50$

  • test batch: mean age $30$, standard deviation $5$

  • person to score: age $50$

Hint 1/4

The model was fitted on one scale of ages; put the new age on that same scale before using the coefficient.

Hint 2/4

$\tilde z=(z-\bar x)/\sigma$ with the training statistics, then $\hat y=\bar y+\hat\beta^R\tilde z$.

Hint 3/4

Training: $\bar x=40$, $\sigma=10$, $\hat\beta^R=2$, $\bar y=50$; the age is $z=50$. The test batch's mean $30$ and spread $5$ play no part.

Hint 4/4

$\tilde z=1$, so $\hat y=52$.

Show solution

The coefficient is per training standard deviation, so the new age must be expressed in exactly that unit.

Standardize with the training statistics

$$\tilde z=\frac{50-40}{10}=1$$

The model knows only the training mean and spread.

Predict

$$\hat y=\bar y+\hat\beta^R\tilde z=50+2\times1=52$$

The fit was made on centered responses, so $\bar y$ comes back.

Answer $$\boxed{\hat y=52}$$
Check

Units check: $2$ per standard deviation of $10$ years is $0.2$ per year, and $50$ is $10$ years above the training mean: $50+0.2\times10=52$.

A new input is always standardized with the training statistics, whatever batch it arrives in.

⚠ Standardizing a new input with its own statistics

each batch of data seems to deserve its own mean and spread

wrong$$\tilde z=\frac{z-\bar z_{\text{test}}}{\sigma_{\text{test}}}$$
right$$\tilde z=\frac{z-\bar x_{\text{train}}}{\sigma_{\text{train}}}$$
⚠ Forgetting to add the mean response back

the fit was made on centered responses, so its raw output already looks like a prediction

wrong$$\hat y=\tilde z^T\hat\beta^R$$
right$$\hat y=\bar y+\tilde z^T\hat\beta^R$$
⚠ Fitting ridge on raw units

the closed form works on any centered data, so standardizing looks optional

wrong$$\text{income in TL},\ \lambda=100:\ \ \hat\beta^R\approx\hat\beta^{\mathrm{LS}}\ \ \text{(almost no shrinkage)}$$
right$$\text{standardize first: the same }\lambda\text{ shrinks the same in every unit}$$

5.4Why shrinking helps: ridge trades a little bias for a large cut in variance

Computes the bias and variance of a ridge coefficient and shows that some $\lambda>0$ always beats least squares on expected test error.

Cross-validation said $\lambda=9$ beats $\lambda=0$ on the kiosk data; with one predictor we can see exactly why.

TheoremBias and variance of ridge with one predictor
Conditions
  • one predictor with fixed standardized inputs, $\sum_ix_i^2=n$

  • $y_i=\beta x_i+\varepsilon_i$ with independent noise of mean $0$ and variance $\sigma^2$, and no intercept to estimate

  • a new pair $(x,Y)$ from the same model, with $E[x^2]=1$, independent of the training noise

$$\boxed{\begin{aligned}c&=\tfrac{n}{n+\lambda},\qquad \hat\beta^R=c\,\hat\beta^{\mathrm{LS}}\\ \text{bias}^2&=(1-c)^2\beta^2,\qquad \text{variance}=c^2\,\frac{\sigma^2}{n}\\ \textcolor{#1f6feb}{E\big[(Y-x\hat\beta^R)^2\big]}&=(1-c)^2\beta^2+c^2\frac{\sigma^2}{n}+\sigma^2\\ \text{smallest at}\ \ \lambda^*&=\sigma^2/\beta^2\end{aligned}}$$

Shrinking by the factor $c$ moves the average fit a fraction $1-c$ of the way to zero, which is the bias, and multiplies the spread of the fit by $c$, which squares into the variance. Near $\lambda=0$ the variance falls faster than the bias² rises, so a little shrinkage always helps; the best amount is the noise-to-signal ratio $\sigma^2/\beta^2$.

Derivation (further reading: the lecture draws these curves without formulas)

$\hat\beta^{\mathrm{LS}}=\frac1n\sum_ix_iy_i=\beta+\frac1n\sum_ix_i\varepsilon_i$, so $E[\hat\beta^{\mathrm{LS}}]=\beta$ and $\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\frac{\sigma^2}{n^2}\sum_ix_i^2=\frac{\sigma^2}{n}$.

Ridge is $c\,\hat\beta^{\mathrm{LS}}$, so $E[\hat\beta^R]=c\beta$ and $\mathrm{Var}(\hat\beta^R)=c^2\sigma^2/n$.

At a new input the error is $Y-x\hat\beta^R=\varepsilon+x(\beta-\hat\beta^R)$. The new noise has mean $0$ and is independent of the rest, so $E[(Y-x\hat\beta^R)^2]=\sigma^2+E[x^2]\,E[(\hat\beta^R-\beta)^2]$: the noise plus the bias² and variance of the coefficient.

With $1-c=\frac{\lambda}{n+\lambda}$, bias² plus variance is $g(\lambda)=\frac{\lambda^2\beta^2+n\sigma^2}{(n+\lambda)^2}$. Its derivative $\frac{2n(\lambda\beta^2-\sigma^2)}{(n+\lambda)^3}$ is negative below $\sigma^2/\beta^2$ and positive above it.

Looks like this, but is not

Since ridge is biased and least squares is unbiased, least squares looks like the more accurate of the two.

Accuracy on new data is bias² plus variance, not bias alone. In the figure's setting least squares is $0.5$ above the noise, all of it variance, while ridge at $\lambda=4$ has $\frac19$ of bias² and $\frac29$ of variance, $\frac13$ in total.

The best λ for one predictor: n = 8, σ² = 4, β = 1

A standardized predictor has $n=8$ fixed inputs, true slope $\beta=1$ and noise variance $\sigma^2=4$. Compute the bias², the variance and the expected test error of least squares and of ridge at $\lambda=4$ and $\lambda=24$.

FindThe three terms and the total for $\lambda=0$, $4$ and $24$.
Given
  • $n=8$, $\sum_ix_i^2=8$, $\beta=1$, $\sigma^2=4$

  • a new input with $E[x^2]=1$

Solution

Everything follows from the shrink factor $c=\frac{n}{n+\lambda}$, so we compute $c$ first for each $\lambda$.

Shrink factors

$$c=\tfrac{8}{8+\lambda}:\qquad \lambda=0\to1,\quad \lambda=4\to\tfrac23,\quad \lambda=24\to\tfrac14$$

One number per $\lambda$ carries the whole computation.

Bias² and variance

$$\lambda=0:\quad 0+1^2\cdot\tfrac48=0.5$$

Least squares is unbiased; its variance is $\sigma^2/n$.

$$\lambda=4:\quad \big(\tfrac13\big)^2\cdot1+\big(\tfrac23\big)^2\cdot0.5=\tfrac19+\tfrac29=\tfrac13$$

A third of the slope is lost, and the spread is cut to two thirds.

$$\lambda=24:\quad \big(\tfrac34\big)^2+\big(\tfrac14\big)^2\cdot0.5=0.5625+0.03125=0.59375$$

Too much shrinkage: the bias now outweighs everything the variance saved.

Add the noise

$$E\big[(Y-x\hat\beta)^2\big]:\qquad 4.5,\quad 4.333,\quad 4.594$$

Every predictor pays $\sigma^2=4$ on top.

Answer $$\boxed{\lambda=4:\ 4.333\ \ <\ \ \lambda=0:\ 4.5\ \ <\ \ \lambda=24:\ 4.594}$$
Check

The formula agrees: $\lambda^*=\sigma^2/\beta^2=4$, and the minimum above the noise, $\frac{\sigma^2\beta^2}{n\beta^2+\sigma^2}=\frac{4}{12}=\frac13$, matches the direct sum.

Shrinkage is a purchase: bias² is the price, the cut in variance is what it buys, and the best $\lambda$ is the noise-to-signal ratio.

Checkpoint
§05.4 — the best λ from the noise and the slope

One standardized predictor is fitted on $n=10$ points. The true slope is $\beta=2$ and the noise variance is $\sigma^2=8$, and the intercept is known to be $0$.

Find(a) Which $\lambda$ gives ridge the smallest expected test error?
Given
  • $n=10$, $\sum_ix_i^2=10$

  • $\beta=2$, $\sigma^2=8$

Hint 1/4

The best λ balances the bias ridge adds against the variance it removes; the box gives it in terms of two model quantities.

Hint 2/4

For one standardized predictor, $\lambda^*=\sigma^2/\beta^2$.

Hint 3/4

Here $\sigma^2=8$ and $\beta=2$; the sample size $n=10$ does not enter.

Hint 4/4

$\lambda^*=8/4=2$.

Show solution

The derivative of bias² plus variance vanishes at one point, so we use the result instead of re-deriving it.

Apply the formula

$$\lambda^*=\frac{\sigma^2}{\beta^2}=\frac{8}{4}=2$$

The sign of $\lambda\beta^2-\sigma^2$ decides whether more shrinkage helps.

Answer $$\boxed{\lambda^*=2}$$
Check

Direct check with $c=\frac{n}{n+\lambda}$: bias² plus variance is $\frac{\lambda^2\cdot4+80}{(10+\lambda)^2}$, which gives $0.8$ at $\lambda=0$, $\frac{84}{121}\approx0.694$ at $\lambda=1$, $\frac{96}{144}\approx0.667$ at $\lambda=2$ and $\frac{116}{169}\approx0.686$ at $\lambda=3$.

More noise calls for more shrinkage and a stronger signal for less; the sample size sets how much each unit of λ shrinks.

⚠ Calling the unbiased estimator the more accurate one

the word unbiased sounds like a guarantee of accuracy

wrong$$\text{bias}=0\ \Rightarrow\ \text{smallest test error}$$
right$$\text{test error}=\text{bias}^2+\text{variance}+\text{noise}$$
⚠ Turning the noise-to-signal ratio upside down

both ratios contain the same two numbers

wrong$$\lambda^*=\beta^2/\sigma^2$$
right$$\lambda^*=\sigma^2/\beta^2\quad(\text{more noise, more shrinkage})$$

5.5Ridge always has one answer: even when least squares has infinitely many

Proves that adding $\lambda I$ makes $X^TX$ invertible, so ridge works with duplicated predictors or $p\ge n$, though it never drops a predictor.

So far $X^TX$ could be inverted; now the two thermometers read exactly alike and it cannot, which always happens once there are at least as many predictors as observations.

TheoremAdding λI makes the matrix invertible
Conditions
  • any $n\times p$ matrix $X$, whatever its rank

  • $\lambda>0$

$$\boxed{\begin{aligned}v^T(X^TX+\lambda I)v&=\lVert Xv\rVert_2^2+\lambda\lVert v\rVert_2^2>0\quad(v\neq0)\\ \Rightarrow\ \ \textcolor{#1f6feb}{\hat\beta^R}&=(X^TX+\lambda I)^{-1}X^Ty\ \ \text{exists and is unique}\end{aligned}}$$

For every nonzero direction $v$ the matrix gives a strictly positive number: a squared length that can be $0$, plus $\lambda$ times a squared length that cannot. A matrix that sends no nonzero vector to $0$ is invertible, so ridge has exactly one solution for every $\lambda>0$, even when least squares has a whole line or plane of them.

Proof

$v^TX^TXv=(Xv)^T(Xv)=\lVert Xv\rVert_2^2\ge0$, and it is $0$ exactly when $Xv=0$. That happens for some $v\neq0$ whenever the columns of $X$ are dependent, which is certain when $p\ge n$: centered columns live in a space of dimension $n-1$.

The identity adds $\lambda v^Tv=\lambda\lVert v\rVert_2^2$, which is positive for every $v\neq0$.

If $(X^TX+\lambda I)v=0$, then $v^T(X^TX+\lambda I)v=0$, which forces $v=0$. So $Mv=0$ has only the solution $v=0$, and $M=X^TX+\lambda I$ is invertible.

Looks like this, but is not

Ridge coefficients are small, so a ridge coefficient of $0.02$ looks like ridge dropping that predictor from the model.

A predictor leaves the model only with a coefficient of exactly $0$. Ridge solves a linear system whose answer is, apart from coincidences, nonzero in every entry, so all $p$ predictors stay in; with many predictors that makes the fit hard to read. Exact zeros need a different penalty, the next block's.

Two identical thermometers: least squares has a line of answers, ridge has one

Suppose the second thermometer had copied the first on all five kiosk days, so both standardized columns are $(-1.5, \allowbreak -0.5, \allowbreak 0, \allowbreak 0.5, \allowbreak 1.5)$ and the centered sales are unchanged. Find all least squares fits, the ridge fit at $\lambda=3$, and its limit as $\lambda\to0$.

FindThe set of least squares fits, $\hat\beta^R$ at $\lambda=3$, and its limit as $\lambda\to0$.
Given
  • $x_1=x_2=(-1.5,\, \allowbreak -0.5,\, \allowbreak 0,\, \allowbreak 0.5,\, \allowbreak 1.5)$

  • $y=(-3,\, \allowbreak 1,\, \allowbreak -4,\, \allowbreak 2,\, \allowbreak 4)$, the centered sales

Solution

$X^TX$ has no inverse here, so we solve the normal equations directly for least squares and use the $2\times2$ inverse only for ridge.

Least squares

$$X^TX=\begin{pmatrix}5&5\\5&5\end{pmatrix},\quad \det=0,\qquad X^Ty=\begin{pmatrix}11\\11\end{pmatrix}$$

Identical columns give identical rows; $\sum x_1y=11$ as before.

$$5\beta_1+5\beta_2=11\ \Rightarrow\ \beta_1+\beta_2=2.2$$

Both normal equations say the same thing, so a whole line solves them.

Ridge at λ = 3

$$X^TX+3I=\begin{pmatrix}8&5\\5&8\end{pmatrix},\qquad \det=64-25=39$$

Positive, as the theorem guarantees.

$$\hat\beta^R=\frac1{39}\begin{pmatrix}8&-5\\-5&8\end{pmatrix}\begin{pmatrix}11\\11\end{pmatrix}=\frac{33}{39}\begin{pmatrix}1\\1\end{pmatrix}\approx\begin{pmatrix}0.846\\0.846\end{pmatrix}$$

Symmetric data give a symmetric answer: the two copies share the weight equally.

The limit as λ → 0

$$\hat\beta^R_\lambda=\frac{11\lambda}{\lambda(10+\lambda)}\begin{pmatrix}1\\1\end{pmatrix}=\frac{11}{10+\lambda}\begin{pmatrix}1\\1\end{pmatrix}\ \to\ \begin{pmatrix}1.1\\1.1\end{pmatrix}$$

In general $\det=(5+\lambda)^2-25=\lambda(10+\lambda)$, and $\lambda$ cancels before the limit.

Answer $$\boxed{\begin{aligned}&\text{LS: every }\beta\text{ with }\beta_1+\beta_2=2.2\\ &\hat\beta^R_3=\big(\tfrac{11}{13},\,\tfrac{11}{13}\big),\qquad \lim_{\lambda\to0}\hat\beta^R_\lambda=(1.1,\,1.1)\end{aligned}}$$
Check

$(1.1,\,1.1)$ lies on the line, and it is the point of the line closest to the origin: the line's normal direction is $(1,1)$, so the foot of the perpendicular from $0$ is $\frac{2.2}{2}(1,1)$.

When least squares cannot decide between fits, ridge decides for the one with the smallest coefficients, and for every $\lambda>0$ it decides uniquely.

Checkpoint
§05.5 — ridge with more predictors than observations

A gene study measures $p=500$ expression levels on $n=60$ patients and fits ridge with $\lambda=0.1$. A classmate says no ridge fit can exist, because $X^TX$ is singular.

Find(a) True or false: the ridge estimate exists and is unique for every $\lambda>0$, including $\lambda=0.1$.
Given
  • $p=500$ predictors, $n=60$ observations

  • $\lambda=0.1$

Hint 1/4

Existence of the ridge estimate hinges on one matrix; ask what could make it singular.

Hint 2/4

$v^T(X^TX+\lambda I)v=\lVert Xv\rVert_2^2+\lambda\lVert v\rVert_2^2$.

Hint 3/4

With $p=500>n=60$, $Xv=0$ for some $v\neq0$, so the first term can vanish; the second is $0.1\lVert v\rVert_2^2$.

Hint 4/4

The form is positive for every $v\neq0$, so the matrix is invertible and the statement is true.

Show solution

The quadratic form argument works for any rank, so we do not need to know anything else about the data.

The least squares matrix

$$p>n\ \Rightarrow\ Xv=0\ \text{for some}\ v\neq0\ \Rightarrow\ \det(X^TX)=0$$

Five hundred columns in a space of dimension at most sixty must be dependent.

The ridge matrix

$$v^T(X^TX+0.1I)v=\lVert Xv\rVert_2^2+0.1\lVert v\rVert_2^2\ \ge\ 0.1\lVert v\rVert_2^2>0$$

Holds for every $v\neq0$, so no nonzero vector is sent to $0$.

Answer $$\boxed{\text{True: }\hat\beta^R\ \text{exists and is unique}}$$
Check

Small case of the same kind: two identical standardized columns with $n=5$ give $\det(X^TX+\lambda I)=\lambda(10+\lambda)$, positive for every $\lambda>0$ and $0$ only at $\lambda=0$.

Any positive λ, however small, rescues invertibility; how much it shrinks is a separate question for cross-validation.

⚠ Believing ridge fails wherever least squares fails

the least squares formula needs that inverse, and ridge looks like the same formula

wrong$$\det(X^TX)=0\ \Rightarrow\ \text{no ridge fit}$$
right$$\det(X^TX+\lambda I)>0\ \ \text{for every}\ \lambda>0$$
⚠ Reading a small ridge coefficient as a dropped predictor

small and zero look alike in a printed table

wrong$$\hat\beta^R_j=0.02\ \Rightarrow\ x_j\ \text{is out of the model}$$
right$$x_j\ \text{is out only if}\ \hat\beta_j=0\ \text{exactly}$$

5.6The lasso: an absolute-value penalty that sets coefficients exactly to zero

Replaces $\beta_j^2$ by $\lvert\beta_j\rvert$ in the penalty; the fit has no closed form in general, but it drops predictors and gives a sparse model.

Ridge kept both thermometers, as it keeps every predictor it is given; the lasso changes one exponent and starts dropping predictors.

DefinitionThe lasso (least absolute shrinkage and selection operator)
Conditions
  • centered response, standardized predictors, $\lambda\ge0$; the intercept is not penalized

  • last line only: one predictor with $\sum_ix_i^2=n$ (further reading)

$$\boxed{\begin{aligned}\mathrm{Loss}_L(\beta,\lambda)&=\mathrm{RSS}(\beta)+\lambda\sum_{j=1}^p\lvert\beta_j\rvert\\ \textcolor{#d1690a}{\hat\beta^L}&=\arg\min_\beta\,\mathrm{Loss}_L(\beta,\lambda)\\ p=1:\ \ \textcolor{#d1690a}{\hat\beta^L}&=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\Big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\frac{\lambda}{2n}\Big)_+\end{aligned}}$$

The lasso charges $\lambda$ per unit of absolute size instead of squared size. As $\lambda\to0$ it becomes least squares, and as $\lambda\to\infty$ every coefficient is $0$. In between, some coefficients sit at exactly $0$ over whole ranges of $\lambda$: that is variable selection. With one predictor the fit moves the least squares value $\frac{\lambda}{2n}$ toward $0$ and stops there.

The one-predictor formula (further reading)

With $\sum_ix_i^2=n$ and $z=\sum_ix_iy_i=n\hat\beta^{\mathrm{LS}}$, the loss is $\sum_iy_i^2-2z\beta+n\beta^2+\lambda\lvert\beta\rvert$.

On $\beta>0$ it is a parabola with derivative $-2z+2n\beta+\lambda$, zero at $\beta=\frac{z-\lambda/2}{n}$; that point is positive only when $z>\lambda/2$. On $\beta<0$ the stationary point is $\frac{z+\lambda/2}{n}$, negative only when $z<-\lambda/2$.

If $\lvert z\rvert\le\lambda/2$, neither piece has a stationary point on its own side: the loss rises in both directions from $\beta=0$, so the minimum sits at the corner $0$. Dividing $z$ by $n$ gives the formula in the box.

Looks like this, but is not

The ridge path of the kiosk fit also passes through $\hat\beta_2=0$, at $\lambda=\frac97$, so ridge looks able to drop a predictor too.

Ridge touches $0$ at that single $\lambda$ and moves on; at every other value thermometer 2 is back in the model. The lasso holds $\hat\beta_2$ at exactly $0$ for every $\lambda\ge2$, a whole range, and that is what selecting a variable means.

least squaresridgelasso

$2.0$

$1.0$

$1.5$

$1.0$

$0.5$

$0.5$

$0.5$

$0.25$

$0$

$0.3$

$0.15$

$0$

$-0.8$

$-0.4$

$-0.3$

Ridge multiplies every entry by $\frac{n}{n+\lambda}=\frac12$ and never reaches $0$. The lasso takes $\frac{\lambda}{2n}=0.5$ off each size and returns exactly $0$ whenever the least squares value is within $0.5$ of zero.

One predictor by cases: the lasso at λ = 4 and at λ = 30

A standardized predictor with $n=10$ has $\sum_ix_iy_i=12$ on centered data. Minimize the lasso loss for $\lambda=4$ and $\lambda=30$ by treating $\beta>0$, $\beta<0$ and $\beta=0$ separately, and compare with ridge at $\lambda=4$.

Find$\hat\beta^L$ for both values and $\hat\beta^R$ at $\lambda=4$.
Given
  • $n=10$, $\sum_ix_i^2=10$, $\sum_ix_iy_i=12$

  • $\lambda\in\{4,\,30\}$

Solution

The absolute value has a corner at $0$, so no single derivative can be set to zero; splitting by the sign of $\beta$ turns each piece into a parabola.

Write the loss

$$\mathrm{Loss}_L(\beta)=\textstyle\sum_iy_i^2-24\beta+10\beta^2+\lambda\lvert\beta\rvert$$

Expand $\sum_i(y_i-\beta x_i)^2$ with $\sum x_i^2=10$ and $\sum x_iy_i=12$.

λ = 4

$$\beta>0:\quad -24+20\beta+4=0\ \Rightarrow\ \beta=1.0>0$$

The stationary point lies in the region assumed, so it is a candidate.

$$\beta<0:\quad -24+20\beta-4=0\ \Rightarrow\ \beta=1.4\not<0$$

No candidate on the negative side.

$$\mathrm{Loss}_L(1)-\mathrm{Loss}_L(0)=-24+10+4=-10<0\ \Rightarrow\ \hat\beta^L=1.0$$

The candidate beats the corner, so it is the minimum.

λ = 30

$$\beta>0:\ \ \beta=\tfrac{24-30}{20}<0;\qquad \beta<0:\ \ \beta=\tfrac{24+30}{20}>0$$

Neither side has a stationary point in its own region.

$$\hat\beta^L=0$$

The loss rises on both sides of $0$, so the corner is the minimum.

Ridge for comparison

$$\hat\beta^R=\tfrac{12}{10+4}\approx0.857$$

Ridge divides by $n+\lambda$ instead of subtracting.

Answer $$\boxed{\hat\beta^L=1.0\ \ (\lambda=4),\qquad \hat\beta^L=0\ \ (\lambda=30),\qquad \hat\beta^R\approx0.857\ \ (\lambda=4)}$$
Check

The formula in the box agrees: $\hat\beta^{\mathrm{LS}}=1.2$ and $\frac{\lambda}{2n}=0.2$ or $1.5$, so $1.2-0.2=1.0$ and $(1.2-1.5)_+=0$.

Split by sign whenever an absolute value is minimized: each piece is a parabola, and the corner at $0$ wins when neither parabola's lowest point lies on its own side.

The kiosk at λ = 3: why the lasso keeps only thermometer 1

For the kiosk data, show that $\hat\beta^L=(1.9,\,0)$ at $\lambda=3$: find the best $\beta_1$ with $\beta_2=0$, then check that moving $\beta_2$ away from $0$ in either direction does not help.

Find$\hat\beta^L$ at $\lambda=3$, and the reason $\hat\beta_2=0$.
Given
  • $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$, $X^Ty=(11,\,7)^T$

  • $\lambda=3$

Solution

The lasso has no closed form here, so we guess which coefficient is zero and then verify the guess. Guess and check can feel like cheating; a checked guess is a proof, and it beats searching all sign patterns.

Best β₁ with β₂ = 0

$$\mathrm{Loss}_L(\beta_1,0)=\textstyle\sum_iy_i^2-22\beta_1+5\beta_1^2+3\lvert\beta_1\rvert$$

Only the first column is in use: $2\cdot11=22$ and $\sum_ix_{i1}^2=5$.

$$\beta_1>0:\quad -22+10\beta_1+3=0\ \Rightarrow\ \beta_1=1.9$$

The positive piece, exactly as in the one-predictor example.

Nudge β₂

$$\frac{\partial\mathrm{RSS}}{\partial\beta_2}\Big|_{(1.9,\,0)}=-2\big(7-4\cdot1.9-5\cdot0\big)=1.2$$

The RSS slope in the $\beta_2$ direction is $-2x_2^T(y-X\beta)$, with $x_2^Ty=7$ and $x_2^Tx_1=4$.

$$\mathrm{Loss}_L(1.9,t)-\mathrm{Loss}_L(1.9,0)=1.2t+5t^2+3\lvert t\rvert\ \ge\ 1.8\lvert t\rvert>0$$

Exact, since the RSS is quadratic. The penalty's slope $3$ beats the RSS slope $1.2$ in both directions.

The rule behind the zero

$$\hat\beta_2=0\ \text{is kept}\iff\Big\lvert\frac{\partial\mathrm{RSS}}{\partial\beta_2}\Big\rvert\le\lambda:\qquad 1.2\le3$$

The same comparison decides every zero coefficient; the penalty is a sum of separate terms, so checking each coordinate is enough.

Answer $$\boxed{\hat\beta^L=(1.9,\ 0)\ \ \text{at}\ \ \lambda=3}$$
Check

Compare losses: $\mathrm{Loss}_L$ is $27.95$ at $(1.9,0)$ against $28$ at $(2,0)$, $30.125$ at the ridge fit $(1.25,0.25)$ and $32$ at the least squares fit $(3,-1)$.

With two nearly equal predictors the lasso tends to keep one and drop the other, and which one it keeps can hinge on a few observations; ridge splits the weight instead.

Checkpoint
§05.6 — soft thresholding with one predictor

One standardized predictor is fitted on $n=25$ centered points, and least squares gives $\hat\beta^{\mathrm{LS}}=-0.6$. You refit with the lasso at $\lambda=20$.

Find(a) What is $\hat\beta^L$?
Given
  • $n=25$, $\sum_ix_i^2=25$

  • $\hat\beta^{\mathrm{LS}}=-0.6$

  • $\lambda=20$

Hint 1/4

With one predictor, the lasso moves the least squares value toward zero by a fixed amount and stops at zero.

Hint 2/4

With $b=\hat\beta^{\mathrm{LS}}$: $\hat\beta^L=\operatorname{sign}(b)\big(\lvert b\rvert-\frac{\lambda}{2n}\big)_+$.

Hint 3/4

Here $\hat\beta^{\mathrm{LS}}=-0.6$, $\lambda=20$ and $n=25$, so the cut is $\frac{20}{50}=0.4$.

Hint 4/4

$\hat\beta^L=-(0.6-0.4)=-0.2$.

Show solution

The one-predictor formula applies directly, so we compute the cut and compare it with the size.

The cut

$$\frac{\lambda}{2n}=\frac{20}{50}=0.4$$

No $\frac12$ in front of the RSS, so the cut carries the factor 2.

Subtract and keep the sign

$$\hat\beta^L=-\big(0.6-0.4\big)_+=-0.2$$

The size is above the cut, so the result is not zero.

Answer $$\boxed{\hat\beta^L=-0.2}$$
Check

By cases: on $\beta<0$ the loss derivative is $-2z+2n\beta-\lambda$ with $z=25(-0.6)=-15$, which vanishes at $\beta=\frac{-30+20}{50}=-0.2$, inside the region.

Soft thresholding is a subtraction on the size; the sign comes back at the end.

⚠ Dropping the factor 2 in the cut

books that put 1/2 in front of the RSS write the threshold without it

wrong$$\hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{n}\big)_+$$
right$$\hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$
⚠ Losing the sign

the formula works on sizes, and the sign has to be put back by hand

wrong$$\hat\beta^L=\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$
right$$\hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$
⚠ Giving the lasso ridge's closed form

ridge has a formula, and the two losses differ by one exponent

wrong$$\hat\beta^L=(X^TX+\lambda I)^{-1}X^Ty$$
right$$\hat\beta^L:\ \text{no closed form in general; solve by cases or numerically}$$

5.7Budgets and shapes: why the lasso's solution sits on a corner

Rewrites both penalties as a budget on the coefficients and shows why the lasso's diamond produces zeros and ridge's disk does not.

The path figure showed the lasso zeroing thermometer 2; the constrained form shows why, in one picture.

TheoremPenalized and constrained forms
Conditions
  • centered response, standardized predictors

  • each $\lambda$ has a matching budget $s$, and each budget below the least squares value of the penalty has a matching $\lambda>0$

$$\boxed{\begin{aligned}\text{ridge: }&\min_\beta\,\mathrm{RSS}(\beta)\ \ \text{subject to}\ \sum_{j=1}^p\beta_j^2\le s\\ \text{lasso: }&\min_\beta\,\mathrm{RSS}(\beta)\ \ \text{subject to}\ \sum_{j=1}^p\lvert\beta_j\rvert\le s\\ \text{matching: }&s=\textstyle\sum_j(\hat\beta^R_{\lambda,j})^2\ \ \text{or}\ \ s=\sum_j\lvert\hat\beta^L_{\lambda,j}\rvert\end{aligned}}$$

Instead of a price $\lambda$ per unit of size, give the coefficients a budget $s$ and take the smallest RSS the budget allows. Both forms trace the same fits: the penalized fit at $\lambda$ spends exactly the budget $s$, and a smaller budget matches a larger $\lambda$. Ridge's budget region is a disk, the lasso's a diamond with its corners on the axes.

Why a penalized fit solves the budget problem

Let $\hat\beta_\lambda$ minimize $\mathrm{RSS}+\lambda P$ and set $s=P(\hat\beta_\lambda)$. For any $\beta$ with $P(\beta)\le s$, optimality gives $\mathrm{RSS}(\beta)+\lambda P(\beta)\ge\mathrm{RSS}(\hat\beta_\lambda)+\lambda s$.

Since $\lambda P(\beta)\le\lambda s$, this leaves $\mathrm{RSS}(\beta)\ge\mathrm{RSS}(\hat\beta_\lambda)$: nothing within the budget fits better, so $\hat\beta_\lambda$ solves the constrained problem.

In the picture, the RSS contours are ellipses around $\hat\beta^{\mathrm{LS}}$, and the constrained solution is where the growing ellipses first reach the region. A diamond's corners stick out toward the contours, so first contact is often at a corner, where a coefficient is $0$; a disk has no corners.

Looks like this, but is not

Every budget looks as if it forces some shrinkage, since the fit now has to obey a constraint.

A budget that the least squares fit already meets costs nothing: the constrained solution is $\hat\beta^{\mathrm{LS}}$ itself, which matches $\lambda=0$. For the kiosk that happens for the lasso once $s\ge\lvert3\rvert+\lvert-1\rvert=4$, and for ridge once $s\ge3^2+(-1)^2=10$.

One predictor with a budget: which λ matches β = 0.8

A standardized predictor with $n=10$ and $\sum_ix_iy_i=12$ has $\hat\beta^{\mathrm{LS}}=1.2$. Solve the ridge problem with budget $\beta^2\le0.64$ and the lasso problem with budget $\lvert\beta\rvert\le0.8$, and find the $\lambda$ that gives each answer in the penalized form.

FindBoth constrained solutions and their matching values of $\lambda$.
Given
  • $n=10$, $\sum_ix_i^2=10$, $\sum_ix_iy_i=12$

  • budgets: $\beta^2\le0.64$ (ridge), $\lvert\beta\rvert\le0.8$ (lasso)

Solution

With one coefficient each budget region is an interval, so we read the constrained solution off the number line and then solve the penalized formulas backwards for $\lambda$.

Constrained solutions

$$\mathrm{RSS}(\beta)=\mathrm{RSS}(1.2)+10(\beta-1.2)^2$$

A parabola in β with its lowest point at the least squares value.

$$\beta^2\le0.64\iff\lvert\beta\rvert\le0.8:\qquad \beta=0.8$$

$1.2$ lies outside $[-0.8,\,0.8]$, and a parabola is smallest at the allowed point nearest its vertex.

$$\lvert\beta\rvert\le0.8:\qquad \beta=0.8$$

In one dimension the disk and the diamond are the same interval.

Matching λ

$$\text{ridge: }\ \frac{12}{10+\lambda}=0.8\ \Rightarrow\ \lambda=5$$

The penalized fit matches when it lands on the same coefficient, which fixes λ.

$$\text{lasso: }\ \frac{12-\lambda/2}{10}=0.8\ \Rightarrow\ \lambda=8$$

The target is positive, so only the positive branch of the soft threshold can produce it.

Answer $$\boxed{\beta=0.8\ \text{in both};\qquad \lambda_{\text{ridge}}=5,\qquad \lambda_{\text{lasso}}=8}$$
Check

Penalized check: $\frac{12}{15}=0.8$ and $\frac{12-4}{10}=0.8$. The budgets are spent exactly, $0.8^2=0.64$ and $\lvert0.8\rvert=0.8$, as the matching rule in the box says.

A budget and a price are two handles on the same family of fits, but the numbers do not carry over between methods: the same coefficient needs $\lambda=5$ for ridge and $\lambda=8$ for the lasso.

The kiosk budgets: which s gives which fit

For the kiosk fit at $\lambda=3$, find the ridge and lasso budgets, and the smallest budget at which each constrained problem returns plain least squares.

FindThe budgets at λ = 3, the budgets beyond which nothing is shrunk, and the lasso's solution for any budget up to 2.
Given
  • $\hat\beta^{\mathrm{LS}}=(3,\,-1)$

  • $\hat\beta^R=(1.25,\,0.25)$ and $\hat\beta^L=(1.9,\,0)$ at $\lambda=3$

  • lasso path: $(3-\frac\lambda2,\,-1+\frac\lambda2)$ for $\lambda\le2$, $(\frac{22-\lambda}{10},\,0)$ for $2\le\lambda\le22$

Solution

Each budget is the penalty of the matching penalized fit, so the work is evaluating penalties.

Budgets at λ = 3

$$s_R=1.25^2+0.25^2=1.625,\qquad s_L=\lvert1.9\rvert+\lvert0\rvert=1.9$$

Plug each penalized fit into its own penalty.

When nothing is shrunk

$$s_R\ge3^2+(-1)^2=10,\qquad s_L\ge\lvert3\rvert+\lvert-1\rvert=4$$

From these budgets on, the least squares fit is allowed, and nothing beats its RSS.

The lasso's corner

$$2\le\lambda\le22:\quad s_L=\tfrac{22-\lambda}{10}\in[0,\,2],\qquad \hat\beta^L=(s_L,\,0)$$

On this stretch of the path $\hat\beta_2=0$, so every budget up to $2$ buys a corner solution.

Answer $$\boxed{s_R=1.625,\quad s_L=1.9;\qquad \text{no shrinkage once}\ s_R\ge10\ \text{or}\ s_L\ge4}$$
Check

The picture agrees: the diamond with $s_L=1.9$ is touched at its corner $(1.9,0)$, and $(3,-1)$ sits on the boundary of the diamond with $s_L=4$, since $3+1=4$.

Budgets and prices run in opposite directions: as $\lambda$ goes from $0$ to $\infty$, the budget runs from its least squares value down to $0$.

Checkpoint
§05.7 — where the zeros of the lasso come from

Two predictors, the same elliptical RSS contours around $\hat\beta^{\mathrm{LS}}$, and two budget regions: the disk $\beta_1^2+\beta_2^2\le s$ and the diamond $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le s$.

Find(a) Why does the lasso often give a coefficient of exactly 0 while ridge does not?
Given
  • ridge region: $\beta_1^2+\beta_2^2\le s$

  • lasso region: $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le s$

  • both solutions: the first point of the region reached by the growing RSS contours

Hint 1/4

Both solutions sit where the smallest RSS ellipse first touches the region; look at what differs between the two regions.

Hint 2/4

The disk is round. The diamond has vertices $(\pm s,0)$ and $(0,\pm s)$, each with one coordinate equal to $0$.

Hint 3/4

The regions are $\beta_1^2+\beta_2^2\le s$ and $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le s$, and the contours are the same ellipses for both.

Hint 4/4

The zeros come from the diamond's corners, which sit on the axes and are often where first contact happens.

Show solution

The two problems differ only in the region, so we compare the regions' shapes.

Where the solution sits

$$\hat\beta=\text{first point of the region on a growing contour}$$

Any other point of the region lies on a larger contour, so it has a larger RSS.

Compare the shapes

$$\text{diamond vertices }(\pm s,0),\ (0,\pm s);\qquad \text{disk: no vertices}$$

A vertex sticks out toward the contours, and it has a zero coordinate.

Answer $$\boxed{\text{the diamond's corners lie on the axes}}$$
Check

The kiosk figure shows it: the diamond is met at its corner $(1.9,0)$, the disk at $(1.25,0.25)$.

When asked why the lasso selects variables, point to the corners of the constraint region.

⚠ Reading the ridge budget as a radius

the budget is written where a radius usually goes

wrong$$\beta_1^2+\beta_2^2\le s\ \Rightarrow\ \text{disk of radius }s$$
right$$\beta_1^2+\beta_2^2\le s\ \Rightarrow\ \text{disk of radius }\sqrt s$$
⚠ Expecting shrinkage from a budget least squares already meets

a constraint sounds as if it always bites

wrong$$s_L=5\ \text{for}\ \hat\beta^{\mathrm{LS}}=(3,-1):\ \ \text{a shrunk fit}$$
right$$s_L\ge\lvert3\rvert+\lvert-1\rvert=4:\ \ \hat\beta=\hat\beta^{\mathrm{LS}}$$
⚠ Carrying λ over between ridge and the lasso

both methods call their tuning constant λ

wrong$$\text{same }\lambda\ \Rightarrow\ \text{same amount of shrinkage}$$
right$$\beta=0.8\ \text{from}\ \hat\beta^{\mathrm{LS}}=1.2:\ \ \lambda=5\ \text{(ridge)},\ \ \lambda=8\ \text{(lasso)}$$
−4−4−2−22244β₁β₂budget s = 1.9solution (1.9, 0)on the corner (s, 0), soβ₂ is exactly 0RSS above its minimum: 2.25least squares (3, −1)

Step the lasso budget $s$ on the kiosk data. Up to $s=2$ the growing $\textcolor{#8250df}{\text{RSS ellipses}}$ first touch the $\textcolor{#d1690a}{\text{diamond}}$ at its corner $(s,0)$, so $\beta_2=0$; from $s=2$ to $s=4$ the contact point slides along the lower edge; from $s=4$ on, least squares itself fits the budget.

At the edges
s = 1.9 (1.9, 0)

The budget that matches λ = 3: thermometer 2 is out of the model.

s = 2 (2, 0)

The last budget with a corner solution; it matches λ = 2, where the lasso path lets thermometer 2 back in.

s = 4 (3, −1)

The least squares fit spends exactly |3| + |−1| = 4, so every larger budget returns it unchanged.

A ridge fit by hand, from raw data to a prediction

A small table of raw data, a given λ, and a question about the coefficients or a new input.

  1. Training statistics

    Compute $\bar x_j$, $\sigma_j=\sqrt{\frac1n\sum_i(x_{ij}-\bar x_j)^2}$ and $\bar y$ from the training rows only.

  2. Standardize and center

    $\tilde x_{ij}=(x_{ij}-\bar x_j)/\sigma_j$ and $y_i-\bar y$; check that $\sum_i\tilde x_{ij}^2=n$.

  3. Two products

    $X^TX$ and $X^Ty$ from the standardized columns and the centered response.

  4. Add λ to the diagonal

    Only the diagonal of $X^TX$ changes; $X^Ty$ stays as it is.

  5. Solve and check

    For two predictors use the $2\times2$ inverse, then multiply back to recover $X^Ty$.

  6. Predict

    $\tilde z_j=(z_j-\bar x_j)/\sigma_j$ with the training statistics, then $\hat y=\bar y+\tilde z^T\hat\beta^R$.

Where it goes wrong
  • Adding $\lambda$ to every entry of $X^TX$, or to $X^Ty$.

  • Standardizing the new input with its own statistics.

  • Stopping at $\tilde z^T\hat\beta^R$ without adding $\bar y$ back.

A lasso fit by cases

One predictor, uncorrelated standardized predictors, or two predictors where you can guess which coefficient is zero.

  1. Reduce

    With one predictor put $z=\sum_ix_iy_i$ and $n=\sum_ix_i^2$; the loss is $-2z\beta+n\beta^2+\lambda\lvert\beta\rvert$ plus a constant.

  2. Compare with the cut

    If $\lvert z\rvert\le\lambda/2$, then $\hat\beta^L=0$.

  3. Otherwise subtract

    $\hat\beta^L=\big(z-\operatorname{sign}(z)\,\lambda/2\big)/n$: the least squares value moved $\frac{\lambda}{2n}$ toward $0$.

  4. Several predictors

    Guess which coefficients are $0$, fit the others by the same rule, then check each zero: $\lvert\partial\mathrm{RSS}/\partial\beta_j\rvert\le\lambda$ there.

  5. Confirm

    Compare the lasso loss at your answer with its value at the least squares fit and at one other candidate.

Where it goes wrong
  • Using the cut $\lambda/n$ instead of $\lambda/(2n)$.

  • Setting a derivative to zero at $\beta=0$, where the absolute value has none.

  • Dropping the sign of the least squares value.

Choosing λ by cross-validation

Any ridge or lasso fit that will be used on new data.

  1. Grid

    Pick values on a log scale, for example $0,1,3,9,27,81$ or powers of ten.

  2. Folds

    Split the data into k folds once and use the same folds for every λ.

  3. Fit inside each fold

    Standardize with the training part of the fold, fit, and score the held-out part.

  4. Average

    Compute $\mathrm{CV}(k)$ for each $\lambda$ and keep the smallest.

  5. Refit

    Fit the chosen λ on all the data, standardizing with all the data.

Where it goes wrong
  • Scoring $\lambda$ on the training data, which always picks $\lambda=0$.

  • Standardizing with all the data before splitting, which lets each held-out fold shape its own scaling.

  • Reporting one fold's fit instead of the refit.

Ridge on three uncorrelated predictors: every coefficient scaled by the same factor

Three standardized, uncorrelated predictors with $n=10$ ($X^TX=10I$) have least squares coefficients $(2.0,\,-0.6,\,0.15)$. Compute ridge with $\lambda=5$.

Find$\hat\beta^R$.
Given
  • $X^TX=10I$, $\hat\beta^{\mathrm{LS}}=(2.0,\,-0.6,\,0.15)$

  • $\lambda=5$

Solution

With $X^TX=nI$ the system splits into three one-predictor problems.

The factor

$$\frac{n}{n+\lambda}=\frac{10}{15}=\frac23$$

The same for all three: each has $\sum_ix_{ij}^2=10$ and there are no cross terms.

Scale

$$\hat\beta^R=\tfrac23\,(2.0,\,-0.6,\,0.15)=(1.333,\,-0.4,\,0.1)$$

Each coefficient keeps its sign and two thirds of its size.

Answer $$\boxed{\hat\beta^R=(1.333,\ -0.4,\ 0.1)}$$
Check

Ratios are preserved: $2.0:0.6:0.15$ and $1.333:0.4:0.1$ are both $40:12:3$.

Uncorrelated predictors make ridge one division per coefficient, and no coefficient reaches zero.

The lasso on the same three predictors: every coefficient moved by the same amount

Same predictors and least squares coefficients $(2.0,\,-0.6,\,0.15)$ with $X^TX=10I$. Compute the lasso with $\lambda=5$.

Find$\hat\beta^L$.
Given
  • $X^TX=10I$, $\hat\beta^{\mathrm{LS}}=(2.0,\,-0.6,\,0.15)$

  • $\lambda=5$

Solution

Uncorrelated predictors split the lasso too, so each coefficient follows the one-predictor rule.

The cut

$$\frac{\lambda}{2n}=\frac{5}{20}=0.25$$

The same cut for all three coefficients.

Subtract and clip

$$\big(2.0-0.25,\ -(0.6-0.25),\ (0.15-0.25)_+\big)=(1.75,\,-0.35,\,0)$$

The third size is below the cut, so it stops at $0$.

Answer $$\boxed{\hat\beta^L=(1.75,\ -0.35,\ 0)}$$
Check

Check the zero: at $\beta_3=0$ the RSS slope is $-2x_3^Ty=-2\cdot10\cdot0.15=-3$, and $\lvert-3\rvert\le5$.

The lasso shifts every size by the same amount, so small coefficients vanish and large ones barely change in proportion.

Same data and the same $\lambda$: ridge multiplies every coefficient by $\frac23$ and keeps all three; the lasso subtracts $0.25$ from every size and drops the third.

How to tell them apart

Ridge divides, the lasso subtracts. Look at the smallest coefficient: if it survives in proportion, the fit is ridge; if it is exactly $0$, it is the lasso.

Least squares after a change of units: the coefficient absorbs the factor

Three people's centered ages in years are $(-10,\,0,\,10)$ and their centered responses $(-4,\,1,\,3)$. Fit least squares, refit with age in decades, and predict for someone $20$ years above the mean age.

FindBoth slopes and both predictions.
Givencentered ages $(-10,\,0,\,10)$ years; centered responses $(-4,\,1,\,3)$
Solution

One predictor, so each fit is one division, done once per unit.

Years

$$\hat\beta=\frac{40+0+30}{100+0+100}=\frac{70}{200}=0.35$$

Centered data let the slope be computed without an intercept.

Decades

$$\hat\beta=\frac{4+0+3}{1+0+1}=\frac72=3.5$$

Every age is now ten times smaller.

Predict

$$0.35\times20=7=3.5\times2$$

$20$ years is $2$ decades.

Answer $$\boxed{0.35\ \text{per year}=3.5\ \text{per decade};\quad \text{prediction }7\ \text{either way}}$$
Check

The coefficient changed by exactly the unit factor $10$, so the product with age did not change.

Least squares answers a change of units with the reverse change in its coefficient.

Ridge with λ = 10 after the same change: the prediction moves

Same data. Fit ridge with $\lambda=10$ on the centered, unstandardized ages, once in years and once in decades, and predict for someone $20$ years above the mean age.

FindBoth ridge slopes and both predictions.
Given
  • centered ages $(-10,\,0,\,10)$ years; centered responses $(-4,\,1,\,3)$

  • $\lambda=10$

Solution

The same two divisions as least squares, each with λ added to the sum of squares.

Years

$$\hat\beta^R=\frac{70}{200+10}=0.333$$

The $10$ is small next to $200$.

Decades

$$\hat\beta^R=\frac{7}{2+10}=0.583$$

The same $10$ now dwarfs the sum of squares $2$.

Predict

$$0.333\times20=6.67\qquad \text{vs}\qquad 0.583\times2=1.17$$

Same person, same λ, different answers.

Answer $$\boxed{\text{prediction }6.67\ \text{(years)}\quad \text{vs}\quad 1.17\ \text{(decades)}}$$
Check

The shrink factors explain it: $\frac{200}{210}\approx0.95$ in years and $\frac{2}{12}\approx0.17$ in decades, multiplying the least squares prediction $7$.

On raw units the amount of ridge shrinkage is decided by the unit, not by the data.

One change of units, two methods: least squares moves its coefficient by exactly the unit factor and keeps the prediction $7$; ridge with a fixed $\lambda$ turns $6.67$ into $1.17$.

How to tell them apart

If a method's predictions survive a change of units, it puts no penalty on raw coefficients; before ridge or the lasso, standardize so that theirs survive too.

Scaffolding comes off
The common skeleton
  1. Put the data on the standard scale: training means and spreads, standardized predictors, centered response.

  2. Form $X^TX$ and $X^Ty$.

  3. Add $\lambda$ to the diagonal of $X^TX$ only.

  4. Solve $(X^TX+\lambda I)\beta=X^Ty$ and check by multiplying back.

  5. For a new input, standardize with the training statistics and add $\bar y$ back.

1 · fully worked

Six wheat plots: ridge at λ = 4 from raw data to a prediction

Six plots got nitrogen $10, \allowbreak 10, \allowbreak 10, \allowbreak 14, \allowbreak 14, \allowbreak 14$ kg and irrigation $3,3,7,3,7,7$ hours a week, and yielded $6,8,9,11,12,14$ kg. Fit ridge with $\lambda=4$ on standardized inputs and centered yield, and predict a plot with $13$ kg of nitrogen and $4$ hours of irrigation.

Find$\hat\beta^R$ and the predicted yield.
Given
  • nitrogen $10, \allowbreak 10, \allowbreak 10, \allowbreak 14, \allowbreak 14, \allowbreak 14$; irrigation $3,3,7,3,7,7$; yield $6,8,9,11,12,14$

  • $\lambda=4$; new plot: nitrogen $13$, irrigation $4$

Solution

Two predictors call for the $2\times2$ inverse, and each input takes only two values, which makes standardizing quick.

Training statistics and standardized columns

$$\bar x_1=12,\ \sigma_1=2;\qquad \bar x_2=5,\ \sigma_2=2;\qquad \bar y=10$$

Each input sits $2$ above or below its mean, so each spread is $2$.

$$\tilde x_1=(-1,-1,-1,1,1,1),\quad \tilde x_2=(-1,-1,1,-1,1,1),\quad y^c=(-4,-2,-1,1,2,4)$$

Divide the deviations by 2 and subtract the mean yield.

Form the products

$$X^TX=\begin{pmatrix}6&2\\2&6\end{pmatrix},\qquad X^Ty=\begin{pmatrix}4+2+1+1+2+4\\4+2-1-1+2+4\end{pmatrix}=\begin{pmatrix}14\\10\end{pmatrix}$$

The two columns agree on four plots and disagree on two, so their product is $4-2=2$.

Add λ to the diagonal

$$X^TX+4I=\begin{pmatrix}10&2\\2&10\end{pmatrix},\qquad \det=100-4=96$$

Only the diagonal changes.

Solve and check

$$\hat\beta^R=\frac1{96}\begin{pmatrix}10&-2\\-2&10\end{pmatrix}\begin{pmatrix}14\\10\end{pmatrix}=\frac1{96}\begin{pmatrix}120\\72\end{pmatrix}=\begin{pmatrix}1.25\\0.75\end{pmatrix}$$

The $2\times2$ inverse: swap, negate, divide by the determinant.

$$10(1.25)+2(0.75)=14,\qquad 2(1.25)+10(0.75)=10$$

Multiplying back reproduces $X^Ty$.

Predict

$$\tilde z=\Big(\tfrac{13-12}{2},\ \tfrac{4-5}{2}\Big)=(0.5,\,-0.5),\qquad \hat y=10+1.25(0.5)+0.75(-0.5)=10.25$$

Training means and spreads for the new plot, then $\bar y$ back.

Answer $$\boxed{\hat\beta^R=(1.25,\ 0.75),\qquad \hat y=10.25\ \text{kg}}$$
Check

Least squares for comparison: $\frac1{32}\begin{pmatrix}6&-2\\-2&6\end{pmatrix}\begin{pmatrix}14\\10\end{pmatrix}=(2,\,1)$, whose prediction is $10+1-0.5=10.5$; ridge shrank both coefficients and landed between that and the mean $10$.

The skeleton never changes: training statistics, products, λ on the diagonal, solve and check, predict on the training scale.

2 · you write the reasoning

Easier, and this time you write the reasons. One standardized predictor from $n=8$ training points with mean $50$ and spread $10$, $\sum_i\tilde x_iy_i^c=6$, mean response $20$ and $\lambda=4$. Predict at a new input $65$, and for each line write why it is allowed.

  1. $\sum_i\tilde x_i^2=8$ for the standardized column.

    reasoning

    Standardizing with the training mean and the lecture's spread makes the sum of squares equal to n.

  2. $X^TX+\lambda I=8+4=12$.

    reasoning

    With one column $X^TX$ is a number, and the penalty adds $\lambda$ to it, the only diagonal entry.

  3. $\hat\beta^R=\frac{6}{12}=0.5$.

    reasoning

    The ridge system $(X^TX+\lambda I)\beta=X^Ty$ is one division here, with $X^Ty=\sum_i\tilde x_iy_i^c=6$.

  4. $\tilde z=\frac{65-50}{10}=1.5$ for a new input $65$.

    reasoning

    A new input is standardized with the training mean and spread, not its own.

  5. $\hat y=20+0.5\times1.5=20.75$.

    reasoning

    The fit was made on centered responses, so the mean response is added back.

3 · find the buried error

Harder, with two errors buried in a worked solution. Six plots got nitrogen $10, \allowbreak 10, \allowbreak 10, \allowbreak 14, \allowbreak 14, \allowbreak 14$ kg and irrigation $3,3,7,3,7,7$ hours, and yielded $16, \allowbreak 17, \allowbreak 19, \allowbreak 21, \allowbreak 23, \allowbreak 24$ kg. The task: ridge with $\lambda=6$, and a prediction for a plot with $15$ kg of nitrogen and $6$ hours of irrigation. Which two steps are wrong?

  1. Step 1. $\bar x=(12,5)$, $\sigma=(2,2)$, $\bar y=20$; $\tilde x_1=(-1, \allowbreak -1, \allowbreak -1, \allowbreak 1, \allowbreak 1, \allowbreak 1)$ and $\tilde x_2=(-1, \allowbreak -1, \allowbreak 1, \allowbreak -1, \allowbreak 1, \allowbreak 1)$.

  2. Step 2. $X^TX=\begin{pmatrix}6&2\\2&6\end{pmatrix}$ and $X^Ty=(16,\,12)^T$, from the centered yields $(-4, \allowbreak -3, \allowbreak -1, \allowbreak 1, \allowbreak 3, \allowbreak 4)$.

  3. Step 3. Add the penalty: $X^TX+6I=\begin{pmatrix}12&8\\8&12\end{pmatrix}$, with determinant $144-64=80$.

  4. Step 4. Invert and multiply: $\hat\beta^R=\frac1{80}(12\cdot16-8\cdot12,\ -8\cdot16+12\cdot12)=(1.2,\ 0.2)$.

  5. Step 5. New plot: $\tilde z=\big(\frac{15-12}{4},\ \frac{6-5}{4}\big)=(0.75,\,0.25)$, so $\hat y=20+1.2(0.75)+0.2(0.25)=20.95$.

the two buried errors (2)
⚠ step 3

The penalty was added to every entry. $\lambda I$ adds $6$ to the diagonal only, so the matrix is $\begin{pmatrix}12&2\\2&12\end{pmatrix}$ with determinant $140$.

'Add λ' sticks in memory without the identity matrix that says where.

right

$\hat\beta^R=\frac{1}{140}(12\cdot16-2\cdot12,\ -2\cdot16+12\cdot12)$, which is $\frac{1}{140}(168,\,112)=(1.2,\ 0.8)$

⚠ step 5

The new plot was divided by the variance $4$ instead of the standard deviation $2$.

$\sigma^2=4$ is computed on the way to $\sigma=2$, and the two swap easily once the training columns are done.

right

$\tilde z=(1.5,\,0.5)$ and $\hat y=20+1.2(1.5)+0.8(0.5)=22.2$

4 · the bare problem
§05.1 — a ridge fit and a prediction, no scaffolding

A shop models daily sales from two standardized predictors measured on $n=6$ days: temperature and hours of rain, which tend to move in opposite directions. The fit is ridge with $\lambda=2$.

Find
  1. (a) Compute $\hat\beta^R$.

  2. (b) Predict sales for the new day.

Given
  • $X^TX=\begin{pmatrix}6&-2\\-2&6\end{pmatrix}$ and $X^Ty=(12,\,-6)^T$ (standardized predictors, centered sales)

  • training means $24$ °C and $3$ hours; spreads $4$ °C and $2$ hours; mean sales $30$

  • new day: $28$ °C and $2$ hours of rain

Hint 1/4

Follow the skeleton: add λ to the diagonal, solve, then put the new day on the training scale.

Hint 2/4

$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$; $\tilde z_j=(z_j-\bar x_j)/\sigma_j$ and $\hat y=\bar y+\tilde z^T\hat\beta^R$.

Hint 3/4

$X^TX+2I=\begin{pmatrix}8&-2\\-2&8\end{pmatrix}$ with determinant $60$ and $X^Ty=(12,-6)^T$; the new day is $\big(\frac{28-24}{4},\ \frac{2-3}{2}\big)$ on the training scale, and mean sales are $30$.

Hint 4/4

$\hat\beta^R=(1.4,\,-0.4)$, $\tilde z=(1,\,-0.5)$ and $\hat y=31.6$.

Show solution

The same skeleton as the plots; only the sign of the off-diagonal entry differs.

Add λ to the diagonal

$$X^TX+2I=\begin{pmatrix}8&-2\\-2&8\end{pmatrix},\qquad \det=64-4=60$$

Only the diagonal changes.

Solve

$$\hat\beta^R=\frac1{60}\begin{pmatrix}8&2\\2&8\end{pmatrix}\begin{pmatrix}12\\-6\end{pmatrix}=\frac1{60}\begin{pmatrix}84\\-24\end{pmatrix}=\begin{pmatrix}1.4\\-0.4\end{pmatrix}$$

Negating the off-diagonal $-2$ gives $+2$ in the inverse.

Predict

$$\tilde z=\Big(\tfrac{28-24}{4},\ \tfrac{2-3}{2}\Big)=(1,\,-0.5),\qquad \hat y=30+1.4(1)-0.4(-0.5)=31.6$$

Training statistics for the new day, then the mean back.

Answer $$\boxed{\hat\beta^R=(1.4,\ -0.4),\qquad \hat y=31.6}$$
Check

Multiply back: $8(1.4)-2(-0.4)=12$ and $-2(1.4)+8(-0.4)=-6$. Least squares would give $\frac1{32}(6\cdot12-2\cdot6,\ 2\cdot12-6\cdot6)=(1.875,\,-0.375)$: larger in the first coefficient ($1.875>1.4$) but smaller in the second ($0.375<0.4$). The vector still shrinks, $\lVert\beta\rVert_2^2$ from $3.66$ to $2.12$, while one coefficient grows, as on the kiosk path.

A negative correlation between the inputs only flips the sign of the off-diagonal entry; the skeleton stays the same.

Full exam-style question

Exam-style: ridge and lasso on two uncorrelated predictorsexam format

Centered responses and two standardized, uncorrelated predictors from $n=4$ observations give $X^TX=4I$ and $X^Ty=(8,\,-1)^T$. (a) Derive the ridge estimate from $\mathrm{Loss}_R$ and evaluate it at $\lambda=4$. (b) Evaluate the lasso estimate at $\lambda=4$. (c) Find the smallest $\lambda$ at which the lasso sets both coefficients to $0$. (d) Give the lasso budget $s$ that matches $\lambda=4$.

Find$\hat\beta^R$, $\hat\beta^L$, the smallest all-zero $\lambda$, and the budget $s$.
Given
  • $n=4$, $X^TX=4I$, $X^Ty=(8,\,-1)^T$

  • $\lambda=4$

Solution

Uncorrelated standardized predictors split both problems into one-predictor problems, which is quicker than any matrix inverse.

(a) Derive and evaluate ridge

$$-2y^TX+2\beta^T(X^TX+\lambda I)=0\ \Rightarrow\ \hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$$

Set the row derivative of $\mathrm{Loss}_R$ to zero; the matrix is invertible for $\lambda>0$.

$$\hat\beta^R=\frac{1}{4+4}\,(8,\,-1)=(1,\ -0.125)$$

$X^TX+4I=8I$, whose inverse is $\frac18I$.

(b) The lasso, coordinate by coordinate

$$\hat\beta^{\mathrm{LS}}=\tfrac14(8,\,-1)=(2,\,-0.25),\qquad \frac{\lambda}{2n}=\frac48=0.5$$

Each coordinate is a one-predictor lasso with the same cut.

$$\hat\beta^L=\big(2-0.5,\ -(0.25-0.5)_+\big)=(1.5,\ 0)$$

The second size, $0.25$, is below the cut.

(c) The empty model

$$\hat\beta^L=0\iff\lambda\ge2\max_j\lvert x_j^Ty\rvert=2\cdot8=16$$

At $\beta=0$ the RSS slopes are $-2x_j^Ty$, and each must be at most $\lambda$ in size.

(d) The budget

$$s=\lvert1.5\rvert+\lvert0\rvert=1.5$$

The matching budget is the penalty of the penalized fit.

Answer $$\boxed{\hat\beta^R=(1,\,-0.125),\quad \hat\beta^L=(1.5,\,0),\quad \lambda\ge16,\quad s=1.5}$$
Check

Check the zero in (b): at $(1.5,0)$ the RSS slope in $\beta_2$ is $-2(x_2^Ty-x_2^Tx_1\cdot1.5)=-2(-1-0)=2\le4$. For (a), multiply back: $8\,(1,\,-0.125)=(8,\,-1)$.

With $X^TX=nI$, ridge divides every coefficient by the same $1+\lambda/n$, while the lasso subtracts the same $\lambda/(2n)$ and drops whatever falls below it.

Practice

A · concept 4 questions
1§05.2 — can one ridge coefficient grow with λ?

A classmate sums ridge up as 'the larger $\lambda$, the closer every coefficient is to zero'. You test the claim on the kiosk fit.

Find(a) True or false: increasing $\lambda$ never increases the absolute value of any single ridge coefficient.
Given
  • $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$, $X^Ty=(11,\,7)^T$

  • ridge fits: $(1.9,\,-0.1)$ at $\lambda=1$ and $(1.25,\,0.25)$ at $\lambda=3$

Hint 1/4

The claim is about each coefficient separately; test it on one coefficient at two values of λ.

Hint 2/4

Ridge guarantees only that $\sum_j\beta_j^2$ never grows with $\lambda$; nothing is promised for a single coefficient.

Hint 3/4

At $\lambda=1$ the fit is $(1.9,\,-0.1)$ and at $\lambda=3$ it is $(1.25,\,0.25)$.

Hint 4/4

$\lvert\hat\beta_2\rvert$ rises from $0.1$ to $0.25$, so the claim is false.

Show solution

One counterexample refutes a claim about every coefficient, so we compare the two fits instead of arguing in general.

Coefficient by coefficient

$$\lvert\hat\beta_1\rvert:\ 1.9\to1.25,\qquad \lvert\hat\beta_2\rvert:\ 0.1\to0.25$$

The second coefficient grows in size.

The whole vector

$$\textstyle\sum_j\hat\beta_j^2:\ 3.62\to1.625$$

This is the quantity the ordering proof controls, and it does fall.

Answer $$\boxed{\text{False}}$$
Check

Recompute both fits: $\frac1{20}(6\cdot11-4\cdot7,\ -4\cdot11+6\cdot7)=(1.9,\,-0.1)$ and $\frac1{48}(8\cdot11-4\cdot7,\ -4\cdot11+8\cdot7)=(1.25,\,0.25)$.

Ridge shrinks the coefficient vector, not each coefficient; with correlated predictors a single coefficient can move away from zero for a while.

2§05.1 — why the intercept is not penalized

A ridge fit predicts exam scores out of $100$. A classmate suggests adding $\lambda\beta_0^2$ to the penalty so that every coefficient is treated alike.

Find(a) What is the main reason ridge leaves $\beta_0$ out of the penalty?
Given
  • the model $\hat y=\beta_0+\sum_{j=1}^p\beta_jx_j$ on standardized predictors

  • the proposal: penalty $\lambda\sum_{j=0}^p\beta_j^2$

Hint 1/4

Ask what the size of $\beta_0$ depends on, apart from the pattern in the data.

Hint 2/4

With standardized predictors, $\hat\beta_0=\bar y$: the intercept is the average response.

Hint 3/4

Recording scores as points above $50$ instead of out of $100$ moves $\bar y$ by $50$ and changes nothing else in the data.

Hint 4/4

The intercept only records where zero sits on the response scale, so charging for its size would tie the fit to an arbitrary choice.

Show solution

Shifting the response is a harmless change of scale, so we test the proposal against it.

What the intercept is

$$\hat\beta_0=\bar y-\textstyle\sum_j\hat\beta_j\bar x_j=\bar y$$

Standardized predictors have mean zero.

Shift the scale

$$y\to y-50:\qquad \hat\beta_0\to\bar y-50,\qquad \hat\beta_1,\dots,\hat\beta_p\ \text{unchanged}$$

Only the intercept notices, so a penalty on it would make the fit depend on the choice of zero.

Answer $$\boxed{\beta_0\ \text{records location, not complexity}}$$
Check

Unpenalized, the shifted fit predicts exactly $50$ less for everyone, as it should; penalized, $\beta_0$ would be pulled toward $0$ by different amounts on the two scales.

Penalize what measures complexity, the slopes, and leave the location of the response alone.

3§05.5 — can ridge remove a small coefficient?

Three standardized, uncorrelated predictors have least squares coefficients $2.0$, $-0.6$ and $0.05$ from $n=10$ observations. A classmate expects a large enough $\lambda$ to remove the third predictor while keeping the other two.

Find(a) True or false: for some $\lambda>0$, ridge sets the third coefficient exactly to $0$ while the first two stay nonzero.
Given
  • $X^TX=10I$

  • $\hat\beta^{\mathrm{LS}}=(2.0,\,-0.6,\,0.05)$

Hint 1/4

With uncorrelated standardized predictors each ridge coefficient is its own one-predictor problem; ask what ridge does to one coefficient.

Hint 2/4

When $X^TX=nI$, $\hat\beta^R_j=\frac{n}{n+\lambda}\,\hat\beta^{\mathrm{LS}}_j$.

Hint 3/4

Here the factor is $\frac{10}{10+\lambda}$ for all three, applied to $2.0$, $-0.6$ and $0.05$.

Hint 4/4

The factor is positive for every λ, so the third coefficient is never exactly 0 and the statement is false.

Show solution

The system splits into three one-predictor problems, so one formula answers for every λ.

Split the system

$$(X^TX+\lambda I)^{-1}X^Ty=\frac{1}{10+\lambda}X^Ty=\frac{10}{10+\lambda}\,\hat\beta^{\mathrm{LS}}$$

Here $X^Ty=10\,\hat\beta^{\mathrm{LS}}$.

Look for a zero

$$\frac{10}{10+\lambda}\times0.05>0\quad\text{for every }\lambda\ge0$$

A positive factor times a nonzero number is never zero.

Answer $$\boxed{\text{False}}$$
Check

Compare the lasso with the same data: its cut $\frac{\lambda}{20}$ passes $0.05$ at $\lambda=1$, and from there on the third coefficient is exactly $0$ while the first two are not, until $\lambda=12$ removes $-0.6$ as well.

Ridge rescales; it never selects. Exact zeros need the lasso.

4§05.2 — the fit as λ grows without bound

A ridge model of house prices uses standardized predictors and a centered response; the mean price in the training data is $\bar y=2.4$ million TL. You refit with larger and larger $\lambda$.

Find(a) What happens to the predicted price of every house as $\lambda\to\infty$?
Given
  • $\bar y=2.4$ million TL

  • standardized predictors, centered response, unpenalized intercept

Hint 1/4

Split the prediction into the part the penalty acts on and the part it leaves alone.

Hint 2/4

$\hat y=\bar y+\sum_j\hat\beta^R_j\tilde x_j$, and $\hat\beta^R_\lambda\approx X^Ty/\lambda\to0$.

Hint 3/4

Here $\bar y=2.4$ million TL, and the sum over $j$ goes to $0$.

Hint 4/4

Every prediction tends to $2.4$ million TL.

Show solution

The prediction splits into an unpenalized part and a penalized part, so we take the limit of each.

The penalized part

$$\hat\beta^R_\lambda=\tfrac1\lambda\big(I+\tfrac1\lambda X^TX\big)^{-1}X^Ty\to0$$

The bracket tends to $I$ while the factor $\frac1\lambda$ goes to $0$.

The prediction

$$\hat y=\bar y+\textstyle\sum_j\hat\beta^R_j\tilde x_j\ \to\ \bar y=2.4$$

The intercept is not in the penalty, so nothing pulls it.

Answer $$\boxed{\hat y\to2.4\ \text{million TL for every house}}$$
Check

The kiosk at $\lambda=1000$ shows the same thing: coefficients about $0.011$ and $0.007$, and every day predicted within $0.03$ of the mean sales $30$.

The heaviest ridge penalty leaves the mean predictor, not the zero predictor.

B · computation 7 questions
1§05.1 — shrinking one standardized coefficient

A clinic predicts recovery time from one standardized blood marker measured on $n=12$ patients. On centered data the marker and the recovery times give the products below.

Find
  1. (a) Compute $\hat\beta^{\mathrm{LS}}$ and $\hat\beta^R$ at $\lambda=6$.

  2. (b) For which $\lambda$ is $\hat\beta^R$ exactly half of $\hat\beta^{\mathrm{LS}}$?

Given$n=12$, $\sum_ix_i^2=12$, $\sum_ix_iy_i=18$
Hint 1/4

One standardized predictor turns every matrix in the ridge formula into a number.

Hint 2/4

$\hat\beta^R=\frac{\sum_ix_iy_i}{n+\lambda}=\frac{n}{n+\lambda}\,\hat\beta^{\mathrm{LS}}$.

Hint 3/4

Here $n=12$ and $\sum_ix_iy_i=18$, so $\hat\beta^{\mathrm{LS}}=\frac{18}{12}$ and the factor is $\frac{12}{12+\lambda}$.

Hint 4/4

$\hat\beta^{\mathrm{LS}}=1.5$, $\hat\beta^R=1$ at $\lambda=6$, and the factor is $\frac12$ at $\lambda=12$.

Show solution

With one column the ridge formula is a division, so we compute the factor and reuse it.

Least squares

$$\hat\beta^{\mathrm{LS}}=\frac{18}{12}=1.5$$

The one-predictor normal equation.

Ridge at λ = 6

$$\hat\beta^R=\frac{18}{12+6}=1$$

With one column, $X^TX+\lambda I$ is the number $12+\lambda$.

Halving

$$\frac{12}{12+\lambda}=\frac12\ \Rightarrow\ \lambda=12$$

Half means the factor $\frac{n}{n+\lambda}$ equals $\frac12$, which happens exactly at $\lambda=n$.

Answer $$\boxed{\hat\beta^{\mathrm{LS}}=1.5,\quad \hat\beta^R_6=1,\quad \lambda_{1/2}=12}$$
Check

Factor check: $\frac{12}{18}\times1.5=1$, and $\frac{18}{24}=0.75$ is half of $1.5$.

For one standardized predictor, λ = n halves the coefficient, which gives a feel for what a given λ does.

2§05.1 — a zero least squares coefficient that ridge makes nonzero

Two positively correlated standardized predictors come from $n=4$ observations. Least squares gives the second one no weight at all; you refit with ridge.

Find
  1. (a) Compute $\hat\beta^{\mathrm{LS}}$.

  2. (b) Compute $\hat\beta^R$ at $\lambda=2$.

  3. (c) Explain in one sentence why the second coefficient moved away from 0.

Given
  • $X^TX=\begin{pmatrix}4&2\\2&4\end{pmatrix}$, $X^Ty=(8,\,4)^T$

  • $\lambda=2$

Hint 1/4

Two linear systems with the same right-hand side; only the diagonal differs.

Hint 2/4

$\hat\beta=(X^TX+\lambda I)^{-1}X^Ty$ with $\lambda=0$ and with $\lambda=2$, using the $2\times2$ inverse.

Hint 3/4

$X^TX=\begin{pmatrix}4&2\\2&4\end{pmatrix}$ has determinant $12$, $X^TX+2I=\begin{pmatrix}6&2\\2&6\end{pmatrix}$ has determinant $32$, and $X^Ty=(8,4)^T$.

Hint 4/4

$\hat\beta^{\mathrm{LS}}=(2,\,0)$ and $\hat\beta^R=(1.25,\,0.25)$: ridge shrinks the difference of the coefficients harder than their sum.

Show solution

The $2\times2$ inverse gives both fits; the directions $(1,1)$ and $(1,-1)$ explain them.

Least squares

$$\frac1{12}\begin{pmatrix}4&-2\\-2&4\end{pmatrix}\begin{pmatrix}8\\4\end{pmatrix}=\frac1{12}\begin{pmatrix}24\\0\end{pmatrix}=\begin{pmatrix}2\\0\end{pmatrix}$$

Determinant $16-4=12$ is nonzero, so least squares has one answer.

Ridge

$$\frac1{32}\begin{pmatrix}6&-2\\-2&6\end{pmatrix}\begin{pmatrix}8\\4\end{pmatrix}=\frac1{32}\begin{pmatrix}40\\8\end{pmatrix}=\begin{pmatrix}1.25\\0.25\end{pmatrix}$$

Adding $2$ to the diagonal raises the determinant to $32$; the larger determinant does the shrinking.

Why the zero disappears

$$X^TX(1,1)^T=6(1,1)^T,\quad X^TX(1,-1)^T=2(1,-1)^T:\quad \tfrac{6}{6+2}=0.75,\ \ \tfrac{2}{2+2}=0.5$$

The sum direction keeps three quarters, the difference direction only half.

$$0.75(1,1)+0.5(1,-1)=(1.25,\,0.25)$$

The cancellation that made the second coefficient 0 is broken.

Answer $$\boxed{\hat\beta^{\mathrm{LS}}=(2,\,0),\qquad \hat\beta^R=(1.25,\,0.25)}$$
Check

Multiply back: $6(1.25)+2(0.25)=8$ and $2(1.25)+6(0.25)=4$.

Ridge shrinks along the data's directions, not coefficient by coefficient, so a zero can turn nonzero.

3§05.3 — from raw heights to a ridge prediction

Four adults have heights $158, \allowbreak 164, \allowbreak 166, \allowbreak 172$ cm and weights $52, 60, 58, 66$ kg. Weight is predicted from standardized height by ridge with $\lambda=4$.

Find
  1. (a) Standardize the heights and center the weights.

  2. (b) Compute $\hat\beta^R$.

  3. (c) Predict the new adult's weight and compare with least squares.

Given
  • heights $158, \allowbreak 164, \allowbreak 166, \allowbreak 172$ cm; weights $52,60,58,66$ kg

  • $\lambda=4$

  • a new adult of height $175$ cm

Hint 1/4

Put the training heights on the standard scale, fit there, and send the new height through the same conversion.

Hint 2/4

$\tilde x=\frac{x-\bar x}{\sigma}$ with $\sigma=\sqrt{\frac1n\sum(x-\bar x)^2}$; $\hat\beta^R=\frac{\sum\tilde x\,y^c}{n+\lambda}$; $\hat y=\bar y+\hat\beta^R\tilde z$.

Hint 3/4

Heights $158, \allowbreak 164, \allowbreak 166, \allowbreak 172$ (mean $165$), weights $52,60,58,66$ (mean $59$), $\lambda=4$, new height $175$.

Hint 4/4

$\tilde x=(-1.4,-0.2,0.2,1.4)$, $\hat\beta^R=2.4$, and the prediction is $63.8$ kg against $68.6$ kg for least squares.

Show solution

The training mean and spread are computed once and reused for the new adult.

Training statistics

$$\bar x=165,\quad \sigma=\sqrt{\tfrac{49+1+1+49}{4}}=5,\quad \bar y=59$$

The spread divides by $n=4$.

Standardize and center

$$\tilde x=(-1.4,\,-0.2,\,0.2,\,1.4),\qquad y^c=(-7,\,1,\,-1,\,7)$$

Centered heights over $5$; weights minus $59$.

Fit

$$\textstyle\sum\tilde x\,y^c=9.8-0.2-0.2+9.8=19.2,\qquad \hat\beta^R=\frac{19.2}{4+4}=2.4$$

One column: divide by $n+\lambda$.

Predict

$$\tilde z=\frac{175-165}{5}=2,\qquad \hat y=59+2.4\times2=63.8$$

Training statistics for the new adult, then the mean back.

Answer $$\boxed{\hat\beta^R=2.4,\qquad \hat y=63.8\ \text{kg}\ \ (\text{least squares: }68.6\ \text{kg})}$$
Check

Shrink check: $\frac{n}{n+\lambda}=\frac48=\frac12$, and $2.4$ is half of the least squares $\frac{19.2}{4}=4.8$.

Every raw-data ridge question runs the same way: training statistics, standardize, fit, standardize the new input with the same numbers, add the mean.

4§05.6 — one lasso coefficient at four values of λ

A standardized predictor from $n=20$ centered observations has $\sum_ix_iy_i=18$. You track its lasso coefficient as $\lambda$ grows.

Find
  1. (a) Compute $\hat\beta^L$ for each $\lambda$.

  2. (b) Find the smallest $\lambda$ at which $\hat\beta^L=0$.

Given
  • $n=20$, $\sum_ix_i^2=20$, $\sum_ix_iy_i=18$

  • $\lambda\in\{8,\,20,\,36,\,40\}$

Hint 1/4

Compare the size of the least squares coefficient with the cut at each λ.

Hint 2/4

With $b=\hat\beta^{\mathrm{LS}}$: $\hat\beta^L=\operatorname{sign}(b)\big(\lvert b\rvert-\frac{\lambda}{2n}\big)_+$, which is $0$ exactly when $\lambda\ge2\lvert\sum_ix_iy_i\rvert$.

Hint 3/4

$\hat\beta^{\mathrm{LS}}=\frac{18}{20}=0.9$ and the cuts are $\frac{\lambda}{40}=0.2,\ \allowbreak 0.5,\ \allowbreak 0.9,\ \allowbreak 1.0$.

Hint 4/4

$\hat\beta^L=0.7,\ \allowbreak 0.4,\ \allowbreak 0,\ \allowbreak 0$, and the coefficient first reaches $0$ at $\lambda=36$.

Show solution

One size and four cuts: a table of subtractions is quicker than solving four losses.

The cuts

$$\frac{\lambda}{2n}=\frac{\lambda}{40}:\quad 0.2,\ \ 0.5,\ \ 0.9,\ \ 1.0$$

The cut is the only part that changes with λ, so we list it once per value.

Subtract and clip

$$(0.9-0.2,\ 0.9-0.5,\ (0.9-0.9)_+,\ (0.9-1.0)_+)=(0.7,\ 0.4,\ 0,\ 0)$$

The size is positive, so the sign stays positive.

First zero

$$\frac{\lambda}{40}=0.9\ \Rightarrow\ \lambda=36=2\textstyle\sum_ix_iy_i$$

The cut reaches the size exactly there.

Answer $$\boxed{\hat\beta^L=0.7,\ 0.4,\ 0,\ 0;\qquad \lambda_{\text{zero}}=36}$$
Check

By cases at $\lambda=20$: on $\beta>0$ the derivative $-36+40\beta+20$ vanishes at $\beta=0.4$, inside the region.

The lasso coefficient falls in a straight line with λ and then sits at exactly zero.

5§05.6 — ridge and the lasso side by side on uncorrelated predictors

Three standardized, uncorrelated predictors from $n=20$ observations have least squares coefficients $1.5$, $-0.4$ and $0.1$. Both penalties are tried with $\lambda=10$.

Find
  1. (a) Compute $\hat\beta^R$.

  2. (b) Compute $\hat\beta^L$.

  3. (c) Which predictors does each method keep?

Given
  • $X^TX=20I$

  • $\hat\beta^{\mathrm{LS}}=(1.5,\,-0.4,\,0.1)$

  • $\lambda=10$

Hint 1/4

Uncorrelated standardized predictors split both problems into three one-predictor problems.

Hint 2/4

Ridge multiplies by $\frac{n}{n+\lambda}$; the lasso subtracts $\frac{\lambda}{2n}$ from each size and stops at $0$.

Hint 3/4

$n=20$ and $\lambda=10$, so the factor is $\frac{20}{30}$ and the cut is $\frac{10}{40}=0.25$; the coefficients are $1.5$, $-0.4$ and $0.1$.

Hint 4/4

$\hat\beta^R=(1,\,-0.267,\,0.067)$ keeps all three; $\hat\beta^L=(1.25,\,-0.15,\,0)$ keeps the first two.

Show solution

No cross terms, so each coefficient follows the one-predictor rules.

Ridge

$$\tfrac{20}{30}(1.5,\,-0.4,\,0.1)=(1,\,-0.267,\,0.067)$$

One common factor, two thirds.

Lasso

$$\big(1.5-0.25,\ -(0.4-0.25),\ (0.1-0.25)_+\big)=(1.25,\,-0.15,\,0)$$

One common cut, a quarter.

Answer $$\boxed{\hat\beta^R=(1,\,-0.267,\,0.067),\qquad \hat\beta^L=(1.25,\,-0.15,\,0)}$$
Check

Zero check for the lasso: the RSS slope at $\beta_3=0$ has size $2\cdot20\cdot0.1=4\le10$.

Ridge keeps proportions and every predictor; the lasso keeps differences of sizes and drops what falls below the cut.

6§05.5 — a duplicated predictor, ridge and its limit

A data set accidentally contains the same standardized predictor twice, from $n=4$ observations. On centered data each copy gives $\sum_ix_iy_i=6$.

Find
  1. (a) Show that least squares has no unique solution and describe all solutions.

  2. (b) Compute $\hat\beta^R$ at $\lambda=2$.

  3. (c) Find $\lim_{\lambda\to0}\hat\beta^R_\lambda$.

Given
  • $X^TX=\begin{pmatrix}4&4\\4&4\end{pmatrix}$, $X^Ty=(6,\,6)^T$

  • $\lambda=2$

Hint 1/4

Ask first whether least squares can have a unique answer here, then look for the direction that two identical columns single out.

Hint 2/4

Least squares solves $X^TX\beta=X^Ty$, ridge solves $(X^TX+\lambda I)\beta=X^Ty$, and $X^TX(1,1)^T=8(1,1)^T$.

Hint 3/4

$X^TX=\begin{pmatrix}4&4\\4&4\end{pmatrix}$ has determinant $0$, $X^Ty=(6,6)^T$ and $\lambda=2$.

Hint 4/4

Least squares: every $\beta$ with $\beta_1+\beta_2=1.5$; ridge: $\frac{6}{8+\lambda}(1,1)=(0.6,\,0.6)$; limit $(0.75,\,0.75)$.

Show solution

$X^TX$ only stretches the direction $(1,1)$, so ridge along it is a single division.

Least squares

$$\det\begin{pmatrix}4&4\\4&4\end{pmatrix}=0,\qquad 4\beta_1+4\beta_2=6$$

Two identical equations, one line of solutions.

Ridge

$$(X^TX+\lambda I)(1,1)^T=(8+\lambda)(1,1)^T\ \Rightarrow\ \hat\beta^R=\frac{6}{8+\lambda}(1,1)$$

$X^Ty=6(1,1)^T$ lies along the same direction.

$$\lambda=2:\ \ (0.6,\,0.6);\qquad \lambda\to0:\ \ (0.75,\,0.75)$$

The formula is continuous at $\lambda=0$ although $X^TX$ is singular, so the limit is a substitution.

Answer $$\boxed{\beta_1+\beta_2=1.5\ \text{(LS)},\qquad \hat\beta^R_2=(0.6,\,0.6),\qquad \lim_{\lambda\to0}=(0.75,\,0.75)}$$
Check

The inverse gives the same: $X^TX+2I=\begin{pmatrix}6&4\\4&6\end{pmatrix}$ with determinant $20$, and $\frac1{20}(36-24,\ -24+36)=(0.6,\,0.6)$.

With duplicated predictors, ridge splits the weight equally and picks the smallest least squares fit in the limit.

7§05.7 — turning a budget into a λ

A standardized predictor from $n=10$ centered observations has $\sum_ix_iy_i=15$, so $\hat\beta^{\mathrm{LS}}=1.5$. Both penalties are rewritten with a budget.

Find
  1. (a) Solve both constrained problems.

  2. (b) Find the λ that gives each solution in penalized form.

  3. (c) Which budgets would leave least squares untouched?

Given
  • $n=10$, $\sum_ix_i^2=10$, $\sum_ix_iy_i=15$

  • budgets: $\beta^2\le1$ for ridge, $\lvert\beta\rvert\le1$ for the lasso

Hint 1/4

With one coefficient each budget is an interval around 0; find the allowed point nearest the least squares value, then match λ.

Hint 2/4

Ridge: $\hat\beta^R=\frac{\sum x_iy_i}{n+\lambda}$; lasso on its positive branch: $\hat\beta^L=\frac{\sum x_iy_i-\lambda/2}{n}$.

Hint 3/4

$\hat\beta^{\mathrm{LS}}=1.5$, budgets $\beta^2\le1$ and $\lvert\beta\rvert\le1$, $n=10$, $\sum x_iy_i=15$.

Hint 4/4

Both constrained solutions are $\beta=1$, matched by $\lambda=5$ for ridge and $\lambda=10$ for the lasso; budgets $s\ge2.25$ and $s\ge1.5$ leave least squares untouched.

Show solution

In one dimension the constrained solution is the allowed point nearest the least squares value, so no calculus is needed.

Constrained solutions

$$\beta^2\le1\iff\lvert\beta\rvert\le1;\qquad 1.5\notin[-1,1]\ \Rightarrow\ \beta=1$$

The RSS is a parabola around 1.5, smallest at the nearest allowed point.

Matching λ

$$\frac{15}{10+\lambda}=1\Rightarrow\lambda=5;\qquad \frac{15-\lambda/2}{10}=1\Rightarrow\lambda=10$$

A penalized fit matches the constrained one when it lands on the same coefficient, which fixes λ.

Idle budgets

$$s_R\ge1.5^2=2.25,\qquad s_L\ge\lvert1.5\rvert=1.5$$

Then least squares is allowed and nothing is shrunk.

Answer $$\boxed{\beta=1;\quad \lambda_R=5,\ \ \lambda_L=10;\quad s_R\ge2.25,\ \ s_L\ge1.5}$$
Check

Penalized check: $\frac{15}{15}=1$ and $\frac{15-5}{10}=1$.

A budget question in one dimension is a clipping question: clip the least squares value to the interval, then read off λ.

C · exam level 5 questions
1§05.1 — deriving the ridge estimator and its uniqueness

An exam asks for the ridge estimator from first principles. The response is centered and the predictors are standardized, so there is no intercept.

Find
  1. (a) Expand the loss and set its derivative to zero.

  2. (b) Show that $X^TX+\lambda I$ is invertible for every $\lambda>0$.

  3. (c) Explain why the stationary point is the unique minimizer.

Given
  • $\mathrm{Loss}_R(\beta,\lambda)=(y-X\beta)^T(y-X\beta)+\lambda\beta^T\beta$ with $\lambda>0$

  • row derivatives: $\frac{\partial}{\partial\beta}(a^T\beta)=a^T$, and $\frac{\partial}{\partial\beta}(\beta^TA\beta)=2\beta^TA$ for symmetric $A$

Hint 1/4

Three pieces are needed: a stationary-point equation, an invertibility argument, and a reason the stationary point is a minimum.

Hint 2/4

Expand $(y-X\beta)^T(y-X\beta)$, apply the two derivative rules, and for (b) look at $v^T(X^TX+\lambda I)v$.

Hint 3/4

$\mathrm{Loss}_R=y^Ty-2y^TX\beta+\beta^T(X^TX+\lambda I)\beta$, with $\lambda>0$.

Hint 4/4

$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$, unique because $v^T(X^TX+\lambda I)v=\lVert Xv\rVert_2^2+\lambda\lVert v\rVert_2^2>0$ for $v\neq0$.

Show solution

Matrix calculus gives the equation in two lines, and the λ term settles invertibility and uniqueness at once.

Expand

$$\mathrm{Loss}_R=y^Ty-2y^TX\beta+\beta^T(X^TX+\lambda I)\beta$$

$\beta^TX^Ty$ is a number, so it equals its transpose $y^TX\beta$.

Differentiate

$$\frac{\partial\mathrm{Loss}_R}{\partial\beta}=-2y^TX+2\beta^T(X^TX+\lambda I)=0$$

Row derivatives; $X^TX+\lambda I$ is symmetric.

$$(X^TX+\lambda I)\beta=X^Ty$$

The transpose carries the same equations in the usual column form.

Invertibility

$$v^T(X^TX+\lambda I)v=\lVert Xv\rVert_2^2+\lambda\lVert v\rVert_2^2>0\quad(v\neq0)$$

If $Mv=0$ the left side would be $0$, which forces $v=0$.

A unique minimum

$$\frac{\partial^2\mathrm{Loss}_R}{\partial\beta\,\partial\beta^T}=2(X^TX+\lambda I)\ \ \text{positive definite}$$

A quadratic that curves upward in every direction has exactly one stationary point, its minimum.

Answer $$\boxed{\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty,\ \text{unique for every}\ \lambda>0}$$
Check

Special cases: with one predictor this is $\frac{\sum_ix_iy_i}{\sum_ix_i^2+\lambda}$, and $\lambda=0$ gives back the least squares normal equations.

Every ridge derivation has the same three beats: expand, differentiate, and argue invertibility from the λ term.

2§05.6 — the smallest λ that empties the lasso model

Three standardized predictors and a centered response give the products below. You want to know when the lasso keeps no predictor at all, and which one enters first as $\lambda$ is lowered.

Find(a) What is the smallest $\lambda$ at which the lasso sets all three coefficients to $0$, and which predictor enters first as $\lambda$ falls below it?
Given
  • $X^Ty=(9,\,-14,\,5)^T$

  • a lasso coefficient at $0$ stays optimal exactly when $\lvert\partial\mathrm{RSS}/\partial\beta_j\rvert\le\lambda$ there

Hint 1/4

Test the all-zero fit with the zero rule, one coefficient at a time.

Hint 2/4

At $\beta=0$, $\frac{\partial\mathrm{RSS}}{\partial\beta_j}=-2x_j^T(y-X\cdot0)=-2x_j^Ty$, so every zero is kept exactly when $\lambda\ge2\max_j\lvert x_j^Ty\rvert$.

Hint 3/4

$x_j^Ty=9,\ -14,\ 5$, so the three slopes have sizes $18$, $28$ and $10$.

Hint 4/4

All three stay at $0$ exactly when $\lambda\ge28$; below $28$ predictor 2 enters first, with a negative coefficient.

Show solution

At the all-zero fit the residual is the response itself, so every slope is a product the data give directly.

Slopes at zero

$$\frac{\partial\mathrm{RSS}}{\partial\beta_j}\Big|_{\beta=0}=-2x_j^Ty:\qquad -18,\ \ 28,\ \ -10$$

With every coefficient at 0 the residual is the centered response.

All zeros kept

$$\hat\beta^L=0\iff2\lvert x_j^Ty\rvert\le\lambda\ \text{for every }j\iff\lambda\ge28$$

Every coordinate must pass the zero rule at once, and the largest slope is the binding one.

First to enter

$$\lambda<28:\ \ \text{coordinate }2\ \text{fails first};\ \ \operatorname{sign}(\hat\beta_2)=\operatorname{sign}(x_2^Ty)<0$$

Its slope is the largest in size, so its zero breaks first, and the fit moves against the slope.

Answer $$\boxed{\hat\beta^L=0\iff\lambda\ge28;\qquad \text{predictor 2 enters first, negative}}$$
Check

The same rule on the kiosk data: $X^Ty=(11,7)$ gives $2\times11=22$, exactly where the lasso path in the figure reaches $(0,0)$.

The largest correlation with the response in absolute value decides both the empty-model threshold and the first predictor in.

3§05.4 — how much ridge beats least squares at its best λ

A sensor's response is predicted from one standardized input with $n=20$ fixed readings. The true slope is $\beta=0.5$, the noise variance is $\sigma^2=4$, and the intercept is known.

Find
  1. (a) Find $\lambda^*$ and the shrink factor $c$ there.

  2. (b) Compute bias², variance and the expected test error at $\lambda^*$, and for least squares.

  3. (c) What fraction of the error above the noise does ridge remove?

Given
  • $n=20$, $\sum_ix_i^2=20$

  • $\beta=0.5$, $\sigma^2=4$

  • a new input with $E[x^2]=1$

Hint 1/4

Find the best λ first; everything else follows from the shrink factor there.

Hint 2/4

$\lambda^*=\sigma^2/\beta^2$, $c=\frac{n}{n+\lambda}$, bias² $=(1-c)^2\beta^2$, variance $=c^2\sigma^2/n$.

Hint 3/4

$\sigma^2=4$, $\beta=0.5$ and $n=20$, so $\lambda^*=16$ and $c=\frac{20}{36}=\frac59$.

Hint 4/4

Bias² $\frac{4}{81}$ and variance $\frac{5}{81}$ make $\frac19$ above the noise, against $0.2$ for least squares: $\frac49$ of it is removed.

Show solution

The shrink factor at the best λ is a clean fraction, so we work in fractions throughout.

Best λ

$$\lambda^*=\frac{4}{0.25}=16,\qquad c=\frac{20}{20+16}=\frac59$$

The derivative of bias² plus variance changes sign exactly at the noise-to-signal ratio.

Ridge terms

$$\text{bias}^2=\big(\tfrac49\big)^2(0.5)^2=\tfrac{4}{81},\qquad \text{variance}=\big(\tfrac59\big)^2\tfrac{4}{20}=\tfrac{5}{81}$$

$1-c=\frac49$ of the slope is lost; the spread is cut to $\frac59$.

$$\tfrac{4}{81}+\tfrac{5}{81}=\tfrac19,\qquad \text{test error}=4+\tfrac19\approx4.111$$

Every predictor pays the noise, and it does not depend on λ.

Least squares and the saving

$$0+\tfrac{4}{20}=0.2;\qquad \frac{0.2-1/9}{0.2}=\frac49$$

Least squares pays only variance, $\sigma^2/n$.

Answer $$\boxed{\lambda^*=16;\quad \tfrac19\ \text{vs}\ 0.2\ \text{above the noise};\quad \tfrac49\ \text{of it removed}}$$
Check

The closed form for the minimum agrees: $\frac{\sigma^2\beta^2}{n\beta^2+\sigma^2}=\frac{1}{5+4}=\frac19$.

The weaker the signal relative to the noise, the larger the best λ and the more ridge saves.

4§05.7 — the lasso solution when the contours are circles

Two standardized, uncorrelated predictors make the RSS contours circles around $\hat\beta^{\mathrm{LS}}=(2,\,0.5)$. The lasso is solved in budget form with $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le1.2$.

Find(a) Where is the constrained lasso solution?
Given
  • $X^TX=nI$, so the RSS contours are circles around $(2,\,0.5)$

  • budget: $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le1.2$

Hint 1/4

With circular contours, the constrained solution is the point of the region closest to the least squares point.

Hint 2/4

Try the edge $\beta_1+\beta_2=1.2$ with $\beta_1,\beta_2\ge0$ first: project onto it by subtracting the same amount from both coordinates.

Hint 3/4

The least squares point is $(2,\,0.5)$ with $2+0.5=2.5$, so the projection subtracts $\frac{2.5-1.2}{2}=0.65$ from each coordinate.

Hint 4/4

The projection $(1.35,\,-0.15)$ leaves the edge, so the closest point is the corner $(1.2,\,0)$.

Show solution

With circular contours, closest in distance means smallest RSS, so the question is a projection.

Project onto the edge

$$(2,\,0.5)-0.65\,(1,1)=(1.35,\,-0.15)$$

Move along the edge's normal $(1,1)$ until the coordinates add up to $1.2$.

Check the edge

$$\beta_2=-0.15<0\ \Rightarrow\ \text{off this edge}$$

This edge covers only $\beta_1,\beta_2\ge0$.

Take the corner

$$(1.2,\,0):\qquad \text{distance}^2=0.8^2+0.5^2=0.89$$

The neighbouring edge's own projection, $(1.85,\,0.65)$, is off that edge too, so both edges end at this corner.

Answer $$\boxed{(1.2,\ 0)}$$
Check

Soft thresholding agrees: circular contours mean subtracting the same $t$ from each size, and $t=0.8$ gives $\big(1.2,\,(0.5-0.8)_+\big)=(1.2,\,0)$, spending exactly the budget $1.2$.

When the least squares point lies far out along one axis, the diamond's corner on that axis catches the solution and the other coefficient is exactly zero.

5§05.6 — find the step that breaks a lasso computation

An exam question gives two standardized, uncorrelated predictors from $n=5$ observations, with $X^TX=5I$ and $X^Ty=(4,\,-9)^T$, and asks for the lasso at $\lambda=6$. A student writes the four steps below; one of them is wrong.

Find
  1. (a) Which step is wrong?

  2. (b) What is the correct lasso estimate?

Given
  • $n=5$, $X^TX=5I$, $X^Ty=(4,\,-9)^T$, $\lambda=6$

  • Step 1: $\hat\beta^{\mathrm{LS}}=\frac15(4,\,-9)=(0.8,\,-1.8)$

  • Step 2: the cut is $\frac{\lambda}{n}=\frac65=1.2$

  • Step 3: $\hat\beta^L_1=(0.8-1.2)_+=0$

  • Step 4: $\hat\beta^L_2=-(1.8-1.2)=-0.6$, so $\hat\beta^L=(0,\,-0.6)$

Hint 1/4

Check each step against the rule it uses, not against the final answer.

Hint 2/4

With $X^TX=nI$: $\hat\beta^{\mathrm{LS}}=X^Ty/n$, and the lasso subtracts $\frac{\lambda}{2n}$ from each size and stops at $0$.

Hint 3/4

Here $n=5$, $\lambda=6$ and $X^Ty=(4,-9)^T$; the student's cut is $\frac65=1.2$.

Hint 4/4

Step 2 is wrong: the cut is $\frac{6}{10}=0.6$, and the lasso gives $(0.2,\,-1.2)$.

Show solution

Each step applies one rule, so we test the rules in order instead of redoing the whole problem.

Step 1

$$\hat\beta^{\mathrm{LS}}=\tfrac15(4,\,-9)=(0.8,\,-1.8)$$

Correct: with $X^TX=5I$ least squares divides $X^Ty$ by $5$.

Step 2

$$\frac{\lambda}{2n}=\frac{6}{10}=0.6\neq1.2$$

Wrong: the lecture's loss has no $\frac12$ in front of the RSS, so the cut carries the factor 2.

Redo steps 3 and 4

$$\hat\beta^L=\big((0.8-0.6)_+,\ -(1.8-0.6)\big)=(0.2,\,-1.2)$$

Both sizes exceed the cut, so neither coefficient reaches 0.

Answer $$\boxed{\text{Step 2};\qquad \hat\beta^L=(0.2,\ -1.2)}$$
Check

By cases for the first coefficient: on $\beta_1>0$ the derivative $-2\cdot4+2\cdot5\,\beta_1+6$ vanishes at $\beta_1=0.2$, inside the region, so the first predictor stays in.

Most lasso slips happen in the cut; write down the cut before touching the coefficients.

D · interleaved 4 questions
1§05.1 — least squares or ridge? read it off the residuals

A report gives a fitted coefficient vector for the kiosk data but not the method. For least squares the residual vector is orthogonal to every column of $X$; you check what these residuals do.

Find
  1. (a) Compute $X^T(y-X\hat\beta)$.

  2. (b) Is the fit least squares? If it is ridge, find $\lambda$.

Given
  • $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$, $X^Ty=(11,\,7)^T$ (standardized predictors, centered sales)

  • reported fit: $\hat\beta=(0.7,\,0.3)$

Hint 1/4

Least squares and ridge satisfy different equations for the same residual; compute the residual's products with the columns.

Hint 2/4

Least squares: $X^T(y-X\hat\beta)=0$. Ridge: $(X^TX+\lambda I)\hat\beta=X^Ty$, that is $X^T(y-X\hat\beta)=\lambda\hat\beta$.

Hint 3/4

$X^Ty=(11,7)^T$ and $X^TX\hat\beta=(5\cdot0.7+4\cdot0.3,\ 4\cdot0.7+5\cdot0.3)$.

Hint 4/4

$X^T(y-X\hat\beta)=(6.3,\,2.7)=9\,(0.7,\,0.3)$: not least squares, but ridge with $\lambda=9$.

Show solution

The two methods differ in one equation for the residual, so computing it decides between them without refitting.

Residual products

$$X^T(y-X\hat\beta)=X^Ty-X^TX\hat\beta=(11,\,7)-(4.7,\,4.3)=(6.3,\,2.7)$$

Row by row: $5(0.7)+4(0.3)=4.7$ and $4(0.7)+5(0.3)=4.3$.

Match an equation

$$(6.3,\,2.7)=9\,(0.7,\,0.3)\ \Rightarrow\ \lambda=9$$

Ridge's equation $X^T(y-X\hat\beta)=\lambda\hat\beta$ holds with one common λ; least squares' would need $(0,0)$.

Answer $$\boxed{\text{ridge with}\ \lambda=9}$$
Check

Refit check: $\frac1{180}\begin{pmatrix}14&-4\\-4&14\end{pmatrix}\begin{pmatrix}11\\7\end{pmatrix}=(0.7,\,0.3)$, the ridge fit at $\lambda=9$.

Ridge residuals are not orthogonal to the columns; they lean toward the fit by exactly λ times the coefficients.

2§05.2 — a validation split between two values of λ

Five training points on one standardized predictor give the products below on centered data. Three held-out points, already on the training scale, decide between least squares and ridge with $\lambda=5$.

Find
  1. (a) Fit both models on the training points.

  2. (b) Compute each validation MSE and choose.

Given
  • training: $n=5$, $\sum_ix_i^2=5$, $\sum_ix_iy_i=10$

  • validation points $(\tilde x,\,y^c)$: $(1,\,1.2)$, $(-0.5,\,-0.2)$, $(2,\,1.8)$

Hint 1/4

Fit on the training part only, then score each fit on points it never saw.

Hint 2/4

$\hat\beta=\frac{\sum x_iy_i}{\sum x_i^2+\lambda}$, and $\mathrm{MSE}_{\text{val}}=\frac13\sum(y^c-\hat\beta\tilde x)^2$ over the held-out points.

Hint 3/4

Training: $\sum x_i^2=5$, $\sum x_iy_i=10$; held-out points $(1,1.2)$, $(-0.5,-0.2)$, $(2,1.8)$.

Hint 4/4

$\hat\beta^{\mathrm{LS}}=2$ scores $2.04$ and $\hat\beta^R=1$ scores $0.057$, so ridge with $\lambda=5$ wins.

Show solution

The validation set approach needs one fit per candidate and one score each, the cheapest honest comparison.

Fit

$$\hat\beta^{\mathrm{LS}}=\frac{10}{5}=2,\qquad \hat\beta^R=\frac{10}{5+5}=1$$

One predictor, so each fit is a division.

Score least squares

$$\tfrac13\big[(1.2-2)^2+(-0.2+1)^2+(1.8-4)^2\big]=\tfrac{0.64+0.64+4.84}{3}=2.04$$

Predictions $2$, $-1$ and $4$.

Score ridge

$$\tfrac13\big[(1.2-1)^2+(-0.2+0.5)^2+(1.8-2)^2\big]=\tfrac{0.04+0.09+0.04}{3}\approx0.057$$

Predictions $1$, $-0.5$ and $2$.

Answer $$\boxed{\mathrm{MSE}_{\text{val}}:\ 2.04\ \text{(LS)}\ \ \text{vs}\ \ 0.057\ \text{(ridge)};\ \ \text{choose}\ \lambda=5}$$
Check

Residual check: ridge's three misses $0.2,\ 0.3,\ -0.2$ are all smaller in size than least squares' $-0.8,\ 0.8,\ -2.2$.

Held-out error chooses λ; the winner is then refitted on all the data.

3§05.4 — a 95% interval built around a ridge estimate

Noise is normal with $\sigma=2$, one standardized predictor has $n=16$ fixed inputs, and the true slope is $\beta=1$. A student reports $\hat\beta^R\pm2\,\mathrm{sd}(\hat\beta^R)$ at $\lambda=16$ as a 95% interval for $\beta$.

Find
  1. (a) Find the mean and the standard deviation of $\hat\beta^R$.

  2. (b) How often does the student's interval contain $\beta=1$?

Given
  • $n=16$, $\sum_ix_i^2=16$, $\sigma=2$, $\beta=1$

  • $\lambda=16$

  • standard normal $Z$: $P(Z\le0)=0.5$ and $P(Z\le4)\approx1.000$

Hint 1/4

An interval of the form estimate ± 2 sd covers the truth 95 times in 100 only when the estimate is centered on the truth; check where this one is centered.

Hint 2/4

$\hat\beta^R=c\,\hat\beta^{\mathrm{LS}}$ with $c=\frac{n}{n+\lambda}$, so $E[\hat\beta^R]=c\beta$ and $\mathrm{sd}(\hat\beta^R)=c\,\sigma/\sqrt n$.

Hint 3/4

$c=\frac{16}{32}=\frac12$, $\beta=1$, and $\sigma/\sqrt n=\frac24=0.5$.

Hint 4/4

$\hat\beta^R$ is normal with mean $0.5$ and sd $0.25$; the interval $\hat\beta^R\pm0.5$ contains $1$ only when $\hat\beta^R\ge0.5$, about $50$ times in $100$.

Show solution

Coverage is a probability about the estimate, so we find its distribution first and then the event.

Distribution of the estimate

$$\hat\beta^{\mathrm{LS}}\sim N\big(1,\,0.5^2\big)\ \Rightarrow\ \hat\beta^R=\tfrac12\hat\beta^{\mathrm{LS}}\sim N\big(0.5,\,0.25^2\big)$$

Normal noise makes the least squares slope normal, and halving it halves mean and sd.

The covering event

$$\lvert\hat\beta^R-1\rvert\le0.5\iff0.5\le\hat\beta^R\le1.5\iff0\le Z\le4$$

Standardize with mean $0.5$ and sd $0.25$.

Its probability

$$P(0\le Z\le4)\approx1.000-0.5=0.5$$

$P(0\le Z\le4)=P(Z\le4)-P(Z\le0)$, and the question gives both values.

Answer $$\boxed{E=0.5,\ \ \mathrm{sd}=0.25;\qquad \text{coverage}\approx50\ \text{in}\ 100}$$
Check

Least squares at the same settings: $\hat\beta^{\mathrm{LS}}\pm2(0.5)$ is centered on $1$ in expectation and covers it about $95$ times in $100$.

A biased estimate with a small spread makes a confidently wrong interval; the ±2 sd recipe needs an unbiased center.

4§05.4 — does ridge contradict Gauss-Markov?

Gauss-Markov says least squares has the smallest variance among linear unbiased estimators. Ridge is also linear in $y$, since $\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$, and its variance is smaller.

Find(a) How do the two facts fit together?
Given
  • Gauss-Markov: among linear unbiased estimators, least squares has the smallest variance

  • one standardized predictor: $\mathrm{Var}(\hat\beta^R)=c^2\sigma^2/n$ with $c=\frac{n}{n+\lambda}<1$

Hint 1/4

Check each condition of the theorem against ridge: linear, and unbiased.

Hint 2/4

An estimator $Ay$ is linear; it is unbiased when $E[Ay]=\beta$ for every $\beta$.

Hint 3/4

For one standardized predictor $E[\hat\beta^R]=\frac{n}{n+\lambda}\beta$, which differs from $\beta$ whenever $\lambda>0$ and $\beta\neq0$.

Hint 4/4

Ridge is linear but biased, so Gauss-Markov says nothing about it, and its smaller variance contradicts nothing.

Show solution

A theorem about a class says nothing about estimators outside it, so we test membership.

Linear?

$$\hat\beta^R=Ay,\qquad A=(X^TX+\lambda I)^{-1}X^T\ \text{fixed}$$

Yes: a fixed matrix times $y$.

Unbiased?

$$E[\hat\beta^R]=\tfrac{n}{n+\lambda}\,\beta\neq\beta\quad(\lambda>0,\ \beta\neq0)$$

No: the shrink factor pulls the mean toward 0.

Answer $$\boxed{\text{ridge is biased, so Gauss-Markov does not cover it}}$$
Check

Numbers from the tradeoff block: at $\lambda=4$ with $n=8$, $\sigma^2=4$, $\beta=1$, ridge has variance $\frac29<0.5$ and bias² $\frac19>0$, and its total $\frac13$ beats least squares' $0.5$.

Gauss-Markov ranks unbiased estimators only; shrinkage wins by leaving that class on purpose.

Mistake ledger (19 entries)
⚠ Adding λ to every entry, not only the diagonal

The phrase 'add λ' is remembered without the identity matrix that says where.

wrong$$X^TX+\lambda\mathbf 1\mathbf 1^T=\begin{pmatrix}5+\lambda&4+\lambda\\4+\lambda&5+\lambda\end{pmatrix}$$
right$$X^TX+\lambda I=\begin{pmatrix}5+\lambda&4\\4&5+\lambda\end{pmatrix}$$
⚠ Penalizing the intercept

The penalty sum is written from j = 0 out of habit.

wrong$$\lambda\sum_{j=0}^p\beta_j^2$$
right$$\lambda\sum_{j=1}^p\beta_j^2\qquad(\hat\beta_0=\bar y\ \text{is left alone})$$
⚠ Inverting first and adding λ afterwards

Both formulas contain the same three pieces, and the order of inverting and adding is easy to swap.

wrong$$\hat\beta^R=\big((X^TX)^{-1}+\lambda I\big)X^Ty$$
right$$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$$
⚠ Choosing λ by the training RSS

It is the error at hand, and for most fitting questions it is the right score.

wrong$$\hat\lambda=\arg\min_\lambda\mathrm{RSS}(\hat\beta^R_\lambda)=0$$
right$$\hat\lambda=\arg\min_\lambda\mathrm{CV}(k)\ \ \text{over a grid}$$
⚠ Thinking a huge λ predicts 0

On centered data the fitted values do go to 0, and the step back to raw units is forgotten.

wrong$$\lambda\to\infty:\ \hat y\to0$$
right$$\lambda\to\infty:\ \hat y\to\bar y\quad(\text{the intercept is not penalized})$$
⚠ Keeping a fold's fit instead of refitting

The fold fits are already computed, and refitting looks like extra work.

wrong$$\hat\beta=\text{one of the }k\text{ fold fits at }\hat\lambda$$
right$$\hat\beta=\hat\beta^R_{\hat\lambda}\ \text{refitted on all }n\text{ points}$$
⚠ Standardizing a new input with its own statistics

Each batch of data seems to deserve its own mean and spread.

wrong$$\tilde z=\frac{z-\bar z_{\text{test}}}{\sigma_{\text{test}}}$$
right$$\tilde z=\frac{z-\bar x_{\text{train}}}{\sigma_{\text{train}}}$$
⚠ Forgetting to add the mean response back

The fit was made on centered responses, so its raw output already looks like a prediction.

wrong$$\hat y=\tilde z^T\hat\beta^R$$
right$$\hat y=\bar y+\tilde z^T\hat\beta^R$$
⚠ Fitting ridge on raw units

The closed form works on any centered data, so standardizing looks optional.

wrong$$\text{income in TL},\ \lambda=100:\ \ \hat\beta^R\approx\hat\beta^{\mathrm{LS}}\ \ \text{(almost no shrinkage)}$$
right$$\text{standardize first: the same }\lambda\text{ shrinks the same in every unit}$$
⚠ Calling the unbiased estimator the more accurate one

The word unbiased sounds like a guarantee of accuracy.

wrong$$\text{bias}=0\ \Rightarrow\ \text{smallest test error}$$
right$$\text{test error}=\text{bias}^2+\text{variance}+\text{noise}$$
⚠ Turning the noise-to-signal ratio upside down

Both ratios contain the same two numbers.

wrong$$\lambda^*=\beta^2/\sigma^2$$
right$$\lambda^*=\sigma^2/\beta^2\quad(\text{more noise, more shrinkage})$$
⚠ Believing ridge fails wherever least squares fails

The least squares formula needs that inverse, and ridge looks like the same formula.

wrong$$\det(X^TX)=0\ \Rightarrow\ \text{no ridge fit}$$
right$$\det(X^TX+\lambda I)>0\ \ \text{for every}\ \lambda>0$$
⚠ Reading a small ridge coefficient as a dropped predictor

Small and zero look alike in a printed table.

wrong$$\hat\beta^R_j=0.02\ \Rightarrow\ x_j\ \text{is out of the model}$$
right$$x_j\ \text{is out only if}\ \hat\beta_j=0\ \text{exactly}$$
⚠ Dropping the factor 2 in the cut

Books that put 1/2 in front of the RSS write the threshold without it.

wrong$$\hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{n}\big)_+$$
right$$\hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$
⚠ Losing the sign

The formula works on sizes, and the sign has to be put back by hand.

wrong$$\hat\beta^L=\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$
right$$\hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$
⚠ Giving the lasso ridge's closed form

Ridge has a formula, and the two losses differ by one exponent.

wrong$$\hat\beta^L=(X^TX+\lambda I)^{-1}X^Ty$$
right$$\hat\beta^L:\ \text{no closed form in general; solve by cases or numerically}$$
⚠ Reading the ridge budget as a radius

The budget is written where a radius usually goes.

wrong$$\beta_1^2+\beta_2^2\le s\ \Rightarrow\ \text{disk of radius }s$$
right$$\beta_1^2+\beta_2^2\le s\ \Rightarrow\ \text{disk of radius }\sqrt s$$
⚠ Expecting shrinkage from a budget least squares already meets

A constraint sounds as if it always bites.

wrong$$s_L=5\ \text{for}\ \hat\beta^{\mathrm{LS}}=(3,-1):\ \ \text{a shrunk fit}$$
right$$s_L\ge\lvert3\rvert+\lvert-1\rvert=4:\ \ \hat\beta=\hat\beta^{\mathrm{LS}}$$
⚠ Carrying λ over between ridge and the lasso

Both methods call their tuning constant λ.

wrong$$\text{same }\lambda\ \Rightarrow\ \text{same amount of shrinkage}$$
right$$\beta=0.8\ \text{from}\ \hat\beta^{\mathrm{LS}}=1.2:\ \ \lambda=5\ \text{(ridge)},\ \ \lambda=8\ \text{(lasso)}$$
Formula card
Ridge loss and estimate
$$\mathrm{Loss}_R=\mathrm{RSS}(\beta)+\lambda\sum_{j=1}^p\beta_j^2,\qquad \hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$$

centered response, standardized predictors; any $\lambda>0$

The two ends of λ, and its choice
$$\lambda\to0:\ \hat\beta^R\to\hat\beta^{\mathrm{LS}};\qquad \lambda\to\infty:\ \hat\beta^R\to0,\ \ \hat y\to\bar y;\qquad \hat\lambda=\arg\min\mathrm{CV}(k)$$

$X^TX$ invertible for the first limit

Standardize and predict
$$\tilde x_{ij}=\frac{x_{ij}-\bar x_j}{\sigma_j},\quad \sigma_j=\sqrt{\tfrac1n\textstyle\sum_i(x_{ij}-\bar x_j)^2},\quad \hat y=\bar y+\tilde z^T\hat\beta^R$$

training means and spreads, also for new inputs

One predictor: bias, variance, best λ
$$c=\tfrac{n}{n+\lambda}:\quad \text{bias}^2=(1-c)^2\beta^2,\quad \text{variance}=c^2\tfrac{\sigma^2}{n},\quad \lambda^*=\tfrac{\sigma^2}{\beta^2}$$

one standardized predictor, fixed inputs, no intercept to estimate

Invertibility
$$v^T(X^TX+\lambda I)v=\lVert Xv\rVert_2^2+\lambda\lVert v\rVert_2^2>0\quad(v\neq0,\ \lambda>0)$$

any $X$, even with $p\ge n$ or duplicated columns

Lasso loss and the one-predictor rule
$$\mathrm{Loss}_L=\mathrm{RSS}+\lambda\sum_j\lvert\beta_j\rvert;\qquad p=1:\ \hat\beta^L=\operatorname{sign}(\hat\beta^{\mathrm{LS}})\big(\lvert\hat\beta^{\mathrm{LS}}\rvert-\tfrac{\lambda}{2n}\big)_+$$

the rule needs $\sum_ix_i^2=n$; coefficient by coefficient when $X^TX=nI$

Zero test and the empty model
$$\hat\beta_j=0\ \text{kept}\iff\Big\lvert\frac{\partial\mathrm{RSS}}{\partial\beta_j}\Big\rvert\le\lambda;\qquad \hat\beta^L=0\iff\lambda\ge2\max_j\lvert x_j^Ty\rvert$$

lasso; the slope is taken at the candidate fit

Constrained forms
$$\min\,\mathrm{RSS}\ \ \text{s.t.}\ \sum_j\beta_j^2\le s\ \ \text{or}\ \sum_j\lvert\beta_j\rvert\le s;\qquad s=\text{penalty of }\hat\beta_\lambda$$

a budget at or above the least squares value of the penalty means $\lambda=0$

Check yourself

Close the page and write down from memory:

  • the ridge and lasso losses, and which coefficient neither of them penalizes;
  • the ridge estimate, where $\lambda$ enters it, and why that matrix is always invertible;
  • the standardization formula and how a new raw input becomes a prediction;
  • the one-predictor rules: ridge multiplies by $\frac{n}{n+\lambda}$, the lasso subtracts $\frac{\lambda}{2n}$;
  • why the diamond gives zeros and the disk does not.

Then reopen the page and compare; whatever is missing is your reread list.

  • Compute $\hat\beta^R$ for the kiosk data at $\lambda=3$ by hand and derive the formula from the gradient?

    c-ridge

  • Say what $\lambda\to0$ and $\lambda\to\infty$ do, and pick $\lambda$ from a table of fold errors?

    c-lambda

  • Show with numbers that recording income in thousands changes a raw ridge fit, and predict at a new raw input?

    c-scale

  • Compute bias² and variance of a one-predictor ridge fit and find $\lambda^*=\sigma^2/\beta^2$?

    c-tradeoff

  • Prove that $X^TX+\lambda I$ is invertible and find the ridge fit for two identical predictors?

    c-invertible

  • Soft-threshold a coefficient, and check a guessed lasso zero with the slope rule?

    c-lasso

  • Match a budget $s$ to a $\lambda$ and explain the corner in the diamond picture?

    c-geometry

Glossary (15 terms)
düzenlileştirme

Adding a penalty on the size of the coefficients to the fitting criterion, to lower the variance of the fit at the cost of some bias.

ridge regression

Least squares with the added penalty $\lambda\sum_j\beta_j^2$; on centered, standardized data its estimate is $(X^TX+\lambda I)^{-1}X^Ty$.

lasso

Least squares with the added penalty $\lambda\sum_j\lvert\beta_j\rvert$; it shrinks coefficients and sets some of them exactly to zero.

tuning parameterayar parametresi

A constant fixed before fitting that sets how strongly a method regularizes; here λ, chosen by cross-validation.

penaltyceza terimi

The term added to the RSS that grows with the size of the coefficients.

standardizationstandartlaştırma

Subtracting a predictor's mean and dividing by its standard deviation, so that it has mean 0 and variance 1.

scale invarianceölçek değişmezliği

The property that a method's predictions do not change when a predictor is recorded in different units.

eşdoğrusallık

Near linear dependence among the predictors, which leaves some combinations of coefficients poorly determined by the data.

positive definitepozitif tanımlı

A symmetric matrix $M$ with $v^TMv>0$ for every nonzero $v$; such a matrix is invertible.

identity matrixbirim matris

The square matrix with ones on the diagonal and zeros elsewhere; $\lambda I$ adds $\lambda$ to each diagonal entry.

coefficient path

The estimated coefficients plotted against the tuning parameter λ.

variable selectiondeğişken seçimi

Deciding which predictors enter a model; the lasso does it by setting coefficients exactly to zero.

sparse modelseyrek model

A model in which many coefficients are exactly zero.

soft thresholdingyumuşak eşikleme

Moving a number toward zero by a fixed amount and stopping at zero; the lasso's rule for one standardized predictor.

kısıtlı minimizasyon

Minimizing a function over the points that satisfy a condition, here a budget on the size of the coefficients.

What comes next
§06 · Linear regression from a Bayesian perspective

Next, the regression coefficients get a probability distribution of their own: a prior before the data and a posterior after, the way the Bayesian section treated an unknown rate.

Sources
  • textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The weekly line names no chapter of the book, so no section numbers are cited here.
  • course materialEEE 485 lecture slides and lecture notes, chapter 5: Regularized regression Scope, order and notation (the two losses, the tuning parameter, the ridge and lasso estimates, standardization with 1/n, the constrained forms) follow these materials; every data set, figure and exercise here is original.
  • course materialEEE 485 syllabus page on STARS, Fall 2026-27, printed 21 September 2026 Assessment weights, the course learning outcomes, the weekly list and the recommended books.
  • textbookG. James, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning, Springer, 2013 Recommended in the syllabus; the lecture's coefficient-path, bias-variance and diamond-and-disk figures come from it. None of them is reproduced here.

Spotted something missing or wrong? tell us · share your own notes or an old exam.

Last updated .