On the kiosk data of this section the RSS rises from $20$ to $25.625$ at $\lambda=3$, while $\sum_j\beta_j^2$ falls from $10$ to $1.625$.
A penalty always costs training fit; whether it pays off is decided on data the fit has not seen.
Two thermometers hang side by side outside a kiosk. On five days they read $18, 20, 21, 22, 24$ and $20, 18, 21, 22, 24$ degrees, and a least squares fit of daily sales on both readings gives the second thermometer a negative coefficient: by the fit, its warmth costs sales. Had the two thermometers swapped their first two readings, the fit would have blamed the first thermometer instead.
By the end you can compute, by hand, a fit that gives both thermometers a small positive weight, choose from the data how strongly to shrink, and say why a second kind of penalty keeps only one thermometer.
In 60 seconds
Ridge and the lasso are least squares plus a price $\lambda$ on the size of the coefficients: ridge's $\lambda\sum_j\beta_j^2$ shrinks every coefficient smoothly, with the closed form $(X^TX+\lambda I)^{-1}X^Ty$, and the lasso's $\lambda\sum_j\lvert\beta_j\rvert$ sets some coefficients exactly to zero, with $\lambda$ chosen by cross-validation.
Ridge estimate
$$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$$
centered response, standardized predictors and any $\lambda>0$, even when $X^TX$ is singular
geometry questions; each budget $s$ matches one $\lambda$
Three most common mistakes
Fitting ridge or the lasso on raw predictors: the penalty then depends on the units, so recording income in thousands of TL instead of TL changes the predictions. Standardize first, with the training means and standard deviations.
Choosing $\lambda$ by the training RSS: it never falls as $\lambda$ grows, so it always picks $\lambda=0$, plain least squares. Use cross-validation.
Expecting zeros from ridge: $(X^TX+\lambda I)^{-1}X^Ty$ shrinks the coefficients but, apart from coincidences, leaves every one of them nonzero. Only the lasso holds coefficients at exactly zero.
Two course documents give different weights:
Chapter 1 slides, undergraduate line: midterm 25, final 25, four quizzes 20, two-phase project 30.
STARS syllabus page printed on 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. It is the later document; confirm which split applies.
The syllabus assesses the outcome 'apply probability and linear algebra knowledge to analyze the performance of statistical learning algorithms' in the midterm, the final, the problem sets and quizzes, and the project.
How much time do you have?
10 minutes
The ridge formula with the hand computation that settles the opening puzzle, and the lasso's one-predictor rule with the reason it produces zeros.
The 60-second card · Ridge regression · The lasso · Formula card
45 minutes
Every block once with its first worked example and checkpoint, then one ladder from a full ridge computation down to a bare problem.
The 60-second card · Ridge regression · The tuning parameter λ · Scale matters for ridge · Why shrinking helps · Ridge always has one answer · The lasso · Budgets and shapes · Scaffolding comes off · Formula card
full read
The derivations, the look-alike pairs and enough mixed practice to decide on your own which tool a question needs.
The opening pages · Recall first · Ridge regression · The tuning parameter λ · Scale matters for ridge · Why shrinking helps · Ridge always has one answer · The lasso · Budgets and shapes · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
Compute the ridge estimate $(X^TX+\lambda I)^{-1}X^Ty$ for two predictors by hand, and derive it by setting the gradient of the ridge loss to zero.
Describe how the ridge coefficients move as $\lambda$ runs from $0$ to $\infty$, and choose $\lambda$ by cross-validation.
Standardize the predictors, show why ridge needs this while least squares does not, and turn a new raw input into a prediction.
Quantify the bias and the variance of a ridge coefficient, and find the $\lambda$ with the smallest expected test error for one predictor.
Prove that $X^TX+\lambda I$ is invertible for every $\lambda>0$, and use ridge when least squares has no unique solution.
Compute lasso estimates for one predictor and for uncorrelated standardized predictors, and explain why the lasso selects variables and ridge does not.
Rewrite ridge and the lasso as constrained problems, match a budget $s$ to a $\lambda$, and explain the lasso's zeros with the diamond picture.
Syllabus coverage
covered
Regularized regression — covered
Why shrink at all: least squares has low bias but can have high variance when predictors are nearly collinear or $p$ is close to $n$, and a penalty on the size of the coefficients trades a little bias for less variance.
The lecturer's chapter title, 'Regularized regression: ridge and lasso'. The motivation opens the first block and the tradeoff is worked out in the bias-variance block.
covered
Ridge regression — covered
The ridge loss and its closed form
the choice of $\lambda$ by cross-validation
bias and variance against $\lambda$
invertibility of $X^TX+\lambda I$
ridge keeps every predictor
Official weekly line. Ridge fills the first five blocks, in the lecture's order.
covered
lasso — covered
The lasso loss, its limits as $\lambda\to0$ and $\lambda\to\infty$, and sparse models, the constrained forms and the diamond picture.
The lasso has no closed form in general; the one-predictor formula is the exception.
deferred
Bayesian linear regression — deferred
Regression with a prior distribution on the coefficients.
The official line groups it with ridge and the lasso; it is the fourth line of the STARS weekly list. The lecturer teaches it as the next chapter, 'Linear regression from a Bayesian perspective', and the next section covers it. Check the course schedule for the week.
off syllabus
The one-predictor lasso formula — off syllabus
: with $b=\hat\beta^{\mathrm{LS}}$, $\hat\beta^L=\operatorname{sign}(b)(\lvert b\rvert-\lambda/2n)_+$ for one standardized predictor, and coefficient by coefficient for uncorrelated ones.
Further reading. The slides say the lasso has no closed form in general; in one dimension it has one, and the formula makes the zeros visible.
off syllabus
Checking a zero coefficient — off syllabus
A lasso coefficient belongs at $0$ exactly when $\lvert\partial\mathrm{RSS}/\partial\beta_j\rvert\le\lambda$ there; all coefficients are $0$ once $\lambda\ge2\max_j\lvert x_j^Ty\rvert$.
Further reading: an elementary check used in two worked examples, the exam example and one practice question.
off syllabus
The best $\lambda$ for one predictor — off syllabus
$\lambda^*=\sigma^2/\beta^2$ from the exact bias and variance of a ridge coefficient.
Further reading: the lecture shows the tradeoff as curves; this block derives them for one predictor so that every number can be checked.
off syllabus
Ridge as $\lambda\to0$ with a singular $X^TX$ — off syllabus
The limit is the least squares fit with the smallest sum of squared coefficients.
Further reading, used in one figure, one worked example and one practice question.
Recall first
Least squares in matrix form
$\mathrm{RSS}(\beta)=(y-X\beta)^T(y-X\beta)$, and the normal equations $X^TX\hat\beta=X^Ty$ give $\hat\beta=(X^TX)^{-1}X^Ty$ when the columns of $X$ are linearly independent.
Ridge changes one term of these equations, and least squares is the case $\lambda=0$ throughout.
Matrix derivatives, row convention
$\frac{\partial}{\partial\beta}(a^T\beta)=a^T$ and, for symmetric $A$, $\frac{\partial}{\partial\beta}(\beta^TA\beta)=2\beta^TA$. Setting a row derivative to zero and transposing gives the same equations as a column gradient.
The ridge formula comes from one such derivative.
The 2 × 2 inverse and invertibility
$\begin{pmatrix}a&b\\c&d\end{pmatrix}^{-1}=\frac{1}{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}$ when $ad-bc\neq0$. A square matrix $M$ is invertible exactly when $Mv=0$ has only the solution $v=0$.
Every hand computation here has two predictors, and ridge's main advantage is an invertibility argument.
Directions a matrix only stretches
If $Av=d\,v$ for some $v\neq0$, then $(A+\lambda I)v=(d+\lambda)v$ and $(A+\lambda I)^{-1}v=\frac{v}{d+\lambda}$. For $A=\begin{pmatrix}a&c\\c&a\end{pmatrix}$ the directions $(1,1)$ and $(1,-1)$ are stretched by $a+c$ and $a-c$.
It shows why ridge shrinks some combinations of the coefficients much harder than others.
Bias-variance decomposition
Expected test error $=\text{bias}^2+\text{variance}+\text{noise}$: bias² compares the average model with the expected label, and the variance measures how far one fit strays from the average model.
Ridge buys a little bias for a large cut in variance.
Cross-validation
$\mathrm{CV}(k)=\frac1k\sum_{i=1}^k\mathrm{MSE}_i$, where fold $i$ is scored by a model fitted without it; with $k=n$ this is leave-one-out, $\mathrm{CV}(n)$.
$\lambda$ is chosen this way.
Scaled estimates
$E[cZ]=c\,E[Z]$ and $\mathrm{Var}(cZ)=c^2\mathrm{Var}(Z)$. For one centered predictor with fixed inputs and noise variance $\sigma^2$, $\mathrm{Var}(\hat\beta_1)=\sigma^2/\sum_ix_i^2$.
With one predictor a ridge coefficient is a scaled least squares coefficient.
Confidence intervals
For an unbiased, normally distributed estimate, $\hat\theta\pm2\,\mathrm{SE}$ covers the true value about $95$ times in $100$ repetitions.
One interleaved question asks what happens to this recipe when the estimate is biased.
Try it yourself first (2 questions)
1§05.5 — adding to the diagonal of a singular matrix
Two predictors give $X^TX=\begin{pmatrix}4&2\\2&1\end{pmatrix}$. Before any regression, check two matrices for an inverse.
Find(a) Which of $X^TX$ and $X^TX+I$ has an inverse?
Given
$X^TX=\begin{pmatrix}4&2\\2&1\end{pmatrix}$
$I=\begin{pmatrix}1&0\\0&1\end{pmatrix}$
Hint 1/4
An inverse exists exactly when a determinant is nonzero; decide which two determinants to compute.
Hint 2/4
$\det\begin{pmatrix}a&b\\c&d\end{pmatrix}=ad-bc$, and adding $I$ adds $1$ to each diagonal entry.
Hint 3/4
The two matrices are $X^TX=\begin{pmatrix}4&2\\2&1\end{pmatrix}$ and $X^TX+I=\begin{pmatrix}5&2\\2&2\end{pmatrix}$.
Hint 4/4
$\det(X^TX)=0$ and $\det(X^TX+I)=6$, so only $X^TX+I$ has an inverse.
Show solution
A $2\times2$ matrix is invertible exactly when its determinant is nonzero, which is quicker to check than attempting the inverse.
The original matrix
$$\det(X^TX)=4\cdot1-2\cdot2=0$$
The second column is half the first, so the columns are dependent.
After adding the identity
$$\det(X^TX+I)=5\cdot2-2\cdot2=6\neq0$$
Only the diagonal grew, and that alone broke the dependence.
Answer $$\boxed{\text{only }X^TX+I\text{ is invertible}}$$
Check
Direct check: $X^TX\,(1,\,-2)^T=(0,\,0)^T$, so $X^TX$ sends a nonzero vector to zero and has no inverse, while $(X^TX+I)(1,\,-2)^T=(1,\,-2)^T\neq0$.
A singular $X^TX$ blocks least squares, and adding $\lambda I$ with $\lambda>0$ removes the block.
2§05.3 — least squares after a change of units
A least squares fit predicts monthly spending from income recorded in TL, and the income coefficient is $0.0004$. The same data are refitted with income recorded in thousands of TL.
Find(a) What is the income coefficient after the refit?
Given
income coefficient with income in TL: $0.0004$
new unit: $1$ thousand TL
Hint 1/4
Least squares keeps its predictions under a change of units; ask what the coefficient must do to keep them.
Hint 2/4
If every value of a predictor is multiplied by $c$, the least squares coefficient is divided by $c$.
Hint 3/4
Income in thousands of TL is income in TL times $c=\frac{1}{1000}$, and the old coefficient is $0.0004$.
Hint 4/4
The new coefficient is $0.0004\times1000=0.4$.
Show solution
of least squares turns this into one multiplication, so we use it instead of refitting.
Least squares sees the same data in new units and returns the same fitted values.
Solve for the new coefficient
$$\hat\beta_{\text{new}}=1000\times0.0004=0.4$$
The equation must hold for every income $x$, so the factors in front of $x$ must match.
Answer $$\boxed{0.4\ \text{per thousand TL}}$$
Check
Units check: $0.0004$ per TL times $1000$ TL per unit is $0.4$ per unit, and an income of $25\,000$ TL gives $0.0004\times25\,000=10=0.4\times25$ either way.
Least squares absorbs any change of units into its coefficients; ridge will not, which is why this section standardizes.
Notation
symbol
reads as
means
watch out
$\lambda$
lambda, the tuning parameter
the price per unit of penalty, fixed before fitting, $\lambda\ge0$
Chosen by cross-validation, never by the training RSS.
The least squares section wrote $\hat\beta_{\mathrm{RSS}}$; the superscript leaves room for an index, as in $\hat\beta^R_j$.
$\hat\beta^R_\lambda$
ridge estimate at lambda
the ridge estimate for one particular $\lambda$
Used when several values of $\lambda$ are compared.
$p,\ n$
p, n
the number of predictors and the number of observations
Least squares gets into trouble when $p$ is close to $n$ or larger.
$\bar x_j,\ \sigma_j,\ \bar y$
x bar j, sigma j, y bar
mean and standard deviation of predictor $j$, and the mean response, all on the training data
$\sigma_j$ divides by $n$, as in the lecture, not by $n-1$.
$\tilde x_{ij},\ \tilde z_j$
x tilde i j, z tilde j
a standardized training value and a standardized new input
The new input uses the training $\bar x_j$ and $\sigma_j$.
$\lVert\beta\rVert_2^2,\ \lVert\beta\rVert_1$
squared 2-norm, 1-norm
$\sum_{j=1}^p\beta_j^2$ and $\sum_{j=1}^p\lvert\beta_j\rvert$
Both leave out $\beta_0$.
$s$
the budget
the bound in the constrained forms, $\lVert\beta\rVert_2^2\le s$ or $\lVert\beta\rVert_1\le s$
A larger budget matches a smaller $\lambda$.
$(u)_+,\ \operatorname{sign}(u)$
positive part, sign
$\max(u,0)$, and $+1$, $-1$ or $0$ by the sign of $u$
Used only in the lasso formula for one predictor.
$I$
the
$p\times p$, ones on the diagonal and zeros elsewhere
$\lambda I$ adds $\lambda$ to the diagonal only.
Conventions used here
The intercept is never penalized.
Every penalty sums over $j=1,\dots,p$. After centering the response and every predictor the fitted intercept is $0$, so $X$ has no column of ones in this section, and on standardized inputs the model reads $\hat y=\bar y+\sum_j\hat\beta_j\tilde x_j$.
Penalizing the intercept would make predictions depend on where zero sits on the response scale.
Standardization divides by n.
$\sigma_j=\sqrt{\frac1n\sum_i(x_{ij}-\bar x_j)^2}$, as in the lecture, so every standardized column has $\sum_i\tilde x_{ij}=0$ and $\sum_i\tilde x_{ij}^2=n$.
Software that divides by $n-1$ produces the same family of fits at slightly different values of $\lambda$.
Our λ is the lecture's λ.
The losses are $\mathrm{RSS}+\lambda\cdot\text{penalty}$, with no $\frac12$ and no $\frac1n$ in front of the RSS. Books and packages that minimize $\frac{1}{2n}\mathrm{RSS}+\lambda'\cdot\text{penalty}$ use $\lambda'=\lambda/(2n)$.
The same fit can carry very different $\lambda$ labels in different sources.
Zero means exactly zero.
A coefficient 'at zero' is exactly $0$, which the lasso produces over whole ranges of $\lambda$. A ridge coefficient can pass through $0$ at one isolated $\lambda$; that does not count as selecting a variable.
Selection is about ranges of λ, not about a curve crossing an axis.
Log scales for λ.
Figures and grids of $\lambda$ use powers of ten or of three, because the interesting range spans several orders of magnitude.
A linear grid from 0 to 1000 would spend almost every value on the flat end.
Rounding.
Intermediate steps keep at least four significant figures; final answers are rounded to three decimals unless an exact fraction is asked for.
Ridge and least squares answers often differ only in the second decimal.
5.1Ridge regression: least squares plus a price on the size of the coefficients
Adds $\lambda\sum_j\beta_j^2$ to the RSS so that unstable coefficients shrink; on centered, standardized data the fit is $(X^TX+\lambda I)^{-1}X^Ty$.
Least squares gave us the smallest RSS; on the kiosk data the smallest RSS is exactly what goes wrong.
Solvable with what we have
Fit $\hat\beta^{\mathrm{LS}}=(X^TX)^{-1}X^Ty$ whenever the columns of $X$ are independent.
Split a model's expected test error into bias², variance and noise.
Choose among a few candidate models by cross-validation.
Not solvable yet
Get stable coefficients when two predictors almost copy each other.
Fit at all when $X^TX$ is singular, for instance when $p>n$.
Turn the complexity of a model up or down smoothly instead of dropping whole predictors.
Fit least squares anyway. The five kiosk days, standardized, give $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$ and $X^Ty=(11,\,7)^T$, so $\hat\beta^{\mathrm{LS}}=(3,\,-1)$: each standardized degree on the second thermometer costs one unit of sales.
Why it fails
The data fix the sum $\beta_1+\beta_2$ from how sales follow the average reading, but the difference $\beta_1-\beta_2$ only from days 1 and 2, the two days the thermometers disagree. Least squares sets $\hat\beta_1-\hat\beta_2$ to the sales gap between those two days, $31-27=4$, noise and all; swapping two readings turns it into $-4$.
DefinitionThe ridge loss and its minimizer
Conditions
$\lambda\ge0$ is fixed before fitting, and the penalty starts at $j=1$: the intercept is not penalized
for the closed form, the response is centered and the predictors are centered and standardized, so $X$ has no column of ones
The ridge loss is the least squares score plus a charge of $\lambda$ for every unit of $\sum_j\beta_j^2$. Its minimizer solves $(X^TX+\lambda I)\beta=X^Ty$, the normal equations with $\lambda$ added to each diagonal entry; that is the only change the penalty makes.
Deriving the closed form
The intercept first. It is not penalized, so setting $\partial\mathrm{Loss}_R/\partial\beta_0=-2\sum_i\big(y_i-\beta_0-\sum_j\beta_jx_{ij}\big)$ to zero gives $\hat\beta_0=\bar y-\sum_j\beta_j\bar x_j$. With centered columns every $\bar x_j=0$, and with a centered response $\hat\beta_0=0$: the column of ones drops out.
Write the rest with vectors; it is the kiosk computation with letters instead of numbers. Starting from $(y-X\beta)^T(y-X\beta)+\lambda\beta^T\beta$:
By the two derivative rules recalled above, the row derivative is $-2y^TX+2\beta^T(X^TX+\lambda I)$. Setting it to zero and transposing gives $(X^TX+\lambda I)\beta=X^Ty$, because $X^TX+\lambda I$ is symmetric.
For $\lambda>0$ the matrix $X^TX+\lambda I$ is invertible, and the loss has the second derivative $2(X^TX+\lambda I)$, which is ; the invertibility block below proves both. The loss is therefore a bowl with one lowest point, $\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$.
Forty simulated weeks on the kiosk's ten thermometer readings, with true coefficients $(1,1)$ and noise standard deviation $2$. $\textcolor{#8250df}{\text{Least squares}}$ estimates spread along the line $\beta_1+\beta_2=2$: the standard deviation of $\hat\beta_1-\hat\beta_2$ is $2\sqrt2\approx2.83$. $\textcolor{#1f6feb}{\text{Ridge}}$ at $\lambda=3$ divides it by $1+\lambda=4$, to $0.71$, and pulls the cloud toward the origin.
Looks like this, but is not
Fitting least squares and then multiplying $\hat\beta^{\mathrm{LS}}$ by one shrink factor, say $0.5$, looks like ridge: every coefficient gets smaller.
Ridge does not shrink all directions alike. On the kiosk data it multiplies the well measured sum $\beta_1+\beta_2$ by $\frac{9}{9+\lambda}$ and the badly measured difference by $\frac{1}{1+\lambda}$; at $\lambda=3$ that is $0.75$ against $0.25$. One common factor equals ridge only when $X^TX$ is a multiple of $I$.
Two thermometers, five days: least squares against ridge at λ = 3
The kiosk's readings are $T_1=18, \allowbreak 20, \allowbreak 21, \allowbreak 22, \allowbreak 24$ and $T_2=20, \allowbreak 18, \allowbreak 21, \allowbreak 22, \allowbreak 24$ degrees, with sales $27,31,26,32,34$. Standardize the readings, center the sales, and fit least squares and ridge with $\lambda=3$. Then predict sales on a day with $T_1=23$ and $T_2=19$.
Find$\hat\beta^{\mathrm{LS}}$, $\hat\beta^R$ and both predictions for $T_1=23$, $T_2=19$.
Multiply back: $(X^TX+3I)\hat\beta^R=(8\cdot1.25+4\cdot0.25,\ 4\cdot1.25+8\cdot0.25)=(11,\,7)=X^Ty$. The sum $1.25+0.25=1.5$ is $\frac{9}{12}$ of the least squares sum $2$, and the difference $1$ is $\frac14$ of $4$: the two shrink factors of the counterexample above.
One $2\times2$ inverse per value of $\lambda$; the products $X^TX$ and $X^Ty$ are computed once.
This settles the opening puzzle: ridge gives both thermometers a positive weight, and on a day when they disagree by four degrees it predicts $31$, not $34$, because the difference least squares read off two noisy days is cut to a quarter.
One standardized predictor: ridge is least squares times n/(n + λ)
A single standardized predictor has $\sum_ix_i^2=n=20$ and $\sum_ix_iy_i=30$ on centered data. Compute $\hat\beta^{\mathrm{LS}}$ and $\hat\beta^R$ for $\lambda=5$ and $\lambda=20$, and find the factor that turns one into the other.
FindBoth ridge estimates and the factor from $\hat\beta^{\mathrm{LS}}$ to $\hat\beta^R$.
Given
$n=20$, $\sum_ix_i^2=20$, $\sum_ix_iy_i=30$
$\lambda\in\{5,\,20\}$
Solution
With one column, $X^TX$ is the number $\sum_ix_i^2$, so the matrix formula collapses to one division.
Factor check: $\frac{20}{25}=0.8$ and $0.8\times1.5=1.2$; at $\lambda=n$ the factor is exactly $\frac12$, and $0.75$ is half of $1.5$.
With one standardized predictor, ridge never changes the sign of the fit and never reaches $0$: it multiplies by $\frac{n}{n+\lambda}$, which lies strictly between $0$ and $1$.
Checkpoint
§05.1 — one ridge solve with two predictors
Two standardized predictors on centered data give the products below. You fit ridge with $\lambda=2$.
Find(a) Which vector is $\hat\beta^R$?
Given
$X^TX=\begin{pmatrix}6&2\\2&6\end{pmatrix}$
$X^Ty=(9,\,6)^T$
$\lambda=2$
Hint 1/4
The ridge estimate solves one linear system; decide which matrix and which right-hand side it uses before computing.
Hint 2/4
$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$, and $\begin{pmatrix}a&b\\b&a\end{pmatrix}^{-1}=\frac{1}{a^2-b^2}\begin{pmatrix}a&-b\\-b&a\end{pmatrix}$.
Hint 3/4
Here $X^TX+2I=\begin{pmatrix}8&2\\2&8\end{pmatrix}$ with determinant $60$, and $X^Ty=(9,\,6)^T$.
5.2The tuning parameter λ: from least squares to the mean, chosen by cross-validation
Shows what each $\lambda$ does, from $\hat\beta^{\mathrm{LS}}$ at $\lambda=0$ to the mean prediction as $\lambda\to\infty$, and picks $\lambda$ by cross-validation.
The kiosk fit used $\lambda=3$ without a reason; this block shows what the other values do and how the data choose one.
At $\lambda=0$ the penalty is free and ridge is least squares; as $\lambda$ grows the coefficients are pulled toward zero, and for a huge $\lambda$ every prediction is the mean response. The training RSS cannot choose between these fits, because it is smallest at $\lambda=0$; cross-validation over a grid of values can.
Why the two ends and the ordering hold
Small $\lambda$: when $X^TX$ is invertible, $(X^TX+\lambda I)^{-1}$ depends continuously on $\lambda$, so it tends to $(X^TX)^{-1}$ and $\hat\beta^R_\lambda$ tends to $\hat\beta^{\mathrm{LS}}$.
Large $\lambda$: $(X^TX+\lambda I)^{-1}=\frac1\lambda\big(I+\frac1\lambda X^TX\big)^{-1}$ and the bracket tends to $I$, so $\hat\beta^R_\lambda\approx\frac1\lambda X^Ty\to0$.
The training RSS never falls as $\lambda$ grows. Take $\lambda_1<\lambda_2$ with fits $\beta_1,\beta_2$, RSS values $R_1,R_2$ and penalties $P_1,P_2$. Each fit is optimal for its own loss: $R_1+\lambda_1P_1\le R_2+\lambda_1P_2$ and $R_2+\lambda_2P_2\le R_1+\lambda_2P_1$.
Adding the two gives $(\lambda_2-\lambda_1)(P_1-P_2)\ge0$, so $P_1\ge P_2$: the penalty never grows. The first inequality then gives $R_1\le R_2+\lambda_1(P_2-P_1)\le R_2$.
The ridge path of the kiosk fit, $\hat\beta^R_\lambda=\frac{9}{9+\lambda}(1,1)+\frac{2}{1+\lambda}(1,-1)$. Both coefficients start at the $\textcolor{#8250df}{\text{least squares}}$ values $3$ and $-1$ and end at $0$. The $\textcolor{#1f6feb}{\text{ridge}}$ coefficient of thermometer 2 changes sign once, at $\lambda=\frac97$, and is zero nowhere else.
Looks like this, but is not
Since a larger $\lambda$ shrinks the fit, every coefficient looks as if it must move toward $0$ as $\lambda$ grows.
Only the whole vector is guaranteed to shrink: $\sum_j\beta_j^2$ never grows with $\lambda$. Single coefficients can grow. In the kiosk fit $\hat\beta_2$ climbs from $-1$ through $0$ to about $0.31$ before it heads back to $0$.
λ
thermometer 1
thermometer 2
sum of squares
$0$
$3$
$-1$
$10$
$1$
$1.9$
$-0.1$
$3.62$
$3$
$1.25$
$0.25$
$1.625$
$9$
$0.7$
$0.3$
$0.58$
$27$
$0.321$
$0.179$
$0.135$
$81$
$0.124$
$0.076$
$0.021$
The sum of squares falls at every step, as the proof in the box says it must; thermometer 2's coefficient does not, since it climbs from $-1$ to $0.3$.
The kiosk at λ = 1000: every prediction lands near the mean
Fit ridge to the kiosk data with $\lambda=1000$, compare the coefficients with $X^Ty/\lambda$, and compute the training RSS.
Find$\hat\beta^R$, its approximation and the training RSS.
Every fitted value is within $0.03$ of the mean, so the residuals are almost the centered sales.
Answer $$\boxed{\hat\beta^R\approx(0.0109,\ 0.0069),\qquad \hat y\approx30\ \text{on every day},\qquad \mathrm{RSS}\approx45.66}$$
Check
The two ends bracket every fit: $\mathrm{RSS}=20$ at $\lambda=0$ (least squares) and $46=\sum_iy_i^2$ as $\lambda\to\infty$ (predict $\bar y$); $45.66$ lies between them, near the top.
A very large $\lambda$ does not predict $0$; it predicts the mean response, because the intercept is never penalized.
Choosing λ for the kiosk by leave-one-out cross-validation
Each kiosk day was left out in turn, ridge was refitted on the other four days (standardizing with their means and spreads), and the left-out day's squared error was recorded. Average the errors for each $\lambda$, choose $\lambda$, and refit on all five days.
Find$\mathrm{CV}(5)$ for each $\lambda$, the chosen $\lambda$ and the refitted coefficients.
all five days: $X^TX=\begin{pmatrix}5&4\\4&5\end{pmatrix}$, $X^Ty=(11,\,7)^T$, mean sales $30$
Solution
With five days, leave-one-out is five-fold cross-validation, the most any split can hold out; the errors are given, so the work is five averages and one refit.
Day 3 contributes $25.00$ to every row: its readings equal the other four days' means, so every fit predicts their mean sales, $31$, against the actual $26$. The ranking comes from the other days, above all days 1 and 2, the two that separate the thermometers.
Five values of $\lambda$ times five folds is $25$ ridge fits, plus one refit.
Cross-validation picks $\lambda$ from errors on held-out days, which the training RSS cannot do; the winner is then refitted on all the data.
Checkpoint
§05.2 — choosing λ by the training RSS
A friend fits ridge with $\lambda=0,1,10,100$ on a training set and keeps the $\lambda$ whose fit has the smallest RSS on that same training set.
Find(a) Which λ will the friend choose?
Given
candidates: $\lambda\in\{0,\,1,\,10,\,100\}$
score: RSS on the training data used for fitting
Hint 1/4
The friend's score is measured on the data the fits were made from; ask which fit is best at exactly that.
Hint 2/4
Least squares minimizes RSS over all $\beta$, and the training RSS of $\hat\beta^R_\lambda$ never falls as $\lambda$ grows.
Hint 3/4
The candidates are $\lambda=0,1,10,100$, and $\lambda=0$ gives $\hat\beta^{\mathrm{LS}}$.
Hint 4/4
The friend always picks $\lambda=0$, plain least squares.
Show solution
The ordering proved in the box answers this for every data set at once, so no fit needs computing.
The kiosk numbers show it: training RSS $20$ at $\lambda=0$, $25.625$ at $\lambda=3$ and $45.66$ at $\lambda=1000$, while leave-one-out preferred $\lambda=9$.
Never choose a tuning parameter by the error on the data the fit was trained on; hold data out.
⚠ Choosing λ by the training RSS
it is the error at hand, and for most fitting questions it is the right score
Measure every predictor in standard deviations from its own mean, center the response, and fit ridge there. A new input goes through the same two operations, with the training means and spreads, before $\bar y$ is added back. Least squares predicts the same with or without this step; ridge does not.
Why least squares ignores units and ridge does not
Record predictor $j$ in a new unit, $x_{ij}\to c\,x_{ij}$ with $c>0$. The RSS depends on $\beta_j$ only through the products $\beta_jx_{ij}$, so least squares answers with $\hat\beta_j\to\hat\beta_j/c$ and every prediction stays the same.
Ridge also charges $\lambda\beta_j^2$. Keeping the old fit would need the coefficient $\beta_j/c$, whose charge is $\lambda\beta_j^2/c^2$: a change of unit acts like the penalty $\lambda/c^2$ on that predictor. Going from TL to thousands of TL, $c=\frac1{1000}$, makes that penalty a million times stronger.
After standardizing, $\tilde x_{ij}=(x_{ij}-\bar x_j)/\sigma_j$ is the same number in every unit, because $c$ cancels between the numerator and $\sigma_j$. The standardized fit and its predictions no longer depend on units.
The four people of the next worked example, predicted at an income of $30\,000$ TL. With $\lambda=100$ held fixed, $\textcolor{#1f6feb}{\text{ridge on raw income}}$ predicts $6.8$ above the mean when income is recorded in TL, $3.4$ in thousands of TL and almost $0$ in tens of thousands. $\textcolor{#8250df}{\text{Least squares}}$, and ridge with $\lambda=4$ on standardized income, do not depend on the unit.
Looks like this, but is not
Rescaling every predictor by the same factor looks harmless for ridge, since the predictors keep their relative sizes.
A common factor $c$ still turns $\lambda$ into $\lambda/c^2$ for every coefficient, so the amount of shrinkage changes. It looks harmless only because least squares, which has no $\lambda$, really is unaffected.
Income in TL or in thousands of TL: least squares shrugs, ridge does not
Four people earn $13, 19, 21, 27$ thousand TL a month and spend $5, 11, 9, 15$ hundred TL. Fit least squares and ridge with $\lambda=100$ on centered data, once with income in thousands of TL and once in TL, and predict spending at $30$ thousand TL.
FindBoth slopes in both units, and the four predictions.
The least squares slopes differ by exactly the unit factor, $0.68=1000\times0.00068$, as scale invariance requires. The ridge slopes do not: $0.34$ is half of $0.68$, while $0.00068$ is essentially the least squares value.
Same people, same $\lambda$, two predictions: on raw data the unit decides how much ridge shrinks, which is why the lecture standardizes first.
Standardize, fit, predict: the same four people
Standardize the four incomes, fit ridge with $\lambda=4$ on standardized income and centered spending, and predict spending at $30$ thousand TL. Then check that recording income in TL changes nothing.
Find$\hat\beta^R$ and the prediction at $30$ thousand TL.
The unit cancels in the ratio, so $\tilde x$, $\hat\beta^R$ and $\hat y$ do not move.
Answer $$\boxed{\hat\beta^R=1.7\ \text{per standard deviation},\qquad \hat y=13.4\ \ (1340\ \text{TL})}$$
Check
The previous example's ridge in thousands of TL used $\lambda=100=4\sigma^2$ and also predicted $13.4$: dividing a column by $\sigma$ has the same effect as multiplying its penalty by $\sigma^2$.
Standardize with the training data, fit, send every new input through the same means and spreads, then add $\bar y$ back.
Checkpoint
§05.3 — predicting with the training statistics
A ridge model was trained on standardized ages and centered responses. It now has to score a test batch whose ages have a different mean and spread from the training ages.
Find(a) What does the model predict for the 50-year-old?
Given
training ages: mean $40$, standard deviation $10$; age coefficient $\hat\beta^R=2$; mean response $\bar y=50$
test batch: mean age $30$, standard deviation $5$
person to score: age $50$
Hint 1/4
The model was fitted on one scale of ages; put the new age on that same scale before using the coefficient.
Hint 2/4
$\tilde z=(z-\bar x)/\sigma$ with the training statistics, then $\hat y=\bar y+\hat\beta^R\tilde z$.
Hint 3/4
Training: $\bar x=40$, $\sigma=10$, $\hat\beta^R=2$, $\bar y=50$; the age is $z=50$. The test batch's mean $30$ and spread $5$ play no part.
Hint 4/4
$\tilde z=1$, so $\hat y=52$.
Show solution
The coefficient is per training standard deviation, so the new age must be expressed in exactly that unit.
Standardize with the training statistics
$$\tilde z=\frac{50-40}{10}=1$$
The model knows only the training mean and spread.
Shrinking by the factor $c$ moves the average fit a fraction $1-c$ of the way to zero, which is the bias, and multiplies the spread of the fit by $c$, which squares into the variance. Near $\lambda=0$ the variance falls faster than the bias² rises, so a little shrinkage always helps; the best amount is the noise-to-signal ratio $\sigma^2/\beta^2$.
Derivation (further reading: the lecture draws these curves without formulas)
$\hat\beta^{\mathrm{LS}}=\frac1n\sum_ix_iy_i=\beta+\frac1n\sum_ix_i\varepsilon_i$, so $E[\hat\beta^{\mathrm{LS}}]=\beta$ and $\mathrm{Var}(\hat\beta^{\mathrm{LS}})=\frac{\sigma^2}{n^2}\sum_ix_i^2=\frac{\sigma^2}{n}$.
Ridge is $c\,\hat\beta^{\mathrm{LS}}$, so $E[\hat\beta^R]=c\beta$ and $\mathrm{Var}(\hat\beta^R)=c^2\sigma^2/n$.
At a new input the error is $Y-x\hat\beta^R=\varepsilon+x(\beta-\hat\beta^R)$. The new noise has mean $0$ and is independent of the rest, so $E[(Y-x\hat\beta^R)^2]=\sigma^2+E[x^2]\,E[(\hat\beta^R-\beta)^2]$: the noise plus the bias² and variance of the coefficient.
With $1-c=\frac{\lambda}{n+\lambda}$, bias² plus variance is $g(\lambda)=\frac{\lambda^2\beta^2+n\sigma^2}{(n+\lambda)^2}$. Its derivative $\frac{2n(\lambda\beta^2-\sigma^2)}{(n+\lambda)^3}$ is negative below $\sigma^2/\beta^2$ and positive above it.
One standardized predictor with $n=8$, noise variance $\sigma^2=4$ and slope $\beta=1$. As $\lambda$ grows, bias² rises from $0$ toward $\beta^2=1$ and the variance falls from $\sigma^2/n=0.5$ toward $0$. Their $\textcolor{#1f6feb}{\text{sum for ridge}}$ dips below the $\textcolor{#8250df}{\text{least squares}}$ level $0.5$ and bottoms out at $\frac13$ when $\lambda=\sigma^2/\beta^2=4$.
Looks like this, but is not
Since ridge is biased and least squares is unbiased, least squares looks like the more accurate of the two.
Accuracy on new data is bias² plus variance, not bias alone. In the figure's setting least squares is $0.5$ above the noise, all of it variance, while ridge at $\lambda=4$ has $\frac19$ of bias² and $\frac29$ of variance, $\frac13$ in total.
The best λ for one predictor: n = 8, σ² = 4, β = 1
A standardized predictor has $n=8$ fixed inputs, true slope $\beta=1$ and noise variance $\sigma^2=4$. Compute the bias², the variance and the expected test error of least squares and of ridge at $\lambda=4$ and $\lambda=24$.
FindThe three terms and the total for $\lambda=0$, $4$ and $24$.
Given
$n=8$, $\sum_ix_i^2=8$, $\beta=1$, $\sigma^2=4$
a new input with $E[x^2]=1$
Solution
Everything follows from the shrink factor $c=\frac{n}{n+\lambda}$, so we compute $c$ first for each $\lambda$.
The formula agrees: $\lambda^*=\sigma^2/\beta^2=4$, and the minimum above the noise, $\frac{\sigma^2\beta^2}{n\beta^2+\sigma^2}=\frac{4}{12}=\frac13$, matches the direct sum.
Shrinkage is a purchase: bias² is the price, the cut in variance is what it buys, and the best $\lambda$ is the noise-to-signal ratio.
Checkpoint
§05.4 — the best λ from the noise and the slope
One standardized predictor is fitted on $n=10$ points. The true slope is $\beta=2$ and the noise variance is $\sigma^2=8$, and the intercept is known to be $0$.
Find(a) Which $\lambda$ gives ridge the smallest expected test error?
Given
$n=10$, $\sum_ix_i^2=10$
$\beta=2$, $\sigma^2=8$
Hint 1/4
The best λ balances the bias ridge adds against the variance it removes; the box gives it in terms of two model quantities.
Hint 2/4
For one standardized predictor, $\lambda^*=\sigma^2/\beta^2$.
Hint 3/4
Here $\sigma^2=8$ and $\beta=2$; the sample size $n=10$ does not enter.
Hint 4/4
$\lambda^*=8/4=2$.
Show solution
The derivative of bias² plus variance vanishes at one point, so we use the result instead of re-deriving it.
The sign of $\lambda\beta^2-\sigma^2$ decides whether more shrinkage helps.
Answer $$\boxed{\lambda^*=2}$$
Check
Direct check with $c=\frac{n}{n+\lambda}$: bias² plus variance is $\frac{\lambda^2\cdot4+80}{(10+\lambda)^2}$, which gives $0.8$ at $\lambda=0$, $\frac{84}{121}\approx0.694$ at $\lambda=1$, $\frac{96}{144}\approx0.667$ at $\lambda=2$ and $\frac{116}{169}\approx0.686$ at $\lambda=3$.
More noise calls for more shrinkage and a stronger signal for less; the sample size sets how much each unit of λ shrinks.
⚠ Calling the unbiased estimator the more accurate one
the word unbiased sounds like a guarantee of accuracy
wrong$$\text{bias}=0\ \Rightarrow\ \text{smallest test error}$$
right$$\lambda^*=\sigma^2/\beta^2\quad(\text{more noise, more shrinkage})$$
5.5Ridge always has one answer: even when least squares has infinitely many
Proves that adding $\lambda I$ makes $X^TX$ invertible, so ridge works with duplicated predictors or $p\ge n$, though it never drops a predictor.
So far $X^TX$ could be inverted; now the two thermometers read exactly alike and it cannot, which always happens once there are at least as many predictors as observations.
TheoremAdding λI makes the matrix invertible
Conditions
any $n\times p$ matrix $X$, whatever its rank
$\lambda>0$
$$\boxed{\begin{aligned}v^T(X^TX+\lambda I)v&=\lVert Xv\rVert_2^2+\lambda\lVert v\rVert_2^2>0\quad(v\neq0)\\ \Rightarrow\ \ \textcolor{#1f6feb}{\hat\beta^R}&=(X^TX+\lambda I)^{-1}X^Ty\ \ \text{exists and is unique}\end{aligned}}$$
For every nonzero direction $v$ the matrix gives a strictly positive number: a squared length that can be $0$, plus $\lambda$ times a squared length that cannot. A matrix that sends no nonzero vector to $0$ is invertible, so ridge has exactly one solution for every $\lambda>0$, even when least squares has a whole line or plane of them.
Proof
$v^TX^TXv=(Xv)^T(Xv)=\lVert Xv\rVert_2^2\ge0$, and it is $0$ exactly when $Xv=0$. That happens for some $v\neq0$ whenever the columns of $X$ are dependent, which is certain when $p\ge n$: centered columns live in a space of dimension $n-1$.
The identity adds $\lambda v^Tv=\lambda\lVert v\rVert_2^2$, which is positive for every $v\neq0$.
If $(X^TX+\lambda I)v=0$, then $v^T(X^TX+\lambda I)v=0$, which forces $v=0$. So $Mv=0$ has only the solution $v=0$, and $M=X^TX+\lambda I$ is invertible.
If the kiosk's thermometers had agreed every day, $X^TX=\begin{pmatrix}5&5\\5&5\end{pmatrix}$ and $X^Ty=(11,\,11)^T$. Every point of the $\textcolor{#8250df}{\text{line }\beta_1+\beta_2=2.2}$ is a least squares fit. $\textcolor{#1f6feb}{\text{Ridge}}$ picks one point, $\frac{11}{10+\lambda}(1,1)$, for each $\lambda$; as $\lambda\to0$ it approaches $(1.1,\,1.1)$, the point of the line closest to the origin.
Looks like this, but is not
Ridge coefficients are small, so a ridge coefficient of $0.02$ looks like ridge dropping that predictor from the model.
A predictor leaves the model only with a coefficient of exactly $0$. Ridge solves a linear system whose answer is, apart from coincidences, nonzero in every entry, so all $p$ predictors stay in; with many predictors that makes the fit hard to read. Exact zeros need a different penalty, the next block's.
Two identical thermometers: least squares has a line of answers, ridge has one
Suppose the second thermometer had copied the first on all five kiosk days, so both standardized columns are $(-1.5, \allowbreak -0.5, \allowbreak 0, \allowbreak 0.5, \allowbreak 1.5)$ and the centered sales are unchanged. Find all least squares fits, the ridge fit at $\lambda=3$, and its limit as $\lambda\to0$.
FindThe set of least squares fits, $\hat\beta^R$ at $\lambda=3$, and its limit as $\lambda\to0$.
In general $\det=(5+\lambda)^2-25=\lambda(10+\lambda)$, and $\lambda$ cancels before the limit.
Answer $$\boxed{\begin{aligned}&\text{LS: every }\beta\text{ with }\beta_1+\beta_2=2.2\\ &\hat\beta^R_3=\big(\tfrac{11}{13},\,\tfrac{11}{13}\big),\qquad \lim_{\lambda\to0}\hat\beta^R_\lambda=(1.1,\,1.1)\end{aligned}}$$
Check
$(1.1,\,1.1)$ lies on the line, and it is the point of the line closest to the origin: the line's normal direction is $(1,1)$, so the foot of the perpendicular from $0$ is $\frac{2.2}{2}(1,1)$.
When least squares cannot decide between fits, ridge decides for the one with the smallest coefficients, and for every $\lambda>0$ it decides uniquely.
Checkpoint
§05.5 — ridge with more predictors than observations
A gene study measures $p=500$ expression levels on $n=60$ patients and fits ridge with $\lambda=0.1$. A classmate says no ridge fit can exist, because $X^TX$ is singular.
Find(a) True or false: the ridge estimate exists and is unique for every $\lambda>0$, including $\lambda=0.1$.
Given
$p=500$ predictors, $n=60$ observations
$\lambda=0.1$
Hint 1/4
Existence of the ridge estimate hinges on one matrix; ask what could make it singular.
Holds for every $v\neq0$, so no nonzero vector is sent to $0$.
Answer $$\boxed{\text{True: }\hat\beta^R\ \text{exists and is unique}}$$
Check
Small case of the same kind: two identical standardized columns with $n=5$ give $\det(X^TX+\lambda I)=\lambda(10+\lambda)$, positive for every $\lambda>0$ and $0$ only at $\lambda=0$.
Any positive λ, however small, rescues invertibility; how much it shrinks is a separate question for cross-validation.
⚠ Believing ridge fails wherever least squares fails
the least squares formula needs that inverse, and ridge looks like the same formula
The lasso charges $\lambda$ per unit of absolute size instead of squared size. As $\lambda\to0$ it becomes least squares, and as $\lambda\to\infty$ every coefficient is $0$. In between, some coefficients sit at exactly $0$ over whole ranges of $\lambda$: that is variable selection. With one predictor the fit moves the least squares value $\frac{\lambda}{2n}$ toward $0$ and stops there.
The one-predictor formula (further reading)
With $\sum_ix_i^2=n$ and $z=\sum_ix_iy_i=n\hat\beta^{\mathrm{LS}}$, the loss is $\sum_iy_i^2-2z\beta+n\beta^2+\lambda\lvert\beta\rvert$.
On $\beta>0$ it is a parabola with derivative $-2z+2n\beta+\lambda$, zero at $\beta=\frac{z-\lambda/2}{n}$; that point is positive only when $z>\lambda/2$. On $\beta<0$ the stationary point is $\frac{z+\lambda/2}{n}$, negative only when $z<-\lambda/2$.
If $\lvert z\rvert\le\lambda/2$, neither piece has a stationary point on its own side: the loss rises in both directions from $\beta=0$, so the minimum sits at the corner $0$. Dividing $z$ by $n$ gives the formula in the box.
The lasso path of the kiosk fit. $\textcolor{#d1690a}{\text{Thermometer 2}}$ climbs from $-1$, hits exactly $0$ at $\lambda=2$ and stays there, so from then on the model uses one thermometer. $\textcolor{#d1690a}{\text{Thermometer 1}}$ falls in straight pieces from $3$ to $0$, reached at $\lambda=22=2\max_j\lvert x_j^Ty\rvert$.
Looks like this, but is not
The ridge path of the kiosk fit also passes through $\hat\beta_2=0$, at $\lambda=\frac97$, so ridge looks able to drop a predictor too.
Ridge touches $0$ at that single $\lambda$ and moves on; at every other value thermometer 2 is back in the model. The lasso holds $\hat\beta_2$ at exactly $0$ for every $\lambda\ge2$, a whole range, and that is what selecting a variable means.
least squares
ridge
lasso
$2.0$
$1.0$
$1.5$
$1.0$
$0.5$
$0.5$
$0.5$
$0.25$
$0$
$0.3$
$0.15$
$0$
$-0.8$
$-0.4$
$-0.3$
Ridge multiplies every entry by $\frac{n}{n+\lambda}=\frac12$ and never reaches $0$. The lasso takes $\frac{\lambda}{2n}=0.5$ off each size and returns exactly $0$ whenever the least squares value is within $0.5$ of zero.
One predictor by cases: the lasso at λ = 4 and at λ = 30
A standardized predictor with $n=10$ has $\sum_ix_iy_i=12$ on centered data. Minimize the lasso loss for $\lambda=4$ and $\lambda=30$ by treating $\beta>0$, $\beta<0$ and $\beta=0$ separately, and compare with ridge at $\lambda=4$.
Find$\hat\beta^L$ for both values and $\hat\beta^R$ at $\lambda=4$.
Given
$n=10$, $\sum_ix_i^2=10$, $\sum_ix_iy_i=12$
$\lambda\in\{4,\,30\}$
Solution
The absolute value has a corner at $0$, so no single derivative can be set to zero; splitting by the sign of $\beta$ turns each piece into a parabola.
The formula in the box agrees: $\hat\beta^{\mathrm{LS}}=1.2$ and $\frac{\lambda}{2n}=0.2$ or $1.5$, so $1.2-0.2=1.0$ and $(1.2-1.5)_+=0$.
Split by sign whenever an absolute value is minimized: each piece is a parabola, and the corner at $0$ wins when neither parabola's lowest point lies on its own side.
The kiosk at λ = 3: why the lasso keeps only thermometer 1
For the kiosk data, show that $\hat\beta^L=(1.9,\,0)$ at $\lambda=3$: find the best $\beta_1$ with $\beta_2=0$, then check that moving $\beta_2$ away from $0$ in either direction does not help.
Find$\hat\beta^L$ at $\lambda=3$, and the reason $\hat\beta_2=0$.
The lasso has no closed form here, so we guess which coefficient is zero and then verify the guess. Guess and check can feel like cheating; a checked guess is a proof, and it beats searching all sign patterns.
Compare losses: $\mathrm{Loss}_L$ is $27.95$ at $(1.9,0)$ against $28$ at $(2,0)$, $30.125$ at the ridge fit $(1.25,0.25)$ and $32$ at the least squares fit $(3,-1)$.
With two nearly equal predictors the lasso tends to keep one and drop the other, and which one it keeps can hinge on a few observations; ridge splits the weight instead.
Checkpoint
§05.6 — soft thresholding with one predictor
One standardized predictor is fitted on $n=25$ centered points, and least squares gives $\hat\beta^{\mathrm{LS}}=-0.6$. You refit with the lasso at $\lambda=20$.
Find(a) What is $\hat\beta^L$?
Given
$n=25$, $\sum_ix_i^2=25$
$\hat\beta^{\mathrm{LS}}=-0.6$
$\lambda=20$
Hint 1/4
With one predictor, the lasso moves the least squares value toward zero by a fixed amount and stops at zero.
Hint 2/4
With $b=\hat\beta^{\mathrm{LS}}$: $\hat\beta^L=\operatorname{sign}(b)\big(\lvert b\rvert-\frac{\lambda}{2n}\big)_+$.
Hint 3/4
Here $\hat\beta^{\mathrm{LS}}=-0.6$, $\lambda=20$ and $n=25$, so the cut is $\frac{20}{50}=0.4$.
Hint 4/4
$\hat\beta^L=-(0.6-0.4)=-0.2$.
Show solution
The one-predictor formula applies directly, so we compute the cut and compare it with the size.
The cut
$$\frac{\lambda}{2n}=\frac{20}{50}=0.4$$
No $\frac12$ in front of the RSS, so the cut carries the factor 2.
Subtract and keep the sign
$$\hat\beta^L=-\big(0.6-0.4\big)_+=-0.2$$
The size is above the cut, so the result is not zero.
Answer $$\boxed{\hat\beta^L=-0.2}$$
Check
By cases: on $\beta<0$ the loss derivative is $-2z+2n\beta-\lambda$ with $z=25(-0.6)=-15$, which vanishes at $\beta=\frac{-30+20}{50}=-0.2$, inside the region.
Soft thresholding is a subtraction on the size; the sign comes back at the end.
⚠ Dropping the factor 2 in the cut
books that put 1/2 in front of the RSS write the threshold without it
Instead of a price $\lambda$ per unit of size, give the coefficients a budget $s$ and take the smallest RSS the budget allows. Both forms trace the same fits: the penalized fit at $\lambda$ spends exactly the budget $s$, and a smaller budget matches a larger $\lambda$. Ridge's budget region is a disk, the lasso's a diamond with its corners on the axes.
Why a penalized fit solves the budget problem
Let $\hat\beta_\lambda$ minimize $\mathrm{RSS}+\lambda P$ and set $s=P(\hat\beta_\lambda)$. For any $\beta$ with $P(\beta)\le s$, optimality gives $\mathrm{RSS}(\beta)+\lambda P(\beta)\ge\mathrm{RSS}(\hat\beta_\lambda)+\lambda s$.
Since $\lambda P(\beta)\le\lambda s$, this leaves $\mathrm{RSS}(\beta)\ge\mathrm{RSS}(\hat\beta_\lambda)$: nothing within the budget fits better, so $\hat\beta_\lambda$ solves the constrained problem.
In the picture, the RSS contours are ellipses around $\hat\beta^{\mathrm{LS}}$, and the constrained solution is where the growing ellipses first reach the region. A diamond's corners stick out toward the contours, so first contact is often at a corner, where a coefficient is $0$; a disk has no corners.
The kiosk fit at $\lambda=3$ as a budget problem. $\textcolor{#8250df}{\text{RSS contours}}$ are ellipses around $(3,-1)$, long in the $(1,-1)$ direction where the data say little. Growing outward, the first ellipse to reach the $\textcolor{#d1690a}{\text{lasso diamond}}$ ($s=1.9$) meets its corner $(1.9,\,0)$; the first to reach the $\textcolor{#1f6feb}{\text{ridge disk}}$ ($s=1.625$) touches it at $(1.25,\,0.25)$.
Looks like this, but is not
Every budget looks as if it forces some shrinkage, since the fit now has to obey a constraint.
A budget that the least squares fit already meets costs nothing: the constrained solution is $\hat\beta^{\mathrm{LS}}$ itself, which matches $\lambda=0$. For the kiosk that happens for the lasso once $s\ge\lvert3\rvert+\lvert-1\rvert=4$, and for ridge once $s\ge3^2+(-1)^2=10$.
One predictor with a budget: which λ matches β = 0.8
A standardized predictor with $n=10$ and $\sum_ix_iy_i=12$ has $\hat\beta^{\mathrm{LS}}=1.2$. Solve the ridge problem with budget $\beta^2\le0.64$ and the lasso problem with budget $\lvert\beta\rvert\le0.8$, and find the $\lambda$ that gives each answer in the penalized form.
FindBoth constrained solutions and their matching values of $\lambda$.
With one coefficient each budget region is an interval, so we read the constrained solution off the number line and then solve the penalized formulas backwards for $\lambda$.
Penalized check: $\frac{12}{15}=0.8$ and $\frac{12-4}{10}=0.8$. The budgets are spent exactly, $0.8^2=0.64$ and $\lvert0.8\rvert=0.8$, as the matching rule in the box says.
A budget and a price are two handles on the same family of fits, but the numbers do not carry over between methods: the same coefficient needs $\lambda=5$ for ridge and $\lambda=8$ for the lasso.
The kiosk budgets: which s gives which fit
For the kiosk fit at $\lambda=3$, find the ridge and lasso budgets, and the smallest budget at which each constrained problem returns plain least squares.
FindThe budgets at λ = 3, the budgets beyond which nothing is shrunk, and the lasso's solution for any budget up to 2.
Given
$\hat\beta^{\mathrm{LS}}=(3,\,-1)$
$\hat\beta^R=(1.25,\,0.25)$ and $\hat\beta^L=(1.9,\,0)$ at $\lambda=3$
lasso path: $(3-\frac\lambda2,\,-1+\frac\lambda2)$ for $\lambda\le2$, $(\frac{22-\lambda}{10},\,0)$ for $2\le\lambda\le22$
Solution
Each budget is the penalty of the matching penalized fit, so the work is evaluating penalties.
The picture agrees: the diamond with $s_L=1.9$ is touched at its corner $(1.9,0)$, and $(3,-1)$ sits on the boundary of the diamond with $s_L=4$, since $3+1=4$.
Budgets and prices run in opposite directions: as $\lambda$ goes from $0$ to $\infty$, the budget runs from its least squares value down to $0$.
Checkpoint
§05.7 — where the zeros of the lasso come from
Two predictors, the same elliptical RSS contours around $\hat\beta^{\mathrm{LS}}$, and two budget regions: the disk $\beta_1^2+\beta_2^2\le s$ and the diamond $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le s$.
Find(a) Why does the lasso often give a coefficient of exactly 0 while ridge does not?
Step the lasso budget $s$ on the kiosk data. Up to $s=2$ the growing $\textcolor{#8250df}{\text{RSS ellipses}}$ first touch the $\textcolor{#d1690a}{\text{diamond}}$ at its corner $(s,0)$, so $\beta_2=0$; from $s=2$ to $s=4$ the contact point slides along the lower edge; from $s=4$ on, least squares itself fits the budget.
At the edges
s = 1.9 (1.9, 0)
The budget that matches λ = 3: thermometer 2 is out of the model.
s = 2 (2, 0)
The last budget with a corner solution; it matches λ = 2, where the lasso path lets thermometer 2 back in.
s = 4 (3, −1)
The least squares fit spends exactly |3| + |−1| = 4, so every larger budget returns it unchanged.
A ridge fit by hand, from raw data to a prediction
A small table of raw data, a given λ, and a question about the coefficients or a new input.
Training statistics
Compute $\bar x_j$, $\sigma_j=\sqrt{\frac1n\sum_i(x_{ij}-\bar x_j)^2}$ and $\bar y$ from the training rows only.
Standardize and center
$\tilde x_{ij}=(x_{ij}-\bar x_j)/\sigma_j$ and $y_i-\bar y$; check that $\sum_i\tilde x_{ij}^2=n$.
Two products
$X^TX$ and $X^Ty$ from the standardized columns and the centered response.
Add λ to the diagonal
Only the diagonal of $X^TX$ changes; $X^Ty$ stays as it is.
Solve and check
For two predictors use the $2\times2$ inverse, then multiply back to recover $X^Ty$.
Predict
$\tilde z_j=(z_j-\bar x_j)/\sigma_j$ with the training statistics, then $\hat y=\bar y+\tilde z^T\hat\beta^R$.
Where it goes wrong
Adding $\lambda$ to every entry of $X^TX$, or to $X^Ty$.
Standardizing the new input with its own statistics.
Stopping at $\tilde z^T\hat\beta^R$ without adding $\bar y$ back.
A lasso fit by cases
One predictor, uncorrelated standardized predictors, or two predictors where you can guess which coefficient is zero.
Reduce
With one predictor put $z=\sum_ix_iy_i$ and $n=\sum_ix_i^2$; the loss is $-2z\beta+n\beta^2+\lambda\lvert\beta\rvert$ plus a constant.
Compare with the cut
If $\lvert z\rvert\le\lambda/2$, then $\hat\beta^L=0$.
Otherwise subtract
$\hat\beta^L=\big(z-\operatorname{sign}(z)\,\lambda/2\big)/n$: the least squares value moved $\frac{\lambda}{2n}$ toward $0$.
Several predictors
Guess which coefficients are $0$, fit the others by the same rule, then check each zero: $\lvert\partial\mathrm{RSS}/\partial\beta_j\rvert\le\lambda$ there.
Confirm
Compare the lasso loss at your answer with its value at the least squares fit and at one other candidate.
Where it goes wrong
Using the cut $\lambda/n$ instead of $\lambda/(2n)$.
Setting a derivative to zero at $\beta=0$, where the absolute value has none.
Dropping the sign of the least squares value.
Choosing λ by cross-validation
Any ridge or lasso fit that will be used on new data.
Grid
Pick values on a log scale, for example $0,1,3,9,27,81$ or powers of ten.
Folds
Split the data into k folds once and use the same folds for every λ.
Fit inside each fold
Standardize with the training part of the fold, fit, and score the held-out part.
Average
Compute $\mathrm{CV}(k)$ for each $\lambda$ and keep the smallest.
Refit
Fit the chosen λ on all the data, standardizing with all the data.
Where it goes wrong
Scoring $\lambda$ on the training data, which always picks $\lambda=0$.
Standardizing with all the data before splitting, which lets each held-out fold shape its own scaling.
Reporting one fold's fit instead of the refit.
Ridge on three uncorrelated predictors: every coefficient scaled by the same factor
Three standardized, uncorrelated predictors with $n=10$ ($X^TX=10I$) have least squares coefficients $(2.0,\,-0.6,\,0.15)$. Compute ridge with $\lambda=5$.
The third size is below the cut, so it stops at $0$.
Answer $$\boxed{\hat\beta^L=(1.75,\ -0.35,\ 0)}$$
Check
Check the zero: at $\beta_3=0$ the RSS slope is $-2x_3^Ty=-2\cdot10\cdot0.15=-3$, and $\lvert-3\rvert\le5$.
The lasso shifts every size by the same amount, so small coefficients vanish and large ones barely change in proportion.
Same data and the same $\lambda$: ridge multiplies every coefficient by $\frac23$ and keeps all three; the lasso subtracts $0.25$ from every size and drops the third.
How to tell them apart
Ridge divides, the lasso subtracts. Look at the smallest coefficient: if it survives in proportion, the fit is ridge; if it is exactly $0$, it is the lasso.
Least squares after a change of units: the coefficient absorbs the factor
Three people's centered ages in years are $(-10,\,0,\,10)$ and their centered responses $(-4,\,1,\,3)$. Fit least squares, refit with age in decades, and predict for someone $20$ years above the mean age.
The coefficient changed by exactly the unit factor $10$, so the product with age did not change.
Least squares answers a change of units with the reverse change in its coefficient.
Ridge with λ = 10 after the same change: the prediction moves
Same data. Fit ridge with $\lambda=10$ on the centered, unstandardized ages, once in years and once in decades, and predict for someone $20$ years above the mean age.
The shrink factors explain it: $\frac{200}{210}\approx0.95$ in years and $\frac{2}{12}\approx0.17$ in decades, multiplying the least squares prediction $7$.
On raw units the amount of ridge shrinkage is decided by the unit, not by the data.
One change of units, two methods: least squares moves its coefficient by exactly the unit factor and keeps the prediction $7$; ridge with a fixed $\lambda$ turns $6.67$ into $1.17$.
How to tell them apart
If a method's predictions survive a change of units, it puts no penalty on raw coefficients; before ridge or the lasso, standardize so that theirs survive too.
Scaffolding comes off
The common skeleton
Put the data on the standard scale: training means and spreads, standardized predictors, centered response.
Form $X^TX$ and $X^Ty$.
Add $\lambda$ to the diagonal of $X^TX$ only.
Solve $(X^TX+\lambda I)\beta=X^Ty$ and check by multiplying back.
For a new input, standardize with the training statistics and add $\bar y$ back.
1 · fully worked
Six wheat plots: ridge at λ = 4 from raw data to a prediction
Six plots got nitrogen $10, \allowbreak 10, \allowbreak 10, \allowbreak 14, \allowbreak 14, \allowbreak 14$ kg and irrigation $3,3,7,3,7,7$ hours a week, and yielded $6,8,9,11,12,14$ kg. Fit ridge with $\lambda=4$ on standardized inputs and centered yield, and predict a plot with $13$ kg of nitrogen and $4$ hours of irrigation.
Least squares for comparison: $\frac1{32}\begin{pmatrix}6&-2\\-2&6\end{pmatrix}\begin{pmatrix}14\\10\end{pmatrix}=(2,\,1)$, whose prediction is $10+1-0.5=10.5$; ridge shrank both coefficients and landed between that and the mean $10$.
The skeleton never changes: training statistics, products, λ on the diagonal, solve and check, predict on the training scale.
2 · you write the reasoning
Easier, and this time you write the reasons. One standardized predictor from $n=8$ training points with mean $50$ and spread $10$, $\sum_i\tilde x_iy_i^c=6$, mean response $20$ and $\lambda=4$. Predict at a new input $65$, and for each line write why it is allowed.
$\sum_i\tilde x_i^2=8$ for the standardized column.
reasoning
Standardizing with the training mean and the lecture's spread makes the sum of squares equal to n.
$X^TX+\lambda I=8+4=12$.
reasoning
With one column $X^TX$ is a number, and the penalty adds $\lambda$ to it, the only diagonal entry.
$\hat\beta^R=\frac{6}{12}=0.5$.
reasoning
The ridge system $(X^TX+\lambda I)\beta=X^Ty$ is one division here, with $X^Ty=\sum_i\tilde x_iy_i^c=6$.
$\tilde z=\frac{65-50}{10}=1.5$ for a new input $65$.
reasoning
A new input is standardized with the training mean and spread, not its own.
$\hat y=20+0.5\times1.5=20.75$.
reasoning
The fit was made on centered responses, so the mean response is added back.
3 · find the buried error
Harder, with two errors buried in a worked solution. Six plots got nitrogen $10, \allowbreak 10, \allowbreak 10, \allowbreak 14, \allowbreak 14, \allowbreak 14$ kg and irrigation $3,3,7,3,7,7$ hours, and yielded $16, \allowbreak 17, \allowbreak 19, \allowbreak 21, \allowbreak 23, \allowbreak 24$ kg. The task: ridge with $\lambda=6$, and a prediction for a plot with $15$ kg of nitrogen and $6$ hours of irrigation. Which two steps are wrong?
Step 2.$X^TX=\begin{pmatrix}6&2\\2&6\end{pmatrix}$ and $X^Ty=(16,\,12)^T$, from the centered yields $(-4, \allowbreak -3, \allowbreak -1, \allowbreak 1, \allowbreak 3, \allowbreak 4)$.
Step 3. Add the penalty: $X^TX+6I=\begin{pmatrix}12&8\\8&12\end{pmatrix}$, with determinant $144-64=80$.
Step 4. Invert and multiply: $\hat\beta^R=\frac1{80}(12\cdot16-8\cdot12,\ -8\cdot16+12\cdot12)=(1.2,\ 0.2)$.
Step 5. New plot: $\tilde z=\big(\frac{15-12}{4},\ \frac{6-5}{4}\big)=(0.75,\,0.25)$, so $\hat y=20+1.2(0.75)+0.2(0.25)=20.95$.
the two buried errors (2)
⚠ step 3
The penalty was added to every entry. $\lambda I$ adds $6$ to the diagonal only, so the matrix is $\begin{pmatrix}12&2\\2&12\end{pmatrix}$ with determinant $140$.
'Add λ' sticks in memory without the identity matrix that says where.
right
$\hat\beta^R=\frac{1}{140}(12\cdot16-2\cdot12,\ -2\cdot16+12\cdot12)$, which is $\frac{1}{140}(168,\,112)=(1.2,\ 0.8)$
⚠ step 5
The new plot was divided by the variance $4$ instead of the standard deviation $2$.
$\sigma^2=4$ is computed on the way to $\sigma=2$, and the two swap easily once the training columns are done.
right
$\tilde z=(1.5,\,0.5)$ and $\hat y=20+1.2(1.5)+0.8(0.5)=22.2$
4 · the bare problem
§05.1 — a ridge fit and a prediction, no scaffolding
A shop models daily sales from two standardized predictors measured on $n=6$ days: temperature and hours of rain, which tend to move in opposite directions. The fit is ridge with $\lambda=2$.
Find
(a) Compute $\hat\beta^R$.
(b) Predict sales for the new day.
Given
$X^TX=\begin{pmatrix}6&-2\\-2&6\end{pmatrix}$ and $X^Ty=(12,\,-6)^T$ (standardized predictors, centered sales)
training means $24$ °C and $3$ hours; spreads $4$ °C and $2$ hours; mean sales $30$
new day: $28$ °C and $2$ hours of rain
Hint 1/4
Follow the skeleton: add λ to the diagonal, solve, then put the new day on the training scale.
Hint 2/4
$\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$; $\tilde z_j=(z_j-\bar x_j)/\sigma_j$ and $\hat y=\bar y+\tilde z^T\hat\beta^R$.
Hint 3/4
$X^TX+2I=\begin{pmatrix}8&-2\\-2&8\end{pmatrix}$ with determinant $60$ and $X^Ty=(12,-6)^T$; the new day is $\big(\frac{28-24}{4},\ \frac{2-3}{2}\big)$ on the training scale, and mean sales are $30$.
Hint 4/4
$\hat\beta^R=(1.4,\,-0.4)$, $\tilde z=(1,\,-0.5)$ and $\hat y=31.6$.
Show solution
The same skeleton as the plots; only the sign of the off-diagonal entry differs.
Multiply back: $8(1.4)-2(-0.4)=12$ and $-2(1.4)+8(-0.4)=-6$. Least squares would give $\frac1{32}(6\cdot12-2\cdot6,\ 2\cdot12-6\cdot6)=(1.875,\,-0.375)$: larger in the first coefficient ($1.875>1.4$) but smaller in the second ($0.375<0.4$). The vector still shrinks, $\lVert\beta\rVert_2^2$ from $3.66$ to $2.12$, while one coefficient grows, as on the kiosk path.
A negative correlation between the inputs only flips the sign of the off-diagonal entry; the skeleton stays the same.
Full exam-style question
Exam-style: ridge and lasso on two uncorrelated predictorsexam format
Centered responses and two standardized, uncorrelated predictors from $n=4$ observations give $X^TX=4I$ and $X^Ty=(8,\,-1)^T$. (a) Derive the ridge estimate from $\mathrm{Loss}_R$ and evaluate it at $\lambda=4$. (b) Evaluate the lasso estimate at $\lambda=4$. (c) Find the smallest $\lambda$ at which the lasso sets both coefficients to $0$. (d) Give the lasso budget $s$ that matches $\lambda=4$.
Find$\hat\beta^R$, $\hat\beta^L$, the smallest all-zero $\lambda$, and the budget $s$.
Given
$n=4$, $X^TX=4I$, $X^Ty=(8,\,-1)^T$
$\lambda=4$
Solution
Uncorrelated standardized predictors split both problems into one-predictor problems, which is quicker than any matrix inverse.
Check the zero in (b): at $(1.5,0)$ the RSS slope in $\beta_2$ is $-2(x_2^Ty-x_2^Tx_1\cdot1.5)=-2(-1-0)=2\le4$. For (a), multiply back: $8\,(1,\,-0.125)=(8,\,-1)$.
With $X^TX=nI$, ridge divides every coefficient by the same $1+\lambda/n$, while the lasso subtracts the same $\lambda/(2n)$ and drops whatever falls below it.
Practice
A · concept 4 questions
1§05.2 — can one ridge coefficient grow with λ?
A classmate sums ridge up as 'the larger $\lambda$, the closer every coefficient is to zero'. You test the claim on the kiosk fit.
Find(a) True or false: increasing $\lambda$ never increases the absolute value of any single ridge coefficient.
This is the quantity the ordering proof controls, and it does fall.
Answer $$\boxed{\text{False}}$$
Check
Recompute both fits: $\frac1{20}(6\cdot11-4\cdot7,\ -4\cdot11+6\cdot7)=(1.9,\,-0.1)$ and $\frac1{48}(8\cdot11-4\cdot7,\ -4\cdot11+8\cdot7)=(1.25,\,0.25)$.
Ridge shrinks the coefficient vector, not each coefficient; with correlated predictors a single coefficient can move away from zero for a while.
2§05.1 — why the intercept is not penalized
A ridge fit predicts exam scores out of $100$. A classmate suggests adding $\lambda\beta_0^2$ to the penalty so that every coefficient is treated alike.
Find(a) What is the main reason ridge leaves $\beta_0$ out of the penalty?
Given
the model $\hat y=\beta_0+\sum_{j=1}^p\beta_jx_j$ on standardized predictors
the proposal: penalty $\lambda\sum_{j=0}^p\beta_j^2$
Hint 1/4
Ask what the size of $\beta_0$ depends on, apart from the pattern in the data.
Hint 2/4
With standardized predictors, $\hat\beta_0=\bar y$: the intercept is the average response.
Hint 3/4
Recording scores as points above $50$ instead of out of $100$ moves $\bar y$ by $50$ and changes nothing else in the data.
Hint 4/4
The intercept only records where zero sits on the response scale, so charging for its size would tie the fit to an arbitrary choice.
Show solution
Shifting the response is a harmless change of scale, so we test the proposal against it.
Only the intercept notices, so a penalty on it would make the fit depend on the choice of zero.
Answer $$\boxed{\beta_0\ \text{records location, not complexity}}$$
Check
Unpenalized, the shifted fit predicts exactly $50$ less for everyone, as it should; penalized, $\beta_0$ would be pulled toward $0$ by different amounts on the two scales.
Penalize what measures complexity, the slopes, and leave the location of the response alone.
3§05.5 — can ridge remove a small coefficient?
Three standardized, uncorrelated predictors have least squares coefficients $2.0$, $-0.6$ and $0.05$ from $n=10$ observations. A classmate expects a large enough $\lambda$ to remove the third predictor while keeping the other two.
Find(a) True or false: for some $\lambda>0$, ridge sets the third coefficient exactly to $0$ while the first two stay nonzero.
Given
$X^TX=10I$
$\hat\beta^{\mathrm{LS}}=(2.0,\,-0.6,\,0.05)$
Hint 1/4
With uncorrelated standardized predictors each ridge coefficient is its own one-predictor problem; ask what ridge does to one coefficient.
Hint 2/4
When $X^TX=nI$, $\hat\beta^R_j=\frac{n}{n+\lambda}\,\hat\beta^{\mathrm{LS}}_j$.
Hint 3/4
Here the factor is $\frac{10}{10+\lambda}$ for all three, applied to $2.0$, $-0.6$ and $0.05$.
Hint 4/4
The factor is positive for every λ, so the third coefficient is never exactly 0 and the statement is false.
Show solution
The system splits into three one-predictor problems, so one formula answers for every λ.
$$\frac{10}{10+\lambda}\times0.05>0\quad\text{for every }\lambda\ge0$$
A positive factor times a nonzero number is never zero.
Answer $$\boxed{\text{False}}$$
Check
Compare the lasso with the same data: its cut $\frac{\lambda}{20}$ passes $0.05$ at $\lambda=1$, and from there on the third coefficient is exactly $0$ while the first two are not, until $\lambda=12$ removes $-0.6$ as well.
Ridge rescales; it never selects. Exact zeros need the lasso.
4§05.2 — the fit as λ grows without bound
A ridge model of house prices uses standardized predictors and a centered response; the mean price in the training data is $\bar y=2.4$ million TL. You refit with larger and larger $\lambda$.
Find(a) What happens to the predicted price of every house as $\lambda\to\infty$?
The intercept is not in the penalty, so nothing pulls it.
Answer $$\boxed{\hat y\to2.4\ \text{million TL for every house}}$$
Check
The kiosk at $\lambda=1000$ shows the same thing: coefficients about $0.011$ and $0.007$, and every day predicted within $0.03$ of the mean sales $30$.
The heaviest ridge penalty leaves the mean predictor, not the zero predictor.
B · computation 7 questions
1§05.1 — shrinking one standardized coefficient
A clinic predicts recovery time from one standardized blood marker measured on $n=12$ patients. On centered data the marker and the recovery times give the products below.
Find
(a) Compute $\hat\beta^{\mathrm{LS}}$ and $\hat\beta^R$ at $\lambda=6$.
(b) For which $\lambda$ is $\hat\beta^R$ exactly half of $\hat\beta^{\mathrm{LS}}$?
Given$n=12$, $\sum_ix_i^2=12$, $\sum_ix_iy_i=18$
Hint 1/4
One standardized predictor turns every matrix in the ridge formula into a number.
Factor check: $\frac{12}{18}\times1.5=1$, and $\frac{18}{24}=0.75$ is half of $1.5$.
For one standardized predictor, λ = n halves the coefficient, which gives a feel for what a given λ does.
2§05.1 — a zero least squares coefficient that ridge makes nonzero
Two positively correlated standardized predictors come from $n=4$ observations. Least squares gives the second one no weight at all; you refit with ridge.
Find
(a) Compute $\hat\beta^{\mathrm{LS}}$.
(b) Compute $\hat\beta^R$ at $\lambda=2$.
(c) Explain in one sentence why the second coefficient moved away from 0.
Two linear systems with the same right-hand side; only the diagonal differs.
Hint 2/4
$\hat\beta=(X^TX+\lambda I)^{-1}X^Ty$ with $\lambda=0$ and with $\lambda=2$, using the $2\times2$ inverse.
Hint 3/4
$X^TX=\begin{pmatrix}4&2\\2&4\end{pmatrix}$ has determinant $12$, $X^TX+2I=\begin{pmatrix}6&2\\2&6\end{pmatrix}$ has determinant $32$, and $X^Ty=(8,4)^T$.
Hint 4/4
$\hat\beta^{\mathrm{LS}}=(2,\,0)$ and $\hat\beta^R=(1.25,\,0.25)$: ridge shrinks the difference of the coefficients harder than their sum.
Show solution
The $2\times2$ inverse gives both fits; the directions $(1,1)$ and $(1,-1)$ explain them.
Multiply back: $6(1.25)+2(0.25)=8$ and $2(1.25)+6(0.25)=4$.
Ridge shrinks along the data's directions, not coefficient by coefficient, so a zero can turn nonzero.
3§05.3 — from raw heights to a ridge prediction
Four adults have heights $158, \allowbreak 164, \allowbreak 166, \allowbreak 172$ cm and weights $52, 60, 58, 66$ kg. Weight is predicted from standardized height by ridge with $\lambda=4$.
Find
(a) Standardize the heights and center the weights.
(b) Compute $\hat\beta^R$.
(c) Predict the new adult's weight and compare with least squares.
Shrink check: $\frac{n}{n+\lambda}=\frac48=\frac12$, and $2.4$ is half of the least squares $\frac{19.2}{4}=4.8$.
Every raw-data ridge question runs the same way: training statistics, standardize, fit, standardize the new input with the same numbers, add the mean.
4§05.6 — one lasso coefficient at four values of λ
A standardized predictor from $n=20$ centered observations has $\sum_ix_iy_i=18$. You track its lasso coefficient as $\lambda$ grows.
Find
(a) Compute $\hat\beta^L$ for each $\lambda$.
(b) Find the smallest $\lambda$ at which $\hat\beta^L=0$.
Given
$n=20$, $\sum_ix_i^2=20$, $\sum_ix_iy_i=18$
$\lambda\in\{8,\,20,\,36,\,40\}$
Hint 1/4
Compare the size of the least squares coefficient with the cut at each λ.
Hint 2/4
With $b=\hat\beta^{\mathrm{LS}}$: $\hat\beta^L=\operatorname{sign}(b)\big(\lvert b\rvert-\frac{\lambda}{2n}\big)_+$, which is $0$ exactly when $\lambda\ge2\lvert\sum_ix_iy_i\rvert$.
Hint 3/4
$\hat\beta^{\mathrm{LS}}=\frac{18}{20}=0.9$ and the cuts are $\frac{\lambda}{40}=0.2,\ \allowbreak 0.5,\ \allowbreak 0.9,\ \allowbreak 1.0$.
Hint 4/4
$\hat\beta^L=0.7,\ \allowbreak 0.4,\ \allowbreak 0,\ \allowbreak 0$, and the coefficient first reaches $0$ at $\lambda=36$.
Show solution
One size and four cuts: a table of subtractions is quicker than solving four losses.
By cases at $\lambda=20$: on $\beta>0$ the derivative $-36+40\beta+20$ vanishes at $\beta=0.4$, inside the region.
The lasso coefficient falls in a straight line with λ and then sits at exactly zero.
5§05.6 — ridge and the lasso side by side on uncorrelated predictors
Three standardized, uncorrelated predictors from $n=20$ observations have least squares coefficients $1.5$, $-0.4$ and $0.1$. Both penalties are tried with $\lambda=10$.
Find
(a) Compute $\hat\beta^R$.
(b) Compute $\hat\beta^L$.
(c) Which predictors does each method keep?
Given
$X^TX=20I$
$\hat\beta^{\mathrm{LS}}=(1.5,\,-0.4,\,0.1)$
$\lambda=10$
Hint 1/4
Uncorrelated standardized predictors split both problems into three one-predictor problems.
Hint 2/4
Ridge multiplies by $\frac{n}{n+\lambda}$; the lasso subtracts $\frac{\lambda}{2n}$ from each size and stops at $0$.
Hint 3/4
$n=20$ and $\lambda=10$, so the factor is $\frac{20}{30}$ and the cut is $\frac{10}{40}=0.25$; the coefficients are $1.5$, $-0.4$ and $0.1$.
Hint 4/4
$\hat\beta^R=(1,\,-0.267,\,0.067)$ keeps all three; $\hat\beta^L=(1.25,\,-0.15,\,0)$ keeps the first two.
Show solution
No cross terms, so each coefficient follows the one-predictor rules.
The inverse gives the same: $X^TX+2I=\begin{pmatrix}6&4\\4&6\end{pmatrix}$ with determinant $20$, and $\frac1{20}(36-24,\ -24+36)=(0.6,\,0.6)$.
With duplicated predictors, ridge splits the weight equally and picks the smallest least squares fit in the limit.
7§05.7 — turning a budget into a λ
A standardized predictor from $n=10$ centered observations has $\sum_ix_iy_i=15$, so $\hat\beta^{\mathrm{LS}}=1.5$. Both penalties are rewritten with a budget.
Find
(a) Solve both constrained problems.
(b) Find the λ that gives each solution in penalized form.
(c) Which budgets would leave least squares untouched?
Given
$n=10$, $\sum_ix_i^2=10$, $\sum_ix_iy_i=15$
budgets: $\beta^2\le1$ for ridge, $\lvert\beta\rvert\le1$ for the lasso
Hint 1/4
With one coefficient each budget is an interval around 0; find the allowed point nearest the least squares value, then match λ.
Hint 2/4
Ridge: $\hat\beta^R=\frac{\sum x_iy_i}{n+\lambda}$; lasso on its positive branch: $\hat\beta^L=\frac{\sum x_iy_i-\lambda/2}{n}$.
Hint 3/4
$\hat\beta^{\mathrm{LS}}=1.5$, budgets $\beta^2\le1$ and $\lvert\beta\rvert\le1$, $n=10$, $\sum x_iy_i=15$.
Hint 4/4
Both constrained solutions are $\beta=1$, matched by $\lambda=5$ for ridge and $\lambda=10$ for the lasso; budgets $s\ge2.25$ and $s\ge1.5$ leave least squares untouched.
Show solution
In one dimension the constrained solution is the allowed point nearest the least squares value, so no calculus is needed.
A quadratic that curves upward in every direction has exactly one stationary point, its minimum.
Answer $$\boxed{\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty,\ \text{unique for every}\ \lambda>0}$$
Check
Special cases: with one predictor this is $\frac{\sum_ix_iy_i}{\sum_ix_i^2+\lambda}$, and $\lambda=0$ gives back the least squares normal equations.
Every ridge derivation has the same three beats: expand, differentiate, and argue invertibility from the λ term.
2§05.6 — the smallest λ that empties the lasso model
Three standardized predictors and a centered response give the products below. You want to know when the lasso keeps no predictor at all, and which one enters first as $\lambda$ is lowered.
Find(a) What is the smallest $\lambda$ at which the lasso sets all three coefficients to $0$, and which predictor enters first as $\lambda$ falls below it?
Given
$X^Ty=(9,\,-14,\,5)^T$
a lasso coefficient at $0$ stays optimal exactly when $\lvert\partial\mathrm{RSS}/\partial\beta_j\rvert\le\lambda$ there
Hint 1/4
Test the all-zero fit with the zero rule, one coefficient at a time.
Hint 2/4
At $\beta=0$, $\frac{\partial\mathrm{RSS}}{\partial\beta_j}=-2x_j^T(y-X\cdot0)=-2x_j^Ty$, so every zero is kept exactly when $\lambda\ge2\max_j\lvert x_j^Ty\rvert$.
Hint 3/4
$x_j^Ty=9,\ -14,\ 5$, so the three slopes have sizes $18$, $28$ and $10$.
Hint 4/4
All three stay at $0$ exactly when $\lambda\ge28$; below $28$ predictor 2 enters first, with a negative coefficient.
Show solution
At the all-zero fit the residual is the response itself, so every slope is a product the data give directly.
The same rule on the kiosk data: $X^Ty=(11,7)$ gives $2\times11=22$, exactly where the lasso path in the figure reaches $(0,0)$.
The largest correlation with the response in absolute value decides both the empty-model threshold and the first predictor in.
3§05.4 — how much ridge beats least squares at its best λ
A sensor's response is predicted from one standardized input with $n=20$ fixed readings. The true slope is $\beta=0.5$, the noise variance is $\sigma^2=4$, and the intercept is known.
Find
(a) Find $\lambda^*$ and the shrink factor $c$ there.
(b) Compute bias², variance and the expected test error at $\lambda^*$, and for least squares.
(c) What fraction of the error above the noise does ridge remove?
Given
$n=20$, $\sum_ix_i^2=20$
$\beta=0.5$, $\sigma^2=4$
a new input with $E[x^2]=1$
Hint 1/4
Find the best λ first; everything else follows from the shrink factor there.
Answer $$\boxed{\lambda^*=16;\quad \tfrac19\ \text{vs}\ 0.2\ \text{above the noise};\quad \tfrac49\ \text{of it removed}}$$
Check
The closed form for the minimum agrees: $\frac{\sigma^2\beta^2}{n\beta^2+\sigma^2}=\frac{1}{5+4}=\frac19$.
The weaker the signal relative to the noise, the larger the best λ and the more ridge saves.
4§05.7 — the lasso solution when the contours are circles
Two standardized, uncorrelated predictors make the RSS contours circles around $\hat\beta^{\mathrm{LS}}=(2,\,0.5)$. The lasso is solved in budget form with $\lvert\beta_1\rvert+\lvert\beta_2\rvert\le1.2$.
Find(a) Where is the constrained lasso solution?
Given
$X^TX=nI$, so the RSS contours are circles around $(2,\,0.5)$
The neighbouring edge's own projection, $(1.85,\,0.65)$, is off that edge too, so both edges end at this corner.
Answer $$\boxed{(1.2,\ 0)}$$
Check
Soft thresholding agrees: circular contours mean subtracting the same $t$ from each size, and $t=0.8$ gives $\big(1.2,\,(0.5-0.8)_+\big)=(1.2,\,0)$, spending exactly the budget $1.2$.
When the least squares point lies far out along one axis, the diamond's corner on that axis catches the solution and the other coefficient is exactly zero.
5§05.6 — find the step that breaks a lasso computation
An exam question gives two standardized, uncorrelated predictors from $n=5$ observations, with $X^TX=5I$ and $X^Ty=(4,\,-9)^T$, and asks for the lasso at $\lambda=6$. A student writes the four steps below; one of them is wrong.
By cases for the first coefficient: on $\beta_1>0$ the derivative $-2\cdot4+2\cdot5\,\beta_1+6$ vanishes at $\beta_1=0.2$, inside the region, so the first predictor stays in.
Most lasso slips happen in the cut; write down the cut before touching the coefficients.
D · interleaved 4 questions
1§05.1 — least squares or ridge? read it off the residuals
A report gives a fitted coefficient vector for the kiosk data but not the method. For least squares the residual vector is orthogonal to every column of $X$; you check what these residuals do.
Find
(a) Compute $X^T(y-X\hat\beta)$.
(b) Is the fit least squares? If it is ridge, find $\lambda$.
Ridge's equation $X^T(y-X\hat\beta)=\lambda\hat\beta$ holds with one common λ; least squares' would need $(0,0)$.
Answer $$\boxed{\text{ridge with}\ \lambda=9}$$
Check
Refit check: $\frac1{180}\begin{pmatrix}14&-4\\-4&14\end{pmatrix}\begin{pmatrix}11\\7\end{pmatrix}=(0.7,\,0.3)$, the ridge fit at $\lambda=9$.
Ridge residuals are not orthogonal to the columns; they lean toward the fit by exactly λ times the coefficients.
2§05.2 — a validation split between two values of λ
Five training points on one standardized predictor give the products below on centered data. Three held-out points, already on the training scale, decide between least squares and ridge with $\lambda=5$.
Residual check: ridge's three misses $0.2,\ 0.3,\ -0.2$ are all smaller in size than least squares' $-0.8,\ 0.8,\ -2.2$.
Held-out error chooses λ; the winner is then refitted on all the data.
3§05.4 — a 95% interval built around a ridge estimate
Noise is normal with $\sigma=2$, one standardized predictor has $n=16$ fixed inputs, and the true slope is $\beta=1$. A student reports $\hat\beta^R\pm2\,\mathrm{sd}(\hat\beta^R)$ at $\lambda=16$ as a 95% interval for $\beta$.
Find
(a) Find the mean and the standard deviation of $\hat\beta^R$.
(b) How often does the student's interval contain $\beta=1$?
Given
$n=16$, $\sum_ix_i^2=16$, $\sigma=2$, $\beta=1$
$\lambda=16$
standard normal $Z$: $P(Z\le0)=0.5$ and $P(Z\le4)\approx1.000$
Hint 1/4
An interval of the form estimate ± 2 sd covers the truth 95 times in 100 only when the estimate is centered on the truth; check where this one is centered.
Hint 2/4
$\hat\beta^R=c\,\hat\beta^{\mathrm{LS}}$ with $c=\frac{n}{n+\lambda}$, so $E[\hat\beta^R]=c\beta$ and $\mathrm{sd}(\hat\beta^R)=c\,\sigma/\sqrt n$.
Hint 3/4
$c=\frac{16}{32}=\frac12$, $\beta=1$, and $\sigma/\sqrt n=\frac24=0.5$.
Hint 4/4
$\hat\beta^R$ is normal with mean $0.5$ and sd $0.25$; the interval $\hat\beta^R\pm0.5$ contains $1$ only when $\hat\beta^R\ge0.5$, about $50$ times in $100$.
Show solution
Coverage is a probability about the estimate, so we find its distribution first and then the event.
Least squares at the same settings: $\hat\beta^{\mathrm{LS}}\pm2(0.5)$ is centered on $1$ in expectation and covers it about $95$ times in $100$.
A biased estimate with a small spread makes a confidently wrong interval; the ±2 sd recipe needs an unbiased center.
4§05.4 — does ridge contradict Gauss-Markov?
Gauss-Markov says least squares has the smallest variance among linear unbiased estimators. Ridge is also linear in $y$, since $\hat\beta^R=(X^TX+\lambda I)^{-1}X^Ty$, and its variance is smaller.
Find(a) How do the two facts fit together?
Given
Gauss-Markov: among linear unbiased estimators, least squares has the smallest variance
one standardized predictor: $\mathrm{Var}(\hat\beta^R)=c^2\sigma^2/n$ with $c=\frac{n}{n+\lambda}<1$
Hint 1/4
Check each condition of the theorem against ridge: linear, and unbiased.
Hint 2/4
An estimator $Ay$ is linear; it is unbiased when $E[Ay]=\beta$ for every $\beta$.
Hint 3/4
For one standardized predictor $E[\hat\beta^R]=\frac{n}{n+\lambda}\beta$, which differs from $\beta$ whenever $\lambda>0$ and $\beta\neq0$.
Hint 4/4
Ridge is linear but biased, so Gauss-Markov says nothing about it, and its smaller variance contradicts nothing.
Show solution
A theorem about a class says nothing about estimators outside it, so we test membership.
Answer $$\boxed{\text{ridge is biased, so Gauss-Markov does not cover it}}$$
Check
Numbers from the tradeoff block: at $\lambda=4$ with $n=8$, $\sigma^2=4$, $\beta=1$, ridge has variance $\frac29<0.5$ and bias² $\frac19>0$, and its total $\frac13$ beats least squares' $0.5$.
Gauss-Markov ranks unbiased estimators only; shrinkage wins by leaving that class on purpose.
Mistake ledger (19 entries)
⚠ Adding λ to every entry, not only the diagonal
The phrase 'add λ' is remembered without the identity matrix that says where.
a budget at or above the least squares value of the penalty means $\lambda=0$
Check yourself
Close the page and write down from memory:
the ridge and lasso losses, and which coefficient neither of them penalizes;
the ridge estimate, where $\lambda$ enters it, and why that matrix is always invertible;
the standardization formula and how a new raw input becomes a prediction;
the one-predictor rules: ridge multiplies by $\frac{n}{n+\lambda}$, the lasso subtracts $\frac{\lambda}{2n}$;
why the diamond gives zeros and the disk does not.
Then reopen the page and compare; whatever is missing is your reread list.
Compute $\hat\beta^R$ for the kiosk data at $\lambda=3$ by hand and derive the formula from the gradient?
c-ridge
Say what $\lambda\to0$ and $\lambda\to\infty$ do, and pick $\lambda$ from a table of fold errors?
c-lambda
Show with numbers that recording income in thousands changes a raw ridge fit, and predict at a new raw input?
c-scale
Compute bias² and variance of a one-predictor ridge fit and find $\lambda^*=\sigma^2/\beta^2$?
c-tradeoff
Prove that $X^TX+\lambda I$ is invertible and find the ridge fit for two identical predictors?
c-invertible
Soft-threshold a coefficient, and check a guessed lasso zero with the slope rule?
c-lasso
Match a budget $s$ to a $\lambda$ and explain the corner in the diamond picture?
c-geometry
Glossary (15 terms)
düzenlileştirme
Adding a penalty on the size of the coefficients to the fitting criterion, to lower the variance of the fit at the cost of some bias.
ridge regression
Least squares with the added penalty $\lambda\sum_j\beta_j^2$; on centered, standardized data its estimate is $(X^TX+\lambda I)^{-1}X^Ty$.
lasso
Least squares with the added penalty $\lambda\sum_j\lvert\beta_j\rvert$; it shrinks coefficients and sets some of them exactly to zero.
tuning parameterayar parametresi
A constant fixed before fitting that sets how strongly a method regularizes; here λ, chosen by cross-validation.
penaltyceza terimi
The term added to the RSS that grows with the size of the coefficients.
standardizationstandartlaştırma
Subtracting a predictor's mean and dividing by its standard deviation, so that it has mean 0 and variance 1.
scale invarianceölçek değişmezliği
The property that a method's predictions do not change when a predictor is recorded in different units.
eşdoğrusallık
Near linear dependence among the predictors, which leaves some combinations of coefficients poorly determined by the data.
positive definitepozitif tanımlı
A symmetric matrix $M$ with $v^TMv>0$ for every nonzero $v$; such a matrix is invertible.
identity matrixbirim matris
The square matrix with ones on the diagonal and zeros elsewhere; $\lambda I$ adds $\lambda$ to each diagonal entry.
coefficient path
The estimated coefficients plotted against the tuning parameter λ.
variable selectiondeğişken seçimi
Deciding which predictors enter a model; the lasso does it by setting coefficients exactly to zero.
sparse modelseyrek model
A model in which many coefficients are exactly zero.
soft thresholdingyumuşak eşikleme
Moving a number toward zero by a fixed amount and stopping at zero; the lasso's rule for one standardized predictor.
kısıtlı minimizasyon
Minimizing a function over the points that satisfy a condition, here a budget on the size of the coefficients.
What comes next
§06 · Linear regression from a Bayesian perspective
Next, the regression coefficients get a probability distribution of their own: a prior before the data and a posterior after, the way the Bayesian section treated an unknown rate.
Sources
textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The weekly line names no chapter of the book, so no section numbers are cited here.
course materialEEE 485 lecture slides and lecture notes, chapter 5: Regularized regression Scope, order and notation (the two losses, the tuning parameter, the ridge and lasso estimates, standardization with 1/n, the constrained forms) follow these materials; every data set, figure and exercise here is original.
course materialEEE 485 syllabus page on STARS, Fall 2026-27, printed 21 September 2026 Assessment weights, the course learning outcomes, the weekly list and the recommended books.
textbookG. James, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning, Springer, 2013 Recommended in the syllabus; the lecture's coefficient-path, bias-variance and diamond-and-disk figures come from it. None of them is reproduced here.