← back to EEE 485
Week 3117 min full read
7 concepts22 worked examples32 exercises5 exam-level7 figures
What are you here for?

03 Linear regression and ordinary least squares

Start with this

One question before you read anything. Getting it wrong is the point: it shows you what this section is for.

§03.1 — one number closest to three

Before any regression: you must summarize the measurements $1$, $2$ and $6$ by a single number $c$, and you are charged $(1-c)^2+(2-c)^2+(6-c)^2$.

Find(a) Which $c$ makes the cost smallest?
Given
  • measurements $1,\ 2,\ 6$

  • cost $(1-c)^2+(2-c)^2+(6-c)^2$

Hint 1/4

Treat the cost as a function of $c$ and look for its lowest point.

Hint 2/4

Set $\frac{d}{dc}\sum_i(y_i-c)^2=-2\sum_i(y_i-c)$ to zero.

Hint 3/4

With the measurements $1,2,6$: $(1-c)+(2-c)+(6-c)=9-3c$.

Hint 4/4

$9-3c=0$ gives $c=3$, the mean; the cost there is $14$.

Show solution

The cost is a smooth parabola in c, so one derivative finds its lowest point.

Differentiate

$$\frac{d}{dc}\sum_i(y_i-c)^2=-2(9-3c)$$

The derivative of $(y_i-c)^2$ is $-2(y_i-c)$; the three terms add to $9-3c$.

Solve

$$9-3c=0\ \Rightarrow\ c=3$$

The zero of the derivative.

Compare the tempting answers

$$\text{cost}(3)=14,\ \ \text{cost}(2)=17,\ \ \text{cost}(3.5)=14.75,\ \ \text{cost}(3.70)\approx 15.5$$

The middle value, the midpoint of the extremes and the root mean square all do worse.

Answer $$\boxed{c=3}$$
Check

The second derivative is $6>0$, so the cost is an upward parabola and $c=3$ is its minimum.

Squared misses pull the best single number to the mean; with an starts from exactly this fact.

A café counts its social media posts and the cups it sells, in hundreds, for five weeks: $(1,3),\, \allowbreak (2,5),\, \allowbreak (3,4),\, \allowbreak (4,7),\, \allowbreak (5,9)$. One barista rules a line through weeks 1 and 5 and reads 150 extra cups per post; another uses weeks 2 and 4 and reads 100. Which of them is right, and how sure can the café be of its number?

By the end you can fit that line by hand from five pairs, put a 95% interval on its , write the same fit as $\hat\beta=(X^TX)^{-1}X^Ty$ for any number of inputs, and say why squared misses are the natural score.

In 60 seconds

Least squares picks the coefficients whose squared vertical misses add up to the least: for one input $\hat\beta_1=S_{xy}/S_{xx}$ and $\hat\beta_0=\bar y-\hat\beta_1\bar x$; for many, $\hat\beta=(X^TX)^{-1}X^Ty$, and $\hat y$ is the of $y$ onto the columns of $X$.

One input
$$\hat\beta_1=\frac{\sum_i(x_i-\bar x)(y_i-\bar y)}{\sum_i(x_i-\bar x)^2},\qquad \hat\beta_0=\bar y-\hat\beta_1\bar x$$

a table of pairs and a model with an intercept

Precision of the slope
$$\operatorname{Var}(\hat\beta_1)=\frac{\sigma^2}{\sum_i(x_i-\bar x)^2},\qquad \hat\beta_1\pm 2\,\frac{\mathrm{RSE}}{\sqrt{S_{xx}}}$$

a 95% interval, with $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$ in place of $\sigma$

$$X^TX\hat\beta=X^Ty\ \Rightarrow\ \hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$$

any number of inputs, when the columns of X are independent

Projection
$$\hat y=Hy,\quad H=X(X^TX)^{-1}X^T,\quad X^T(y-\hat y)=0$$

testing a fit, or any question about residuals

Three most common mistakes
  1. Raw sums instead of centered ones: $\sum x_iy_i/\sum x_i^2$ is the slope of a line forced through the origin, not the least squares slope of a model with an intercept.

  2. Treating residuals as perpendicular distances to the line. Least squares charges vertical misses; the right angle lives in $\mathbb R^n$, between the residual vector and the columns of $X$.

  3. Dividing the RSS by $n$ instead of $n-2$ for the RSE, or stepping $\pm 2$ variances instead of $\pm 2$ standard errors.

The weights differ between two course documents:

  • Chapter 1 slides, undergraduate line, issued on 15 September and again on 24 September: midterm 25, final 25, four quizzes 20, two-phase project 30.
  • STARS syllabus page for Fall 2026-27, printed 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. Confirm with the course which split applies.
  • To sit the final, both require every project report on time, no disciplinary penalty, and midterm plus quiz points worth at least a fifth of what those two components can give.
How much time do you have?
10 minutes

The two formulas most regression questions start from, each with its first worked example.

The 60-second card · Simple linear regression · Many inputs at once · Formula card
45 minutes

Every block once with its first example and checkpoint, then one ladder from a full solution to a bare problem.

The 60-second card · Simple linear regression · How accurate are the estimates · Many inputs at once · The geometry of least squares · Why squares · Classification with linear regression · Curves with a linear model · Scaffolding comes off · Formula card
full read

Where each formula comes from, how uncertain each estimate is, and enough mixed practice to choose the method yourself.

The opening pages · Recall first · Simple linear regression · How accurate are the estimates · Many inputs at once · The geometry of least squares · Why squares · Classification with linear regression · Curves with a linear model · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
  1. Compute the least squares intercept and slope from a table of pairs, and check the fit through its residuals.

  2. Estimate $\sigma$ by the and build 95% intervals $\hat\beta_j\pm 2\,\mathrm{SE}$ for the intercept and the slope.

  3. Derive the normal equations with the matrix derivative rules and solve $X^TX\beta=X^Ty$ when $X$ has .

  4. Interpret $\hat y=Hy$ as an orthogonal projection and use $X^Te=0$ to test whether a fit is the least squares fit.

  5. Justify least squares twice: as the best linear unbiased estimator and as the maximum likelihood estimator under Gaussian noise.

  6. Turn a least squares fit on 0/1 labels into a , and name the two weaknesses of doing so.

  7. Fit curves with the same machinery by adding interaction, polynomial or basis-function columns, and decide when the fit is unique.

Syllabus coverage

Linear regression — covered

The lecture's chapter

  • the model $Y=f(X,\beta^{\mathrm{true}})+\varepsilon$
  • simple linear regression
  • accuracy of the estimates and 95% intervals
  • classification with linear regression
  • interactions, and basis functions

Seven blocks follow the lecture's order; the model and simple regression open the first block.

ordinary least squares — covered

The fitting method

  • RSS and its minimizer
  • the closed form for one input
  • the normal equations and $\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$ from matrix derivatives
  • the projection $H$
  • Gauss-Markov and the maximum likelihood reading

Proof of the Gauss-Markov theorem — off syllabus

A four-line sketch that splits the weights of any linear unbiased estimator.

Further reading. The lecture states the theorem without proof; nothing later depends on the sketch.

Covariance of the two coefficients — off syllabus

$\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-\bar x\,\sigma^2/S_{xx}$, built from the covariance rules of the first section.

Further reading, used in one interleaved practice question to explain the pivot in the sampling figure.

Maximum likelihood estimate of the noise variance — off syllabus

Maximizing the Gaussian log-likelihood over $\sigma$ gives $\mathrm{RSS}/n$, set against $\mathrm{RSE}^2=\mathrm{RSS}/(n-2)$.

Further reading, used in one interleaved practice question.

Recall first
Minimizing a smooth function

At an interior minimum of a differentiable function every partial derivative is zero; a quadratic $at^2+bt+c$ with $a>0$ is smallest at $t=-b/(2a)$.

Least squares sets two, or p + 1, partial derivatives of the RSS to zero.

Transpose, inverse and rank

$(AB)^T=B^TA^T$; a square matrix is invertible exactly when its columns are independent; $\begin{bmatrix}a&b\\c&d\end{bmatrix}^{-1}=\frac{1}{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix}$.

Every matrix step of the normal equations uses them.

Expectation and variance of weighted sums

$E[\sum_ia_iY_i]=\sum_ia_iE[Y_i]$ always. If the $Y_i$ are uncorrelated with variance $\sigma^2$, $\operatorname{Var}(\sum_ia_iY_i)=\sigma^2\sum_ia_i^2$; in vector form $\operatorname{Var}(a^TY)=a^T\Sigma a$.

Unbiasedness and both variance formulas are these rules applied to the least squares weights.

Gaussian density

$Y\sim\mathcal N(\mu,\sigma^2)$ has density $\frac{1}{\sqrt{2\pi}\,\sigma}e^{-(y-\mu)^2/(2\sigma^2)}$.

The likelihood of the regression model is a product of these.

Maximum likelihood and the 95% recipe

The MLE maximizes $L$, or equivalently $\log L$. A 95% confidence interval steps $2$ standard errors each side of the estimate, and its 95% refers to repeated samples.

Both return here: the MLE of $\beta$, and intervals for $\beta_0$ and $\beta_1$.

Orthogonality and length

$u^Tv=0$ means $u$ and $v$ are perpendicular; $\lVert u\rVert^2=u^Tu$; if $u^Tv=0$ then $\lVert u+v\rVert^2=\lVert u\rVert^2+\lVert v\rVert^2$.

The geometric reading of least squares is Pythagoras in n dimensions.

Roots of polynomials

A nonzero polynomial of degree at most $p$ has at most $p$ real roots.

It decides when polynomial regression has a unique fit.

Try it yourself first (2 questions)
1§03.3 — transposing a product

Matrix warm-up: $A$ is $3\times 2$ and $B$ is $2\times 4$.

Find(a) Which expression equals $(AB)^T$, and what size is it?
Given
  • $A$: $3\times 2$

  • $B$: $2\times 4$

Hint 1/4

Transposing a product reverses the order of the factors; the sizes confirm it.

Hint 2/4

$(AB)^T=B^TA^T$.

Hint 3/4

$AB$ is $3\times 4$, so its transpose is $4\times 3$; $B^T$ is $4\times 2$ and $A^T$ is $2\times 3$.

Hint 4/4

$(AB)^T=B^TA^T$, a $4\times 3$ matrix.

Show solution

The size bookkeeping catches a wrong order immediately.

Rule

$$(AB)^T=B^TA^T$$

Entry $(i,j)$ of $(AB)^T$ is row $j$ of $A$ times column $i$ of $B$.

Sizes

$$(4\times 2)(2\times 3)=4\times 3$$

Inner dimensions match, outer ones give the size.

Answer $$\boxed{B^TA^T,\ 4\times 3}$$
Check

$AB$ is $3\times 4$, and transposing swaps the two sizes, giving $4\times 3$ again.

This reversal is how $(y-X\beta)^T$ becomes $y^T-\beta^TX^T$ in the least squares derivation.

2§03.2 — a constant plus scaled noise

Probability warm-up: $\varepsilon$ has mean $0$ and variance $\sigma^2$, and $Y=2+3\varepsilon$.

Find(a) What are $E[Y]$ and $\operatorname{Var}(Y)$?
Given
  • $E[\varepsilon]=0$, $\operatorname{Var}(\varepsilon)=\sigma^2$

  • $Y=2+3\varepsilon$

Hint 1/4

The mean and the variance react differently to adding a constant and to scaling.

Hint 2/4

$E[a+bX]=a+bE[X]$ and $\operatorname{Var}(a+bX)=b^2\operatorname{Var}(X)$.

Hint 3/4

Here $a=2$, $b=3$, $E[\varepsilon]=0$ and $\operatorname{Var}(\varepsilon)=\sigma^2$.

Hint 4/4

$E[Y]=2$ and $\operatorname{Var}(Y)=9\sigma^2$.

Show solution

Two one-line rules from the first section.

Mean

$$E[Y]=2+3\cdot 0=2$$

Linearity of expectation.

Variance

$$\operatorname{Var}(Y)=3^2\sigma^2=9\sigma^2$$

The constant 2 does not spread anything.

Answer $$\boxed{E[Y]=2,\quad \operatorname{Var}(Y)=9\sigma^2}$$
Check

Standard deviation check: $\sqrt{9\sigma^2}=3\sigma$, three times the spread of $\varepsilon$, as scaling by $3$ should give.

This is the model of the section in miniature: a fixed part plus noise, with the noise alone carrying the variance.

Notation
symbolreads asmeanswatch out
$Y,\ X=[X_1,\dots,X_p]^T$

Y; the input vector X

the response and the inputs, also called

Capitals are random quantities; the observed values are $y_i$ and $x_i$.

$\beta^{\mathrm{true}}$

beta true

the unknown coefficients of the model that produced the data

Never observed; everything with a hat is computed from data.

$\hat\beta_0,\ \hat\beta_1$

beta zero hat, beta one hat

the least squares intercept and slope

Random: a new sample gives new values.

$D=\{(x_i,y_i)\}_{i=1}^n$

the data set D

$n$ observed pairs

In multiple regression each $x_i$ is a vector.

$S_{xx},\ S_{xy}$

S x x, S x y

$\sum_i(x_i-\bar x)^2$ and $\sum_i(x_i-\bar x)(y_i-\bar y)$

Centered sums; the raw sums $\sum x_i^2$ and $\sum x_iy_i$ are different numbers.

$\hat y_i,\ e_i$

y i hat, e i

the fitted value and the residual $y_i-\hat y_i$

A residual is computed from the fit; the error $\varepsilon_i$ is not observable.

$\varepsilon,\ \sigma^2$

epsilon, sigma squared

the zero-mean noise and its variance

$\sigma$ is estimated by the RSE.

$\mathrm{RSS},\ \mathrm{RSE}$

residual sum of squares, residual standard error

$\sum_ie_i^2$ and $\sqrt{\mathrm{RSS}/(n-2)}$

The $n-2$ belongs to simple regression, with two fitted coefficients.

$x_i=[1,x_{i1},\dots,x_{ip}]^T$

x i

the $i$-th input vector with a leading $1$ for the intercept

Forgetting the 1 drops the intercept from the model.

$X,\ y,\ \hat y$

X, y, y hat

the $n\times(p+1)$ data matrix, the response vector and the fitted vector

$X$ has one row per observation and one column per coefficient.

$\hat\beta_{\mathrm{RSS}}$

beta hat RSS

the minimizer of $\mathrm{RSS}(\beta)$, equal to $(X^TX)^{-1}X^Ty$

Needs full column rank.

$H$

the

$X(X^TX)^{-1}X^T$, the projection onto the column space of $X$

$n\times n$, symmetric, and $H^2=H$.

$\partial h/\partial g$

the Jacobian of h with respect to g

the $m\times n$ matrix of partial derivatives $\partial h_i/\partial g_j$

For a scalar $\alpha$, $\partial\alpha/\partial g$ is a row vector in this course.

$L(\beta),\ l(\beta)$

the likelihood, the log-likelihood

the density of the data as a function of $\beta$, and its natural logarithm

Both have the same maximizer.

$\phi_j(x)$

phi j of x

the $j$-th basis function, a column computed from the inputs

$\phi_0(x)=1$ carries the intercept.

Conventions used here
Input rows and the column of ones.

Every $x_i$ is a column $[1,x_{i1},\dots,x_{ip}]^T$, so $x_i^T\beta$ is a number and the rows of $X$ are the $x_i^T$. With an intercept, the first column of $X$ is all ones.

Most dimension errors come from a missing column of ones or a transpose in the wrong place.

Derivatives of scalars are rows.

Following the lecture, $\partial\alpha/\partial\beta$ of a scalar $\alpha$ is a $1\times(p+1)$ row, and $\partial h/\partial g$ is $m\times n$ with one row per output. Setting a row derivative to zero and transposing gives the same equations as a column gradient.

Some books write gradients as columns; the equations you end with are identical.

Interval multiplier and small samples.

As in the previous section, a 95% interval steps $2$ standard errors each way, with $\sigma$ replaced by the RSE. With only a handful of points this interval is on the narrow side; we still use $2$ so that every example can be done by hand.

Exact small-sample multipliers exist but are not part of this chapter; answers here use 2 unless a question says otherwise.

Linear means linear in the coefficients.

A model is linear regression when $y$ is a linear combination of the coefficients, whatever the columns are: $x$, $x^2$, $\log x$ or $x_1x_2$.

This is what lets polynomial and basis-function models use the same formula.

Residuals versus errors.

$e_i=y_i-\hat y_i$ is computed from the fit; $\varepsilon_i=y_i-(\beta^{\mathrm{true}})^Tx_i$ needs the unknown truth. Facts such as $\sum_ie_i=0$ are about residuals only.

Mixing the two produces false claims such as the noise summing to zero.

Logarithms and rounding.

$\log$ is the natural logarithm. Intermediate steps keep at least four significant figures; coefficients and interval ends are rounded to two or three decimals at the end.

Rounding the RSE early can move an interval end in the second decimal.

3.1Simple linear regression: the line with the smallest residual sum of squares

Picks the intercept and slope that make the squared vertical misses as small as possible; both come from two short formulas.

In the previous section an estimate was one number; now the data are pairs $(x_i,y_i)$ and we need two numbers at once, an intercept and a slope.

Solvable with what we have
  • Estimate one unknown number from data, such as a rate by its MLE $N_1/N$.

  • Minimize a function of one variable by setting its derivative to zero.

  • Average a column of numbers into $\bar x$ or $\bar y$.

Not solvable yet
  • Decide which of two lines through the same points is better.

  • Find a slope and an intercept together from one data set.

  • Predict sales for a number of posts the café has not tried.

A tempting shortcut: average the slopes between neighbouring weeks. For the café they are $2,\,-1,\,3,\,2$, with mean $1.5$. But the sum telescopes: $$\tfrac14\big[(y_2-y_1)+(y_3-y_2)+(y_4-y_3)+(y_5-y_4)\big]=\tfrac14(y_5-y_1).$$

Why it fails

Only the first and last weeks survive. Change week 3's sales from $4$ to $40$ and the average slope is still $1.5$. We need one score that charges a line for missing every point.

TheoremLeast squares for one predictor
Conditions
  • data $D=\{(x_i,y_i)\}_{i=1}^n$ with at least two different $x_i$

  • model $Y\approx\beta_0+\beta_1X$; a candidate line is scored by its residual sum of squares

$$\boxed{\begin{aligned}\mathrm{RSS}(\beta_0,\beta_1)&=\sum_{i=1}^n\big(y_i-(\beta_0+\beta_1x_i)\big)^2\\ \textcolor{#1f6feb}{\hat\beta_1}&=\frac{\sum_i(x_i-\bar x)(y_i-\bar y)}{\sum_i(x_i-\bar x)^2}\\ \textcolor{#1f6feb}{\hat\beta_0}&=\bar y-\hat\beta_1\bar x\end{aligned}}$$

Among all lines, keep the one whose squared vertical misses add up to the least. Its slope is how much $x$ and $y$ move together, divided by how much $x$ moves on its own. The intercept then puts the line through the point of means $(\bar x,\bar y)$.

Derivation: two partial derivatives set to zero

Set the derivative in $\beta_0$ to zero: $$\frac{\partial\,\mathrm{RSS}}{\partial\beta_0}=-2\sum_i\big(y_i-\beta_0-\beta_1x_i\big)=0\ \Longrightarrow\ \beta_0=\bar y-\beta_1\bar x.$$ So the best line passes through $(\bar x,\bar y)$, whatever its slope.

Put this $\beta_0$ back. Each residual becomes $(y_i-\bar y)-\beta_1(x_i-\bar x)$, a function of $\beta_1$ alone.

Set the derivative in $\beta_1$ to zero: $$\sum_i(x_i-\bar x)\big[(y_i-\bar y)-\beta_1(x_i-\bar x)\big]=0\ \Longrightarrow\ \beta_1=\frac{S_{xy}}{S_{xx}}.$$

It is a minimum: after the substitution, $\mathrm{RSS}=S_{yy}-2\beta_1S_{xy}+\beta_1^2S_{xx}$ with $S_{yy}=\sum_i(y_i-\bar y)^2$, an upward parabola in $\beta_1$ because $S_{xx}>0$ when the $x_i$ are not all equal.

Looks like this, but is not

A line whose residuals sum to zero looks like the least squares line: the misses above and below cancel.

Every line through $(\bar x,\bar y)$ has residuals that sum to zero. The café line $2+1.2x$ passes through $(3,5.6)$ and its residuals $-0.2,\, \allowbreak 0.6,\, \allowbreak -1.6,\, \allowbreak 0.2,\, \allowbreak 1$ cancel, yet its RSS is $4.0$, not $3.6$. A zero sum is necessary, not sufficient.

weekxyx − 3y − 5.6product(x − 3)²

1

$1$

$3$

$-2$

$-2.6$

$5.2$

$4$

2

$2$

$5$

$-1$

$-0.6$

$0.6$

$1$

3

$3$

$4$

$0$

$-1.6$

$0$

$0$

4

$4$

$7$

$1$

$1.4$

$1.4$

$1$

5

$5$

$9$

$2$

$3.4$

$6.8$

$4$

sum

$15$

$28$

$0$

$0$

$14$

$10$

The last two column sums are $S_{xy}=14$ and $S_{xx}=10$, so $\hat\beta_1=1.4$. Both centered columns sum to zero, which is a free check on the two means.

Café posts and cups: the least squares line and its RSS

Five weeks of café data: posts per week $x=1,2,3,4,5$ and cups sold, in hundreds, $y=3,5,4,7,9$. Fit the least squares line, then score it and the two lines drawn by eye, $1.5+1.5x$ and $3+x$, by their RSS.

Find$\hat\beta_0$, $\hat\beta_1$ and the RSS of all three lines.
Given
  • $x=1,2,3,4,5$

  • $y=3,5,4,7,9$

  • eye-drawn lines: $1.5+1.5x$ (weeks 1 and 5) and $3+x$ (weeks 2 and 4)

Solution

We center first, so every product stays small; the raw-sum version subtracts two large, nearly equal numbers and invites slips.

Find the means

$$\bar x=\frac{15}{5}=3,\qquad \bar y=\frac{28}{5}=5.6$$

The slope formula works with distances from the means, so they come first.

Build the centered sums

$$S_{xy}=(-2)(-2.6)+(-1)(-0.6)+0+(1)(1.4)+(2)(3.4)=14$$

Each term pairs one week's distance from $\bar x$ with its distance from $\bar y$.

$$S_{xx}=4+1+0+1+4=10$$

Squared distances of the posts from their mean; this is the spread of the input.

Slope, then intercept

$$\hat\beta_1=\frac{S_{xy}}{S_{xx}}=\frac{14}{10}=1.4$$

Co-movement divided by the spread of the input.

$$\hat\beta_0=\bar y-\hat\beta_1\bar x=5.6-1.4\cdot 3=1.4$$

The intercept that puts the line through $(\bar x,\bar y)$.

Score all three lines

$$\mathrm{RSS}=0.2^2+0.8^2+1.6^2+0^2+0.6^2=3.6$$

The residuals of $1.4+1.4x$ are $0.2,\ \allowbreak 0.8,\ \allowbreak -1.6,\ \allowbreak 0,\ \allowbreak 0.6$.

$$\mathrm{RSS}(1.5+1.5x)=4.5,\qquad \mathrm{RSS}(3+x)=6$$

Their residuals are $0,\ \allowbreak 0.5,\ \allowbreak -2,\ \allowbreak -0.5,\ \allowbreak 0$ and $-1,\ 0,\ -2,\ 0,\ 1$.

Answer $$\boxed{\hat y=1.4+1.4x,\qquad \mathrm{RSS}=3.6<4.5<6}$$
Check

The five residuals $0.2+0.8-1.6+0+0.6$ sum to $0$, as they must for a line through $(\bar x,\bar y)$, and the slope $1.4$ sits between the two eye-drawn slopes $1$ and $1.5$.

This answers the opening question: each extra post goes with about $140$ more cups, and no straight line scores below $3.6$ on these five weeks.

Counting posts from the average week: what centering changes

Refit the café data with the input measured from its mean, $u=x-3$, so $u=-2,-1,0,1,2$ and still $y=3,5,4,7,9$. Compare the fit with $\hat y=1.4+1.4x$.

FindThe least squares line in $u$ and its meaning.
Given
  • $u=-2,-1,0,1,2$

  • $y=3,5,4,7,9$

Solution

With $\sum_i u_i=0$ the intercept formula collapses to $\bar y$, so we get the fit almost without arithmetic.

Use the zero mean of u

$$\bar u=0\ \Rightarrow\ \hat\beta_0=\bar y-\hat\beta_1\cdot 0=5.6$$

The slope term drops out of the intercept when the input has mean zero.

Slope from raw sums

$$\hat\beta_1=\frac{\sum_i u_iy_i}{\sum_i u_i^2}=\frac{-6-5+0+7+18}{10}=1.4$$

Centered and raw sums agree when $\bar u=0$, so the raw products are safe here.

Translate back to x

$$5.6+1.4u=5.6+1.4(x-3)=1.4+1.4x$$

Substituting the definition of $u$ recovers the line in the original units.

Answer $$\boxed{\hat y=5.6+1.4u=1.4+1.4x}$$
Check

At $u=0$, the average week, the line gives $5.6$, which is exactly $\bar y$; and the slope matches the previous example to every digit.

Shifting the input moves only the intercept: the slope and every fitted value stay the same, and the intercept now means sales in an average week.

Checkpoint
§03.1 — slope and intercept from summary sums

An engineer logged outdoor temperature $x$ (°C) and a building's heating load $y$ (kW) on four days, and kept only the summary numbers below.

Find(a) Which line is the least squares fit?
Given
  • $n=4$, $\bar x=2$, $\bar y=10$

  • $\sum_i(x_i-\bar x)(y_i-\bar y)=-7.5$

  • $\sum_i(x_i-\bar x)^2=5$

Hint 1/4

Two numbers are needed: the slope from the two sums, then the intercept from the means.

Hint 2/4

$\hat\beta_1=\dfrac{\sum(x_i-\bar x)(y_i-\bar y)}{\sum(x_i-\bar x)^2}$ and $\hat\beta_0=\bar y-\hat\beta_1\bar x$.

Hint 3/4

The sums are $-7.5$ and $5$, with $\bar x=2$ and $\bar y=10$: $\hat\beta_1=-7.5/5$ and $\hat\beta_0=10-\hat\beta_1\cdot 2$.

Hint 4/4

$\hat\beta_1=-1.5$ and $\hat\beta_0=13$, so the line is $\hat y=13-1.5x$.

Show solution

The summary sums are exactly what the two formulas need, so no raw data are missing.

Slope

$$\hat\beta_1=\frac{-7.5}{5}=-1.5$$

Co-movement over spread; the negative sign says load falls as it gets warmer.

Intercept

$$\hat\beta_0=10-(-1.5)(2)=13$$

Subtracting a negative slope times $\bar x$ adds $3$.

Answer $$\boxed{\hat y=13-1.5x}$$
Check

At $x=\bar x=2$ the line gives $13-3=10=\bar y$, as every least squares line with an intercept must.

Slope first, then the intercept through $(\bar x,\bar y)$; the sign of $-\hat\beta_1\bar x$ is where the slips happen.

⚠ Raw sums in place of centered sums

the raw products $x_iy_i$ are quicker to add up than the deviations

wrong$$\hat\beta_1=\frac{\sum_i x_iy_i}{\sum_i x_i^2}=\frac{98}{55}\approx 1.78$$
right$$\hat\beta_1=\frac{\sum_i(x_i-\bar x)(y_i-\bar y)}{\sum_i(x_i-\bar x)^2}=\frac{14}{10}=1.4$$
⚠ Sign slip in the intercept

the formula has a minus sign and $\hat\beta_1$ may itself be negative

wrong$$\hat\beta_0=\bar y+\hat\beta_1\bar x$$
right$$\hat\beta_0=\bar y-\hat\beta_1\bar x$$
01234560246810posts per weekcups (hundreds)0.81.4202468slope bRSSslope 1.4 · RSS 3.6

Every line here passes through the point of means $(3,\,5.6)$; only the slope $b$ changes. The right panel is $\textcolor{#d1690a}{\mathrm{RSS}(b)=3.6+10\,(b-1.4)^2}$, lowest at the least squares slope $\textcolor{#1f6feb}{1.4}$.

At the edges
slope 0.8 RSS 7.2

Too flat: week 1 falls 1.0 below the line and week 5 sits 1.8 above it.

slope 1.4 RSS 3.6

The least squares slope. A step of 0.2 either way costs the same 0.4, because the curve is a parabola.

slope 2.0 RSS 7.2

Too steep: weeks 1 and 2 sit 1.4 above the line and the last two weeks fall below it. It costs exactly as much as 0.8.

3.2How accurate are the estimates: unbiasedness, variance and a 95% interval

Says how far $\hat\beta_0$ and $\hat\beta_1$ would move with new data, and turns that into a $\pm 2$ standard error interval.

The café line came from five particular weeks; five other weeks with the same true relationship would give a different line.

TheoremMean and variance of the least squares estimates
Conditions
  • true model $Y=\beta_0^{\mathrm{true}}+\beta_1^{\mathrm{true}}X+\varepsilon$, with the $x_i$ fixed and known

  • $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$, and the $\varepsilon_i$ are uncorrelated

  • $\sigma$ is estimated by the residual standard error $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$

$$\boxed{\begin{aligned}E[\hat\beta_0]&=\beta_0^{\mathrm{true}},\qquad E[\hat\beta_1]=\beta_1^{\mathrm{true}}\\ \operatorname{Var}(\hat\beta_1)&=\frac{\sigma^2}{\sum_i(x_i-\bar x)^2}\\ \operatorname{Var}(\hat\beta_0)&=\sigma^2\Big[\frac1n+\frac{\bar x^2}{\sum_i(x_i-\bar x)^2}\Big]\\ \text{95\% interval: }&\ \hat\beta_j\pm 2\sqrt{\operatorname{Var}(\hat\beta_j)}\end{aligned}}$$

Averaged over repeated samples, the estimates land on the true values. The slope wobbles less when the inputs are spread out; the intercept wobbles least when the inputs are centered at zero. Replace $\sigma$ by the RSE and step two standard errors each way for a 95% interval.

Why the slope is unbiased, and where its variance comes from

Write the slope as a weighted sum of responses: $$\hat\beta_1=\sum_i k_iy_i,\qquad k_i=\frac{x_i-\bar x}{S_{xx}}.$$ The $\bar y$ term drops out because $\sum_i(x_i-\bar x)=0$.

The weights satisfy $\sum_ik_i=0$ and $\sum_ik_ix_i=1$, so $E[\hat\beta_1]=\sum_ik_i(\beta_0+\beta_1x_i)=\beta_1$.

Uncorrelated noise with one variance gives $\operatorname{Var}(\hat\beta_1)=\sigma^2\sum_ik_i^2=\sigma^2/S_{xx}$.

For the intercept, $\hat\beta_0=\bar y-\hat\beta_1\bar x$ and $\operatorname{Cov}(\bar y,\hat\beta_1)=\frac{\sigma^2}{n}\sum_ik_i=0$, so $\operatorname{Var}(\hat\beta_0)=\frac{\sigma^2}{n}+\bar x^2\,\frac{\sigma^2}{S_{xx}}$.

Looks like this, but is not

Unbiased sounds like accurate: since $E[\hat\beta_1]=\beta_1^{\mathrm{true}}$, the café slope $1.4$ should sit close to the truth.

Unbiased is a promise about the average over many samples, not about your one sample. The café slope has a standard error of about $0.35$, so a miss of that size is ordinary; only more data or more spread in $x$ make a single estimate close.

designinputs xsum of (x − 5)²Var of slopeSE of slope

A

$4,5,5,6$

$2$

$0.5$

$0.707$

B

$3,4,6,7$

$10$

$0.1$

$0.316$

C

$1,3,7,9$

$40$

$0.025$

$0.158$

D

$1,1,9,9$

$64$

$0.0156$

$0.125$

Same four runs, same noise: moving the inputs apart cuts the standard error from $0.707$ to $0.125$, a factor of about $5.7$. Design D bets everything on the line being straight between $1$ and $9$.

How sure is the café about 140 cups per post? 95% intervals for both coefficients

The café fit is $\hat y=1.4+1.4x$ with $\mathrm{RSS}=3.6$ from $n=5$ weeks, $\bar x=3$ and $S_{xx}=10$. Estimate $\sigma$ and give 95% intervals for $\beta_1^{\mathrm{true}}$ and $\beta_0^{\mathrm{true}}$.

FindThe RSE and the two 95% intervals.
Given
  • $\hat\beta_0=1.4$, $\hat\beta_1=1.4$, $\mathrm{RSS}=3.6$

  • $n=5$, $\bar x=3$, $S_{xx}=10$

Solution

$\sigma$ is unknown, so the RSE stands in for it inside the variance formulas; everything else is already on the page.

Estimate the noise level

$$\mathrm{RSE}=\sqrt{\frac{\mathrm{RSS}}{n-2}}=\sqrt{\frac{3.6}{3}}=\sqrt{1.2}\approx 1.095$$

Two coefficients were fitted, so $n-2=3$ residuals are free to vary.

Standard errors

$$\mathrm{SE}(\hat\beta_1)=\frac{1.095}{\sqrt{10}}\approx 0.346$$

The square root of $\sigma^2/S_{xx}$ with $\sigma$ replaced by the RSE.

$$\mathrm{SE}(\hat\beta_0)=1.095\sqrt{\tfrac15+\tfrac{9}{10}}\approx 1.149$$

The intercept pays twice: the $1/n$ term and $\bar x^2/S_{xx}$, because $\bar x=3$ is far from $0$.

Step two standard errors each way

$$1.4\pm 2(0.346)=[0.707,\ 2.093]$$

The lecture's 95% recipe for the slope.

$$1.4\pm 2(1.149)=[-0.898,\ 3.698]$$

The same recipe for the intercept.

Answer $$\boxed{\beta_1^{\mathrm{true}}\in[0.71,\ 2.09],\qquad \beta_0^{\mathrm{true}}\in[-0.90,\ 3.70]}$$
Check

Ratio check: $\mathrm{SE}(\hat\beta_0)/\mathrm{SE}(\hat\beta_1)=\sqrt{S_{xx}/n+\bar x^2}=\sqrt{2+9}\approx 3.32$, and indeed $1.149/0.346\approx 3.32$.

The slope interval stays above $0$, so the five weeks do tie posts to sales; the intercept interval straddles $0$, so they say little about a week with no posts.

Four calibration runs: where to put them to pin down a slope

A sensor's output is calibrated with four runs at temperatures of your choice between $1$ and $9$ °C. The noise is known, $\sigma=0.5$. Compare the standard error of the slope for design A, $x=4,5,5,6$, and design C, $x=1,3,7,9$.

Find$\mathrm{SE}(\hat\beta_1)$ for both designs and their ratio.
Given
  • $\sigma=0.5$ known

  • design A: $x=4,5,5,6$

  • design C: $x=1,3,7,9$

Solution

$\operatorname{Var}(\hat\beta_1)=\sigma^2/S_{xx}$ does not contain the responses, so we can compare designs before running a single measurement.

Spread of each design

$$S_{xx}^{A}=1+0+0+1=2,\qquad S_{xx}^{C}=16+4+4+16=40$$

Both designs have mean $5$; only the distances from $5$ differ.

Standard errors

$$\mathrm{SE}_A=\frac{0.5}{\sqrt 2}\approx 0.354,\qquad \mathrm{SE}_C=\frac{0.5}{\sqrt{40}}\approx 0.079$$

Square root of $\sigma^2/S_{xx}$ for each design.

Compare

$$\frac{\mathrm{SE}_A}{\mathrm{SE}_C}=\sqrt{\frac{40}{2}}=\sqrt{20}\approx 4.47$$

The ratio of standard errors is the square root of the ratio of spreads.

Answer $$\boxed{\mathrm{SE}_A\approx 0.354,\quad \mathrm{SE}_C\approx 0.079,\quad \text{ratio}\approx 4.47}$$
Check

Units check: $\sigma$ is in output units and $\sqrt{S_{xx}}$ in °C, so both standard errors are in output units per °C, the units of a slope.

Spreading the inputs is free precision, as long as the relationship stays straight over the wider range; check that before trusting the far ends.

Checkpoint
§03.2 — spreading the inputs

A study measures reaction time $y$ at four noise levels $x=3,4,6,7$. A second study uses $x=1,3,7,9$: the same mean, with every distance from it doubled. The noise level $\sigma$ is the same.

Find(a) How does $\mathrm{SE}(\hat\beta_1)$ in the second study compare with the first?
Given
  • study 1: $x=3,4,6,7$

  • study 2: $x=1,3,7,9$

  • same $\sigma$ in both

Hint 1/4

The standard error of the slope depends on the inputs only through their spread, $S_{xx}$.

Hint 2/4

$\mathrm{SE}(\hat\beta_1)=\sigma/\sqrt{S_{xx}}$.

Hint 3/4

Study 1 has deviations $-2,-1,1,2$, so $S_{xx}=10$; study 2 has $-4,-2,2,4$, so $S_{xx}=40$.

Hint 4/4

$S_{xx}$ is four times larger, so the standard error is $\sqrt{1/4}=1/2$ as large.

Show solution

Only $S_{xx}$ changes between the studies, so we compare it and nothing else.

Spreads

$$S_{xx}^{(1)}=4+1+1+4=10,\qquad S_{xx}^{(2)}=16+4+4+16=40$$

Both means are $5$; the second set of deviations is twice the first.

Ratio

$$\frac{\mathrm{SE}^{(2)}}{\mathrm{SE}^{(1)}}=\sqrt{\frac{10}{40}}=\frac12$$

$\sigma$ cancels because it is the same in both studies.

Answer $$\boxed{\text{half as large}}$$
Check

Scaling every deviation by $2$ scales $S_{xx}$ by $2^2=4$, and the square root turns that into $2$.

Variances scale with $S_{xx}$, standard errors with its square root: keep track of which one a question asks for.

⚠ Dividing RSS by n instead of n − 2

$n$ is the divisor of an ordinary average, and the $-2$ is easy to forget

wrong$$\mathrm{RSE}=\sqrt{\mathrm{RSS}/n}=\sqrt{3.6/5}\approx 0.85$$
right$$\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}=\sqrt{3.6/3}\approx 1.10$$
⚠ The variance in place of the standard error

the formula box gives variances, and the square root gets dropped on the way to the interval

wrong$$\hat\beta_1\pm 2\operatorname{Var}(\hat\beta_1)=1.4\pm 0.24$$
right$$\hat\beta_1\pm 2\sqrt{\operatorname{Var}(\hat\beta_1)}=1.4\pm 0.69$$

3.3Many inputs at once: least squares in matrix form

Stacks the data into a matrix $X$ and solves all the coefficient equations together: $X^TX\hat\beta=X^Ty$.

With one input we solved two equations by hand; with $p$ inputs there are $p+1$ of them, and matrices keep the bookkeeping to one line.

TheoremLeast squares solution for multiple linear regression
Conditions
  • rows $x_i^T=[1,\,x_{i1},\dots,x_{ip}]$ stacked into the $n\times(p+1)$ matrix $X$, responses into $y$

  • $X$ has full column rank: no column is a combination of the others, which needs $n\ge p+1$

$$\boxed{\begin{aligned}\mathrm{RSS}(\beta)&=(y-X\beta)^T(y-X\beta)\\ X^TX\,\textcolor{#1f6feb}{\hat\beta_{\mathrm{RSS}}}&=X^Ty\\ \textcolor{#1f6feb}{\hat\beta_{\mathrm{RSS}}}&=(X^TX)^{-1}X^Ty\end{aligned}}$$

Stack every residual into one vector; the RSS is its squared length. Setting its derivative to zero gives $p+1$ linear equations, the normal equations, and when the columns of $X$ are independent they have exactly one solution.

Derivation with the course's matrix derivative rules

The rules, where a derivative of a scalar by a column vector is a row: (1) $h=Ag\Rightarrow \partial h/\partial g=A$; (2) $\alpha=y^TAg\Rightarrow \partial\alpha/\partial g=y^TA$ and $\partial\alpha/\partial y=g^TA^T$; (3) $\alpha=y^TAy\Rightarrow \partial\alpha/\partial y=y^T(A+A^T)$. Here $y$ is any vector.

Expand: $\mathrm{RSS}(\beta)=y^Ty-2y^TX\beta+\beta^TX^TX\beta$. The two cross terms $y^TX\beta$ and $\beta^TX^Ty$ are equal, since a number is its own transpose.

Rule 2 on the middle term and rule 3 with $A=X^TX$ on the last: $$\frac{\partial\,\mathrm{RSS}}{\partial\beta}=-2y^TX+\beta^T\big(X^TX+(X^TX)^T\big)=-2y^TX+2\beta^TX^TX.$$

Set it to zero and transpose: $X^TX\beta=X^Ty$. If $X^TXv=0$ then $\lVert Xv\rVert^2=v^TX^TXv=0$, so $Xv=0$ and, with independent columns, $v=0$. So $X^TX$ is invertible and the solution is unique.

It is the minimum: for any $\beta$, $\mathrm{RSS}(\beta)=\mathrm{RSS}(\hat\beta)+\lVert X(\beta-\hat\beta)\rVert^2$, because the cross term is $2(\hat\beta-\beta)^TX^T(y-X\hat\beta)=0$ by the normal equations.

Looks like this, but is not

Cancel the inverse: $(X^TX)^{-1}X^Ty=X^{-1}(X^T)^{-1}X^Ty=X^{-1}y$, so least squares is just $X^{-1}y$.

$X$ is $n\times(p+1)$ with more rows than columns, so it has no inverse, and $(AB)^{-1}=B^{-1}A^{-1}$ needs square, invertible factors. Only when $n=p+1$ is $X^{-1}y$ legal, and then the fit passes through every point with nothing averaged.

The café fit again, in matrix form

Write the café data ($x=1,\dots,5$, $y=3,5,4,7,9$) as $X$ and $y$, and compute $\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$.

Find$\hat\beta_{\mathrm{RSS}}=[\hat\beta_0,\ \hat\beta_1]^T$
Given
  • $x=1,2,3,4,5$

  • $y=3,5,4,7,9$

  • model with an intercept

Solution

With $p=1$ the matrix $X^TX$ is only $2\times 2$, so the closed-form $2\times 2$ inverse is quicker than elimination.

Build X and y

$$X=\begin{bmatrix}1&1\\1&2\\1&3\\1&4\\1&5\end{bmatrix},\qquad y=\begin{bmatrix}3\\5\\4\\7\\9\end{bmatrix}$$

The first column of ones carries the intercept; the second holds the posts.

Form the two products

$$X^TX=\begin{bmatrix}5&15\\15&55\end{bmatrix},\qquad X^Ty=\begin{bmatrix}28\\98\end{bmatrix}$$

The entries are $n$, $\sum x_i$, $\sum x_i^2$ and $\sum y_i$, $\sum x_iy_i$.

Invert and multiply

$$(X^TX)^{-1}=\frac{1}{50}\begin{bmatrix}55&-15\\-15&5\end{bmatrix}$$

The determinant is $5\cdot 55-15^2=50$; swap the diagonal and negate the off-diagonal.

$$\hat\beta=\frac{1}{50}\begin{bmatrix}55\cdot 28-15\cdot 98\\-15\cdot 28+5\cdot 98\end{bmatrix}=\frac{1}{50}\begin{bmatrix}70\\70\end{bmatrix}$$

Row times column, one entry at a time.

Answer $$\boxed{\hat\beta_{\mathrm{RSS}}=\begin{bmatrix}1.4\\1.4\end{bmatrix}}$$
Check

The scalar formulas gave $\hat\beta_1=S_{xy}/S_{xx}=14/10$ and $\hat\beta_0=5.6-1.4\cdot 3$, the same two numbers by a different road.

For one input the matrix route and the centered-sum route are the same computation; the matrix route is the one that keeps working when more columns arrive.

A bakery's two-factor test: three coefficients from one diagonal system

A bakery codes oven temperature as $x_1=-1$ (low) or $+1$ (high) and proofing time as $x_2=-1$ (short) or $+1$ (long), plus one run at the middle settings $(0,0)$. Loaf heights in cm: $(-1,-1)\to 6.0$, $(1,-1)\to 7.2$, $(-1,1)\to 6.8$, $(1,1)\to 8.4$, $(0,0)\to 7.1$. Fit $\hat y=\hat\beta_0+\hat\beta_1x_1+\hat\beta_2x_2$.

Find$\hat\beta$, the residuals and the RSS.
Givenruns $(x_1,x_2,y)$: $(-1,-1,6.0)$, $(1,-1,7.2)$, $(-1,1,6.8)$, $(1,1,8.4)$, $(0,0,7.1)$
Solution

This design makes the three columns of $X$ orthogonal, so $X^TX$ is diagonal and the $3\times 3$ system splits into three one-line divisions.

Form the normal equations

$$X^TX=\begin{bmatrix}5&0&0\\0&4&0\\0&0&4\end{bmatrix},\qquad X^Ty=\begin{bmatrix}35.5\\2.8\\2.0\end{bmatrix}$$

Off-diagonal sums such as $\sum x_{i1}x_{i2}=1-1-1+1+0$ vanish by design.

$$\textstyle\sum x_{i1}y_i=-6.0+7.2-6.8+8.4=2.8$$

The second entry of the right side, written out because its signs are easy to slip.

Solve one equation per row

$$\hat\beta_0=\frac{35.5}{5}=7.1,\quad \hat\beta_1=\frac{2.8}{4}=0.7,\quad \hat\beta_2=\frac{2.0}{4}=0.5$$

A diagonal system is solved entry by entry.

Residuals

$$e=(0.1,\ -0.1,\ -0.1,\ 0.1,\ 0),\qquad \mathrm{RSS}=0.04$$

For example $6.0-(7.1-0.7-0.5)=0.1$ at the first run.

Answer $$\boxed{\hat y=7.1+0.7x_1+0.5x_2,\qquad \mathrm{RSS}=0.04}$$
Check

$X^Te=0$ holds: $\sum e_i=0$, $\sum x_{i1}e_i=-0.1-0.1+0.1+0.1=0$ and $\sum x_{i2}e_i=-0.1+0.1-0.1+0.1=0$.

Going from low to high temperature is $2$ coded units, so it adds about $2\times 0.7=1.4$ cm; coded inputs make that reading immediate.

Rule 3 on a 2 × 2 matrix that is not symmetric

Take $A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$ and $\alpha=y^TAy$ for $y=[y_1,y_2]^T$. Compute $\partial\alpha/\partial y$ directly and compare it with the rule $y^T(A+A^T)$ and with the tempting shortcut $2y^TA$.

Find$\partial\alpha/\partial y$, a $1\times 2$ row.
Given
  • $A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$

  • $\alpha=y^TAy$

Solution

Expanding a 2 × 2 quadratic form takes one line, and it tests the rule instead of trusting it.

Expand the quadratic form

$$\alpha=y_1^2+2y_1y_2+3y_2^2$$

Only $a_{12}=2$ couples the two coordinates, because $a_{21}=0$.

Differentiate entry by entry

$$\frac{\partial\alpha}{\partial y}=\big[\,2y_1+2y_2,\ \ 2y_1+6y_2\,\big]$$

In this course the derivative of a scalar by a column is a row.

$$y^T(A+A^T)=[y_1,\ y_2]\begin{bmatrix}2&2\\2&6\end{bmatrix}=\big[\,2y_1+2y_2,\ \ 2y_1+6y_2\,\big]$$

The rule reproduces the direct answer.

Test the shortcut

$$2y^TA=\big[\,2y_1,\ \ 4y_1+6y_2\,\big]$$

Different entries: $2y^TA$ is only valid when $A=A^T$.

Answer $$\boxed{\frac{\partial}{\partial y}\big(y^TAy\big)=y^T(A+A^T)=\big[\,2y_1+2y_2,\ 2y_1+6y_2\,\big]}$$
Check

At $y=[1,1]^T$: $\alpha(1.001,1)-\alpha(1,1)\approx 0.004$, matching the first entry $2+2=4$ times the step $0.001$.

In the least squares derivation $A=X^TX$ is symmetric, which is the only reason $2\beta^TX^TX$ is allowed there.

Checkpoint
§03.3 — sizes in the normal equations

A data set has $n=200$ customers and $p=3$ inputs (age, income, visits per month), and the model has an intercept.

Find(a) What are the sizes of $X^TX$ and $X^Ty$?
Given
  • $n=200$ rows

  • $p=3$ inputs plus an intercept

Hint 1/4

First fix the size of $X$ itself; the products follow from it.

Hint 2/4

$X$ is $n\times(p+1)$, and an $a\times b$ times a $b\times c$ matrix is $a\times c$.

Hint 3/4

$X$ is $200\times 4$, so $X^T$ is $4\times 200$, and $y$ is $200\times 1$.

Hint 4/4

$X^TX$ is $4\times 4$ and $X^Ty$ is $4\times 1$.

Show solution

Tracking the inner dimensions of each product is faster than picturing the matrices.

Size of X

$$X:\ 200\times(3+1)=200\times 4$$

Three input columns plus the column of ones.

Products

$$X^TX:\ (4\times 200)(200\times 4)=4\times 4$$

The inner dimension 200 is summed away.

$$X^Ty:\ (4\times 200)(200\times 1)=4\times 1$$

Same inner dimension, one column.

Answer $$\boxed{4\times 4\ \text{and}\ 4\times 1}$$
Check

$\hat\beta$ has $p+1=4$ entries, and the normal equations need one equation per entry: a $4\times 4$ system.

The system's size is set by the number of coefficients, never by the number of rows.

⚠ Forgetting the column of ones

the data file lists only the inputs; the intercept column is not in it

wrong$$X=\begin{bmatrix}x_{11}&\cdots&x_{1p}\\ \vdots&&\vdots\\ x_{n1}&\cdots&x_{np}\end{bmatrix}$$
right$$X=\begin{bmatrix}1&x_{11}&\cdots&x_{1p}\\ \vdots&\vdots&&\vdots\\ 1&x_{n1}&\cdots&x_{np}\end{bmatrix}$$
⚠ Transposes in the wrong place

both $X^TX$ and $XX^T$ exist, but $XX^T$ is $n\times n$ and cannot multiply $X^Ty$

wrong$$\hat\beta=(XX^T)^{-1}X^Ty$$
right$$\hat\beta=(X^TX)^{-1}X^Ty$$
⚠ Doubling a quadratic form whose matrix is not symmetric

the shortcut is right for the symmetric $X^TX$ and gets reused where it is not

wrong$$\frac{\partial}{\partial y}\big(y^TAy\big)=2y^TA$$
right$$\frac{\partial}{\partial y}\big(y^TAy\big)=y^T(A+A^T)$$

3.4The geometry of least squares: projecting y onto the columns of X

Shows that $\hat y$ is the point of the column space closest to $y$, so the residual is perpendicular to every column of $X$.

The normal equations $X^T(y-X\hat\beta)=0$ can be read as a picture, and the picture explains why residuals sum to zero.

TheoremLeast squares as an orthogonal projection
Conditions
  • $X$ has full column rank, with columns $w_0=\mathbf 1,\ w_1,\dots,w_p$ in $\mathbb R^n$

  • $e=y-\hat y$ is the residual vector

$$\boxed{\begin{aligned}\textcolor{#1f6feb}{\hat y}&=Hy,\qquad H=X(X^TX)^{-1}X^T\\ X^T\textcolor{#d1690a}{e}&=0\quad\big(w_j^T\textcolor{#d1690a}{e}=0\ \text{for every column}\big)\\ H^T&=H,\qquad H^2=H\end{aligned}}$$

Of all vectors $X\beta$, the fitted vector is the one closest to $y$, reached by dropping a perpendicular. The residual is perpendicular to every column, so with an intercept the residuals sum to zero. Projecting twice changes nothing, which is what $H^2=H$ says.

Why the residual is perpendicular, and why H is a projection

Perpendicular: $X^Te=X^Ty-X^TX\hat\beta=0$ by the normal equations; row $j$ reads $w_j^Te=0$.

Symmetric and idempotent: $H^T=X\big((X^TX)^{-1}\big)^TX^T=H$ because $X^TX$ is symmetric, and $$H^2=X(X^TX)^{-1}\underbrace{X^TX(X^TX)^{-1}}_{I}X^T=H.$$

Closest point: for any $\beta$, $\lVert y-X\beta\rVert^2=\lVert e\rVert^2+\lVert X(\hat\beta-\beta)\rVert^2$ by Pythagoras, since $e$ is perpendicular to $X(\hat\beta-\beta)$.

With an intercept, $w_0=\mathbf 1$ gives $\sum_ie_i=0$. Averaging $y_i=\hat y_i+e_i$ then shows that the fitted line passes through $(\bar x,\bar y)$.

Looks like this, but is not

If the residual is perpendicular to the columns, then in the scatter plot each residual should be perpendicular to the fitted line.

In the scatter plot residuals are vertical segments, and least squares measures them vertically. The right angle is between two vectors in $\mathbb R^n$: the residual vector and a column of $X$. Minimizing perpendicular distances to a line is a different method with a different answer.

Three points and the hat matrix written out in full

Fit $y=\beta_0+\beta_1x$ to the points $(-1,1)$, $(0,2)$, $(1,6)$. Compute $\hat\beta$, $\hat y$, $e$ and the $3\times 3$ matrix $H$, and check $X^Te=0$.

Find$\hat\beta$, $\hat y$, $e$, $H$
Given
  • $x=(-1,0,1)$

  • $y=(1,2,6)$

Solution

The inputs already have mean $0$, so $X^TX$ is diagonal and $H$ can be built from two outer products instead of a full matrix inverse.

Normal equations

$$X^TX=\begin{bmatrix}3&0\\0&2\end{bmatrix},\qquad X^Ty=\begin{bmatrix}9\\5\end{bmatrix}$$

$\sum x_i=0$ kills the off-diagonal; $\sum x_iy_i=-1+0+6=5$.

$$\hat\beta=\begin{bmatrix}3\\2.5\end{bmatrix},\qquad \hat y=\begin{bmatrix}0.5\\3\\5.5\end{bmatrix},\qquad e=\begin{bmatrix}0.5\\-1\\0.5\end{bmatrix}$$

A diagonal system: divide each entry by its diagonal element.

Check perpendicularity

$$X^Te=\begin{bmatrix}0.5-1+0.5\\-0.5+0+0.5\end{bmatrix}=\begin{bmatrix}0\\0\end{bmatrix}$$

The residual is orthogonal to both columns, the ones and the inputs.

Build H

$$H=X\begin{bmatrix}\tfrac13&0\\0&\tfrac12\end{bmatrix}X^T=\tfrac13\mathbf 1\mathbf 1^T+\tfrac12\,xx^T$$

With diagonal $(X^TX)^{-1}$, $H$ splits into one outer product per column.

$$H=\begin{bmatrix}\tfrac56&\tfrac13&-\tfrac16\\ \tfrac13&\tfrac13&\tfrac13\\ -\tfrac16&\tfrac13&\tfrac56\end{bmatrix}$$

Add $\tfrac13$ to every entry, then $\tfrac12x_ix_j$ with $x=(-1,0,1)$.

Answer $$\boxed{\hat y=Hy=(0.5,\ 3,\ 5.5),\qquad e=(0.5,\ -1,\ 0.5)}$$
Check

Every row of $H$ sums to $1$: a constant vector is already a column of $X$, so it projects to itself. And $\operatorname{tr}H=\tfrac56+\tfrac13+\tfrac56=2$, the number of columns.

To test a fit, multiply the residuals by $X^T$; to project anything onto the same columns, reuse $H$ without refitting.

Is this the least squares line? Test the residuals instead of refitting

For the café data ($x=1,\dots,5$, $y=3,5,4,7,9$) three lines are proposed: $1.5+1.5x$, $2+1.2x$ and $1.4+1.4x$. Decide which one is the least squares line using only $X^Te=0$.

FindThe candidate with $\sum e_i=0$ and $\sum x_ie_i=0$.
Given
  • $x=1,2,3,4,5$

  • $y=3,5,4,7,9$

  • candidates $1.5+1.5x$, $2+1.2x$, $1.4+1.4x$

Solution

Two sums per candidate are cheaper than solving the normal equations, and they test the exact property that defines the fit.

First candidate

$$e=(0,\ 0.5,\ -2,\ -0.5,\ 0),\qquad \textstyle\sum e_i=-2\ne 0$$

It fails the first equation already, so it is not the least squares line.

Second candidate

$$e=(-0.2,\ 0.6,\ -1.6,\ 0.2,\ 1),\qquad \textstyle\sum e_i=0$$

It passes through $(\bar x,\bar y)$, so the first test cannot catch it.

$$\textstyle\sum x_ie_i=-0.2+1.2-4.8+0.8+5=2\ne 0$$

The residual is not orthogonal to the input column: the slope is off.

Third candidate

$$e=(0.2,\ 0.8,\ -1.6,\ 0,\ 0.6),\qquad \textstyle\sum e_i=0,\ \ \sum x_ie_i=0.2+1.6-4.8+0+3=0$$

Both normal equations hold, and they have a unique solution here.

Answer $$\boxed{\hat y=1.4+1.4x\ \text{is the least squares line}}$$
Check

Its RSS is $3.6$, below the $4.5$ and $4.0$ of the other two, as the minimum must be.

A zero residual sum only fixes the height of the line; orthogonality to each input column is what fixes the slopes.

Checkpoint
§03.4 — what the normal equations force

A least squares line is fitted to $n$ points $(x_i,y_i)$ with an intercept, giving residuals $e_i=y_i-\hat y_i$. The true errors are $\varepsilon_i=y_i-\beta_0^{\mathrm{true}}-\beta_1^{\mathrm{true}}x_i$.

Find(a) Which of these sums is zero for every data set?
Givenleast squares fit of $y$ on one input $x$, with an intercept
Hint 1/4

Ask what the normal equations say about the residual vector, not about the data.

Hint 2/4

$X^Te=0$: the residuals are orthogonal to every column of $X$.

Hint 3/4

Here the columns are $\mathbf 1$ and $x=(x_1,\dots,x_n)$, so $\sum_ie_i=0$ and $\sum_ix_ie_i=0$.

Hint 4/4

$\sum_ix_ie_i$ is always zero; the other three are not.

Show solution

The normal equations are exactly the statement $X^Te=0$, so we read off its rows.

Write the rows

$$X^Te=\begin{bmatrix}\sum_ie_i\\ \sum_ix_ie_i\end{bmatrix}=\begin{bmatrix}0\\0\end{bmatrix}$$

One row per column of $X$.

Rule out the rest

$$\textstyle\sum_iy_ie_i=\sum_i(\hat y_i+e_i)e_i=\sum_ie_i^2$$

$\hat y$ lies in the column space, so $\sum\hat y_ie_i=0$ and what is left is the RSS.

Answer $$\boxed{\textstyle\sum_i x_ie_i=0}$$
Check

On the café fit: $\sum x_ie_i=0.2+1.6-4.8+0+3=0$, while $\sum e_i^2=3.6$.

Residuals are orthogonal to the columns of $X$, not to $y$, and the unobserved errors obey no such rule.

⚠ Measuring residuals perpendicular to the line

distance to a line in school geometry is the perpendicular one

wrong$$d_i=\frac{\vert y_i-\hat\beta_0-\hat\beta_1x_i\vert}{\sqrt{1+\hat\beta_1^2}}$$
right$$e_i=y_i-\hat\beta_0-\hat\beta_1x_i$$
⚠ Simplifying H to the identity

the inverse of a product looks as if it should split into two inverses

wrong$$H=X(X^TX)^{-1}X^T=XX^{-1}(X^T)^{-1}X^T=I$$
right$$H=X(X^TX)^{-1}X^T,\qquad \operatorname{tr}H=p+1<n$$

3.5Why squares: Gauss-Markov and maximum likelihood

Gives two reasons for least squares: the smallest variance among linear unbiased estimators, and the maximum likelihood answer when the noise is Gaussian.

So far squaring the misses was our choice; this block shows two settings in which that choice is forced on us.

TheoremTwo justifications of least squares
Conditions
  • $y_i=(\beta^{\mathrm{true}})^Tx_i+\varepsilon_i$, with $\beta^{\mathrm{true}}$ fixed and unknown, the $x_i$ fixed and observed

  • Gauss-Markov: $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$, $\operatorname{Cov}(\varepsilon_i,\varepsilon_j)=0$ for $i\ne j$

  • likelihood: in addition, the $\varepsilon_i$ are i.i.d. $\mathcal N(0,\sigma^2)$

$$\boxed{\begin{aligned}&\text{(i) Gauss-Markov: }\operatorname{Var}(\tilde\beta_j)\ge\operatorname{Var}(\textcolor{#1f6feb}{\hat\beta_j})\\ &\qquad\text{for every linear unbiased }\tilde\beta_j=\textstyle\sum_k c_{kj}y_k\\ &\text{(ii) Gaussian noise: }\hat\beta_{\mathrm{MLE}}=\textcolor{#1f6feb}{\hat\beta_{\mathrm{RSS}}}\end{aligned}}$$

Among estimators that are linear in $y$ and right on average, least squares wobbles least: it is the best linear unbiased estimator, . If the noise is also Gaussian, the log-likelihood is a constant minus RSS over $2\sigma^2$, so the most likely $\beta$ is the least squares one, whatever $\sigma$ is.

The likelihood calculation, and a sketch of Gauss-Markov

Each $y_i$ has density $\frac{1}{\sqrt{2\pi}\,\sigma}\exp\!\big(-\frac{(y_i-\beta^Tx_i)^2}{2\sigma^2}\big)$, and independence multiplies them. Taking logs: $$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\,\sigma}-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\beta^Tx_i)^2.$$

The first term has no $\beta$ in it and $\frac{1}{2\sigma^2}>0$, so maximizing $l$ is minimizing RSS. Since $\log$ is increasing, maximizing $L$ and $l$ is the same thing.

Gauss-Markov, sketched as further reading: a linear estimator of $a^T\beta$ is $c^Ty$. It is unbiased for every $\beta$ exactly when $X^Tc=a$.

Split $c=c_0+d$ with $c_0=X(X^TX)^{-1}a$, the least squares weights. Then $X^Td=0$, so $c_0^Td=0$ and $$\operatorname{Var}(c^Ty)=\sigma^2\big(\lVert c_0\rVert^2+\lVert d\rVert^2\big)\ge\sigma^2\lVert c_0\rVert^2.$$

Looks like this, but is not

Gauss-Markov says least squares is the best estimator of $\beta$, so no estimator can have a smaller variance.

The theorem only compares estimators that are linear in $y$ and unbiased. The estimator that always answers $0$ has variance $0$; it is excluded because it is biased, not beaten.

Two café lines scored by log-likelihood, at two noise levels

Assume the café data come from $y_i=\beta_0+\beta_1x_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,\sigma^2)$. Compare $l(\beta)$ for the least squares line $1.4+1.4x$ (RSS $3.6$) and the line $1.5+1.5x$ (RSS $4.5$), first with $\sigma=1$, then with $\sigma=2$.

FindThe four log-likelihoods and which line wins at each $\sigma$.
Given
  • $n=5$

  • $\mathrm{RSS}(1.4,1.4)=3.6$, $\mathrm{RSS}(1.5,1.5)=4.5$

  • $\sigma=1$, then $\sigma=2$

Solution

The log-likelihood depends on $\beta$ only through the RSS, so we never need the individual densities; two RSS values and one constant per $\sigma$ are enough.

The constant term

$$n\log\frac{1}{\sqrt{2\pi}\,\sigma}=5\log\frac{1}{\sqrt{2\pi}}\approx -4.595\quad(\sigma=1)$$

It does not contain $\beta$, so it is shared by every line.

$$5\log\frac{1}{2\sqrt{2\pi}}\approx -8.060\quad(\sigma=2)$$

Doubling $\sigma$ subtracts $5\log 2\approx 3.466$.

Subtract RSS over 2σ²

$$\sigma=1:\ \ l_{\mathrm{LS}}\approx -4.595-1.8=-6.395,\qquad l_{1.5}\approx -4.595-2.25=-6.845$$

Divide each RSS by $2\sigma^2=2$.

$$\sigma=2:\ \ l_{\mathrm{LS}}\approx -8.060-0.45=-8.510,\qquad l_{1.5}\approx -8.060-0.5625=-8.623$$

Now divide by $2\sigma^2=8$.

Answer $$\boxed{\text{least squares wins at both: } -6.395>-6.845,\ \ -8.510>-8.623}$$
Check

The gap in log-likelihood is $(4.5-3.6)/(2\sigma^2)$: $0.45$ at $\sigma=1$ and $0.1125$ at $\sigma=2$, matching the differences above.

$\sigma$ changes how sharply the likelihood prefers one line, never which line it prefers: the maximizer is the RSS minimizer.

Two unbiased slopes for the café weeks, and which one wobbles less

For the café design $x=1,\dots,5$ compare the least squares slope with the end-point slope $\tilde\beta_1=(y_5-y_1)/(x_5-x_1)$ under the Gauss-Markov assumptions.

FindWhether $\tilde\beta_1$ is linear and unbiased, and its variance against $\sigma^2/S_{xx}$.
Given
  • $x=1,2,3,4,5$

  • $\tilde\beta_1=(y_5-y_1)/4$

  • $\operatorname{Var}(\varepsilon_i)=\sigma^2$, uncorrelated

Solution

$\tilde\beta_1$ is the average of neighbouring slopes from the opening, so this is the test that settles the naive attempt.

Linear and unbiased

$$\tilde\beta_1=-\tfrac14y_1+\tfrac14y_5$$

A fixed weighted sum of the $y_k$, so it is linear.

$$E[\tilde\beta_1]=\tfrac14\big[(\beta_0+5\beta_1)-(\beta_0+\beta_1)\big]=\beta_1$$

The intercepts cancel and four slopes remain.

Variances

$$\operatorname{Var}(\tilde\beta_1)=\tfrac{1}{16}(\sigma^2+\sigma^2)=\tfrac{\sigma^2}{8}=0.125\,\sigma^2$$

Uncorrelated terms: variances add, weights enter squared.

$$\operatorname{Var}(\hat\beta_1)=\frac{\sigma^2}{S_{xx}}=\frac{\sigma^2}{10}=0.1\,\sigma^2$$

$S_{xx}=4+1+0+1+4=10$ for this design.

Answer $$\boxed{\operatorname{Var}(\tilde\beta_1)=0.125\,\sigma^2>0.1\,\sigma^2=\operatorname{Var}(\hat\beta_1)}$$
Check

Gauss-Markov predicts exactly this ordering, and on the actual data the two estimates are $1.5$ and $1.4$: close, but only one of them uses all five weeks.

Being unbiased is cheap; among the unbiased linear rules, least squares is the one with the smallest spread.

Checkpoint
§03.5 — what Gauss-Markov needs

A lab fits $y_i=\beta^Tx_i+\varepsilon_i$ by least squares and wants to call the estimate BLUE. It is not sure which properties of its noise it has to check first.

Find(a) Which of these assumptions is not needed for the ?
Givenmodel $y_i=(\beta^{\mathrm{true}})^Tx_i+\varepsilon_i$ with fixed inputs
Hint 1/4

List what the theorem's proof uses: means, variances and covariances of the noise.

Hint 2/4

Gauss-Markov assumes $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$ and $\operatorname{Cov}(\varepsilon_i,\varepsilon_j)=0$ for $i\ne j$.

Hint 3/4

Compare each option with those three conditions; skewed noise can still satisfy all of them.

Hint 4/4

The shape of the distribution is never used, so Gaussian noise is the assumption it does not need.

Show solution

The theorem is stated through first and second moments only, so we check each option against those.

Moments used

$$E[\varepsilon_i]=0,\quad \operatorname{Var}(\varepsilon_i)=\sigma^2,\quad \operatorname{Cov}(\varepsilon_i,\varepsilon_j)=0$$

These give $E[c^Ty]$ and $\operatorname{Var}(c^Ty)=\sigma^2c^Tc$, all the proof needs.

The odd one out

$$\varepsilon_i\sim\mathcal N(0,\sigma^2)\ \text{is not used}$$

Normality enters only the maximum likelihood argument.

Answer $$\boxed{\text{Gaussian noise is not needed}}$$
Check

A skewed, far from Gaussian noise can still have mean $0$, one variance and no correlations, and then least squares is BLUE for it too.

Gauss-Markov lives on means and variances; the Gaussian shape is what the likelihood argument adds.

⚠ Thinking Gauss-Markov needs Gaussian noise

the two names sound alike and both results appear in the same lecture

wrong$$\text{BLUE}\ \Leftarrow\ \varepsilon_i\sim\mathcal N(0,\sigma^2)$$
right$$\text{BLUE}\ \Leftarrow\ E[\varepsilon_i]=0,\ \operatorname{Var}(\varepsilon_i)=\sigma^2,\ \text{uncorrelated}$$
⚠ Letting σ change the maximum likelihood estimate of β

$\sigma$ appears in every term of the likelihood

wrong$$\hat\beta_{\mathrm{MLE}}\ \text{depends on}\ \sigma^2$$
right$$\hat\beta_{\mathrm{MLE}}=\arg\min_\beta\mathrm{RSS}(\beta)\ \text{for every}\ \sigma>0$$
⚠ Losing the minus sign

the RSS term carries a negative coefficient that is easy to drop when copying

wrong$$\arg\max_\beta l(\beta)=\arg\max_\beta\mathrm{RSS}(\beta)$$
right$$\arg\max_\beta l(\beta)=\arg\min_\beta\mathrm{RSS}(\beta)$$

3.6Classification with linear regression: code the classes 0 and 1, cut at 0.5

Fits least squares to 0/1 labels and predicts class 1 where the fitted value exceeds $0.5$; quick, but the outputs are not probabilities.

Least squares does not care whether $y$ is a sales figure or a label, so we can point it at a yes/no problem and see what comes out.

MethodDecision rule from a least squares fit
Conditions
  • labels coded $y_i\in\{0,1\}$

  • $\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$ fitted as before, with rows $x=[1,x_1,\dots,x_p]^T$

$$\boxed{\begin{aligned}\widehat{\text{class}}(x)&=\begin{cases}1,&\textcolor{#1f6feb}{\hat\beta_{\mathrm{RSS}}^Tx}>0.5\\0,&\textcolor{#1f6feb}{\hat\beta_{\mathrm{RSS}}^Tx}\le 0.5\end{cases}\\ \text{boundary: }&\ \{x:\ \hat\beta_{\mathrm{RSS}}^Tx=\textcolor{#d1690a}{0.5}\}\end{aligned}}$$

Fit the line or plane to the zeros and ones, and predict class 1 wherever it lies above one half. Where it equals one half is the boundary, flat in input space. The lecture advises against this in general: fitted values can leave $[0,1]$, and three classes coded $1,2,3$ get an order they may not have.

Looks like this, but is not

A fitted value of $0.8$ reads like an 80% chance of passing.

The fit is a straight line, not a probability model. The same line gives $-0.2$ at $0$ hours and $1.4$ at $8$ hours, values no probability can take. Only the side of $0.5$ is used.

Pass or fail from hours of study: a boundary at 3.5 hours

Six students studied $x=1,2,3,4,5,6$ hours; the outcomes, coded pass $=1$ and fail $=0$, were $y=0,0,1,0,1,1$. Fit least squares, find the decision boundary and count the training mistakes.

FindThe fitted line, the boundary and the number of misclassified students.
Given
  • $x=1,2,3,4,5,6$

  • $y=0,0,1,0,1,1$

  • rule: pass if $\hat y>0.5$

Solution

One input, so the centered-sum formulas are the quickest route; the labels are just numbers to them.

Fit the line

$$\bar x=3.5,\quad \bar y=0.5,\quad S_{xx}=17.5,\quad S_{xy}=3.5$$

$S_{xy}=1.25+0.75-0.25-0.25+0.75+1.25$, pairing each $x_i-3.5$ with $y_i-0.5$.

$$\hat\beta_1=\frac{3.5}{17.5}=0.2,\qquad \hat\beta_0=0.5-0.2\cdot 3.5=-0.2$$

The same two formulas as for any response.

Solve for the boundary

$$-0.2+0.2x=0.5\ \Longrightarrow\ x=3.5$$

The boundary is where the fitted value equals the cut.

Count training mistakes

$$\hat y=(0,\ 0.2,\ 0.4,\ 0.6,\ 0.8,\ 1.0)\ \Rightarrow\ \text{predicted }(0,0,0,1,1,1)$$

Compare each fitted value with $0.5$.

$$\text{actual }(0,0,1,0,1,1):\ \text{students 3 and 4 are misclassified}$$

A straight boundary cannot separate a pass at 3 hours from a fail at 4 hours.

Answer $$\boxed{\hat y=-0.2+0.2x,\qquad \text{pass if } x>3.5,\qquad 2 \text{ of } 6 \text{ wrong}}$$
Check

The boundary sits exactly at $\bar x=3.5$, as it must here: $\bar y=0.5$ and the line passes through $(\bar x,\bar y)$.

The regression only decides which side of the cut a student falls on; its height above or below the cut is not a probability.

Two inputs: the boundary becomes a straight line in the plane

A fit on 0/1 labels for support tickets (urgent $=1$) gave $\hat y=-0.4+0.15x_1+0.1x_2$, with $x_1$ the number of customer replies and $x_2$ the hours open. Write the boundary as a line and classify the tickets $(2,5)$, $(4,4)$ and $(6,0)$.

FindThe boundary line and three classifications.
Given
  • $\hat y=-0.4+0.15x_1+0.1x_2$

  • tickets $(x_1,x_2)$: $(2,5)$, $(4,4)$, $(6,0)$

  • urgent if $\hat y>0.5$

Solution

Setting the fitted value to $0.5$ and solving for $x_2$ gives a line we can sketch and check points against.

Boundary

$$-0.4+0.15x_1+0.1x_2=0.5\ \Longrightarrow\ x_2=9-1.5x_1$$

Move the constant across and divide by the coefficient of $x_2$.

Classify

$$(2,5):\ -0.4+0.3+0.5=0.4\le 0.5\ \Rightarrow\ 0$$

Below the cut, so not urgent; in the plane, $5<9-3=6$.

$$(4,4):\ -0.4+0.6+0.4=0.6>0.5\ \Rightarrow\ 1$$

Above the cut; in the plane, $4>9-6=3$.

$$(6,0):\ -0.4+0.9+0=0.5\ \Rightarrow\ 0$$

Exactly on the boundary; the rule's $\le$ sends it to class $0$.

Answer $$\boxed{x_2=9-1.5x_1;\quad (2,5)\to 0,\ \ (4,4)\to 1,\ \ (6,0)\to 0}$$
Check

Both routes agree for every ticket: the sign of $\hat y-0.5$ and the side of the line $x_2=9-1.5x_1$ give the same labels.

With $p$ inputs the boundary $\hat\beta^Tx=0.5$ is flat: a line for two inputs, a plane for three.

Checkpoint
§03.6 — one student, one prediction

A least squares fit to pass $=1$ and fail $=0$ labels gave $\hat y=-0.2+0.2x$, where $x$ is hours of study. A new student studied $3$ hours.

Find(a) What are the fitted value and the predicted class?
Given
  • $\hat y=-0.2+0.2x$

  • $x=3$ hours

  • pass if $\hat y>0.5$

Hint 1/4

Compute the fitted value first, then compare it with the cut.

Hint 2/4

Predict class $1$ exactly when $\hat\beta^Tx>0.5$.

Hint 3/4

With $x=3$: $\hat y=-0.2+0.2\cdot 3$.

Hint 4/4

$\hat y=0.4\le 0.5$, so the prediction is fail, class $0$.

Show solution

One substitution and one comparison; nothing else is needed.

Substitute

$$\hat y=-0.2+0.6=0.4$$

The intercept is negative, so it lowers the value.

Compare

$$0.4\le 0.5\ \Rightarrow\ \text{class }0$$

Below the cut means fail.

Answer $$\boxed{\hat y=0.4,\ \text{class }0}$$
Check

The boundary of this fit is $x=3.5$ hours, and $3<3.5$ lies on the fail side.

Always compare with the cut that matches the coding: 0.5 for labels 0 and 1.

⚠ Cutting at 0 instead of 0.5

labels coded $-1$ and $+1$ would be cut at $0$; with $0$ and $1$ the cut is halfway

wrong$$\hat\beta^Tx>0\ \Rightarrow\ \text{class }1$$
right$$\hat\beta^Tx>0.5\ \Rightarrow\ \text{class }1$$
⚠ Reading fitted values as probabilities

the labels are 0 and 1, so numbers near them look like probabilities

wrong$$P(\text{pass}\mid x=8)=1.4$$
right$$\hat y(8)=1.4>0.5\ \Rightarrow\ \text{predict pass}$$

3.7Curves with a linear model: interactions, polynomials and basis functions

Adds new columns such as $x^2$, $x_1x_2$ or a bump $\phi_j(x)$ to $X$; the model stays linear in $\beta$, so the same formula fits curves.

So far the model bends nowhere and lets each input act on its own; both limits disappear once we may invent new columns.

RuleLinear basis function models
Conditions
  • basis functions $\phi_0(x)=1,\ \allowbreak \phi_1(x), \allowbreak \dots, \allowbreak \phi_q(x)$ are fixed before fitting

  • polynomial regression of degree $p$: $\phi_j(x)=x^j$; the fit is unique when at least $p+1$ of the $x_i$ are distinct

  • the lecture also lists a sigmoidal basis, $\phi_j(x)=1/\big(1+e^{-\lVert x-\mu_j\rVert_2/s}\big)$

$$\boxed{\begin{aligned}\textcolor{#1f6feb}{\hat y}&=\hat\beta_0+\sum_{j=1}^{q}\hat\beta_j\,\phi_j(x)\\ X_{ij}&=\phi_j(x_i),\qquad \hat\beta=(X^TX)^{-1}X^Ty\\ \text{Gaussian: }\ \phi_j(x)&=\exp\!\Big(-\frac{\lVert x-\mu_j\rVert_2^2}{2s^2}\Big)\end{aligned}}$$

Choose the columns first, then fit exactly as before: each column is a function of the inputs, and $\beta$ still enters linearly. Powers of $x$ give polynomial regression, a product $x_1x_2$ lets one input change the effect of another, and bumps centred at $\mu_j$ with width $s$ give local shapes.

Why distinct inputs decide uniqueness

For polynomial regression $X$ is the with rows $[1,\, \allowbreak x_i,\, \allowbreak x_i^2, \allowbreak \dots, \allowbreak x_i^p]$. If $Xv=0$, the polynomial $v_0+v_1t+\dots+v_pt^p$ vanishes at every $x_i$.

A nonzero polynomial of degree at most $p$ has at most $p$ roots. With $p+1$ distinct $x_i$ it must be the zero polynomial, so $v=0$: full column rank and a unique fit.

With only $p$ distinct inputs, the polynomial with exactly those roots gives $Xv=0$ for some $v\ne 0$: $X^TX$ is singular and many coefficient vectors share the same fitted values.

Looks like this, but is not

$y=\beta_0+\beta_1x+\beta_2x^2$ draws a curve, so it is a nonlinear model and needs a new fitting method.

Linear refers to $\beta$, not to $x$: once $x_i^2$ is computed it is one more number in the table. In $y=\beta_0+e^{\beta_1x}$ the coefficient sits inside the exponential, no column can be computed in advance, and $(X^TX)^{-1}X^Ty$ does not apply.

degreecolumns in XRSS

$0$

$1$

$15.2$

$1$

$2$

$14.8$

$2$

$3$

$0.8$

$3$

$4$

$0.7$

$4$

$5$

$0$

$5$

$6$

no unique fit

RSS never goes up when a column is added, because the smaller model is the bigger one with a coefficient set to $0$. At degree $4$ the curve passes through all five points; at degree $5$ there are six coefficients and only five distinct inputs.

A quadratic through five U-shaped points

Fit $\hat y=\hat\beta_0+\hat\beta_1x+\hat\beta_2x^2$ to $(-2,4)$, $(-1,1)$, $(0,1)$, $(1,1)$, $(2,5)$ and compare its RSS with the straight line's $14.8$.

Find$\hat\beta$ and the RSS of the quadratic.
Given
  • $x=-2,-1,0,1,2$

  • $y=4,1,1,1,5$

  • straight-line fit: $2.4+0.2x$, RSS $14.8$

Solution

The inputs are symmetric about $0$, so every odd power sums to zero and the $3\times 3$ system splits into a $1\times 1$ and a $2\times 2$ piece.

Build the normal equations

$$X^TX=\begin{bmatrix}5&0&10\\0&10&0\\10&0&34\end{bmatrix},\qquad X^Ty=\begin{bmatrix}12\\2\\38\end{bmatrix}$$

Entries are $\sum x_i^k$ for $k=0,\dots,4$: $5,\,0,\,10,\,0,\,34$; the right side is $\sum y_i,\ \sum x_iy_i,\ \sum x_i^2y_i$.

Solve the split system

$$10\hat\beta_1=2\ \Rightarrow\ \hat\beta_1=0.2$$

The middle row stands alone because its off-diagonal entries are $0$.

$$\begin{aligned}5\hat\beta_0+10\hat\beta_2&=12\\10\hat\beta_0+34\hat\beta_2&=38\end{aligned}\ \Rightarrow\ \hat\beta_2=1,\ \ \hat\beta_0=0.4$$

Doubling the first equation and subtracting leaves $14\hat\beta_2=14$.

Residuals

$$\hat y=(4,\ 1.2,\ 0.4,\ 1.6,\ 4.8),\qquad e=(0,\ -0.2,\ 0.6,\ -0.6,\ 0.2)$$

Evaluate $0.4+0.2x+x^2$ at each input.

$$\mathrm{RSS}=0+0.04+0.36+0.36+0.04=0.8$$

Square and add.

Answer $$\boxed{\hat y=0.4+0.2x+x^2,\qquad \mathrm{RSS}=0.8}$$
Check

$X^Te=0$ for all three columns: $\sum e_i=0$, $\sum x_ie_i=0.2-0.6+0.4=0$ and $\sum x_i^2e_i=-0.2-0.6+0.8=0$.

A curve cost one extra column and nothing else: the fitting method did not change.

An interaction term: when proofing time changes what temperature does

Return to the bakery runs $(x_1,x_2,y)$: $(-1,-1,6.0)$, $(1,-1,7.2)$, $(-1,1,6.8)$, $(1,1,8.4)$, $(0,0,7.1)$. Add the column $x_3=x_1x_2$ and fit $\hat y=\hat\beta_0+\hat\beta_1x_1+\hat\beta_2x_2+\hat\beta_3x_1x_2$. How much does one coded unit of temperature add at short and at long proofing?

Find$\hat\beta_3$ and the effect of $x_1$ at $x_2=-1$ and $x_2=+1$.
Given
  • runs $(x_1,x_2,y)$: $(-1,-1,6.0)$, $(1,-1,7.2)$, $(-1,1,6.8)$, $(1,1,8.4)$, $(0,0,7.1)$

  • new column $x_3=x_1x_2$

Solution

The product column is orthogonal to the other three in this design, so the old coefficients stay and only $\hat\beta_3$ is new.

The new column

$$x_3=x_1x_2=(1,\ -1,\ -1,\ 1,\ 0)$$

Multiply the two codes run by run; the centre run gives 0.

Its coefficient

$$\hat\beta_3=\frac{\sum x_{i3}y_i}{\sum x_{i3}^2}=\frac{6.0-7.2-6.8+8.4}{4}=\frac{0.4}{4}=0.1$$

$X^TX$ stays diagonal, $\mathrm{diag}(5,4,4,4)$, so this row solves on its own.

Effect of temperature

$$\frac{\partial\hat y}{\partial x_1}=\hat\beta_1+\hat\beta_3x_2=0.7+0.1x_2$$

The interaction makes the slope in $x_1$ depend on $x_2$.

$$x_2=-1:\ 0.6,\qquad x_2=+1:\ 0.8$$

Short proofing, then long proofing.

Answer $$\boxed{\hat\beta_3=0.1:\ \ 0.6\ \text{cm per unit at short proofing},\ 0.8\ \text{at long}}$$
Check

The four corners are now fitted exactly, for example $7.1-0.7-0.5+0.1=6.0$ at $(-1,-1)$, and the centre run gives $7.1$: RSS falls from $0.04$ to $0$.

With an interaction in the model, never read $\hat\beta_1$ alone as the effect of $x_1$; it is the effect where $x_2=0$.

Two Gaussian bumps as columns: building X and predicting

Use basis functions $\phi_1(x)=e^{-(x-1)^2/2}$ and $\phi_2(x)=e^{-(x-3)^2/2}$ (centres $\mu_1=1$, $\mu_2=3$, $s=1$) plus the constant $\phi_0=1$. Build $X$ for inputs $x=0,1,2,3,4$, then evaluate the fitted model $\hat y=0.5+2\phi_1(x)-\phi_2(x)$ at $x=1.5$.

Find$X$ and $\hat y(1.5)$.
Given
  • $\mu_1=1$, $\mu_2=3$, $s=1$

  • inputs $x=0,1,2,3,4$

  • fit $\hat\beta=(0.5,\ 2,\ -1)$

Solution

Each column is one basis function evaluated at every input, so we fill the matrix column by column.

First bump column

$$\phi_1(0,1,2,3,4)=(0.607,\ 1,\ 0.607,\ 0.135,\ 0.011)$$

$e^{-1/2}\approx 0.607$, $e^{-2}\approx 0.135$, $e^{-9/2}\approx 0.011$.

Second bump column

$$\phi_2(0,1,2,3,4)=(0.011,\ 0.135,\ 0.607,\ 1,\ 0.607)$$

The same values, mirrored about the centre $3$.

Assemble and predict

$$X=\begin{bmatrix}1&0.607&0.011\\1&1&0.135\\1&0.607&0.607\\1&0.135&1\\1&0.011&0.607\end{bmatrix}$$

Column of ones first, then one column per bump.

$$\hat y(1.5)=0.5+2e^{-0.125}-e^{-1.125}\approx 0.5+1.765-0.325=1.940$$

A new input gets its own row $[1,\ \phi_1(1.5),\ \phi_2(1.5)]$.

Answer $$\boxed{\hat y(1.5)\approx 1.94}$$
Check

Order check: $x=1.5$ is close to the first centre and far from the second, so the first bump dominates and $\hat y$ should sit near $0.5+2(0.9)$ minus a little, as it does.

The bumps are fixed before fitting; only their heights $\hat\beta_j$ are learned, which is why the model stays linear.

Checkpoint
§03.7 — repeated inputs and a unique fit

A chemist measured a reaction at temperatures $x=1, \allowbreak 1, \allowbreak 2, \allowbreak 2, \allowbreak 3, \allowbreak 3, \allowbreak 3$: seven runs, but only three different temperatures. She wants to fit a polynomial of degree $p$ by least squares.

Find(a) For which degrees $p$ is the least squares fit unique?
Given
  • $x=1, \allowbreak 1, \allowbreak 2, \allowbreak 2, \allowbreak 3, \allowbreak 3, \allowbreak 3$

  • polynomial model of degree $p$ with an intercept

Hint 1/4

Uniqueness is about how many different input values there are, not how many rows.

Hint 2/4

The Vandermonde matrix has full column rank when at least $p+1$ of the $x_i$ are distinct.

Hint 3/4

There are $3$ distinct values, $1$, $2$ and $3$, so we need $p+1\le 3$.

Hint 4/4

The fit is unique for $p\le 2$.

Show solution

The rank condition for polynomial regression is stated in distinct inputs, so we count those.

Count

$$\{1,2,3\}:\ 3\ \text{distinct values}$$

Repeats add rows but no new information about the shape.

Apply the rule

$$p+1\le 3\ \Rightarrow\ p\le 2$$

At $p=3$ the cubic $(x-1)(x-2)(x-3)$ vanishes at every input, so $X$ loses rank.

Answer $$\boxed{p\le 2}$$
Check

For $p=3$ the columns $1,x,x^2,x^3$ satisfy $x^3=6x^2-11x+6$ at $x=1,2,3$, a linear dependence.

Count distinct inputs before choosing a degree; repeated runs sharpen the estimates but do not raise the ceiling.

⚠ Counting rows instead of distinct inputs

the rank condition sounds like a condition on the sample size

wrong$$n=7\ \text{rows}\ \Rightarrow\ p\le 6$$
right$$3\ \text{distinct}\ x_i\ \Rightarrow\ p\le 2$$
⚠ Reading a main effect alone when an interaction is present

without the product term, $\hat\beta_1$ was the effect of $x_1$, and the habit stays

wrong$$\text{effect of }x_1=\hat\beta_1$$
right$$\text{effect of }x_1=\hat\beta_1+\hat\beta_3x_2$$
−2−101202468input xresponse ydegree 2 · RSS 0.8

Step the degree from $0$ to $4$ on the same five points. Each step adds one column to $X$, and $\textcolor{#d1690a}{\mathrm{RSS}}$ can only go down: $15.2,\ \allowbreak 14.8,\ \allowbreak 0.8,\ \allowbreak 0.7,\ \allowbreak 0$.

At the edges
degree 0 RSS 15.2

Only the column of ones: the fit is the flat line at the mean, 2.4.

degree 4 RSS 0

Five coefficients for five distinct points: the curve passes through all of them. A zero RSS here says nothing about a sixth point.

degree 5 no unique fit

Six coefficients but only five distinct inputs: the normal equations have infinitely many solutions.

Simple regression by hand

A table of pairs, one input, and a question about the line, a prediction or the slope's precision.

  1. Means

    $\bar x$ and $\bar y$.

  2. Centered sums

    $S_{xx}$ and $S_{xy}$ from the deviation columns; each deviation column sums to zero.

  3. Coefficients

    $\hat\beta_1=S_{xy}/S_{xx}$, then $\hat\beta_0=\bar y-\hat\beta_1\bar x$.

  4. Check

    $\sum e_i=0$ and $\sum x_ie_i=0$; then $\mathrm{RSS}=\sum e_i^2$.

  5. Uncertainty

    $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$ and $\hat\beta_1\pm 2\,\mathrm{RSE}/\sqrt{S_{xx}}$.

Where it goes wrong
  • Raw sums $\sum x_iy_i$ where centered ones belong.

  • Dividing the RSS by $n$ instead of $n-2$.

Least squares with a

Two or more inputs, a polynomial or basis model, or any question phrased with $X$ and $y$.

  1. Build X

    A column of ones first, then one column per input or basis function.

  2. Two products

    $X^TX$ and $X^Ty$; the entries are sums such as $\sum_ix_{ij}x_{ik}$ and $\sum_ix_{ij}y_i$.

  3. Rank

    Check that the columns are independent: a nonzero determinant, or enough distinct inputs for a polynomial.

  4. Solve

    Use structure first: orthogonal columns make $X^TX$ diagonal; otherwise invert or eliminate.

  5. Check

    $X^Te=0$, one equation per column.

Where it goes wrong
  • Forgetting the column of ones.

  • Writing $XX^T$ where $X^TX$ belongs.

  • Reading the entries of $\hat\beta$ in a different order from the columns.

Classifying with a least squares fit

Two classes coded 0 and 1, and a question about a boundary or a prediction.

  1. Code

    One class $1$, the other $0$.

  2. Fit

    $\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$ as for any response.

  3. Boundary

    Solve $\hat\beta^Tx=0.5$.

  4. Classify

    Class 1 where $\hat\beta^Tx>0.5$, class 0 otherwise.

  5. Sanity

    Fitted values outside $[0,1]$ can occur and are not probabilities.

Where it goes wrong
  • Cutting at $0$ instead of $0.5$.

  • Coding three classes as $1,2,3$ and trusting the order.

With an intercept: the café slope is 1.4

Fit $y=\beta_0+\beta_1x$ to the café data $x=1,\dots,5$, $y=3,5,4,7,9$.

Find$\hat\beta_1$, $\sum e_i$ and the RSS.
Given
  • $x=1,2,3,4,5$

  • $y=3,5,4,7,9$

Solution

A model with an intercept uses centered sums.

Slope

$$\hat\beta_1=\frac{S_{xy}}{S_{xx}}=\frac{14}{10}=1.4$$

Centered sums, because the intercept absorbs the level.

Residuals

$$e=(0.2,\ 0.8,\ -1.6,\ 0,\ 0.6),\quad \textstyle\sum e_i=0,\quad \mathrm{RSS}=3.6$$

With $\hat\beta_0=1.4$ the line passes through $(3,5.6)$.

Answer $$\boxed{\hat y=1.4+1.4x,\quad \textstyle\sum e_i=0}$$
Check

The line passes through $(\bar x,\bar y)$: $1.4+1.4\cdot 3=5.6=\bar y$.

Through the origin: the same weeks give slope 1.78

Fit $y=\beta x$, with no intercept, to the same café data.

Find$\hat\beta$, $\sum e_i$ and the RSS.
Given
  • $x=1,2,3,4,5$

  • $y=3,5,4,7,9$

Solution

Without an intercept the single normal equation uses raw sums.

Slope

$$\hat\beta=\frac{\sum x_iy_i}{\sum x_i^2}=\frac{98}{55}\approx 1.782$$

$\sum x_iy_i=3+10+12+28+45$ and $\sum x_i^2=55$.

Residuals

$$e\approx(1.22,\ 1.44,\ -1.35,\ -0.13,\ 0.09),\quad \textstyle\sum e_i\approx 1.27,\quad \mathrm{RSS}\approx 5.38$$

Nothing forces the residuals to cancel now.

Answer $$\boxed{\hat y\approx 1.78x,\quad \textstyle\sum e_i\approx 1.27,\quad \mathrm{RSS}\approx 5.38}$$
Check

Orthogonality to the one column still holds: $\sum x_ie_i=98-\hat\beta\cdot 55=0$.

Same five weeks: slope $1.4$ with an intercept, $1.78$ without one. Forcing the line through the origin makes the slope do the intercept's job, and the RSS rises from $3.6$ to $5.38$.

How to tell them apart

Read the model before the formula. A column of ones means centered sums, $S_{xy}/S_{xx}$; no intercept means raw sums, $\sum x_iy_i/\sum x_i^2$, and residuals that need not sum to zero.

Residuals: misses from the fitted line

For the café data and the fitted line $1.4+1.4x$, compute the residuals, their sum and their sum of squares.

Find$\sum e_i$ and $\sum e_i^2$.
Given
  • $x=1,\dots,5$, $y=3,5,4,7,9$

  • fitted line $1.4+1.4x$

Solution

Residuals are computed from quantities we have: the data and the fit.

Residuals

$$e=(0.2,\ 0.8,\ -1.6,\ 0,\ 0.6)$$

Observed minus fitted.

Sums

$$\textstyle\sum e_i=0,\qquad \sum e_i^2=3.6$$

The first normal equation forces the zero sum.

Answer $$\boxed{\textstyle\sum e_i=0,\quad \sum e_i^2=3.6}$$
Check

Both normal equations hold: $\sum e_i=0$ and $\sum x_ie_i=0$.

Errors: misses from the true line

Suppose, as in a simulation, the true line is known to be $1+1.5x$. For the same café data compute $\varepsilon_i=y_i-(1+1.5x_i)$, their sum and their sum of squares.

Find$\sum\varepsilon_i$ and $\sum\varepsilon_i^2$.
Given
  • $x=1,\dots,5$, $y=3,5,4,7,9$

  • true line $1+1.5x$

Solution

Errors are measured from the true line, which only a simulation lets us see.

Errors

$$\varepsilon=(3-2.5,\ 5-4,\ 4-5.5,\ 7-7,\ 9-8.5)=(0.5,\ 1,\ -1.5,\ 0,\ 0.5)$$

Observed minus true.

Sums

$$\textstyle\sum\varepsilon_i=0.5,\qquad \sum\varepsilon_i^2=3.75$$

No equation forces these errors to cancel.

Answer $$\boxed{\textstyle\sum\varepsilon_i=0.5,\quad \sum\varepsilon_i^2=3.75}$$
Check

$3.75\ge 3.6$, as it must be: the least squares line has the smallest RSS of all lines, the true one included.

Residuals come from the fitted line and sum to zero; errors come from the true line and need not. Because least squares minimizes, $\sum e_i^2\le\sum\varepsilon_i^2$ for every data set.

How to tell them apart

If you can compute it from the data alone, it is a residual $e_i$; if it needs $\beta^{\mathrm{true}}$, it is an error $\varepsilon_i$. Squared residuals run small, which is why the RSE divides by $n-2$ rather than $n$.

Scaffolding comes off
The common skeleton
  1. Means: $\bar x$ and $\bar y$.

  2. Centered sums: $S_{xx}=\sum(x_i-\bar x)^2$ and $S_{xy}=\sum(x_i-\bar x)(y_i-\bar y)$.

  3. Coefficients: $\hat\beta_1=S_{xy}/S_{xx}$, then $\hat\beta_0=\bar y-\hat\beta_1\bar x$.

  4. Residual checks: $\sum e_i=0$ and $\sum x_ie_i=0$, then $\mathrm{RSS}=\sum e_i^2$.

  5. Uncertainty: $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$ and $\hat\beta_1\pm 2\,\mathrm{RSE}/\sqrt{S_{xx}}$.

1 · fully worked

Heating load against outdoor temperature: the full skeleton on six days

An office building's daily peak heating load $y$ (kW) and the day's mean outdoor temperature $x$ (°C) over six autumn days: $x=10, \allowbreak 12, \allowbreak 14, \allowbreak 16, \allowbreak 18, \allowbreak 20$ and $y=52, \allowbreak 50, \allowbreak 46, \allowbreak 46, \allowbreak 42, \allowbreak 40$. Fit the line, check it, and give a 95% interval for the slope.

Find$\hat\beta_0$, $\hat\beta_1$, the residual checks and a 95% interval for $\beta_1^{\mathrm{true}}$.
Given
  • $x=10, \allowbreak 12, \allowbreak 14, \allowbreak 16, \allowbreak 18, \allowbreak 20$

  • $y=52, \allowbreak 50, \allowbreak 46, \allowbreak 46, \allowbreak 42, \allowbreak 40$

Solution

All five skeleton steps are needed, and centered deviations keep the numbers small.

Means

$$\bar x=15,\qquad \bar y=46$$

The deviations in the next step are measured from these.

Centered sums

$$x-\bar x=(-5,-3,-1,1,3,5),\qquad y-\bar y=(6,4,0,0,-4,-6)$$

Both columns sum to zero, a check on the means.

$$S_{xy}=-30-12+0+0-12-30=-84,\qquad S_{xx}=25+9+1+1+9+25=70$$

Products of the two deviation columns, and squares of the first.

Coefficients

$$\hat\beta_1=\frac{-84}{70}=-1.2,\qquad \hat\beta_0=46+1.2\cdot 15=64$$

Slope first; the intercept puts the line through the means.

Residual checks

$$\hat y=(52,\ 49.6,\ 47.2,\ 44.8,\ 42.4,\ 40),\qquad e=(0,\ 0.4,\ -1.2,\ 1.2,\ -0.4,\ 0)$$

Observed minus fitted.

$$\textstyle\sum e_i=0,\quad \sum x_ie_i=4.8-16.8+19.2-7.2=0,\quad \mathrm{RSS}=3.2$$

Both normal equations hold, so the coefficients are right.

Uncertainty

$$\mathrm{RSE}=\sqrt{3.2/4}\approx 0.894,\qquad \mathrm{SE}(\hat\beta_1)=\frac{0.894}{\sqrt{70}}\approx 0.107$$

$n-2=4$ degrees of freedom; the slope's spread uses $S_{xx}$.

$$-1.2\pm 2(0.107)=[-1.414,\ -0.986]$$

Two standard errors each way.

Answer $$\boxed{\hat y=64-1.2x,\qquad \beta_1^{\mathrm{true}}\in[-1.41,\ -0.99]}$$
Check

Units: the slope is kW per °C, and a day $10$ °C warmer predicts $12$ kW less, from $52$ at $10$ °C to $40$ at $20$ °C, matching the end points of the table.

Five steps in the same order every time; the two residual sums in step 4 catch slips in steps 1 to 3.

2 · you write the reasoning

An easier one, and this time you write the reasons. Three points: $(0,1)$, $(1,3)$, $(2,2)$. Fit the least squares line and check it.

  1. $\bar x=1,\qquad \bar y=2$

    reasoning

    The slope formula uses deviations from the means, so the means come first.

  2. $S_{xx}=1+0+1=2,\qquad S_{xy}=(-1)(-1)+0\cdot 1+1\cdot 0=1$

    reasoning

    The deviations are $x-\bar x=(-1,0,1)$ and $y-\bar y=(-1,1,0)$: square the first list, multiply the two lists pairwise.

  3. $\hat\beta_1=\tfrac12,\qquad \hat\beta_0=2-\tfrac12\cdot 1=1.5$

    reasoning

    Slope is co-movement over spread; the intercept puts the line through $(\bar x,\bar y)=(1,2)$.

  4. $e=(-0.5,\ 1,\ -0.5)$, $\sum e_i=0$, $\sum x_ie_i=0+1-1=0$, $\mathrm{RSS}=1.5$

    reasoning

    A line with an intercept must leave residuals orthogonal to $\mathbf 1$ and to $x$; if either sum were not zero, a slip happened above.

3 · find the buried error

Harder, with the work done for you and two errors buried in it. A motor is run at voltages $x=2,4,6,8,10$ V and its speed is $y=11, \allowbreak 19, \allowbreak 32, \allowbreak 38, \allowbreak 50$ hundred rpm. Fit the line and give a 95% interval for the slope.

  1. Step 1. $\bar x=6,\qquad \bar y=30$. Means of the two columns.

  2. Step 2. $S_{xx}=40,\qquad S_{xy}=194$. Centered sums from the deviations $(-4,-2,0,2,4)$ and $(-19, \allowbreak -11, \allowbreak 2, \allowbreak 8, \allowbreak 20)$.

  3. Step 3. $\hat\beta_1=194/40=4.85,\qquad \hat\beta_0=30-4.85\cdot 6=0.9$. Slope, then the intercept through the means.

  4. Step 4. $e=(0.4,\ \allowbreak -1.3,\ \allowbreak 2.0,\ \allowbreak -1.7,\ \allowbreak 0.6)$ and $\mathrm{RSS}=9.1$. Both normal equations hold: $\sum e_i=0$ and $\sum x_ie_i=0.8-5.2+12-13.6+6=0$.

  5. Step 5. $\mathrm{RSE}=\sqrt{9.1/5}\approx 1.349$. The typical size of a residual.

  6. Step 6. $\mathrm{SE}(\hat\beta_1)=1.349/\sqrt 5\approx 0.603$. Standard error: spread over the square root of the sample size.

  7. Step 7. $4.85\pm 2(0.603)=[3.64,\ 6.06]$. Two standard errors each way.

the two buried errors (2)
⚠ step 5

The RSE divides the RSS by $n-2=3$, not by $n=5$: two coefficients were fitted from the same five points.

An average divides by $n$, and the RSE looks like the root of an average squared residual.

right

$\mathrm{RSE}=\sqrt{9.1/3}\approx 1.742$.

⚠ step 6

The slope's standard error divides by $\sqrt{S_{xx}}=\sqrt{40}$, not by $\sqrt n$.

$\sigma/\sqrt n$ is the standard error of a sample mean from the previous section, and it gets reused for a slope.

right

$\mathrm{SE}(\hat\beta_1)=1.742/\sqrt{40}\approx 0.275$, so the interval is $4.85\pm 0.551=[4.30,\ 5.40]$.

4 · the bare problem
§03.2 — a courier's minutes per kilometre

A courier logs four deliveries: the distance $x$ in km and the time $y$ in minutes.

Find
  1. (a) Fit the least squares line.

  2. (b) Give a 95% interval for the minutes per km.

Given
  • $x=1,3,5,7$

  • $y=12,17,26,29$

Hint 1/4

Run the whole skeleton: means, centered sums, coefficients, residual checks, then the interval.

Hint 2/4

$\hat\beta_1=S_{xy}/S_{xx}$, $\hat\beta_0=\bar y-\hat\beta_1\bar x$, $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$, interval $\hat\beta_1\pm 2\,\mathrm{RSE}/\sqrt{S_{xx}}$.

Hint 3/4

For $x=1,3,5,7$ and $y=12,17,26,29$: $\bar x=4$, $\bar y=21$, deviations $(-3,-1,1,3)$ and $(-9,-4,5,8)$.

Hint 4/4

$\hat y=9+3x$, $\mathrm{RSS}=6$, $\mathrm{RSE}=\sqrt 3$, interval $3\pm 0.775=[2.23,\ 3.77]$ minutes per km.

Show solution

The same five steps as the ladder, with nothing filled in.

Means and sums

$$\bar x=4,\ \bar y=21,\qquad S_{xy}=27+4+5+24=60,\ \ S_{xx}=20$$

Deviations $(-3,-1,1,3)$ and $(-9,-4,5,8)$.

Coefficients

$$\hat\beta_1=3,\qquad \hat\beta_0=21-12=9$$

Slope, then the intercept through the means.

Residual checks

$$e=(0,\ -1,\ 2,\ -1),\quad \textstyle\sum e_i=0,\quad \sum x_ie_i=0-3+10-7=0$$

Both normal equations hold.

Interval

$$\mathrm{RSE}=\sqrt{6/2}=\sqrt 3,\qquad \mathrm{SE}=\frac{\sqrt 3}{\sqrt{20}}\approx 0.387$$

$n-2=2$ and $S_{xx}=20$.

$$3\pm 0.775=[2.23,\ 3.77]$$

Two standard errors each way.

Answer $$\boxed{\hat y=9+3x,\qquad \beta_1^{\mathrm{true}}\in[2.23,\ 3.77]}$$
Check

Size check: $7$ km at $3$ minutes per km plus $9$ minutes of fixed time is $30$ minutes, next to the $29$ observed.

The intercept reads as fixed time per delivery, the slope as minutes per km; name both in units when you report them.

Full exam-style question

A spring through the origin: derive, test and use the no-intercept slopeexam format

Loads $x=1,2,3,4$ N stretch a spring by $y=2.1, \allowbreak 3.9, \allowbreak 6.2, \allowbreak 7.8$ mm. Zero load gives zero stretch, so fit $y_i=\beta x_i+\varepsilon_i$ with $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$ and uncorrelated noise.

  • (a) Derive the least squares $\hat\beta$.
  • (b) Show that it is unbiased and find its variance.
  • (c) Compute $\hat\beta$ and a 95% interval, taking $\sigma=0.2$ mm as known.
  • (d) Do the residuals sum to zero?
Find$\hat\beta$, its mean and variance, a 95% interval and $\sum_ie_i$.
Given
  • $x=1,2,3,4$ N

  • $y=2.1,\ \allowbreak 3.9,\ \allowbreak 6.2,\ \allowbreak 7.8$ mm

  • $\sigma=0.2$ mm, known, for part (c)

Solution

One derivative gives the estimator, and writing it as a weighted sum of the $y_i$ gives its mean and variance with the rules of the first section.

(a) Derive the estimator

$$\frac{d}{d\beta}\sum_i(y_i-\beta x_i)^2=-2\sum_ix_i(y_i-\beta x_i)=0$$

One coefficient, so one derivative set to zero.

$$\hat\beta=\frac{\sum_ix_iy_i}{\sum_ix_i^2}$$

The second derivative $2\sum_ix_i^2>0$ makes it a minimum.

(b) Mean and variance

$$\hat\beta=\sum_ic_iy_i,\qquad c_i=\frac{x_i}{\sum_jx_j^2}$$

A fixed weighted sum of the responses.

$$E[\hat\beta]=\sum_ic_i\beta x_i=\beta,\qquad \operatorname{Var}(\hat\beta)=\sigma^2\sum_ic_i^2=\frac{\sigma^2}{\sum_ix_i^2}$$

$\sum_ic_ix_i=1$, and uncorrelated terms add their variances.

(c) Numbers

$$\hat\beta=\frac{2.1+7.8+18.6+31.2}{30}=\frac{59.7}{30}=1.99$$

$\sum x_i^2=1+4+9+16=30$.

$$\mathrm{SE}=\frac{0.2}{\sqrt{30}}\approx 0.0365,\qquad 1.99\pm 0.073=[1.917,\ 2.063]$$

$\sigma$ is known here, so no RSE is needed.

(d) Residuals

$$e=(0.11,\ -0.08,\ 0.23,\ -0.16),\qquad \textstyle\sum e_i=0.10\ne 0$$

There is no column of ones, so nothing forces a zero sum.

$$\textstyle\sum x_ie_i=0.11-0.16+0.69-0.64=0$$

The residual is still orthogonal to the one column the model has.

Answer $$\boxed{\hat\beta=\frac{\sum x_iy_i}{\sum x_i^2}=1.99\ \text{mm/N},\quad \operatorname{Var}(\hat\beta)=\frac{\sigma^2}{\sum x_i^2},\quad [1.917,\ 2.063]}$$
Check

Units and size: $1.99$ mm per N predicts $7.96$ mm at $4$ N against $7.8$ observed, and $\sum x_ie_i=0$ confirms the normal equation.

Four parts, one idea each: a derivative, a weighted sum, the arithmetic, and orthogonality.

Drop the intercept and you lose $\sum e_i=0$; what survives is orthogonality to the columns you kept.

Practice

A · concept 4 questions
1§03.1 — a zero residual sum

A classmate offers a quick test for any fitted line: if its residuals add up to zero, it must be the least squares line. You try the claim on the café data.

Find(a) Is the claim true or false? Use the test line.
Given
  • claim: $\sum_ie_i=0$ implies the least squares line

  • café data $x=1,\dots,5$, $y=3,5,4,7,9$; least squares line $1.4+1.4x$ with RSS $3.6$

  • test line: $2+1.2x$

Hint 1/4

One line whose residuals sum to zero but whose RSS exceeds $3.6$ would break the claim.

Hint 2/4

Every line through $(\bar x,\bar y)$ has $\sum e_i=0$; least squares also needs $\sum x_ie_i=0$.

Hint 3/4

The test line gives $\hat y=3.2,\ \allowbreak 4.4,\ \allowbreak 5.6,\ \allowbreak 6.8,\ \allowbreak 8$ against $y=3,5,4,7,9$.

Hint 4/4

Its residuals $-0.2,\ \allowbreak 0.6,\ \allowbreak -1.6,\ \allowbreak 0.2,\ \allowbreak 1$ sum to $0$, yet its RSS is $4.0>3.6$: the claim is false.

Show solution

A claim about every line falls to a single counterexample, so we compute one.

Residuals

$$e=(3-3.2,\ 5-4.4,\ 4-5.6,\ 7-6.8,\ 9-8)=(-0.2,\ 0.6,\ -1.6,\ 0.2,\ 1)$$

Observed minus fitted at each week.

Sum and RSS

$$\textstyle\sum e_i=0,\qquad \mathrm{RSS}=0.04+0.36+2.56+0.04+1=4.0$$

The line passes through $(3,5.6)$, so the sum vanishes; the squares do not shrink with it.

Verdict

$$4.0>3.6\ \Rightarrow\ \text{not least squares}$$

A smaller RSS exists, so this line is not the minimizer.

Answer $$\boxed{\text{False}}$$
Check

The second normal equation fails too: $\sum x_ie_i=-0.2+1.2-4.8+0.8+5=2\ne 0$.

A zero residual sum only says the line passes through the point of means; the slope needs its own equation.

2§03.4 — which distance least squares measures

A figure in a blog post draws a short segment from each data point to the fitted line, at a right angle to the line, and says least squares makes these segments as short as possible.

Find(a) Is the claim true or false?
Givenclaim: least squares minimizes the sum of squared perpendicular distances from the points to the line
Hint 1/4

Check what RSS subtracts: which coordinate of a point is compared with the line?

Hint 2/4

$\mathrm{RSS}=\sum_i\big(y_i-(\beta_0+\beta_1x_i)\big)^2$ compares $y_i$ with the line at the same $x_i$.

Hint 3/4

For the café point $(3,4)$ and the line $1.4+1.4x$ the vertical miss is $-1.6$, while the perpendicular distance is $1.6/\sqrt{1+1.4^2}\approx 0.93$.

Hint 4/4

Least squares squares vertical misses, so the claim is false.

Show solution

One concrete point makes the difference between the two distances visible.

Vertical miss

$$e=4-(1.4+1.4\cdot 3)=-1.6$$

This is the term RSS squares.

Perpendicular distance

$$d=\frac{\vert -1.6\vert}{\sqrt{1+1.4^2}}\approx 0.93$$

The school formula for the distance from a point to a line.

Verdict

$$e^2=2.56\ne d^2\approx 0.86$$

The two methods score the same point differently.

Answer $$\boxed{\text{False}}$$
Check

The vertical and perpendicular distances agree only for a flat line, $\hat\beta_1=0$, where $\sqrt{1+\hat\beta_1^2}=1$.

Least squares treats $x$ as known exactly and charges only the misses in $y$.

3§03.6 — three segments coded 1, 2, 3

A shop codes three customer segments as $y=1$ (students), $y=2$ (families) and $y=3$ (retirees), fits least squares on age, and rounds $\hat y$ to the nearest code.

Find(a) What is the real problem with this plan?
Given
  • codes: students $1$, families $2$, retirees $3$

  • input: customer age

Hint 1/4

Ask what the numbers 1, 2 and 3 claim about the segments beyond naming them.

Hint 2/4

Coding more than two classes on the number line forces an order and a spacing; the lecture names this as a reason not to classify this way.

Hint 3/4

Families are coded as the midpoint of students and retirees, so the fit treats them as halfway between.

Hint 4/4

The coding, not the fitting, is the problem: it invents an order and equal gaps.

Show solution

If codes were only names, relabelling could not change the fit; computing both slopes tests that directly.

Spread of ages

$$\bar x=\tfrac{130}{3},\qquad S_{xx}=6900-3\bar x^2\approx 1266.7$$

$\sum x_i^2=400+1600+4900=6900$.

Two codings

$$(1,2,3):\ S_{xy}=310-260=50,\qquad \hat\beta_1\approx 0.039$$

$\sum x_iy_i=20+80+210=310$ and $n\bar x\bar y=260$.

$$(1,3,2):\ S_{xy}=280-260=20,\qquad \hat\beta_1\approx 0.016$$

Same people, same ages, only the codes of two segments swapped.

Answer $$\boxed{\text{the coding imposes an order; relabelling changes the fit}}$$
Check

The slope moves from $0.039$ to $0.016$ although no data point changed, which is impossible if the codes were mere names.

With more than two classes, number codes smuggle in an order; coding two classes as 0 and 1 is the safe case.

4§03.7 — what linear means

A model for an app's daily active users uses the days since launch $x$: $y=\beta_0+\beta_1\log x+\beta_2x^2+\varepsilon$.

Find(a) True or false: this is a linear regression model that $(X^TX)^{-1}X^Ty$ can fit.
Givenmodel $y=\beta_0+\beta_1\log x+\beta_2x^2+\varepsilon$
Hint 1/4

Look at how the coefficients enter the model, not how $x$ enters.

Hint 2/4

A model is linear regression when $y=\sum_j\beta_j\phi_j(x)+\varepsilon$ with known functions $\phi_j$.

Hint 3/4

Here $\phi_0=1$, $\phi_1(x)=\log x$ and $\phi_2(x)=x^2$, each computable before fitting.

Hint 4/4

Each coefficient multiplies a known column, so the statement is true.

Show solution

Writing one row of $X$ shows at once whether the model fits the linear template.

Identify the columns

$$\phi_0(x)=1,\quad \phi_1(x)=\log x,\quad \phi_2(x)=x^2$$

Each is a fixed function of the data.

One row

$$x=10:\ [\,1,\ \log 10,\ 100\,]\approx[\,1,\ 2.303,\ 100\,]$$

Plain numbers, available before any coefficient is known.

Answer $$\boxed{\text{True}}$$
Check

Linearity test: doubling every coefficient doubles $\beta_0+\beta_1\log x+\beta_2x^2$ at every $x$, which is exactly what linear in $\beta$ means.

Transform the inputs as much as you like; keep the coefficients outside the functions.

B · computation 8 questions
1§03.1 — typing speed after practice

A student logs hours of typing practice $x$ and speed $y$ in words per minute on four days.

Find
  1. (a) Fit the least squares line.

  2. (b) Predict the speed after $3$ hours of practice.

Given
  • $x=1,2,4,5$

  • $y=31,33,43,45$

Hint 1/4

Two coefficients from two centered sums, then one substitution.

Hint 2/4

$\hat\beta_1=S_{xy}/S_{xx}$, $\hat\beta_0=\bar y-\hat\beta_1\bar x$, prediction $\hat\beta_0+\hat\beta_1x$.

Hint 3/4

With $x=1,2,4,5$ and $y=31,33,43,45$: $\bar x=3$, $\bar y=38$, deviations $(-2,-1,1,2)$ and $(-7,-5,5,7)$.

Hint 4/4

$S_{xy}=38$ and $S_{xx}=10$, so $\hat y=26.6+3.8x$, and at $3$ hours the prediction is $38$ wpm.

Show solution

The centered sums are small whole numbers here, so they are the fastest route.

Means

$$\bar x=3,\qquad \bar y=38$$

Needed for the deviations.

Centered sums

$$S_{xy}=14+5+5+14=38,\qquad S_{xx}=4+1+1+4=10$$

Products of $(-2,-1,1,2)$ with $(-7,-5,5,7)$, and squares of the first list.

Coefficients

$$\hat\beta_1=3.8,\qquad \hat\beta_0=38-3.8\cdot 3=26.6$$

Slope, then the intercept through the means.

Predict

$$\hat y(3)=26.6+11.4=38$$

Substitute the new input.

Answer $$\boxed{\hat y=26.6+3.8x,\qquad \hat y(3)=38}$$
Check

$x=3$ is $\bar x$, and a least squares line passes through $(\bar x,\bar y)=(3,38)$, so the prediction must be $38$.

A prediction at the mean input needs no fit at all: it is $\bar y$.

2§03.2 — how sure is the typing slope

Same four days of practice: hours $x=1,2,4,5$, speeds $y=31,33,43,45$, and the fitted line $\hat y=26.6+3.8x$.

Find
  1. (a) Compute the residuals, the RSS and the RSE.

  2. (b) Give 95% intervals for the slope and for the intercept.

Given
  • $x=1,2,4,5$, $y=31,33,43,45$

  • $\hat y=26.6+3.8x$

Hint 1/4

Residuals give the RSS, the RSS gives the estimate of $\sigma$, and that feeds both standard errors.

Hint 2/4

$\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$, $\mathrm{SE}(\hat\beta_1)=\mathrm{RSE}/\sqrt{S_{xx}}$, $\mathrm{SE}(\hat\beta_0)=\mathrm{RSE}\sqrt{1/n+\bar x^2/S_{xx}}$, interval $\pm 2\,\mathrm{SE}$.

Hint 3/4

Fitted values $30.4,\ \allowbreak 34.2,\ \allowbreak 41.8,\ \allowbreak 45.6$ against $31,33,43,45$; $n=4$, $\bar x=3$, $S_{xx}=10$.

Hint 4/4

RSS $=3.6$, RSE $\approx 1.342$; slope $3.8\pm 0.849$, intercept $26.6\pm 2.877$.

Show solution

$\sigma$ is unknown, so the RSE replaces it in both variance formulas.

Residuals and RSS

$$e=(0.6,\ -1.2,\ 1.2,\ -0.6),\qquad \mathrm{RSS}=0.36+1.44+1.44+0.36=3.6$$

Observed minus fitted, then squared and added.

RSE

$$\mathrm{RSE}=\sqrt{3.6/2}=\sqrt{1.8}\approx 1.342$$

$n-2=2$: four points, two fitted coefficients.

Standard errors

$$\mathrm{SE}(\hat\beta_1)=\frac{1.342}{\sqrt{10}}\approx 0.424$$

$\sigma/\sqrt{S_{xx}}$ with the RSE in place of $\sigma$.

$$\mathrm{SE}(\hat\beta_0)=1.342\sqrt{0.25+0.9}\approx 1.439$$

$\bar x^2/S_{xx}=9/10$ dominates the $1/n=0.25$.

Intervals

$$3.8\pm 0.849=[2.95,\ 4.65],\qquad 26.6\pm 2.877=[23.72,\ 29.48]$$

Two standard errors each way.

Answer $$\boxed{\beta_1^{\mathrm{true}}\in[2.95,\ 4.65],\qquad \beta_0^{\mathrm{true}}\in[23.72,\ 29.48]}$$
Check

Ratio check: $\mathrm{SE}(\hat\beta_0)/\mathrm{SE}(\hat\beta_1)=\sqrt{S_{xx}/n+\bar x^2}=\sqrt{2.5+9}\approx 3.39$, and $1.439/0.424\approx 3.39$.

With only two degrees of freedom the RSE is itself shaky; treat both intervals as rough.

3§03.3 — coefficients from the two matrix products

A report on a simple regression with an intercept gives only the two matrix products below.

Find
  1. (a) Find $\hat\beta_{\mathrm{RSS}}$.

  2. (b) Read off $n$, $\bar x$ and $\bar y$, and check the answer with the centered-sum formulas.

Given
  • $X^TX=\begin{bmatrix}4&12\\12&46\end{bmatrix}$

  • $X^Ty=\begin{bmatrix}20\\70\end{bmatrix}$

Hint 1/4

Solve the 2 × 2 normal equations, then decode what the entries of the products mean.

Hint 2/4

$\hat\beta=(X^TX)^{-1}X^Ty$, and the entries are $n,\ \sum x_i,\ \sum x_i^2$ and $\sum y_i,\ \sum x_iy_i$.

Hint 3/4

The determinant is $4\cdot 46-12^2=40$, and $n=4$, $\sum x_i=12$, $\sum x_i^2=46$, $\sum y_i=20$, $\sum x_iy_i=70$.

Hint 4/4

$\hat\beta=(2,\ 1)^T$: the line $\hat y=2+x$.

Show solution

The closed-form 2 × 2 inverse is the shortest path.

Invert

$$(X^TX)^{-1}=\frac{1}{40}\begin{bmatrix}46&-12\\-12&4\end{bmatrix}$$

Swap the diagonal, negate the off-diagonal, divide by the determinant 40.

Multiply

$$\hat\beta=\frac{1}{40}\begin{bmatrix}920-840\\-240+280\end{bmatrix}=\begin{bmatrix}2\\1\end{bmatrix}$$

$46\cdot 20=920$, $12\cdot 70=840$, $12\cdot 20=240$, $4\cdot 70=280$.

Check with centered sums

$$S_{xx}=46-4\cdot 3^2=10,\quad S_{xy}=70-4\cdot 3\cdot 5=10$$

Raw sums minus $n$ times the product of the means.

Answer $$\boxed{\hat\beta_{\mathrm{RSS}}=(2,\ 1)^T}$$
Check

The centered route gives $\hat\beta_1=10/10=1$ and $\hat\beta_0=5-1\cdot 3=2$, the same line.

The first row of $X^TX$ holds $n$ and $\sum x_i$; that alone often decodes a report.

4§03.4 — the hat matrix for three inputs

Three measurements at inputs $x=0,1,2$ gave $y=2,2,5$. The model has an intercept.

Find
  1. (a) Compute $H=X(X^TX)^{-1}X^T$.

  2. (b) Compute $\hat y=Hy$ and the residuals.

  3. (c) Check $\operatorname{tr}H$.

Given
  • $x=0,1,2$

  • $y=2,2,5$

Hint 1/4

Build $X$, invert the $2\times 2$ matrix $X^TX$, then sandwich it between $X$ and $X^T$.

Hint 2/4

$h_{ij}=x_i^T(X^TX)^{-1}x_j$ with rows $x_i=[1,\ x_i]^T$.

Hint 3/4

$X^TX=\begin{bmatrix}3&3\\3&5\end{bmatrix}$, determinant $6$, so $h_{ij}=\frac16(5-3x_i-3x_j+3x_ix_j)$ for $x=0,1,2$.

Hint 4/4

$H$ has rows $(\tfrac56,\tfrac13,-\tfrac16)$, $(\tfrac13,\tfrac13,\tfrac13)$, $(-\tfrac16,\tfrac13,\tfrac56)$; $\hat y=(1.5,3,4.5)$; $\operatorname{tr}H=2$.

Show solution

With only two coefficients, one formula for $h_{ij}$ gives all nine entries.

Invert the 2 × 2 matrix

$$(X^TX)^{-1}=\frac16\begin{bmatrix}5&-3\\-3&3\end{bmatrix}$$

$X^TX$ has entries $n=3$, $\sum x_i=3$, $\sum x_i^2=5$.

Fill H

$$h_{ij}=\tfrac16(5-3x_i-3x_j+3x_ix_j)$$

Expand $[1,\ x_i]\,(X^TX)^{-1}\,[1,\ x_j]^T$.

$$h_{11}=\tfrac56,\ h_{12}=\tfrac13,\ h_{13}=-\tfrac16,\ h_{22}=\tfrac13,\ h_{23}=\tfrac13,\ h_{33}=\tfrac56$$

$H$ is symmetric, so these six entries fix all nine.

Project and check

$$\hat y=Hy=(1.5,\ 3,\ 4.5),\qquad e=(0.5,\ -1,\ 0.5)$$

Row 1: $\tfrac56\cdot 2+\tfrac13\cdot 2-\tfrac16\cdot 5=1.5$.

$$\operatorname{tr}H=\tfrac56+\tfrac13+\tfrac56=2$$

The trace of a projection counts the columns it projects onto.

Answer $$\boxed{\hat y=(1.5,\ 3,\ 4.5),\quad \operatorname{tr}H=2}$$
Check

The centered formulas give $S_{xy}=3$, $S_{xx}=2$, so $\hat y=1.5+1.5x$, which is $(1.5,3,4.5)$ at $x=0,1,2$.

$H$ depends only on the inputs: for $x=-1,0,1$ it is the same matrix, because shifting $x$ leaves the column space unchanged.

5§03.5 — likelihood of two candidate lines

For the typing data ($x=1,2,4,5$, $y=31,33,43,45$) assume i.i.d. Gaussian noise. Two candidate lines are $26.6+3.8x$ and $26+4x$.

Find
  1. (a) Compute $l(\beta)$ for both lines with $\sigma=0.5$.

  2. (b) Repeat with $\sigma=1$. Does the better line change?

Given
  • $x=1,2,4,5$, $y=31,33,43,45$

  • candidates $26.6+3.8x$ and $26+4x$

  • $\sigma=0.5$, then $\sigma=1$

Hint 1/4

The log-likelihood is a constant minus RSS over $2\sigma^2$, so find both RSS values first.

Hint 2/4

$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\,\sigma}-\frac{\mathrm{RSS}(\beta)}{2\sigma^2}$.

Hint 3/4

RSS is $3.6$ for $26.6+3.8x$ and $4$ for $26+4x$; with $n=4$ the constant is $\approx -0.903$ at $\sigma=0.5$ and $\approx -3.676$ at $\sigma=1$.

Hint 4/4

At $\sigma=0.5$: $-8.103$ against $-8.903$; at $\sigma=1$: $-5.476$ against $-5.676$. The least squares line wins both times.

Show solution

$l$ depends on $\beta$ only through the RSS, so two RSS values and two constants do the job.

RSS of the second line

$$e=(1,\ -1,\ 1,\ -1),\qquad \mathrm{RSS}=4$$

$26+4x$ gives $30,34,42,46$.

Constants

$$4\log\frac{1}{0.5\sqrt{2\pi}}\approx -0.903,\qquad 4\log\frac{1}{\sqrt{2\pi}}\approx -3.676$$

$\log\frac{1}{\sqrt{2\pi}}\approx -0.919$ and $\log 2\approx 0.693$.

Log-likelihoods

$$\sigma=0.5:\ -0.903-\tfrac{3.6}{0.5}=-8.103,\quad -0.903-\tfrac{4}{0.5}=-8.903$$

$2\sigma^2=0.5$.

$$\sigma=1:\ -3.676-1.8=-5.476,\quad -3.676-2=-5.676$$

$2\sigma^2=2$.

Answer $$\boxed{26.6+3.8x\ \text{wins at both}\ \sigma}$$
Check

The gap is $(4-3.6)/(2\sigma^2)$: $0.8$ at $\sigma=0.5$ and $0.2$ at $\sigma=1$, matching the differences.

Only the RSS ranks lines; $\sigma$ sets how large the gap looks.

6§03.7 — checking a quadratic fit

A battery's capacity loss $y$ (percent) was measured at four temperature settings coded $x=0,1,2,3$, giving $y=3,1,1,5$. A proposed quadratic fit is $\hat y=3.1-3.9x+1.5x^2$.

Find
  1. (a) Write the Vandermonde matrix $X$ and compute $X^TX$ and $X^Ty$.

  2. (b) Show that the proposed $\hat\beta$ solves the normal equations.

  3. (c) Compute the residuals and the RSS.

Given
  • $x=0,1,2,3$

  • $y=3,1,1,5$

  • proposed $\hat\beta=(3.1,\ -3.9,\ 1.5)$

Hint 1/4

Checking a proposed answer is cheaper than solving: multiply $X^TX$ by it and compare with $X^Ty$.

Hint 2/4

Rows of $X$ are $[1,\ x_i,\ x_i^2]$; the normal equations are $X^TX\hat\beta=X^Ty$.

Hint 3/4

$X^TX=\begin{bmatrix}4&6&14\\6&14&36\\14&36&98\end{bmatrix}$ and $X^Ty=(10,\ 18,\ 50)^T$ for $x=0,1,2,3$, $y=3,1,1,5$.

Hint 4/4

$X^TX\hat\beta=(10,\ 18,\ 50)^T$; residuals $(-0.1,\ 0.3,\ -0.3,\ 0.1)$; RSS $=0.2$.

Show solution

One matrix-vector product settles whether the proposal is the least squares answer.

Build the products

$$X=\begin{bmatrix}1&0&0\\1&1&1\\1&2&4\\1&3&9\end{bmatrix},\qquad X^Ty=\begin{bmatrix}10\\18\\50\end{bmatrix}$$

$\sum y_i=10$, $\sum x_iy_i=1+2+15=18$, $\sum x_i^2y_i=1+4+45=50$.

Multiply

$$4(3.1)+6(-3.9)+14(1.5)=12.4-23.4+21=10$$

First row; the other two rows give 18.6 − 54.6 + 54 = 18 and 43.4 − 140.4 + 147 = 50.

Residuals

$$\hat y=(3.1,\ 0.7,\ 1.3,\ 4.9),\qquad \mathrm{RSS}=0.01+0.09+0.09+0.01=0.2$$

Evaluate the quadratic at each input and subtract.

Answer $$\boxed{X^TX\hat\beta=X^Ty,\qquad \mathrm{RSS}=0.2}$$
Check

The residuals are orthogonal to all three columns: $\sum e_i=0$, $\sum x_ie_i=0.3-0.6+0.3=0$, $\sum x_i^2e_i=0.3-1.2+0.9=0$.

Any claimed least squares answer can be verified with one product; do it before copying the answer into a report.

7§03.7 — one row of a Gaussian basis model

A model uses three Gaussian bumps with centres $0$, $2$, $4$ and width $s=1$, plus a constant: $\hat y=1+2\phi_1(x)-\phi_2(x)+0.5\phi_3(x)$.

Find
  1. (a) Write the row of $X$ for $x=3$.

  2. (b) Compute $\hat y(3)$.

Given
  • $\phi_j(x)=\exp\big(-(x-\mu_j)^2/2\big)$ with $\mu_1=0$, $\mu_2=2$, $\mu_3=4$

  • $\hat\beta=(1,\ 2,\ -1,\ 0.5)$

  • new input $x=3$

Hint 1/4

A new input needs its own row of basis values; the prediction is that row times $\hat\beta$.

Hint 2/4

Row $=[1,\ \phi_1(x),\ \phi_2(x),\ \phi_3(x)]$ and $\hat y=\text{row}\cdot\hat\beta$.

Hint 3/4

At $x=3$: $\phi_1=e^{-9/2}\approx 0.011$, $\phi_2=e^{-1/2}\approx 0.607$, $\phi_3=e^{-1/2}\approx 0.607$.

Hint 4/4

$\hat y(3)\approx 1+0.022-0.607+0.303=0.719$.

Show solution

The model is linear in its coefficients, so prediction is one row times the coefficient vector.

Basis values

$$\phi_1(3)=e^{-9/2}\approx 0.0111,\quad \phi_2(3)=\phi_3(3)=e^{-1/2}\approx 0.6065$$

Squared distances to the centres are $9$, $1$ and $1$.

Dot product

$$\hat y(3)=1+2(0.0111)-0.6065+0.5(0.6065)\approx 0.719$$

Multiply entry by entry and add.

Answer $$\boxed{\hat y(3)\approx 0.719}$$
Check

$x=3$ is equally far from the centres $2$ and $4$, so $\phi_2(3)=\phi_3(3)$, and the far bump at $0$ adds almost nothing.

Basis models predict like any linear model: build the row, then take a dot product.

8§03.7 — an interaction in an ad budget

A fitted sales model with an interaction is $\hat y=6+0.02x_1+0.03x_2+0.001x_1x_2$, where $x_1$ is online ad spend and $x_2$ radio ad spend, both in thousand TL, and $y$ is sales in thousand units.

Find
  1. (a) How much do sales change when $x_1$ rises by $10$ at $x_2=10$?

  2. (b) The same change at $x_2=50$?

  3. (c) What does $\hat\beta_1=0.02$ mean on its own?

Given$\hat y=6+0.02x_1+0.03x_2+0.001x_1x_2$
Hint 1/4

With an interaction the effect of $x_1$ depends on where $x_2$ is, so compute it at each $x_2$.

Hint 2/4

The change in $\hat y$ for $\Delta x_1$ at fixed $x_2$ is $(\hat\beta_1+\hat\beta_3x_2)\,\Delta x_1$.

Hint 3/4

$\hat\beta_1=0.02$, $\hat\beta_3=0.001$, $\Delta x_1=10$, and $x_2=10$ or $50$.

Hint 4/4

$0.3$ thousand units at $x_2=10$, $0.7$ at $x_2=50$; $\hat\beta_1$ alone is the effect when $x_2=0$.

Show solution

The slope in $x_1$ is a formula in $x_2$ here, so we evaluate the formula instead of reading one coefficient.

Slope in x1

$$\frac{\partial\hat y}{\partial x_1}=0.02+0.001\,x_2$$

Differentiate the product term too.

Two radio levels

$$x_2=10:\ 0.03\cdot 10=0.3,\qquad x_2=50:\ 0.07\cdot 10=0.7$$

Multiply the slope by $\Delta x_1=10$.

Meaning of the main effect

$$x_2=0:\ \partial\hat y/\partial x_1=0.02$$

The coefficient alone is the slope where the other input is zero.

Answer $$\boxed{0.3\ \text{at}\ x_2=10,\qquad 0.7\ \text{at}\ x_2=50}$$
Check

Direct check at $x_2=10$: $\hat y(20,10)-\hat y(10,10)=6.9-6.6=0.3$.

Report interaction effects at stated values of the other input, never as one number.

C · exam level 5 questions
1§03.1 — a centred input

An exam asks what happens to the least squares solution when the input is centred, $\sum_ix_i=0$, and what changes when the response is centred too.

Find
  1. (a) Show that $\hat\beta_0=\bar y$ and $\hat\beta_1=\sum_ix_iy_i/\sum_ix_i^2$.

  2. (b) Show that if also $\sum_iy_i=0$, then $\hat\beta_0=0$.

  3. (c) Fit the data and give $\operatorname{Var}(\hat\beta_0)$ in terms of $\sigma^2$.

Given
  • $\sum_ix_i=0$

  • data for part (c): $x=-2,-1,0,1,2$, $y=1,2,2,4,6$

Hint 1/4

Start from the two general formulas and see which terms vanish when $\bar x=0$.

Hint 2/4

$\hat\beta_1=\frac{\sum(x_i-\bar x)(y_i-\bar y)}{\sum(x_i-\bar x)^2}$, $\hat\beta_0=\bar y-\hat\beta_1\bar x$, $\operatorname{Var}(\hat\beta_0)=\sigma^2\big[\frac1n+\frac{\bar x^2}{S_{xx}}\big]$.

Hint 3/4

With $\bar x=0$: $\sum x_i(y_i-\bar y)=\sum x_iy_i-\bar y\sum x_i=\sum x_iy_i$. For the data, $\sum x_iy_i=12$, $\sum x_i^2=10$, $\bar y=3$.

Hint 4/4

$\hat\beta_0=\bar y$ and $\hat\beta_1=\sum x_iy_i/\sum x_i^2$; $\hat\beta_0=0$ when $\bar y=0$; here $\hat y=3+1.2x$ and $\operatorname{Var}(\hat\beta_0)=\sigma^2/5$.

Show solution

Setting $\bar x=0$ in the general formulas is quicker than re-deriving from the RSS.

Intercept

$$\hat\beta_0=\bar y-\hat\beta_1\cdot 0=\bar y$$

The slope term carries a factor $\bar x=0$.

Slope

$$\textstyle\sum x_i(y_i-\bar y)=\sum x_iy_i-\bar y\sum x_i=\sum x_iy_i$$

The $\bar y$ term dies because $\sum x_i=0$; the denominator is $\sum x_i^2$ for the same reason.

Both centred

$$\bar y=0\ \Rightarrow\ \hat\beta_0=0$$

The line then passes through the origin.

The data

$$\hat\beta_1=\frac{-2-2+0+4+12}{10}=1.2,\qquad \hat\beta_0=3$$

$\sum x_iy_i=12$, $\sum x_i^2=10$, $\bar y=15/5$.

$$\operatorname{Var}(\hat\beta_0)=\sigma^2\Big[\frac15+\frac{0}{10}\Big]=\frac{\sigma^2}{5}$$

The $\bar x^2$ term vanishes: the intercept is estimated as well as a plain average.

Answer $$\boxed{\hat y=3+1.2x,\qquad \operatorname{Var}(\hat\beta_0)=\sigma^2/5}$$
Check

Recentre $y$ as $y-3=(-2, \allowbreak -1, \allowbreak -1, \allowbreak 1, \allowbreak 3)$: the slope stays $\sum x_i(y_i-3)/10=12/10$ and the intercept becomes $0$, as (b) says.

Centring decouples intercept and slope and makes the intercept an average, with the smallest possible variance $\sigma^2/n$.

2§03.3 — two columns and no intercept

A sensor's output $y$ is modelled as $y=\beta_1x+\beta_2x^2+\varepsilon$ with no intercept, because zero input gives zero output. Calibration inputs are $x=-2,-1,1,2$ with outputs $y=1.1, \allowbreak -0.8, \allowbreak 2.4, \allowbreak 7.0$.

Find
  1. (a) Write $X$ and the normal equations.

  2. (b) Solve for $\hat\beta$.

  3. (c) Check that the residuals are orthogonal to both columns. Must they sum to zero?

Given
  • model $y=\beta_1x+\beta_2x^2+\varepsilon$

  • $x=-2,-1,1,2$

  • $y=1.1,\ \allowbreak -0.8,\ \allowbreak 2.4,\ \allowbreak 7.0$

Hint 1/4

Two columns, $x$ and $x^2$, and no column of ones: the normal equations are a $2\times 2$ system.

Hint 2/4

$X^TX\hat\beta=X^Ty$ with $X^TX=\begin{bmatrix}\sum x_i^2&\sum x_i^3\\ \sum x_i^3&\sum x_i^4\end{bmatrix}$.

Hint 3/4

For $x=-2,-1,1,2$: $\sum x_i^2=10$, $\sum x_i^3=0$, $\sum x_i^4=34$; with $y=1.1, \allowbreak -0.8, \allowbreak 2.4, \allowbreak 7.0$: $\sum x_iy_i=15$, $\sum x_i^2y_i=34$.

Hint 4/4

$\hat\beta=(1.5,\ 1)$; residuals $(0.1,-0.3,-0.1,0)$ are orthogonal to $x$ and $x^2$ but sum to $-0.3$.

Show solution

The symmetric inputs make $\sum x_i^3=0$, so the system is diagonal and each coefficient is one division.

Design matrix

$$X=\begin{bmatrix}-2&4\\-1&1\\1&1\\2&4\end{bmatrix},\qquad X^TX=\begin{bmatrix}10&0\\0&34\end{bmatrix}$$

Columns $x$ and $x^2$ only; the model has no intercept.

Right-hand side

$$X^Ty=\begin{bmatrix}-2.2+0.8+2.4+14\\4.4-0.8+2.4+28\end{bmatrix}=\begin{bmatrix}15\\34\end{bmatrix}$$

Rows of $X^T$ times $y$.

Solve

$$\hat\beta_1=\frac{15}{10}=1.5,\qquad \hat\beta_2=\frac{34}{34}=1$$

One division per diagonal entry.

Residuals

$$\hat y=(1,\ -0.5,\ 2.5,\ 7),\qquad e=(0.1,\ -0.3,\ -0.1,\ 0)$$

For example $1.5(-2)+1\cdot 4=1$ at the first input.

Answer $$\boxed{\hat\beta=(1.5,\ 1),\qquad \textstyle\sum e_i=-0.3\ne 0}$$
Check

Orthogonality holds: $\sum x_ie_i=-0.2+0.3-0.1+0=0$ and $\sum x_i^2e_i=0.4-0.3-0.1+0=0$.

The normal equations promise orthogonality to the columns you included, and nothing more.

3§03.6 — a boundary in two inputs

A fit on 0/1 labels gave $\hat y=-1+0.5x_1+0.25x_2$, and the rule predicts class 1 when $\hat y>0.5$.

Find(a) Which line in the $(x_1,x_2)$ plane is the decision boundary?
Given
  • $\hat y=-1+0.5x_1+0.25x_2$

  • class 1 if $\hat y>0.5$

Hint 1/4

The boundary is where the fitted value equals the cut; solve that equation for $x_2$.

Hint 2/4

Boundary: $\hat\beta^Tx=0.5$.

Hint 3/4

$-1+0.5x_1+0.25x_2=0.5$.

Hint 4/4

$0.25x_2=1.5-0.5x_1$, so $x_2=6-2x_1$.

Show solution

Putting the cut on the right-hand side first keeps the constant's sign straight.

Move the constant

$$0.25x_2=0.5+1-0.5x_1=1.5-0.5x_1$$

Add 1 to both sides and subtract the $x_1$ term.

Divide

$$x_2=6-2x_1$$

Divide every term by 0.25.

Answer $$\boxed{x_2=6-2x_1}$$
Check

The point $(2,2)$ lies on it and gives $-1+1+0.5=0.5$, exactly the cut.

Write the boundary equation with the cut on the right-hand side before moving anything.

4§03.5 — a two-group slope against least squares

With six equally spaced inputs $x=1,\dots,6$, a quick estimator splits the data into a low half ($x=1,2,3$) and a high half ($x=4,5,6$) and uses $\tilde\beta_1=(\bar y_{\mathrm{high}}-\bar y_{\mathrm{low}})/(5-2)$.

Find
  1. (a) Show that $\tilde\beta_1$ is linear and unbiased.

  2. (b) Compute $\operatorname{Var}(\tilde\beta_1)$ and $\operatorname{Var}(\hat\beta_1)$.

  3. (c) Which is smaller, and which theorem predicts it?

Given
  • $x=1,2,3,4,5,6$

  • $\tilde\beta_1=(\bar y_{\mathrm{high}}-\bar y_{\mathrm{low}})/3$

  • Gauss-Markov assumptions with variance $\sigma^2$

Hint 1/4

Write $\tilde\beta_1$ as a weighted sum of the six $y_i$; linearity and unbiasedness follow from the weights.

Hint 2/4

$E[\sum c_iy_i]=\sum c_i(\beta_0+\beta_1x_i)$ and $\operatorname{Var}(\sum c_iy_i)=\sigma^2\sum c_i^2$.

Hint 3/4

The weights are $-\tfrac19$ on $x=1,2,3$ and $+\tfrac19$ on $x=4,5,6$; for least squares $S_{xx}=17.5$.

Hint 4/4

$\operatorname{Var}(\tilde\beta_1)=\tfrac{2}{27}\sigma^2\approx 0.074\sigma^2$, larger than $\sigma^2/17.5\approx 0.057\sigma^2$, as Gauss-Markov predicts.

Show solution

Both estimators are weighted sums of the responses, so the same two rules settle everything.

Weights

$$\tilde\beta_1=\sum_ic_iy_i,\qquad c=\tfrac19(-1,-1,-1,1,1,1)$$

$\bar y_{\mathrm{high}}-\bar y_{\mathrm{low}}$ divided by $3$.

Unbiased

$$\textstyle\sum c_i=0,\qquad \sum c_ix_i=\tfrac19(15-6)=1$$

So $E[\tilde\beta_1]=0\cdot\beta_0+1\cdot\beta_1$.

Variances

$$\operatorname{Var}(\tilde\beta_1)=\sigma^2\cdot\tfrac{6}{81}=\tfrac{2}{27}\sigma^2\approx 0.0741\,\sigma^2$$

Six weights of size one ninth, squared and added.

$$S_{xx}=6.25+2.25+0.25+0.25+2.25+6.25=17.5,\quad \operatorname{Var}(\hat\beta_1)\approx 0.0571\,\sigma^2$$

Deviations from $\bar x=3.5$.

Answer $$\boxed{\tfrac{2}{27}\sigma^2>\tfrac{\sigma^2}{17.5}}$$
Check

Ratio $0.0741/0.0571\approx 1.30$: the quick estimator needs about $1.3$ times as many observations for the same precision.

Many sensible slopes are unbiased; Gauss-Markov picks the one with the least spread, and it is least squares.

5§03.3 — a matrix solution with one slip

A student fits a line to $x=1,2,3,4$ and $y=3,5,4,6$ in matrix form. The steps are listed below; exactly one of them is wrong.

Find(a) Which step is wrong?
Given
  • Step 1. $X^TX=\begin{bmatrix}4&10\\10&30\end{bmatrix}$, $X^Ty=\begin{bmatrix}18\\49\end{bmatrix}$

  • Step 2. $\det(X^TX)=4\cdot 30-10\cdot 10=20$

  • Step 3. $(X^TX)^{-1}=\frac{1}{20}\begin{bmatrix}30&-10\\-10&4\end{bmatrix}$

  • Step 4. $\hat\beta=\tfrac{1}{20}\,(540-490,\ -180+196)^T=(2.5,\ 0.8)^T$

  • Step 5. So the fitted line is $\hat y=0.8+2.5x$.

Hint 1/4

Recompute each step on its own; one of them misreads what was computed rather than miscalculating.

Hint 2/4

$\hat\beta=[\hat\beta_0,\ \hat\beta_1]^T$ in the order of the columns of $X$: ones first, then $x$.

Hint 3/4

Step 4 gives $\hat\beta=(2.5,\ 0.8)$ for the columns $(\mathbf 1,\ x)$.

Hint 4/4

The line is $\hat y=2.5+0.8x$: Step 5 swapped the intercept and the slope.

Show solution

Recomputing each step independently isolates the slip without redoing the whole fit twice.

Steps 1 to 4

$$n=4,\ \textstyle\sum x_i=10,\ \sum x_i^2=30,\ \sum y_i=18,\ \sum x_iy_i=49$$

The products, the determinant 20, the inverse and the product all check out.

Step 5

$$\hat\beta_0=2.5,\ \ \hat\beta_1=0.8\ \Rightarrow\ \hat y=2.5+0.8x$$

The first entry multiplies the column of ones, so it is the intercept.

Answer $$\boxed{\text{Step 5};\ \ \hat y=2.5+0.8x}$$
Check

Centered route: $\bar x=2.5$, $\bar y=4.5$, $S_{xx}=5$, $S_{xy}=4$, so $\hat\beta_1=0.8$ and $\hat\beta_0=4.5-2=2.5$.

Keep the order of $\hat\beta$ tied to the order of the columns of $X$.

D · interleaved 4 questions
1§03.2 — five weeks, one interval

From five weeks of data a café computes the 95% interval $[0.71,\ 2.09]$ for $\beta_1^{\mathrm{true}}$, in hundreds of cups per post, using $\hat\beta_1\pm 2\,\mathrm{SE}$.

Find(a) Which sentence reads the interval correctly?
Given
  • interval $[0.71,\ 2.09]$ from $\hat\beta_1=1.4$ and $\mathrm{SE}\approx 0.35$

  • $\beta_1^{\mathrm{true}}$ is a fixed, unknown number

Hint 1/4

Ask what is random in this recipe: the interval, or the true slope?

Hint 2/4

A 95% confidence interval is a recipe that, over repeated samples, covers the fixed true value about 95 times in 100.

Hint 3/4

Here the recipe is $\hat\beta_1\pm 2\,\mathrm{SE}$ with $\hat\beta_1=1.4$ and $\mathrm{SE}\approx 0.35$, and $\beta_1^{\mathrm{true}}$ does not move.

Hint 4/4

About 95 in 100 intervals built this way catch the true slope; that is the correct reading.

Show solution

Pinning down which quantity is random settles which sentence can carry the 95%.

What is random

$$\hat\beta_1,\ \mathrm{SE}\ \text{vary with the sample};\quad \beta_1^{\mathrm{true}}\ \text{does not}$$

The interval moves from study to study; the target stays put.

Frequentist reading

$$\Pr\big(\hat\beta_1-2\,\mathrm{SE}\le\beta_1^{\mathrm{true}}\le\hat\beta_1+2\,\mathrm{SE}\big)\approx 0.95$$

The probability is over repeated samples, as in the previous section.

Answer $$\boxed{\text{about 95 in 100 such intervals contain }\beta_1^{\mathrm{true}}}$$
Check

A probability statement about $\beta_1^{\mathrm{true}}$ given this one data set would need a posterior over it, which this recipe never builds.

In regression, as with rates, the 95% belongs to the procedure, not to one interval.

2§03.2 — why the fitted lines pivot

Fitted lines from repeated samples tilt around a pivot: when the slope comes out too steep, the intercept tends to come out too low. You want the number behind that picture.

Find
  1. (a) Show that $\operatorname{Cov}(\bar y,\hat\beta_1)=0$.

  2. (b) Show that $\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-\bar x\,\sigma^2/S_{xx}$.

  3. (c) Evaluate the covariance and the correlation for the café design.

Given
  • $\hat\beta_1=\sum_ik_iy_i$ with $k_i=(x_i-\bar x)/S_{xx}$

  • $\hat\beta_0=\bar y-\hat\beta_1\bar x$

  • $y_i$ uncorrelated with variance $\sigma^2$

  • café design $x=1,2,3,4,5$

Hint 1/4

Everything is a linear combination of the $y_i$, so the covariance rule for sums does all the work.

Hint 2/4

For uncorrelated $y_i$ with variance $\sigma^2$: $\operatorname{Cov}(\sum a_iy_i,\sum b_iy_i)=\sigma^2\sum a_ib_i$.

Hint 3/4

$\bar y$ has weights $\tfrac1n$ and $\hat\beta_1$ has weights $k_i$ with $\sum k_i=0$; for the café, $\bar x=3$, $S_{xx}=10$, $n=5$.

Hint 4/4

$\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-0.3\,\sigma^2$ and the correlation is $-3/\sqrt{11}\approx -0.90$.

Show solution

Writing both estimators as weighted sums turns every covariance into one sum of weight products.

Mean and slope are uncorrelated

$$\operatorname{Cov}(\bar y,\hat\beta_1)=\sigma^2\sum_i\tfrac1n\,k_i=\tfrac{\sigma^2}{n}\cdot 0=0$$

The slope weights sum to zero.

Intercept and slope

$$\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=\operatorname{Cov}(\bar y,\hat\beta_1)-\bar x\operatorname{Var}(\hat\beta_1)=-\frac{\bar x\,\sigma^2}{S_{xx}}$$

Covariance is linear in each argument.

Café numbers

$$-\frac{3\sigma^2}{10}=-0.3\,\sigma^2,\qquad \rho=\frac{-\bar x}{\sqrt{S_{xx}/n+\bar x^2}}=\frac{-3}{\sqrt{11}}\approx -0.90$$

Divide by $\sqrt{\operatorname{Var}(\hat\beta_0)\operatorname{Var}(\hat\beta_1)}$; $\sigma^2$ cancels.

Answer $$\boxed{\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-0.3\,\sigma^2,\qquad \rho\approx -0.90}$$
Check

Sign check: with $\bar x>0$ every fitted line passes near $(\bar x,\bar y)$, so a steeper line must start lower at $x=0$, which is a negative covariance.

The pivot of the sampling picture is the point of means; centring the input ($\bar x=0$) removes the covariance altogether.

3§03.5 — the noise level's own estimate

Under i.i.d. Gaussian noise the log-likelihood also depends on $\sigma$. Maximizing it over $\sigma$ as well as $\beta$ gives a second estimate of $\sigma^2$, to compare with the RSE.

Find
  1. (a) With $\beta=\hat\beta_{\mathrm{RSS}}$ fixed, find the $\sigma^2$ that maximizes $l$.

  2. (b) Evaluate it for the café fit and compare with $\mathrm{RSE}^2=\mathrm{RSS}/(n-2)$.

Given
  • $l(\beta,\sigma)=n\log\frac{1}{\sqrt{2\pi}\,\sigma}-\frac{\mathrm{RSS}(\beta)}{2\sigma^2}$

  • café fit: $n=5$, $\mathrm{RSS}=3.6$

Hint 1/4

$\beta$ is already settled by the RSS; what remains is a maximization in one variable, $\sigma$.

Hint 2/4

Write $l=-n\log\sigma-\tfrac n2\log(2\pi)-\mathrm{RSS}/(2\sigma^2)$, differentiate in $\sigma$ and set the result to zero.

Hint 3/4

$\frac{\partial l}{\partial\sigma}=-\frac n\sigma+\frac{\mathrm{RSS}}{\sigma^3}=0$ with $n=5$ and $\mathrm{RSS}=3.6$.

Hint 4/4

$\hat\sigma^2_{\mathrm{MLE}}=\mathrm{RSS}/n=0.72$, smaller than $\mathrm{RSE}^2=1.2$.

Show solution

Differentiating in $\sigma$ directly is as easy as in $\sigma^2$, and the answer is the same point.

Derivative

$$\frac{\partial l}{\partial\sigma}=-\frac{n}{\sigma}+\frac{\mathrm{RSS}}{\sigma^3}$$

$\log\frac{1}{\sqrt{2\pi}\sigma}=-\log\sigma-\tfrac12\log(2\pi)$.

Solve

$$\sigma^2=\frac{\mathrm{RSS}}{n}$$

Multiply the zero-derivative equation by $\sigma^3/n$.

Café numbers

$$\hat\sigma^2_{\mathrm{MLE}}=\frac{3.6}{5}=0.72,\qquad \mathrm{RSE}^2=\frac{3.6}{3}=1.2$$

Same RSS, different divisors.

Answer $$\boxed{\hat\sigma^2_{\mathrm{MLE}}=\mathrm{RSS}/n=0.72}$$
Check

Second derivative at that point: $\frac{n}{\sigma^2}-\frac{3\,\mathrm{RSS}}{\sigma^4}=\frac{n}{\sigma^2}-\frac{3n}{\sigma^2}<0$, so it is a maximum.

The MLE of the noise level runs small because residuals are smaller than errors; the course's intervals use the RSE.

4§03.6 — a filter trained on counts

A spam filter's first training set has $20$ e-mails, $7$ of them spam (coded $1$) and $13$ not (coded $0$). With no inputs yet, it fits $y=\beta_0+\varepsilon$ by least squares and flags an e-mail when $\hat y>0.5$.

Find(a) What is $\hat\beta_0$, and what does the filter do with every new e-mail?
Given
  • $20$ labels: $7$ ones and $13$ zeros

  • model $y=\beta_0+\varepsilon$

  • flag if $\hat y>0.5$

Hint 1/4

With only an intercept, least squares looks for the single number closest to all twenty labels.

Hint 2/4

Minimizing $\sum_i(y_i-\beta_0)^2$ gives $\hat\beta_0=\bar y$; for 0/1 labels that is the fraction of ones, the rate MLE $N_1/N$.

Hint 3/4

$7$ ones among $20$ labels: $\bar y=7/20$.

Hint 4/4

$\hat\beta_0=0.35\le 0.5$, so no e-mail is ever flagged.

Show solution

The one-parameter least squares problem is the warm-up minimization of this section, so its answer is the mean.

Minimize

$$\frac{d}{d\beta_0}\sum_i(y_i-\beta_0)^2=-2\sum_i(y_i-\beta_0)=0\ \Rightarrow\ \hat\beta_0=\bar y$$

Setting the derivative to zero gives the sample mean.

Evaluate

$$\bar y=\frac{7\cdot 1+13\cdot 0}{20}=0.35$$

The mean of 0/1 labels counts the ones.

Apply the rule

$$0.35\le 0.5\ \Rightarrow\ \text{never flag}$$

The same fitted value for every e-mail, always below the cut.

Answer $$\boxed{\hat\beta_0=0.35,\ \text{no e-mail flagged}}$$
Check

The same $0.35$ is the maximum likelihood estimate of a Bernoulli rate from $7$ successes in $20$ trials.

An intercept-only fit predicts the majority class; inputs are what let a classifier do better than that.

Mistake ledger (16 entries)
⚠ Raw sums in place of centered sums

the raw products $x_iy_i$ are quicker to add up than the deviations

wrong$$\hat\beta_1=\frac{\sum_i x_iy_i}{\sum_i x_i^2}=\frac{98}{55}\approx 1.78$$
right$$\hat\beta_1=\frac{\sum_i(x_i-\bar x)(y_i-\bar y)}{\sum_i(x_i-\bar x)^2}=\frac{14}{10}=1.4$$
⚠ Sign slip in the intercept

the formula has a minus sign and $\hat\beta_1$ may itself be negative

wrong$$\hat\beta_0=\bar y+\hat\beta_1\bar x$$
right$$\hat\beta_0=\bar y-\hat\beta_1\bar x$$
⚠ Dividing RSS by n instead of n − 2

$n$ is the divisor of an ordinary average, and the $-2$ is easy to forget

wrong$$\mathrm{RSE}=\sqrt{\mathrm{RSS}/n}=\sqrt{3.6/5}\approx 0.85$$
right$$\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}=\sqrt{3.6/3}\approx 1.10$$
⚠ The variance in place of the standard error

the formula box gives variances, and the square root gets dropped on the way to the interval

wrong$$\hat\beta_1\pm 2\operatorname{Var}(\hat\beta_1)=1.4\pm 0.24$$
right$$\hat\beta_1\pm 2\sqrt{\operatorname{Var}(\hat\beta_1)}=1.4\pm 0.69$$
⚠ Forgetting the column of ones

the data file lists only the inputs; the intercept column is not in it

wrong$$X=\begin{bmatrix}x_{11}&\cdots&x_{1p}\\ \vdots&&\vdots\\ x_{n1}&\cdots&x_{np}\end{bmatrix}$$
right$$X=\begin{bmatrix}1&x_{11}&\cdots&x_{1p}\\ \vdots&\vdots&&\vdots\\ 1&x_{n1}&\cdots&x_{np}\end{bmatrix}$$
⚠ Transposes in the wrong place

both $X^TX$ and $XX^T$ exist, but $XX^T$ is $n\times n$ and cannot multiply $X^Ty$

wrong$$\hat\beta=(XX^T)^{-1}X^Ty$$
right$$\hat\beta=(X^TX)^{-1}X^Ty$$
⚠ Doubling a quadratic form whose matrix is not symmetric

the shortcut is right for the symmetric $X^TX$ and gets reused where it is not

wrong$$\frac{\partial}{\partial y}\big(y^TAy\big)=2y^TA$$
right$$\frac{\partial}{\partial y}\big(y^TAy\big)=y^T(A+A^T)$$
⚠ Measuring residuals perpendicular to the line

distance to a line in school geometry is the perpendicular one

wrong$$d_i=\frac{\vert y_i-\hat\beta_0-\hat\beta_1x_i\vert}{\sqrt{1+\hat\beta_1^2}}$$
right$$e_i=y_i-\hat\beta_0-\hat\beta_1x_i$$
⚠ Simplifying H to the identity

the inverse of a product looks as if it should split into two inverses

wrong$$H=X(X^TX)^{-1}X^T=XX^{-1}(X^T)^{-1}X^T=I$$
right$$H=X(X^TX)^{-1}X^T,\qquad \operatorname{tr}H=p+1<n$$
⚠ Thinking Gauss-Markov needs Gaussian noise

the two names sound alike and both results appear in the same lecture

wrong$$\text{BLUE}\ \Leftarrow\ \varepsilon_i\sim\mathcal N(0,\sigma^2)$$
right$$\text{BLUE}\ \Leftarrow\ E[\varepsilon_i]=0,\ \operatorname{Var}(\varepsilon_i)=\sigma^2,\ \text{uncorrelated}$$
⚠ Letting σ change the maximum likelihood estimate of β

$\sigma$ appears in every term of the likelihood

wrong$$\hat\beta_{\mathrm{MLE}}\ \text{depends on}\ \sigma^2$$
right$$\hat\beta_{\mathrm{MLE}}=\arg\min_\beta\mathrm{RSS}(\beta)\ \text{for every}\ \sigma>0$$
⚠ Losing the minus sign

the RSS term carries a negative coefficient that is easy to drop when copying

wrong$$\arg\max_\beta l(\beta)=\arg\max_\beta\mathrm{RSS}(\beta)$$
right$$\arg\max_\beta l(\beta)=\arg\min_\beta\mathrm{RSS}(\beta)$$
⚠ Cutting at 0 instead of 0.5

labels coded $-1$ and $+1$ would be cut at $0$; with $0$ and $1$ the cut is halfway

wrong$$\hat\beta^Tx>0\ \Rightarrow\ \text{class }1$$
right$$\hat\beta^Tx>0.5\ \Rightarrow\ \text{class }1$$
⚠ Reading fitted values as probabilities

the labels are 0 and 1, so numbers near them look like probabilities

wrong$$P(\text{pass}\mid x=8)=1.4$$
right$$\hat y(8)=1.4>0.5\ \Rightarrow\ \text{predict pass}$$
⚠ Counting rows instead of distinct inputs

the rank condition sounds like a condition on the sample size

wrong$$n=7\ \text{rows}\ \Rightarrow\ p\le 6$$
right$$3\ \text{distinct}\ x_i\ \Rightarrow\ p\le 2$$
⚠ Reading a main effect alone when an interaction is present

without the product term, $\hat\beta_1$ was the effect of $x_1$, and the habit stays

wrong$$\text{effect of }x_1=\hat\beta_1$$
right$$\text{effect of }x_1=\hat\beta_1+\hat\beta_3x_2$$
Formula card
Least squares, one input
$$\hat\beta_1=\frac{\sum_i(x_i-\bar x)(y_i-\bar y)}{\sum_i(x_i-\bar x)^2},\qquad \hat\beta_0=\bar y-\hat\beta_1\bar x$$

an intercept in the model and at least two distinct $x_i$

Variances of the estimates
$$\operatorname{Var}(\hat\beta_1)=\frac{\sigma^2}{S_{xx}},\qquad \operatorname{Var}(\hat\beta_0)=\sigma^2\Big[\frac1n+\frac{\bar x^2}{S_{xx}}\Big]$$

fixed $x_i$; noise with mean $0$, one variance, uncorrelated

Residual standard error and 95% interval
$$\mathrm{RSE}=\sqrt{\frac{\mathrm{RSS}}{n-2}},\qquad \hat\beta_j\pm 2\sqrt{\operatorname{Var}(\hat\beta_j)}$$

$\sigma$ replaced by the RSE; the multiplier $2$ is the lecture's large-sample value

Normal equations
$$X^TX\hat\beta=X^Ty,\qquad \hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$$

$X$ is $n\times(p+1)$ with full column rank

Matrix derivative rules
$$\frac{\partial(Ag)}{\partial g}=A,\quad \frac{\partial(y^TAg)}{\partial g}=y^TA,\quad \frac{\partial(y^TAy)}{\partial y}=y^T(A+A^T)$$

derivatives of scalars by column vectors are rows

Projection
$$\hat y=Hy,\quad H=X(X^TX)^{-1}X^T,\quad X^Te=0,\quad H^2=H=H^T$$

full column rank

Gaussian log-likelihood
$$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\,\sigma}-\frac{\mathrm{RSS}(\beta)}{2\sigma^2}\ \Rightarrow\ \hat\beta_{\mathrm{MLE}}=\hat\beta_{\mathrm{RSS}}$$

i.i.d. $\mathcal N(0,\sigma^2)$ noise; Gauss-Markov needs only mean, variance and zero correlation

Decision rule from a regression fit
$$\widehat{\text{class}}(x)=1\ \text{if}\ \hat\beta_{\mathrm{RSS}}^Tx>0.5,\ \ \text{else}\ 0$$

labels coded $0$ and $1$; outputs are not probabilities

Basis function model
$$\hat y=\hat\beta_0+\sum_{j=1}^q\hat\beta_j\phi_j(x),\qquad X_{ij}=\phi_j(x_i)$$

polynomial of degree $p$: unique fit when at least $p+1$ inputs are distinct

Gaussian and sigmoidal bases
$$\phi_j(x)=\exp\!\Big(-\frac{\lVert x-\mu_j\rVert_2^2}{2s^2}\Big),\qquad \phi_j(x)=\frac{1}{1+e^{-\lVert x-\mu_j\rVert_2/s}}$$

centres $\mu_j$ and scale $s$ are chosen before fitting

Check yourself

Close the page and write from memory: the RSS; the formulas for $\hat\beta_1$ and $\hat\beta_0$; the two variances and the 95% interval; the normal equations with their rank condition; what $X^Te=0$ says; the two reasons for squares; the $0.5$ decision rule; and when a polynomial fit is unique. Then compare with the formula card.

  • Fit a line from a table of pairs and check it with $\sum e_i=0$ and $\sum x_ie_i=0$?

    c-simple-ls

  • Compute the RSE and a 95% interval for the slope and the intercept, and say which one is wider and why?

    c-accuracy

  • Derive $X^TX\beta=X^Ty$ with the course's derivative rules and solve a $2\times 2$ or diagonal case?

    c-matrix-ls

  • Explain why $e$ is perpendicular to every column of $X$, and use that to test a proposed fit?

    c-projection

  • State what Gauss-Markov assumes, and show that Gaussian noise turns the MLE into least squares?

    c-why-squares

  • Find a decision boundary from a fit on 0/1 labels and name its two weaknesses?

    c-classification

  • Build a design matrix with polynomial, interaction or Gaussian columns and decide whether the fit is unique?

    c-basis

Glossary (29 terms)
linear regressiondoğrusal regresyon

A model that predicts the response by a linear combination of the coefficients, $\hat y=\hat\beta_0+\sum_j\hat\beta_jx_j$.

simple linear regressionbasit doğrusal regresyon

Linear regression with one input: $Y\approx\beta_0+\beta_1X$.

multiple linear regressionçoklu doğrusal regresyon

Linear regression with several inputs, written $y\approx X\beta$.

bağımlı değişken

The quantity being predicted, Y.

predictoraçıklayıcı değişken

An input variable used to predict the response.

interceptsabit terim

The coefficient $\beta_0$, the prediction when every input is zero; the lecture also calls it the bias.

slopeeğim

The change in the prediction per unit change of an input, $\beta_1$ in simple regression.

residualartık

The observed response minus the fitted one, $e_i=y_i-\hat y_i$.

residual sum of squaresartık kareler toplamı

$\mathrm{RSS}=\sum_ie_i^2$, the score that least squares minimizes.

least squaresen küçük kareler

The method that chooses the coefficients minimizing the residual sum of squares.

sıradan en küçük kareler

Least squares with every observation weighted equally, giving $\hat\beta=(X^TX)^{-1}X^Ty$.

fitted value

The model's prediction at an observed input, $\hat y_i=\hat\beta^Tx_i$.

merkezleme

Subtracting the mean from a variable so that its values sum to zero.

residual standard error

$\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$, the estimate of the noise standard deviation in simple regression.

unbiased estimatoryansız kestirici

An estimator whose mean over repeated samples equals the true value, $E[\hat\beta]=\beta^{\mathrm{true}}$.

design matrixtasarım matrisi

The matrix $X$ with one row per observation and one column per coefficient.

full column rank

No column of $X$ is a linear combination of the others; it makes $X^TX$ invertible.

normal equationsnormal denklemler

The system $X^TX\beta=X^Ty$ that the least squares coefficients solve.

Jacobi matrisi

The matrix of partial derivatives $\partial h_i/\partial g_j$ of a vector function $h$ of a vector $g$.

hat matrixşapka matrisi

$H=X(X^TX)^{-1}X^T$, which maps $y$ to $\hat y$; the lecture calls it the projection matrix.

orthogonal projectiondik izdüşüm

The closest point of a subspace to a given vector, reached along a direction perpendicular to the subspace.

column spacesütun uzayı

All vectors of the form $X\beta$: the combinations of the columns of $X$.

Gauss-Markov theoremGauss Markov teoremi

Under zero-mean, equal-variance, uncorrelated noise, least squares has the smallest variance among linear unbiased estimators.

BLUE

Best linear unbiased estimator; under the Gauss-Markov assumptions it is the least squares estimator.

decision boundarykarar sınırı

The set of inputs where a classifier switches class; for a regression fit on 0/1 labels, $\hat\beta^Tx=0.5$.

etkileşim terimi

A product column such as $x_1x_2$ that lets the effect of one input depend on another.

polynomial regressionpolinom regresyonu

Linear regression on the columns $1,x,x^2,\dots,x^p$.

Vandermonde matrixVandermonde matrisi

The design matrix of polynomial regression, with rows $[1, \allowbreak x_i, \allowbreak x_i^2, \allowbreak \dots, \allowbreak x_i^p]$.

basis functiontaban fonksiyonu

A fixed function $\phi_j(x)$ of the inputs used as a column of the design matrix.

What comes next
§04 · Measuring performance: bias-variance and cross-validation

Here every fit was judged on the same points it was fitted to, and the RSS only went down as columns were added. Next comes the question that number cannot answer: how well does a fitted model do on data it has never seen?

Sources
  • textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer, 2003 The textbook named in the syllabus. This week's syllabus line gives no section numbers, so none are cited here.
  • course materialEEE 485/585 chapter 3 lecture slides and lecture notes, Fall 2026 Topic order and notation follow them: beta true, RSS, beta hat RSS, the row convention for matrix derivatives, the projection matrix H and the 0.5 decision rule. All wording, data, examples and exercises here are original.
  • course materialEEE 485 syllabus page on STARS, printed 21 September 2026 Source of the weekly line, the assessment weights quoted in the card and the conditions for sitting the final exam.
  • textbookOther books the syllabus recommends: G. James et al., An Introduction to Statistical Learning (2013); K. P. Murphy, Machine Learning: A Probabilistic Perspective (2012); C. M. Bishop, Pattern Recognition and Machine Learning (2011) Recommended, not required. The lecture slides borrow several of their regression figures from James et al.

Spotted something missing or wrong? tell us · share your own notes or an old exam.

Last updated .