The middle value, the midpoint of the extremes and the root mean square all do worse.
Answer $$\boxed{c=3}$$
Check
The second derivative is $6>0$, so the cost is an upward parabola and $c=3$ is its minimum.
Squared misses pull the best single number to the mean; with an starts from exactly this fact.
A café counts its social media posts and the cups it sells, in hundreds, for five weeks: $(1,3),\, \allowbreak (2,5),\, \allowbreak (3,4),\, \allowbreak (4,7),\, \allowbreak (5,9)$. One barista rules a line through weeks 1 and 5 and reads 150 extra cups per post; another uses weeks 2 and 4 and reads 100. Which of them is right, and how sure can the café be of its number?
By the end you can fit that line by hand from five pairs, put a 95% interval on its , write the same fit as $\hat\beta=(X^TX)^{-1}X^Ty$ for any number of inputs, and say why squared misses are the natural score.
In 60 seconds
Least squares picks the coefficients whose squared vertical misses add up to the least: for one input $\hat\beta_1=S_{xy}/S_{xx}$ and $\hat\beta_0=\bar y-\hat\beta_1\bar x$; for many, $\hat\beta=(X^TX)^{-1}X^Ty$, and $\hat y$ is the of $y$ onto the columns of $X$.
Raw sums instead of centered ones: $\sum x_iy_i/\sum x_i^2$ is the slope of a line forced through the origin, not the least squares slope of a model with an intercept.
Treating residuals as perpendicular distances to the line. Least squares charges vertical misses; the right angle lives in $\mathbb R^n$, between the residual vector and the columns of $X$.
Dividing the RSS by $n$ instead of $n-2$ for the RSE, or stepping $\pm 2$ variances instead of $\pm 2$ standard errors.
The weights differ between two course documents:
Chapter 1 slides, undergraduate line, issued on 15 September and again on 24 September: midterm 25, final 25, four quizzes 20, two-phase project 30.
STARS syllabus page for Fall 2026-27, printed 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. Confirm with the course which split applies.
To sit the final, both require every project report on time, no disciplinary penalty, and midterm plus quiz points worth at least a fifth of what those two components can give.
How much time do you have?
10 minutes
The two formulas most regression questions start from, each with its first worked example.
The 60-second card · Simple linear regression · Many inputs at once · Formula card
45 minutes
Every block once with its first example and checkpoint, then one ladder from a full solution to a bare problem.
The 60-second card · Simple linear regression · How accurate are the estimates · Many inputs at once · The geometry of least squares · Why squares · Classification with linear regression · Curves with a linear model · Scaffolding comes off · Formula card
full read
Where each formula comes from, how uncertain each estimate is, and enough mixed practice to choose the method yourself.
The opening pages · Recall first · Simple linear regression · How accurate are the estimates · Many inputs at once · The geometry of least squares · Why squares · Classification with linear regression · Curves with a linear model · Look-alike pairs · Method boxes · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
Compute the least squares intercept and slope from a table of pairs, and check the fit through its residuals.
Estimate $\sigma$ by the and build 95% intervals $\hat\beta_j\pm 2\,\mathrm{SE}$ for the intercept and the slope.
Derive the normal equations with the matrix derivative rules and solve $X^TX\beta=X^Ty$ when $X$ has .
Interpret $\hat y=Hy$ as an orthogonal projection and use $X^Te=0$ to test whether a fit is the least squares fit.
Justify least squares twice: as the best linear unbiased estimator and as the maximum likelihood estimator under Gaussian noise.
Turn a least squares fit on 0/1 labels into a , and name the two weaknesses of doing so.
Fit curves with the same machinery by adding interaction, polynomial or basis-function columns, and decide when the fit is unique.
Syllabus coverage
covered
Linear regression — covered
The lecture's chapter
the model $Y=f(X,\beta^{\mathrm{true}})+\varepsilon$
simple linear regression
accuracy of the estimates and 95% intervals
classification with linear regression
interactions, and basis functions
Seven blocks follow the lecture's order; the model and simple regression open the first block.
covered
ordinary least squares — covered
The fitting method
RSS and its minimizer
the closed form for one input
the normal equations and $\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$ from matrix derivatives
the projection $H$
Gauss-Markov and the maximum likelihood reading
off syllabus
Proof of the Gauss-Markov theorem — off syllabus
A four-line sketch that splits the weights of any linear unbiased estimator.
Further reading. The lecture states the theorem without proof; nothing later depends on the sketch.
off syllabus
Covariance of the two coefficients — off syllabus
$\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-\bar x\,\sigma^2/S_{xx}$, built from the covariance rules of the first section.
Further reading, used in one interleaved practice question to explain the pivot in the sampling figure.
off syllabus
Maximum likelihood estimate of the noise variance — off syllabus
Maximizing the Gaussian log-likelihood over $\sigma$ gives $\mathrm{RSS}/n$, set against $\mathrm{RSE}^2=\mathrm{RSS}/(n-2)$.
Further reading, used in one interleaved practice question.
Recall first
Minimizing a smooth function
At an interior minimum of a differentiable function every partial derivative is zero; a quadratic $at^2+bt+c$ with $a>0$ is smallest at $t=-b/(2a)$.
Least squares sets two, or p + 1, partial derivatives of the RSS to zero.
Transpose, inverse and rank
$(AB)^T=B^TA^T$; a square matrix is invertible exactly when its columns are independent; $\begin{bmatrix}a&b\\c&d\end{bmatrix}^{-1}=\frac{1}{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix}$.
Every matrix step of the normal equations uses them.
Expectation and variance of weighted sums
$E[\sum_ia_iY_i]=\sum_ia_iE[Y_i]$ always. If the $Y_i$ are uncorrelated with variance $\sigma^2$, $\operatorname{Var}(\sum_ia_iY_i)=\sigma^2\sum_ia_i^2$; in vector form $\operatorname{Var}(a^TY)=a^T\Sigma a$.
Unbiasedness and both variance formulas are these rules applied to the least squares weights.
Gaussian density
$Y\sim\mathcal N(\mu,\sigma^2)$ has density $\frac{1}{\sqrt{2\pi}\,\sigma}e^{-(y-\mu)^2/(2\sigma^2)}$.
The likelihood of the regression model is a product of these.
Maximum likelihood and the 95% recipe
The MLE maximizes $L$, or equivalently $\log L$. A 95% confidence interval steps $2$ standard errors each side of the estimate, and its 95% refers to repeated samples.
Both return here: the MLE of $\beta$, and intervals for $\beta_0$ and $\beta_1$.
Orthogonality and length
$u^Tv=0$ means $u$ and $v$ are perpendicular; $\lVert u\rVert^2=u^Tu$; if $u^Tv=0$ then $\lVert u+v\rVert^2=\lVert u\rVert^2+\lVert v\rVert^2$.
The geometric reading of least squares is Pythagoras in n dimensions.
Roots of polynomials
A nonzero polynomial of degree at most $p$ has at most $p$ real roots.
It decides when polynomial regression has a unique fit.
Try it yourself first (2 questions)
1§03.3 — transposing a product
Matrix warm-up: $A$ is $3\times 2$ and $B$ is $2\times 4$.
Find(a) Which expression equals $(AB)^T$, and what size is it?
Given
$A$: $3\times 2$
$B$: $2\times 4$
Hint 1/4
Transposing a product reverses the order of the factors; the sizes confirm it.
Hint 2/4
$(AB)^T=B^TA^T$.
Hint 3/4
$AB$ is $3\times 4$, so its transpose is $4\times 3$; $B^T$ is $4\times 2$ and $A^T$ is $2\times 3$.
Hint 4/4
$(AB)^T=B^TA^T$, a $4\times 3$ matrix.
Show solution
The size bookkeeping catches a wrong order immediately.
Rule
$$(AB)^T=B^TA^T$$
Entry $(i,j)$ of $(AB)^T$ is row $j$ of $A$ times column $i$ of $B$.
Sizes
$$(4\times 2)(2\times 3)=4\times 3$$
Inner dimensions match, outer ones give the size.
Answer $$\boxed{B^TA^T,\ 4\times 3}$$
Check
$AB$ is $3\times 4$, and transposing swaps the two sizes, giving $4\times 3$ again.
This reversal is how $(y-X\beta)^T$ becomes $y^T-\beta^TX^T$ in the least squares derivation.
2§03.2 — a constant plus scaled noise
Probability warm-up: $\varepsilon$ has mean $0$ and variance $\sigma^2$, and $Y=2+3\varepsilon$.
Find(a) What are $E[Y]$ and $\operatorname{Var}(Y)$?
Standard deviation check: $\sqrt{9\sigma^2}=3\sigma$, three times the spread of $\varepsilon$, as scaling by $3$ should give.
This is the model of the section in miniature: a fixed part plus noise, with the noise alone carrying the variance.
Notation
symbol
reads as
means
watch out
$Y,\ X=[X_1,\dots,X_p]^T$
Y; the input vector X
the response and the inputs, also called
Capitals are random quantities; the observed values are $y_i$ and $x_i$.
$\beta^{\mathrm{true}}$
beta true
the unknown coefficients of the model that produced the data
Never observed; everything with a hat is computed from data.
$\hat\beta_0,\ \hat\beta_1$
beta zero hat, beta one hat
the least squares intercept and slope
Random: a new sample gives new values.
$D=\{(x_i,y_i)\}_{i=1}^n$
the data set D
$n$ observed pairs
In multiple regression each $x_i$ is a vector.
$S_{xx},\ S_{xy}$
S x x, S x y
$\sum_i(x_i-\bar x)^2$ and $\sum_i(x_i-\bar x)(y_i-\bar y)$
Centered sums; the raw sums $\sum x_i^2$ and $\sum x_iy_i$ are different numbers.
$\hat y_i,\ e_i$
y i hat, e i
the fitted value and the residual $y_i-\hat y_i$
A residual is computed from the fit; the error $\varepsilon_i$ is not observable.
$\varepsilon,\ \sigma^2$
epsilon, sigma squared
the zero-mean noise and its variance
$\sigma$ is estimated by the RSE.
$\mathrm{RSS},\ \mathrm{RSE}$
residual sum of squares, residual standard error
$\sum_ie_i^2$ and $\sqrt{\mathrm{RSS}/(n-2)}$
The $n-2$ belongs to simple regression, with two fitted coefficients.
$x_i=[1,x_{i1},\dots,x_{ip}]^T$
x i
the $i$-th input vector with a leading $1$ for the intercept
Forgetting the 1 drops the intercept from the model.
$X,\ y,\ \hat y$
X, y, y hat
the $n\times(p+1)$ data matrix, the response vector and the fitted vector
$X$ has one row per observation and one column per coefficient.
$\hat\beta_{\mathrm{RSS}}$
beta hat RSS
the minimizer of $\mathrm{RSS}(\beta)$, equal to $(X^TX)^{-1}X^Ty$
Needs full column rank.
$H$
the
$X(X^TX)^{-1}X^T$, the projection onto the column space of $X$
$n\times n$, symmetric, and $H^2=H$.
$\partial h/\partial g$
the Jacobian of h with respect to g
the $m\times n$ matrix of partial derivatives $\partial h_i/\partial g_j$
For a scalar $\alpha$, $\partial\alpha/\partial g$ is a row vector in this course.
$L(\beta),\ l(\beta)$
the likelihood, the log-likelihood
the density of the data as a function of $\beta$, and its natural logarithm
Both have the same maximizer.
$\phi_j(x)$
phi j of x
the $j$-th basis function, a column computed from the inputs
$\phi_0(x)=1$ carries the intercept.
Conventions used here
Input rows and the column of ones.
Every $x_i$ is a column $[1,x_{i1},\dots,x_{ip}]^T$, so $x_i^T\beta$ is a number and the rows of $X$ are the $x_i^T$. With an intercept, the first column of $X$ is all ones.
Most dimension errors come from a missing column of ones or a transpose in the wrong place.
Derivatives of scalars are rows.
Following the lecture, $\partial\alpha/\partial\beta$ of a scalar $\alpha$ is a $1\times(p+1)$ row, and $\partial h/\partial g$ is $m\times n$ with one row per output. Setting a row derivative to zero and transposing gives the same equations as a column gradient.
Some books write gradients as columns; the equations you end with are identical.
Interval multiplier and small samples.
As in the previous section, a 95% interval steps $2$ standard errors each way, with $\sigma$ replaced by the RSE. With only a handful of points this interval is on the narrow side; we still use $2$ so that every example can be done by hand.
Exact small-sample multipliers exist but are not part of this chapter; answers here use 2 unless a question says otherwise.
Linear means linear in the coefficients.
A model is linear regression when $y$ is a linear combination of the coefficients, whatever the columns are: $x$, $x^2$, $\log x$ or $x_1x_2$.
This is what lets polynomial and basis-function models use the same formula.
Residuals versus errors.
$e_i=y_i-\hat y_i$ is computed from the fit; $\varepsilon_i=y_i-(\beta^{\mathrm{true}})^Tx_i$ needs the unknown truth. Facts such as $\sum_ie_i=0$ are about residuals only.
Mixing the two produces false claims such as the noise summing to zero.
Logarithms and rounding.
$\log$ is the natural logarithm. Intermediate steps keep at least four significant figures; coefficients and interval ends are rounded to two or three decimals at the end.
Rounding the RSE early can move an interval end in the second decimal.
3.1Simple linear regression: the line with the smallest residual sum of squares
Picks the intercept and slope that make the squared vertical misses as small as possible; both come from two short formulas.
In the previous section an estimate was one number; now the data are pairs $(x_i,y_i)$ and we need two numbers at once, an intercept and a slope.
Solvable with what we have
Estimate one unknown number from data, such as a rate by its MLE $N_1/N$.
Minimize a function of one variable by setting its derivative to zero.
Average a column of numbers into $\bar x$ or $\bar y$.
Not solvable yet
Decide which of two lines through the same points is better.
Find a slope and an intercept together from one data set.
Predict sales for a number of posts the café has not tried.
A tempting shortcut: average the slopes between neighbouring weeks. For the café they are $2,\,-1,\,3,\,2$, with mean $1.5$. But the sum telescopes: $$\tfrac14\big[(y_2-y_1)+(y_3-y_2)+(y_4-y_3)+(y_5-y_4)\big]=\tfrac14(y_5-y_1).$$
Why it fails
Only the first and last weeks survive. Change week 3's sales from $4$ to $40$ and the average slope is still $1.5$. We need one score that charges a line for missing every point.
TheoremLeast squares for one predictor
Conditions
data $D=\{(x_i,y_i)\}_{i=1}^n$ with at least two different $x_i$
model $Y\approx\beta_0+\beta_1X$; a candidate line is scored by its residual sum of squares
Among all lines, keep the one whose squared vertical misses add up to the least. Its slope is how much $x$ and $y$ move together, divided by how much $x$ moves on its own. The intercept then puts the line through the point of means $(\bar x,\bar y)$.
Derivation: two partial derivatives set to zero
Set the derivative in $\beta_0$ to zero: $$\frac{\partial\,\mathrm{RSS}}{\partial\beta_0}=-2\sum_i\big(y_i-\beta_0-\beta_1x_i\big)=0\ \Longrightarrow\ \beta_0=\bar y-\beta_1\bar x.$$ So the best line passes through $(\bar x,\bar y)$, whatever its slope.
Put this $\beta_0$ back. Each residual becomes $(y_i-\bar y)-\beta_1(x_i-\bar x)$, a function of $\beta_1$ alone.
Set the derivative in $\beta_1$ to zero: $$\sum_i(x_i-\bar x)\big[(y_i-\bar y)-\beta_1(x_i-\bar x)\big]=0\ \Longrightarrow\ \beta_1=\frac{S_{xy}}{S_{xx}}.$$
It is a minimum: after the substitution, $\mathrm{RSS}=S_{yy}-2\beta_1S_{xy}+\beta_1^2S_{xx}$ with $S_{yy}=\sum_i(y_i-\bar y)^2$, an upward parabola in $\beta_1$ because $S_{xx}>0$ when the $x_i$ are not all equal.
The café weeks with the least squares line $\textcolor{#1f6feb}{\hat y=1.4+1.4x}$. The vertical segments are the residuals $\textcolor{#d1690a}{e_i=y_i-\hat y_i}$; their squares add to $\textcolor{#d1690a}{\mathrm{RSS}=3.6}$, and the line passes through the point of means $(3,\,5.6)$.
Looks like this, but is not
A line whose residuals sum to zero looks like the least squares line: the misses above and below cancel.
Every line through $(\bar x,\bar y)$ has residuals that sum to zero. The café line $2+1.2x$ passes through $(3,5.6)$ and its residuals $-0.2,\, \allowbreak 0.6,\, \allowbreak -1.6,\, \allowbreak 0.2,\, \allowbreak 1$ cancel, yet its RSS is $4.0$, not $3.6$. A zero sum is necessary, not sufficient.
week
x
y
x − 3
y − 5.6
product
(x − 3)²
1
$1$
$3$
$-2$
$-2.6$
$5.2$
$4$
2
$2$
$5$
$-1$
$-0.6$
$0.6$
$1$
3
$3$
$4$
$0$
$-1.6$
$0$
$0$
4
$4$
$7$
$1$
$1.4$
$1.4$
$1$
5
$5$
$9$
$2$
$3.4$
$6.8$
$4$
sum
$15$
$28$
$0$
$0$
$14$
$10$
The last two column sums are $S_{xy}=14$ and $S_{xx}=10$, so $\hat\beta_1=1.4$. Both centered columns sum to zero, which is a free check on the two means.
Café posts and cups: the least squares line and its RSS
Five weeks of café data: posts per week $x=1,2,3,4,5$ and cups sold, in hundreds, $y=3,5,4,7,9$. Fit the least squares line, then score it and the two lines drawn by eye, $1.5+1.5x$ and $3+x$, by their RSS.
Find$\hat\beta_0$, $\hat\beta_1$ and the RSS of all three lines.
Given
$x=1,2,3,4,5$
$y=3,5,4,7,9$
eye-drawn lines: $1.5+1.5x$ (weeks 1 and 5) and $3+x$ (weeks 2 and 4)
Solution
We center first, so every product stays small; the raw-sum version subtracts two large, nearly equal numbers and invites slips.
The five residuals $0.2+0.8-1.6+0+0.6$ sum to $0$, as they must for a line through $(\bar x,\bar y)$, and the slope $1.4$ sits between the two eye-drawn slopes $1$ and $1.5$.
This answers the opening question: each extra post goes with about $140$ more cups, and no straight line scores below $3.6$ on these five weeks.
Counting posts from the average week: what centering changes
Refit the café data with the input measured from its mean, $u=x-3$, so $u=-2,-1,0,1,2$ and still $y=3,5,4,7,9$. Compare the fit with $\hat y=1.4+1.4x$.
FindThe least squares line in $u$ and its meaning.
Given
$u=-2,-1,0,1,2$
$y=3,5,4,7,9$
Solution
With $\sum_i u_i=0$ the intercept formula collapses to $\bar y$, so we get the fit almost without arithmetic.
the formula has a minus sign and $\hat\beta_1$ may itself be negative
wrong$$\hat\beta_0=\bar y+\hat\beta_1\bar x$$
right$$\hat\beta_0=\bar y-\hat\beta_1\bar x$$
Every line here passes through the point of means $(3,\,5.6)$; only the slope $b$ changes. The right panel is $\textcolor{#d1690a}{\mathrm{RSS}(b)=3.6+10\,(b-1.4)^2}$, lowest at the least squares slope $\textcolor{#1f6feb}{1.4}$.
At the edges
slope 0.8 RSS 7.2
Too flat: week 1 falls 1.0 below the line and week 5 sits 1.8 above it.
slope 1.4 RSS 3.6
The least squares slope. A step of 0.2 either way costs the same 0.4, because the curve is a parabola.
slope 2.0 RSS 7.2
Too steep: weeks 1 and 2 sit 1.4 above the line and the last two weeks fall below it. It costs exactly as much as 0.8.
3.2How accurate are the estimates: unbiasedness, variance and a 95% interval
Says how far $\hat\beta_0$ and $\hat\beta_1$ would move with new data, and turns that into a $\pm 2$ standard error interval.
The café line came from five particular weeks; five other weeks with the same true relationship would give a different line.
TheoremMean and variance of the least squares estimates
Conditions
true model $Y=\beta_0^{\mathrm{true}}+\beta_1^{\mathrm{true}}X+\varepsilon$, with the $x_i$ fixed and known
$E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$, and the $\varepsilon_i$ are uncorrelated
$\sigma$ is estimated by the residual standard error $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$
Averaged over repeated samples, the estimates land on the true values. The slope wobbles less when the inputs are spread out; the intercept wobbles least when the inputs are centered at zero. Replace $\sigma$ by the RSE and step two standard errors each way for a 95% interval.
Why the slope is unbiased, and where its variance comes from
Write the slope as a weighted sum of responses: $$\hat\beta_1=\sum_i k_iy_i,\qquad k_i=\frac{x_i-\bar x}{S_{xx}}.$$ The $\bar y$ term drops out because $\sum_i(x_i-\bar x)=0$.
The weights satisfy $\sum_ik_i=0$ and $\sum_ik_ix_i=1$, so $E[\hat\beta_1]=\sum_ik_i(\beta_0+\beta_1x_i)=\beta_1$.
Uncorrelated noise with one variance gives $\operatorname{Var}(\hat\beta_1)=\sigma^2\sum_ik_i^2=\sigma^2/S_{xx}$.
For the intercept, $\hat\beta_0=\bar y-\hat\beta_1\bar x$ and $\operatorname{Cov}(\bar y,\hat\beta_1)=\frac{\sigma^2}{n}\sum_ik_i=0$, so $\operatorname{Var}(\hat\beta_0)=\frac{\sigma^2}{n}+\bar x^2\,\frac{\sigma^2}{S_{xx}}$.
Ten data sets drawn from the same true line $\textcolor{#8250df}{y=1+0.8x}$, eleven inputs $x=0,1,\dots,10$ and noise with $\sigma=1.5$, each fitted by least squares. The $\textcolor{#1f6feb}{\text{fitted lines}}$ have slopes from $0.58$ to $0.94$ around the true $0.8$. At $x=5$ their heights span about $1.3$; at $x=0$ and $x=10$ they span about $2.3$ and $2.2$.
Looks like this, but is not
Unbiased sounds like accurate: since $E[\hat\beta_1]=\beta_1^{\mathrm{true}}$, the café slope $1.4$ should sit close to the truth.
Unbiased is a promise about the average over many samples, not about your one sample. The café slope has a standard error of about $0.35$, so a miss of that size is ordinary; only more data or more spread in $x$ make a single estimate close.
design
inputs x
sum of (x − 5)²
Var of slope
SE of slope
A
$4,5,5,6$
$2$
$0.5$
$0.707$
B
$3,4,6,7$
$10$
$0.1$
$0.316$
C
$1,3,7,9$
$40$
$0.025$
$0.158$
D
$1,1,9,9$
$64$
$0.0156$
$0.125$
Same four runs, same noise: moving the inputs apart cuts the standard error from $0.707$ to $0.125$, a factor of about $5.7$. Design D bets everything on the line being straight between $1$ and $9$.
How sure is the café about 140 cups per post? 95% intervals for both coefficients
The café fit is $\hat y=1.4+1.4x$ with $\mathrm{RSS}=3.6$ from $n=5$ weeks, $\bar x=3$ and $S_{xx}=10$. Estimate $\sigma$ and give 95% intervals for $\beta_1^{\mathrm{true}}$ and $\beta_0^{\mathrm{true}}$.
Ratio check: $\mathrm{SE}(\hat\beta_0)/\mathrm{SE}(\hat\beta_1)=\sqrt{S_{xx}/n+\bar x^2}=\sqrt{2+9}\approx 3.32$, and indeed $1.149/0.346\approx 3.32$.
The slope interval stays above $0$, so the five weeks do tie posts to sales; the intercept interval straddles $0$, so they say little about a week with no posts.
Four calibration runs: where to put them to pin down a slope
A sensor's output is calibrated with four runs at temperatures of your choice between $1$ and $9$ °C. The noise is known, $\sigma=0.5$. Compare the standard error of the slope for design A, $x=4,5,5,6$, and design C, $x=1,3,7,9$.
Find$\mathrm{SE}(\hat\beta_1)$ for both designs and their ratio.
Given
$\sigma=0.5$ known
design A: $x=4,5,5,6$
design C: $x=1,3,7,9$
Solution
$\operatorname{Var}(\hat\beta_1)=\sigma^2/S_{xx}$ does not contain the responses, so we can compare designs before running a single measurement.
Units check: $\sigma$ is in output units and $\sqrt{S_{xx}}$ in °C, so both standard errors are in output units per °C, the units of a slope.
Spreading the inputs is free precision, as long as the relationship stays straight over the wider range; check that before trusting the far ends.
Checkpoint
§03.2 — spreading the inputs
A study measures reaction time $y$ at four noise levels $x=3,4,6,7$. A second study uses $x=1,3,7,9$: the same mean, with every distance from it doubled. The noise level $\sigma$ is the same.
Find(a) How does $\mathrm{SE}(\hat\beta_1)$ in the second study compare with the first?
Given
study 1: $x=3,4,6,7$
study 2: $x=1,3,7,9$
same $\sigma$ in both
Hint 1/4
The standard error of the slope depends on the inputs only through their spread, $S_{xx}$.
Hint 2/4
$\mathrm{SE}(\hat\beta_1)=\sigma/\sqrt{S_{xx}}$.
Hint 3/4
Study 1 has deviations $-2,-1,1,2$, so $S_{xx}=10$; study 2 has $-4,-2,2,4$, so $S_{xx}=40$.
Hint 4/4
$S_{xx}$ is four times larger, so the standard error is $\sqrt{1/4}=1/2$ as large.
Show solution
Only $S_{xx}$ changes between the studies, so we compare it and nothing else.
Stack every residual into one vector; the RSS is its squared length. Setting its derivative to zero gives $p+1$ linear equations, the normal equations, and when the columns of $X$ are independent they have exactly one solution.
Derivation with the course's matrix derivative rules
The rules, where a derivative of a scalar by a column vector is a row: (1) $h=Ag\Rightarrow \partial h/\partial g=A$; (2) $\alpha=y^TAg\Rightarrow \partial\alpha/\partial g=y^TA$ and $\partial\alpha/\partial y=g^TA^T$; (3) $\alpha=y^TAy\Rightarrow \partial\alpha/\partial y=y^T(A+A^T)$. Here $y$ is any vector.
Expand: $\mathrm{RSS}(\beta)=y^Ty-2y^TX\beta+\beta^TX^TX\beta$. The two cross terms $y^TX\beta$ and $\beta^TX^Ty$ are equal, since a number is its own transpose.
Rule 2 on the middle term and rule 3 with $A=X^TX$ on the last: $$\frac{\partial\,\mathrm{RSS}}{\partial\beta}=-2y^TX+\beta^T\big(X^TX+(X^TX)^T\big)=-2y^TX+2\beta^TX^TX.$$
Set it to zero and transpose: $X^TX\beta=X^Ty$. If $X^TXv=0$ then $\lVert Xv\rVert^2=v^TX^TXv=0$, so $Xv=0$ and, with independent columns, $v=0$. So $X^TX$ is invertible and the solution is unique.
It is the minimum: for any $\beta$, $\mathrm{RSS}(\beta)=\mathrm{RSS}(\hat\beta)+\lVert X(\beta-\hat\beta)\rVert^2$, because the cross term is $2(\hat\beta-\beta)^TX^T(y-X\hat\beta)=0$ by the normal equations.
Shapes in the normal equations for $n=200$ rows and $p=3$ inputs. $X$ is $200\times 4$, but $\textcolor{#1f6feb}{X^TX}$ is only $4\times 4$ and $\textcolor{#1f6feb}{X^Ty}$ is $4\times 1$: adding rows never makes the system to solve bigger.
Looks like this, but is not
Cancel the inverse: $(X^TX)^{-1}X^Ty=X^{-1}(X^T)^{-1}X^Ty=X^{-1}y$, so least squares is just $X^{-1}y$.
$X$ is $n\times(p+1)$ with more rows than columns, so it has no inverse, and $(AB)^{-1}=B^{-1}A^{-1}$ needs square, invertible factors. Only when $n=p+1$ is $X^{-1}y$ legal, and then the fit passes through every point with nothing averaged.
The café fit again, in matrix form
Write the café data ($x=1,\dots,5$, $y=3,5,4,7,9$) as $X$ and $y$, and compute $\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$.
The scalar formulas gave $\hat\beta_1=S_{xy}/S_{xx}=14/10$ and $\hat\beta_0=5.6-1.4\cdot 3$, the same two numbers by a different road.
For one input the matrix route and the centered-sum route are the same computation; the matrix route is the one that keeps working when more columns arrive.
A bakery's two-factor test: three coefficients from one diagonal system
A bakery codes oven temperature as $x_1=-1$ (low) or $+1$ (high) and proofing time as $x_2=-1$ (short) or $+1$ (long), plus one run at the middle settings $(0,0)$. Loaf heights in cm: $(-1,-1)\to 6.0$, $(1,-1)\to 7.2$, $(-1,1)\to 6.8$, $(1,1)\to 8.4$, $(0,0)\to 7.1$. Fit $\hat y=\hat\beta_0+\hat\beta_1x_1+\hat\beta_2x_2$.
$X^Te=0$ holds: $\sum e_i=0$, $\sum x_{i1}e_i=-0.1-0.1+0.1+0.1=0$ and $\sum x_{i2}e_i=-0.1+0.1-0.1+0.1=0$.
Going from low to high temperature is $2$ coded units, so it adds about $2\times 0.7=1.4$ cm; coded inputs make that reading immediate.
Rule 3 on a 2 × 2 matrix that is not symmetric
Take $A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$ and $\alpha=y^TAy$ for $y=[y_1,y_2]^T$. Compute $\partial\alpha/\partial y$ directly and compare it with the rule $y^T(A+A^T)$ and with the tempting shortcut $2y^TA$.
Find$\partial\alpha/\partial y$, a $1\times 2$ row.
Given
$A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$
$\alpha=y^TAy$
Solution
Expanding a 2 × 2 quadratic form takes one line, and it tests the rule instead of trusting it.
Expand the quadratic form
$$\alpha=y_1^2+2y_1y_2+3y_2^2$$
Only $a_{12}=2$ couples the two coordinates, because $a_{21}=0$.
3.4The geometry of least squares: projecting y onto the columns of X
Shows that $\hat y$ is the point of the column space closest to $y$, so the residual is perpendicular to every column of $X$.
The normal equations $X^T(y-X\hat\beta)=0$ can be read as a picture, and the picture explains why residuals sum to zero.
TheoremLeast squares as an orthogonal projection
Conditions
$X$ has full column rank, with columns $w_0=\mathbf 1,\ w_1,\dots,w_p$ in $\mathbb R^n$
$e=y-\hat y$ is the residual vector
$$\boxed{\begin{aligned}\textcolor{#1f6feb}{\hat y}&=Hy,\qquad H=X(X^TX)^{-1}X^T\\ X^T\textcolor{#d1690a}{e}&=0\quad\big(w_j^T\textcolor{#d1690a}{e}=0\ \text{for every column}\big)\\ H^T&=H,\qquad H^2=H\end{aligned}}$$
Of all vectors $X\beta$, the fitted vector is the one closest to $y$, reached by dropping a perpendicular. The residual is perpendicular to every column, so with an intercept the residuals sum to zero. Projecting twice changes nothing, which is what $H^2=H$ says.
Why the residual is perpendicular, and why H is a projection
Perpendicular: $X^Te=X^Ty-X^TX\hat\beta=0$ by the normal equations; row $j$ reads $w_j^Te=0$.
Symmetric and idempotent: $H^T=X\big((X^TX)^{-1}\big)^TX^T=H$ because $X^TX$ is symmetric, and $$H^2=X(X^TX)^{-1}\underbrace{X^TX(X^TX)^{-1}}_{I}X^T=H.$$
Closest point: for any $\beta$, $\lVert y-X\beta\rVert^2=\lVert e\rVert^2+\lVert X(\hat\beta-\beta)\rVert^2$ by Pythagoras, since $e$ is perpendicular to $X(\hat\beta-\beta)$.
With an intercept, $w_0=\mathbf 1$ gives $\sum_ie_i=0$. Averaging $y_i=\hat y_i+e_i$ then shows that the fitted line passes through $(\bar x,\bar y)$.
This picture lives in $\mathbb R^n$, one axis per observation, not in the $(x,y)$ scatter plot. The plane holds every vector $X\beta$; $\textcolor{#1f6feb}{\hat y=Hy}$ is its point closest to $y$, and the residual $\textcolor{#d1690a}{e=y-\hat y}$ meets the plane at a right angle.
Looks like this, but is not
If the residual is perpendicular to the columns, then in the scatter plot each residual should be perpendicular to the fitted line.
In the scatter plot residuals are vertical segments, and least squares measures them vertically. The right angle is between two vectors in $\mathbb R^n$: the residual vector and a column of $X$. Minimizing perpendicular distances to a line is a different method with a different answer.
Three points and the hat matrix written out in full
Fit $y=\beta_0+\beta_1x$ to the points $(-1,1)$, $(0,2)$, $(1,6)$. Compute $\hat\beta$, $\hat y$, $e$ and the $3\times 3$ matrix $H$, and check $X^Te=0$.
Find$\hat\beta$, $\hat y$, $e$, $H$
Given
$x=(-1,0,1)$
$y=(1,2,6)$
Solution
The inputs already have mean $0$, so $X^TX$ is diagonal and $H$ can be built from two outer products instead of a full matrix inverse.
Every row of $H$ sums to $1$: a constant vector is already a column of $X$, so it projects to itself. And $\operatorname{tr}H=\tfrac56+\tfrac13+\tfrac56=2$, the number of columns.
To test a fit, multiply the residuals by $X^T$; to project anything onto the same columns, reuse $H$ without refitting.
Is this the least squares line? Test the residuals instead of refitting
For the café data ($x=1,\dots,5$, $y=3,5,4,7,9$) three lines are proposed: $1.5+1.5x$, $2+1.2x$ and $1.4+1.4x$. Decide which one is the least squares line using only $X^Te=0$.
FindThe candidate with $\sum e_i=0$ and $\sum x_ie_i=0$.
Given
$x=1,2,3,4,5$
$y=3,5,4,7,9$
candidates $1.5+1.5x$, $2+1.2x$, $1.4+1.4x$
Solution
Two sums per candidate are cheaper than solving the normal equations, and they test the exact property that defines the fit.
Both normal equations hold, and they have a unique solution here.
Answer $$\boxed{\hat y=1.4+1.4x\ \text{is the least squares line}}$$
Check
Its RSS is $3.6$, below the $4.5$ and $4.0$ of the other two, as the minimum must be.
A zero residual sum only fixes the height of the line; orthogonality to each input column is what fixes the slopes.
Checkpoint
§03.4 — what the normal equations force
A least squares line is fitted to $n$ points $(x_i,y_i)$ with an intercept, giving residuals $e_i=y_i-\hat y_i$. The true errors are $\varepsilon_i=y_i-\beta_0^{\mathrm{true}}-\beta_1^{\mathrm{true}}x_i$.
Find(a) Which of these sums is zero for every data set?
Givenleast squares fit of $y$ on one input $x$, with an intercept
Hint 1/4
Ask what the normal equations say about the residual vector, not about the data.
Hint 2/4
$X^Te=0$: the residuals are orthogonal to every column of $X$.
Hint 3/4
Here the columns are $\mathbf 1$ and $x=(x_1,\dots,x_n)$, so $\sum_ie_i=0$ and $\sum_ix_ie_i=0$.
Hint 4/4
$\sum_ix_ie_i$ is always zero; the other three are not.
Show solution
The normal equations are exactly the statement $X^Te=0$, so we read off its rows.
3.5Why squares: Gauss-Markov and maximum likelihood
Gives two reasons for least squares: the smallest variance among linear unbiased estimators, and the maximum likelihood answer when the noise is Gaussian.
So far squaring the misses was our choice; this block shows two settings in which that choice is forced on us.
TheoremTwo justifications of least squares
Conditions
$y_i=(\beta^{\mathrm{true}})^Tx_i+\varepsilon_i$, with $\beta^{\mathrm{true}}$ fixed and unknown, the $x_i$ fixed and observed
Gauss-Markov: $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$, $\operatorname{Cov}(\varepsilon_i,\varepsilon_j)=0$ for $i\ne j$
likelihood: in addition, the $\varepsilon_i$ are i.i.d. $\mathcal N(0,\sigma^2)$
$$\boxed{\begin{aligned}&\text{(i) Gauss-Markov: }\operatorname{Var}(\tilde\beta_j)\ge\operatorname{Var}(\textcolor{#1f6feb}{\hat\beta_j})\\ &\qquad\text{for every linear unbiased }\tilde\beta_j=\textstyle\sum_k c_{kj}y_k\\ &\text{(ii) Gaussian noise: }\hat\beta_{\mathrm{MLE}}=\textcolor{#1f6feb}{\hat\beta_{\mathrm{RSS}}}\end{aligned}}$$
Among estimators that are linear in $y$ and right on average, least squares wobbles least: it is the best linear unbiased estimator, . If the noise is also Gaussian, the log-likelihood is a constant minus RSS over $2\sigma^2$, so the most likely $\beta$ is the least squares one, whatever $\sigma$ is.
The likelihood calculation, and a sketch of Gauss-Markov
Each $y_i$ has density $\frac{1}{\sqrt{2\pi}\,\sigma}\exp\!\big(-\frac{(y_i-\beta^Tx_i)^2}{2\sigma^2}\big)$, and independence multiplies them. Taking logs: $$l(\beta)=n\log\frac{1}{\sqrt{2\pi}\,\sigma}-\frac{1}{2\sigma^2}\sum_{i=1}^n(y_i-\beta^Tx_i)^2.$$
The first term has no $\beta$ in it and $\frac{1}{2\sigma^2}>0$, so maximizing $l$ is minimizing RSS. Since $\log$ is increasing, maximizing $L$ and $l$ is the same thing.
Gauss-Markov, sketched as further reading: a linear estimator of $a^T\beta$ is $c^Ty$. It is unbiased for every $\beta$ exactly when $X^Tc=a$.
Split $c=c_0+d$ with $c_0=X(X^TX)^{-1}a$, the least squares weights. Then $X^Td=0$, so $c_0^Td=0$ and $$\operatorname{Var}(c^Ty)=\sigma^2\big(\lVert c_0\rVert^2+\lVert d\rVert^2\big)\ge\sigma^2\lVert c_0\rVert^2.$$
Gaussian noise drawn sideways around the fitted line $\textcolor{#1f6feb}{\hat y=1.4+1.4x}$ at $x=1,3,5$, with $\textcolor{#8250df}{\sigma\approx 1.1}$. Each $\textcolor{#d1690a}{\text{bar}}$ is the density at the observed point. The miss $e=-1.6$ lands in the thin tail, and its log-density drops by $e^2/(2\sigma^2)$: squares enter the likelihood on their own.
Looks like this, but is not
Gauss-Markov says least squares is the best estimator of $\beta$, so no estimator can have a smaller variance.
The theorem only compares estimators that are linear in $y$ and unbiased. The estimator that always answers $0$ has variance $0$; it is excluded because it is biased, not beaten.
Two café lines scored by log-likelihood, at two noise levels
Assume the café data come from $y_i=\beta_0+\beta_1x_i+\varepsilon_i$ with i.i.d. $\varepsilon_i\sim\mathcal N(0,\sigma^2)$. Compare $l(\beta)$ for the least squares line $1.4+1.4x$ (RSS $3.6$) and the line $1.5+1.5x$ (RSS $4.5$), first with $\sigma=1$, then with $\sigma=2$.
FindThe four log-likelihoods and which line wins at each $\sigma$.
The log-likelihood depends on $\beta$ only through the RSS, so we never need the individual densities; two RSS values and one constant per $\sigma$ are enough.
Answer $$\boxed{\text{least squares wins at both: } -6.395>-6.845,\ \ -8.510>-8.623}$$
Check
The gap in log-likelihood is $(4.5-3.6)/(2\sigma^2)$: $0.45$ at $\sigma=1$ and $0.1125$ at $\sigma=2$, matching the differences above.
$\sigma$ changes how sharply the likelihood prefers one line, never which line it prefers: the maximizer is the RSS minimizer.
Two unbiased slopes for the café weeks, and which one wobbles less
For the café design $x=1,\dots,5$ compare the least squares slope with the end-point slope $\tilde\beta_1=(y_5-y_1)/(x_5-x_1)$ under the Gauss-Markov assumptions.
FindWhether $\tilde\beta_1$ is linear and unbiased, and its variance against $\sigma^2/S_{xx}$.
Gauss-Markov predicts exactly this ordering, and on the actual data the two estimates are $1.5$ and $1.4$: close, but only one of them uses all five weeks.
Being unbiased is cheap; among the unbiased linear rules, least squares is the one with the smallest spread.
Checkpoint
§03.5 — what Gauss-Markov needs
A lab fits $y_i=\beta^Tx_i+\varepsilon_i$ by least squares and wants to call the estimate BLUE. It is not sure which properties of its noise it has to check first.
Find(a) Which of these assumptions is not needed for the ?
Givenmodel $y_i=(\beta^{\mathrm{true}})^Tx_i+\varepsilon_i$ with fixed inputs
Hint 1/4
List what the theorem's proof uses: means, variances and covariances of the noise.
Hint 2/4
Gauss-Markov assumes $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$ and $\operatorname{Cov}(\varepsilon_i,\varepsilon_j)=0$ for $i\ne j$.
Hint 3/4
Compare each option with those three conditions; skewed noise can still satisfy all of them.
Hint 4/4
The shape of the distribution is never used, so Gaussian noise is the assumption it does not need.
Show solution
The theorem is stated through first and second moments only, so we check each option against those.
Fit the line or plane to the zeros and ones, and predict class 1 wherever it lies above one half. Where it equals one half is the boundary, flat in input space. The lecture advises against this in general: fitted values can leave $[0,1]$, and three classes coded $1,2,3$ get an order they may not have.
Six students, fail $=0$ as hollow circles and pass $=1$ as filled ones. The least squares line $\textcolor{#1f6feb}{\hat y=-0.2+0.2x}$ crosses the $\textcolor{#d1690a}{0.5}$ cut at $\textcolor{#d1690a}{x=3.5}$ hours, the decision boundary; the students at $3$ and $4$ hours end up on the wrong side.
Looks like this, but is not
A fitted value of $0.8$ reads like an 80% chance of passing.
The fit is a straight line, not a probability model. The same line gives $-0.2$ at $0$ hours and $1.4$ at $8$ hours, values no probability can take. Only the side of $0.5$ is used.
Pass or fail from hours of study: a boundary at 3.5 hours
Six students studied $x=1,2,3,4,5,6$ hours; the outcomes, coded pass $=1$ and fail $=0$, were $y=0,0,1,0,1,1$. Fit least squares, find the decision boundary and count the training mistakes.
FindThe fitted line, the boundary and the number of misclassified students.
Given
$x=1,2,3,4,5,6$
$y=0,0,1,0,1,1$
rule: pass if $\hat y>0.5$
Solution
One input, so the centered-sum formulas are the quickest route; the labels are just numbers to them.
$$\text{actual }(0,0,1,0,1,1):\ \text{students 3 and 4 are misclassified}$$
A straight boundary cannot separate a pass at 3 hours from a fail at 4 hours.
Answer $$\boxed{\hat y=-0.2+0.2x,\qquad \text{pass if } x>3.5,\qquad 2 \text{ of } 6 \text{ wrong}}$$
Check
The boundary sits exactly at $\bar x=3.5$, as it must here: $\bar y=0.5$ and the line passes through $(\bar x,\bar y)$.
The regression only decides which side of the cut a student falls on; its height above or below the cut is not a probability.
Two inputs: the boundary becomes a straight line in the plane
A fit on 0/1 labels for support tickets (urgent $=1$) gave $\hat y=-0.4+0.15x_1+0.1x_2$, with $x_1$ the number of customer replies and $x_2$ the hours open. Write the boundary as a line and classify the tickets $(2,5)$, $(4,4)$ and $(6,0)$.
FindThe boundary line and three classifications.
Given
$\hat y=-0.4+0.15x_1+0.1x_2$
tickets $(x_1,x_2)$: $(2,5)$, $(4,4)$, $(6,0)$
urgent if $\hat y>0.5$
Solution
Setting the fitted value to $0.5$ and solving for $x_2$ gives a line we can sketch and check points against.
Choose the columns first, then fit exactly as before: each column is a function of the inputs, and $\beta$ still enters linearly. Powers of $x$ give polynomial regression, a product $x_1x_2$ lets one input change the effect of another, and bumps centred at $\mu_j$ with width $s$ give local shapes.
Why distinct inputs decide uniqueness
For polynomial regression $X$ is the with rows $[1,\, \allowbreak x_i,\, \allowbreak x_i^2, \allowbreak \dots, \allowbreak x_i^p]$. If $Xv=0$, the polynomial $v_0+v_1t+\dots+v_pt^p$ vanishes at every $x_i$.
A nonzero polynomial of degree at most $p$ has at most $p$ roots. With $p+1$ distinct $x_i$ it must be the zero polynomial, so $v=0$: full column rank and a unique fit.
With only $p$ distinct inputs, the polynomial with exactly those roots gives $Xv=0$ for some $v\ne 0$: $X^TX$ is singular and many coefficient vectors share the same fitted values.
Five points with a U shape. The dashed straight line $\textcolor{#1f6feb}{2.4+0.2x}$ leaves $\mathrm{RSS}=14.8$; adding the column $x^2$ gives $\textcolor{#1f6feb}{0.4+0.2x+x^2}$ with $\mathrm{RSS}=0.8$. Same formula, one more column.
Looks like this, but is not
$y=\beta_0+\beta_1x+\beta_2x^2$ draws a curve, so it is a nonlinear model and needs a new fitting method.
Linear refers to $\beta$, not to $x$: once $x_i^2$ is computed it is one more number in the table. In $y=\beta_0+e^{\beta_1x}$ the coefficient sits inside the exponential, no column can be computed in advance, and $(X^TX)^{-1}X^Ty$ does not apply.
degree
columns in X
RSS
$0$
$1$
$15.2$
$1$
$2$
$14.8$
$2$
$3$
$0.8$
$3$
$4$
$0.7$
$4$
$5$
$0$
$5$
$6$
no unique fit
RSS never goes up when a column is added, because the smaller model is the bigger one with a coefficient set to $0$. At degree $4$ the curve passes through all five points; at degree $5$ there are six coefficients and only five distinct inputs.
A quadratic through five U-shaped points
Fit $\hat y=\hat\beta_0+\hat\beta_1x+\hat\beta_2x^2$ to $(-2,4)$, $(-1,1)$, $(0,1)$, $(1,1)$, $(2,5)$ and compare its RSS with the straight line's $14.8$.
Find$\hat\beta$ and the RSS of the quadratic.
Given
$x=-2,-1,0,1,2$
$y=4,1,1,1,5$
straight-line fit: $2.4+0.2x$, RSS $14.8$
Solution
The inputs are symmetric about $0$, so every odd power sums to zero and the $3\times 3$ system splits into a $1\times 1$ and a $2\times 2$ piece.
$X^Te=0$ for all three columns: $\sum e_i=0$, $\sum x_ie_i=0.2-0.6+0.4=0$ and $\sum x_i^2e_i=-0.2-0.6+0.8=0$.
A curve cost one extra column and nothing else: the fitting method did not change.
An interaction term: when proofing time changes what temperature does
Return to the bakery runs $(x_1,x_2,y)$: $(-1,-1,6.0)$, $(1,-1,7.2)$, $(-1,1,6.8)$, $(1,1,8.4)$, $(0,0,7.1)$. Add the column $x_3=x_1x_2$ and fit $\hat y=\hat\beta_0+\hat\beta_1x_1+\hat\beta_2x_2+\hat\beta_3x_1x_2$. How much does one coded unit of temperature add at short and at long proofing?
Find$\hat\beta_3$ and the effect of $x_1$ at $x_2=-1$ and $x_2=+1$.
The interaction makes the slope in $x_1$ depend on $x_2$.
$$x_2=-1:\ 0.6,\qquad x_2=+1:\ 0.8$$
Short proofing, then long proofing.
Answer $$\boxed{\hat\beta_3=0.1:\ \ 0.6\ \text{cm per unit at short proofing},\ 0.8\ \text{at long}}$$
Check
The four corners are now fitted exactly, for example $7.1-0.7-0.5+0.1=6.0$ at $(-1,-1)$, and the centre run gives $7.1$: RSS falls from $0.04$ to $0$.
With an interaction in the model, never read $\hat\beta_1$ alone as the effect of $x_1$; it is the effect where $x_2=0$.
Two Gaussian bumps as columns: building X and predicting
Use basis functions $\phi_1(x)=e^{-(x-1)^2/2}$ and $\phi_2(x)=e^{-(x-3)^2/2}$ (centres $\mu_1=1$, $\mu_2=3$, $s=1$) plus the constant $\phi_0=1$. Build $X$ for inputs $x=0,1,2,3,4$, then evaluate the fitted model $\hat y=0.5+2\phi_1(x)-\phi_2(x)$ at $x=1.5$.
Find$X$ and $\hat y(1.5)$.
Given
$\mu_1=1$, $\mu_2=3$, $s=1$
inputs $x=0,1,2,3,4$
fit $\hat\beta=(0.5,\ 2,\ -1)$
Solution
Each column is one basis function evaluated at every input, so we fill the matrix column by column.
A new input gets its own row $[1,\ \phi_1(1.5),\ \phi_2(1.5)]$.
Answer $$\boxed{\hat y(1.5)\approx 1.94}$$
Check
Order check: $x=1.5$ is close to the first centre and far from the second, so the first bump dominates and $\hat y$ should sit near $0.5+2(0.9)$ minus a little, as it does.
The bumps are fixed before fitting; only their heights $\hat\beta_j$ are learned, which is why the model stays linear.
Checkpoint
§03.7 — repeated inputs and a unique fit
A chemist measured a reaction at temperatures $x=1, \allowbreak 1, \allowbreak 2, \allowbreak 2, \allowbreak 3, \allowbreak 3, \allowbreak 3$: seven runs, but only three different temperatures. She wants to fit a polynomial of degree $p$ by least squares.
Find(a) For which degrees $p$ is the least squares fit unique?
⚠ Reading a main effect alone when an interaction is present
without the product term, $\hat\beta_1$ was the effect of $x_1$, and the habit stays
wrong$$\text{effect of }x_1=\hat\beta_1$$
right$$\text{effect of }x_1=\hat\beta_1+\hat\beta_3x_2$$
Step the degree from $0$ to $4$ on the same five points. Each step adds one column to $X$, and $\textcolor{#d1690a}{\mathrm{RSS}}$ can only go down: $15.2,\ \allowbreak 14.8,\ \allowbreak 0.8,\ \allowbreak 0.7,\ \allowbreak 0$.
At the edges
degree 0 RSS 15.2
Only the column of ones: the fit is the flat line at the mean, 2.4.
degree 4 RSS 0
Five coefficients for five distinct points: the curve passes through all of them. A zero RSS here says nothing about a sixth point.
degree 5 no unique fit
Six coefficients but only five distinct inputs: the normal equations have infinitely many solutions.
Simple regression by hand
A table of pairs, one input, and a question about the line, a prediction or the slope's precision.
Means
$\bar x$ and $\bar y$.
Centered sums
$S_{xx}$ and $S_{xy}$ from the deviation columns; each deviation column sums to zero.
Coefficients
$\hat\beta_1=S_{xy}/S_{xx}$, then $\hat\beta_0=\bar y-\hat\beta_1\bar x$.
Check
$\sum e_i=0$ and $\sum x_ie_i=0$; then $\mathrm{RSS}=\sum e_i^2$.
Uncertainty
$\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$ and $\hat\beta_1\pm 2\,\mathrm{RSE}/\sqrt{S_{xx}}$.
Where it goes wrong
Raw sums $\sum x_iy_i$ where centered ones belong.
Dividing the RSS by $n$ instead of $n-2$.
Least squares with a
Two or more inputs, a polynomial or basis model, or any question phrased with $X$ and $y$.
Build X
A column of ones first, then one column per input or basis function.
Two products
$X^TX$ and $X^Ty$; the entries are sums such as $\sum_ix_{ij}x_{ik}$ and $\sum_ix_{ij}y_i$.
Rank
Check that the columns are independent: a nonzero determinant, or enough distinct inputs for a polynomial.
Solve
Use structure first: orthogonal columns make $X^TX$ diagonal; otherwise invert or eliminate.
Check
$X^Te=0$, one equation per column.
Where it goes wrong
Forgetting the column of ones.
Writing $XX^T$ where $X^TX$ belongs.
Reading the entries of $\hat\beta$ in a different order from the columns.
Classifying with a least squares fit
Two classes coded 0 and 1, and a question about a boundary or a prediction.
Code
One class $1$, the other $0$.
Fit
$\hat\beta_{\mathrm{RSS}}=(X^TX)^{-1}X^Ty$ as for any response.
Boundary
Solve $\hat\beta^Tx=0.5$.
Classify
Class 1 where $\hat\beta^Tx>0.5$, class 0 otherwise.
Sanity
Fitted values outside $[0,1]$ can occur and are not probabilities.
Where it goes wrong
Cutting at $0$ instead of $0.5$.
Coding three classes as $1,2,3$ and trusting the order.
With an intercept: the café slope is 1.4
Fit $y=\beta_0+\beta_1x$ to the café data $x=1,\dots,5$, $y=3,5,4,7,9$.
Orthogonality to the one column still holds: $\sum x_ie_i=98-\hat\beta\cdot 55=0$.
Same five weeks: slope $1.4$ with an intercept, $1.78$ without one. Forcing the line through the origin makes the slope do the intercept's job, and the RSS rises from $3.6$ to $5.38$.
How to tell them apart
Read the model before the formula. A column of ones means centered sums, $S_{xy}/S_{xx}$; no intercept means raw sums, $\sum x_iy_i/\sum x_i^2$, and residuals that need not sum to zero.
Residuals: misses from the fitted line
For the café data and the fitted line $1.4+1.4x$, compute the residuals, their sum and their sum of squares.
Find$\sum e_i$ and $\sum e_i^2$.
Given
$x=1,\dots,5$, $y=3,5,4,7,9$
fitted line $1.4+1.4x$
Solution
Residuals are computed from quantities we have: the data and the fit.
Both normal equations hold: $\sum e_i=0$ and $\sum x_ie_i=0$.
Errors: misses from the true line
Suppose, as in a simulation, the true line is known to be $1+1.5x$. For the same café data compute $\varepsilon_i=y_i-(1+1.5x_i)$, their sum and their sum of squares.
Find$\sum\varepsilon_i$ and $\sum\varepsilon_i^2$.
Given
$x=1,\dots,5$, $y=3,5,4,7,9$
true line $1+1.5x$
Solution
Errors are measured from the true line, which only a simulation lets us see.
$3.75\ge 3.6$, as it must be: the least squares line has the smallest RSS of all lines, the true one included.
Residuals come from the fitted line and sum to zero; errors come from the true line and need not. Because least squares minimizes, $\sum e_i^2\le\sum\varepsilon_i^2$ for every data set.
How to tell them apart
If you can compute it from the data alone, it is a residual $e_i$; if it needs $\beta^{\mathrm{true}}$, it is an error $\varepsilon_i$. Squared residuals run small, which is why the RSE divides by $n-2$ rather than $n$.
Scaffolding comes off
The common skeleton
Means: $\bar x$ and $\bar y$.
Centered sums: $S_{xx}=\sum(x_i-\bar x)^2$ and $S_{xy}=\sum(x_i-\bar x)(y_i-\bar y)$.
Coefficients: $\hat\beta_1=S_{xy}/S_{xx}$, then $\hat\beta_0=\bar y-\hat\beta_1\bar x$.
Residual checks: $\sum e_i=0$ and $\sum x_ie_i=0$, then $\mathrm{RSS}=\sum e_i^2$.
Uncertainty: $\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$ and $\hat\beta_1\pm 2\,\mathrm{RSE}/\sqrt{S_{xx}}$.
1 · fully worked
Heating load against outdoor temperature: the full skeleton on six days
An office building's daily peak heating load $y$ (kW) and the day's mean outdoor temperature $x$ (°C) over six autumn days: $x=10, \allowbreak 12, \allowbreak 14, \allowbreak 16, \allowbreak 18, \allowbreak 20$ and $y=52, \allowbreak 50, \allowbreak 46, \allowbreak 46, \allowbreak 42, \allowbreak 40$. Fit the line, check it, and give a 95% interval for the slope.
Find$\hat\beta_0$, $\hat\beta_1$, the residual checks and a 95% interval for $\beta_1^{\mathrm{true}}$.
Units: the slope is kW per °C, and a day $10$ °C warmer predicts $12$ kW less, from $52$ at $10$ °C to $40$ at $20$ °C, matching the end points of the table.
Five steps in the same order every time; the two residual sums in step 4 catch slips in steps 1 to 3.
2 · you write the reasoning
An easier one, and this time you write the reasons. Three points: $(0,1)$, $(1,3)$, $(2,2)$. Fit the least squares line and check it.
$\bar x=1,\qquad \bar y=2$
reasoning
The slope formula uses deviations from the means, so the means come first.
A line with an intercept must leave residuals orthogonal to $\mathbf 1$ and to $x$; if either sum were not zero, a slip happened above.
3 · find the buried error
Harder, with the work done for you and two errors buried in it. A motor is run at voltages $x=2,4,6,8,10$ V and its speed is $y=11, \allowbreak 19, \allowbreak 32, \allowbreak 38, \allowbreak 50$ hundred rpm. Fit the line and give a 95% interval for the slope.
Step 1. $\bar x=6,\qquad \bar y=30$. Means of the two columns.
Step 2. $S_{xx}=40,\qquad S_{xy}=194$. Centered sums from the deviations $(-4,-2,0,2,4)$ and $(-19, \allowbreak -11, \allowbreak 2, \allowbreak 8, \allowbreak 20)$.
Step 3.$\hat\beta_1=194/40=4.85,\qquad \hat\beta_0=30-4.85\cdot 6=0.9$. Slope, then the intercept through the means.
Step 4.$e=(0.4,\ \allowbreak -1.3,\ \allowbreak 2.0,\ \allowbreak -1.7,\ \allowbreak 0.6)$ and $\mathrm{RSS}=9.1$. Both normal equations hold: $\sum e_i=0$ and $\sum x_ie_i=0.8-5.2+12-13.6+6=0$.
Step 5.$\mathrm{RSE}=\sqrt{9.1/5}\approx 1.349$. The typical size of a residual.
Step 6.$\mathrm{SE}(\hat\beta_1)=1.349/\sqrt 5\approx 0.603$. Standard error: spread over the square root of the sample size.
Step 7.$4.85\pm 2(0.603)=[3.64,\ 6.06]$. Two standard errors each way.
the two buried errors (2)
⚠ step 5
The RSE divides the RSS by $n-2=3$, not by $n=5$: two coefficients were fitted from the same five points.
An average divides by $n$, and the RSE looks like the root of an average squared residual.
right
$\mathrm{RSE}=\sqrt{9.1/3}\approx 1.742$.
⚠ step 6
The slope's standard error divides by $\sqrt{S_{xx}}=\sqrt{40}$, not by $\sqrt n$.
$\sigma/\sqrt n$ is the standard error of a sample mean from the previous section, and it gets reused for a slope.
right
$\mathrm{SE}(\hat\beta_1)=1.742/\sqrt{40}\approx 0.275$, so the interval is $4.85\pm 0.551=[4.30,\ 5.40]$.
4 · the bare problem
§03.2 — a courier's minutes per kilometre
A courier logs four deliveries: the distance $x$ in km and the time $y$ in minutes.
Find
(a) Fit the least squares line.
(b) Give a 95% interval for the minutes per km.
Given
$x=1,3,5,7$
$y=12,17,26,29$
Hint 1/4
Run the whole skeleton: means, centered sums, coefficients, residual checks, then the interval.
Size check: $7$ km at $3$ minutes per km plus $9$ minutes of fixed time is $30$ minutes, next to the $29$ observed.
The intercept reads as fixed time per delivery, the slope as minutes per km; name both in units when you report them.
Full exam-style question
A spring through the origin: derive, test and use the no-intercept slopeexam format
Loads $x=1,2,3,4$ N stretch a spring by $y=2.1, \allowbreak 3.9, \allowbreak 6.2, \allowbreak 7.8$ mm. Zero load gives zero stretch, so fit $y_i=\beta x_i+\varepsilon_i$ with $E[\varepsilon_i]=0$, $\operatorname{Var}(\varepsilon_i)=\sigma^2$ and uncorrelated noise.
(a) Derive the least squares $\hat\beta$.
(b) Show that it is unbiased and find its variance.
(c) Compute $\hat\beta$ and a 95% interval, taking $\sigma=0.2$ mm as known.
(d) Do the residuals sum to zero?
Find$\hat\beta$, its mean and variance, a 95% interval and $\sum_ie_i$.
Given
$x=1,2,3,4$ N
$y=2.1,\ \allowbreak 3.9,\ \allowbreak 6.2,\ \allowbreak 7.8$ mm
$\sigma=0.2$ mm, known, for part (c)
Solution
One derivative gives the estimator, and writing it as a weighted sum of the $y_i$ gives its mean and variance with the rules of the first section.
Units and size: $1.99$ mm per N predicts $7.96$ mm at $4$ N against $7.8$ observed, and $\sum x_ie_i=0$ confirms the normal equation.
Four parts, one idea each: a derivative, a weighted sum, the arithmetic, and orthogonality.
Drop the intercept and you lose $\sum e_i=0$; what survives is orthogonality to the columns you kept.
Practice
A · concept 4 questions
1§03.1 — a zero residual sum
A classmate offers a quick test for any fitted line: if its residuals add up to zero, it must be the least squares line. You try the claim on the café data.
Find(a) Is the claim true or false? Use the test line.
Given
claim: $\sum_ie_i=0$ implies the least squares line
café data $x=1,\dots,5$, $y=3,5,4,7,9$; least squares line $1.4+1.4x$ with RSS $3.6$
test line: $2+1.2x$
Hint 1/4
One line whose residuals sum to zero but whose RSS exceeds $3.6$ would break the claim.
Hint 2/4
Every line through $(\bar x,\bar y)$ has $\sum e_i=0$; least squares also needs $\sum x_ie_i=0$.
Hint 3/4
The test line gives $\hat y=3.2,\ \allowbreak 4.4,\ \allowbreak 5.6,\ \allowbreak 6.8,\ \allowbreak 8$ against $y=3,5,4,7,9$.
Hint 4/4
Its residuals $-0.2,\ \allowbreak 0.6,\ \allowbreak -1.6,\ \allowbreak 0.2,\ \allowbreak 1$ sum to $0$, yet its RSS is $4.0>3.6$: the claim is false.
Show solution
A claim about every line falls to a single counterexample, so we compute one.
The line passes through $(3,5.6)$, so the sum vanishes; the squares do not shrink with it.
Verdict
$$4.0>3.6\ \Rightarrow\ \text{not least squares}$$
A smaller RSS exists, so this line is not the minimizer.
Answer $$\boxed{\text{False}}$$
Check
The second normal equation fails too: $\sum x_ie_i=-0.2+1.2-4.8+0.8+5=2\ne 0$.
A zero residual sum only says the line passes through the point of means; the slope needs its own equation.
2§03.4 — which distance least squares measures
A figure in a blog post draws a short segment from each data point to the fitted line, at a right angle to the line, and says least squares makes these segments as short as possible.
Find(a) Is the claim true or false?
Givenclaim: least squares minimizes the sum of squared perpendicular distances from the points to the line
Hint 1/4
Check what RSS subtracts: which coordinate of a point is compared with the line?
Hint 2/4
$\mathrm{RSS}=\sum_i\big(y_i-(\beta_0+\beta_1x_i)\big)^2$ compares $y_i$ with the line at the same $x_i$.
Hint 3/4
For the café point $(3,4)$ and the line $1.4+1.4x$ the vertical miss is $-1.6$, while the perpendicular distance is $1.6/\sqrt{1+1.4^2}\approx 0.93$.
Hint 4/4
Least squares squares vertical misses, so the claim is false.
Show solution
One concrete point makes the difference between the two distances visible.
The school formula for the distance from a point to a line.
Verdict
$$e^2=2.56\ne d^2\approx 0.86$$
The two methods score the same point differently.
Answer $$\boxed{\text{False}}$$
Check
The vertical and perpendicular distances agree only for a flat line, $\hat\beta_1=0$, where $\sqrt{1+\hat\beta_1^2}=1$.
Least squares treats $x$ as known exactly and charges only the misses in $y$.
3§03.6 — three segments coded 1, 2, 3
A shop codes three customer segments as $y=1$ (students), $y=2$ (families) and $y=3$ (retirees), fits least squares on age, and rounds $\hat y$ to the nearest code.
Find(a) What is the real problem with this plan?
Given
codes: students $1$, families $2$, retirees $3$
input: customer age
Hint 1/4
Ask what the numbers 1, 2 and 3 claim about the segments beyond naming them.
Hint 2/4
Coding more than two classes on the number line forces an order and a spacing; the lecture names this as a reason not to classify this way.
Hint 3/4
Families are coded as the midpoint of students and retirees, so the fit treats them as halfway between.
Hint 4/4
The coding, not the fitting, is the problem: it invents an order and equal gaps.
Show solution
If codes were only names, relabelling could not change the fit; computing both slopes tests that directly.
Answer $$\boxed{26.6+3.8x\ \text{wins at both}\ \sigma}$$
Check
The gap is $(4-3.6)/(2\sigma^2)$: $0.8$ at $\sigma=0.5$ and $0.2$ at $\sigma=1$, matching the differences.
Only the RSS ranks lines; $\sigma$ sets how large the gap looks.
6§03.7 — checking a quadratic fit
A battery's capacity loss $y$ (percent) was measured at four temperature settings coded $x=0,1,2,3$, giving $y=3,1,1,5$. A proposed quadratic fit is $\hat y=3.1-3.9x+1.5x^2$.
Find
(a) Write the Vandermonde matrix $X$ and compute $X^TX$ and $X^Ty$.
(b) Show that the proposed $\hat\beta$ solves the normal equations.
(c) Compute the residuals and the RSS.
Given
$x=0,1,2,3$
$y=3,1,1,5$
proposed $\hat\beta=(3.1,\ -3.9,\ 1.5)$
Hint 1/4
Checking a proposed answer is cheaper than solving: multiply $X^TX$ by it and compare with $X^Ty$.
Hint 2/4
Rows of $X$ are $[1,\ x_i,\ x_i^2]$; the normal equations are $X^TX\hat\beta=X^Ty$.
Hint 3/4
$X^TX=\begin{bmatrix}4&6&14\\6&14&36\\14&36&98\end{bmatrix}$ and $X^Ty=(10,\ 18,\ 50)^T$ for $x=0,1,2,3$, $y=3,1,1,5$.
$x=3$ is equally far from the centres $2$ and $4$, so $\phi_2(3)=\phi_3(3)$, and the far bump at $0$ adds almost nothing.
Basis models predict like any linear model: build the row, then take a dot product.
8§03.7 — an interaction in an ad budget
A fitted sales model with an interaction is $\hat y=6+0.02x_1+0.03x_2+0.001x_1x_2$, where $x_1$ is online ad spend and $x_2$ radio ad spend, both in thousand TL, and $y$ is sales in thousand units.
Find
(a) How much do sales change when $x_1$ rises by $10$ at $x_2=10$?
(b) The same change at $x_2=50$?
(c) What does $\hat\beta_1=0.02$ mean on its own?
Given$\hat y=6+0.02x_1+0.03x_2+0.001x_1x_2$
Hint 1/4
With an interaction the effect of $x_1$ depends on where $x_2$ is, so compute it at each $x_2$.
Hint 2/4
The change in $\hat y$ for $\Delta x_1$ at fixed $x_2$ is $(\hat\beta_1+\hat\beta_3x_2)\,\Delta x_1$.
Hint 3/4
$\hat\beta_1=0.02$, $\hat\beta_3=0.001$, $\Delta x_1=10$, and $x_2=10$ or $50$.
Hint 4/4
$0.3$ thousand units at $x_2=10$, $0.7$ at $x_2=50$; $\hat\beta_1$ alone is the effect when $x_2=0$.
Show solution
The slope in $x_1$ is a formula in $x_2$ here, so we evaluate the formula instead of reading one coefficient.
With $\bar x=0$: $\sum x_i(y_i-\bar y)=\sum x_iy_i-\bar y\sum x_i=\sum x_iy_i$. For the data, $\sum x_iy_i=12$, $\sum x_i^2=10$, $\bar y=3$.
Hint 4/4
$\hat\beta_0=\bar y$ and $\hat\beta_1=\sum x_iy_i/\sum x_i^2$; $\hat\beta_0=0$ when $\bar y=0$; here $\hat y=3+1.2x$ and $\operatorname{Var}(\hat\beta_0)=\sigma^2/5$.
Show solution
Setting $\bar x=0$ in the general formulas is quicker than re-deriving from the RSS.
Recentre $y$ as $y-3=(-2, \allowbreak -1, \allowbreak -1, \allowbreak 1, \allowbreak 3)$: the slope stays $\sum x_i(y_i-3)/10=12/10$ and the intercept becomes $0$, as (b) says.
Centring decouples intercept and slope and makes the intercept an average, with the smallest possible variance $\sigma^2/n$.
2§03.3 — two columns and no intercept
A sensor's output $y$ is modelled as $y=\beta_1x+\beta_2x^2+\varepsilon$ with no intercept, because zero input gives zero output. Calibration inputs are $x=-2,-1,1,2$ with outputs $y=1.1, \allowbreak -0.8, \allowbreak 2.4, \allowbreak 7.0$.
Find
(a) Write $X$ and the normal equations.
(b) Solve for $\hat\beta$.
(c) Check that the residuals are orthogonal to both columns. Must they sum to zero?
Orthogonality holds: $\sum x_ie_i=-0.2+0.3-0.1+0=0$ and $\sum x_i^2e_i=0.4-0.3-0.1+0=0$.
The normal equations promise orthogonality to the columns you included, and nothing more.
3§03.6 — a boundary in two inputs
A fit on 0/1 labels gave $\hat y=-1+0.5x_1+0.25x_2$, and the rule predicts class 1 when $\hat y>0.5$.
Find(a) Which line in the $(x_1,x_2)$ plane is the decision boundary?
Given
$\hat y=-1+0.5x_1+0.25x_2$
class 1 if $\hat y>0.5$
Hint 1/4
The boundary is where the fitted value equals the cut; solve that equation for $x_2$.
Hint 2/4
Boundary: $\hat\beta^Tx=0.5$.
Hint 3/4
$-1+0.5x_1+0.25x_2=0.5$.
Hint 4/4
$0.25x_2=1.5-0.5x_1$, so $x_2=6-2x_1$.
Show solution
Putting the cut on the right-hand side first keeps the constant's sign straight.
Move the constant
$$0.25x_2=0.5+1-0.5x_1=1.5-0.5x_1$$
Add 1 to both sides and subtract the $x_1$ term.
Divide
$$x_2=6-2x_1$$
Divide every term by 0.25.
Answer $$\boxed{x_2=6-2x_1}$$
Check
The point $(2,2)$ lies on it and gives $-1+1+0.5=0.5$, exactly the cut.
Write the boundary equation with the cut on the right-hand side before moving anything.
4§03.5 — a two-group slope against least squares
With six equally spaced inputs $x=1,\dots,6$, a quick estimator splits the data into a low half ($x=1,2,3$) and a high half ($x=4,5,6$) and uses $\tilde\beta_1=(\bar y_{\mathrm{high}}-\bar y_{\mathrm{low}})/(5-2)$.
Find
(a) Show that $\tilde\beta_1$ is linear and unbiased.
(b) Compute $\operatorname{Var}(\tilde\beta_1)$ and $\operatorname{Var}(\hat\beta_1)$.
(c) Which is smaller, and which theorem predicts it?
Write $\tilde\beta_1$ as a weighted sum of the six $y_i$; linearity and unbiasedness follow from the weights.
Hint 2/4
$E[\sum c_iy_i]=\sum c_i(\beta_0+\beta_1x_i)$ and $\operatorname{Var}(\sum c_iy_i)=\sigma^2\sum c_i^2$.
Hint 3/4
The weights are $-\tfrac19$ on $x=1,2,3$ and $+\tfrac19$ on $x=4,5,6$; for least squares $S_{xx}=17.5$.
Hint 4/4
$\operatorname{Var}(\tilde\beta_1)=\tfrac{2}{27}\sigma^2\approx 0.074\sigma^2$, larger than $\sigma^2/17.5\approx 0.057\sigma^2$, as Gauss-Markov predicts.
Show solution
Both estimators are weighted sums of the responses, so the same two rules settle everything.
Centered route: $\bar x=2.5$, $\bar y=4.5$, $S_{xx}=5$, $S_{xy}=4$, so $\hat\beta_1=0.8$ and $\hat\beta_0=4.5-2=2.5$.
Keep the order of $\hat\beta$ tied to the order of the columns of $X$.
D · interleaved 4 questions
1§03.2 — five weeks, one interval
From five weeks of data a café computes the 95% interval $[0.71,\ 2.09]$ for $\beta_1^{\mathrm{true}}$, in hundreds of cups per post, using $\hat\beta_1\pm 2\,\mathrm{SE}$.
Find(a) Which sentence reads the interval correctly?
Given
interval $[0.71,\ 2.09]$ from $\hat\beta_1=1.4$ and $\mathrm{SE}\approx 0.35$
$\beta_1^{\mathrm{true}}$ is a fixed, unknown number
Hint 1/4
Ask what is random in this recipe: the interval, or the true slope?
Hint 2/4
A 95% confidence interval is a recipe that, over repeated samples, covers the fixed true value about 95 times in 100.
Hint 3/4
Here the recipe is $\hat\beta_1\pm 2\,\mathrm{SE}$ with $\hat\beta_1=1.4$ and $\mathrm{SE}\approx 0.35$, and $\beta_1^{\mathrm{true}}$ does not move.
Hint 4/4
About 95 in 100 intervals built this way catch the true slope; that is the correct reading.
Show solution
Pinning down which quantity is random settles which sentence can carry the 95%.
What is random
$$\hat\beta_1,\ \mathrm{SE}\ \text{vary with the sample};\quad \beta_1^{\mathrm{true}}\ \text{does not}$$
The interval moves from study to study; the target stays put.
The probability is over repeated samples, as in the previous section.
Answer $$\boxed{\text{about 95 in 100 such intervals contain }\beta_1^{\mathrm{true}}}$$
Check
A probability statement about $\beta_1^{\mathrm{true}}$ given this one data set would need a posterior over it, which this recipe never builds.
In regression, as with rates, the 95% belongs to the procedure, not to one interval.
2§03.2 — why the fitted lines pivot
Fitted lines from repeated samples tilt around a pivot: when the slope comes out too steep, the intercept tends to come out too low. You want the number behind that picture.
Find
(a) Show that $\operatorname{Cov}(\bar y,\hat\beta_1)=0$.
(b) Show that $\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-\bar x\,\sigma^2/S_{xx}$.
(c) Evaluate the covariance and the correlation for the café design.
Given
$\hat\beta_1=\sum_ik_iy_i$ with $k_i=(x_i-\bar x)/S_{xx}$
$\hat\beta_0=\bar y-\hat\beta_1\bar x$
$y_i$ uncorrelated with variance $\sigma^2$
café design $x=1,2,3,4,5$
Hint 1/4
Everything is a linear combination of the $y_i$, so the covariance rule for sums does all the work.
Hint 2/4
For uncorrelated $y_i$ with variance $\sigma^2$: $\operatorname{Cov}(\sum a_iy_i,\sum b_iy_i)=\sigma^2\sum a_ib_i$.
Hint 3/4
$\bar y$ has weights $\tfrac1n$ and $\hat\beta_1$ has weights $k_i$ with $\sum k_i=0$; for the café, $\bar x=3$, $S_{xx}=10$, $n=5$.
Hint 4/4
$\operatorname{Cov}(\hat\beta_0,\hat\beta_1)=-0.3\,\sigma^2$ and the correlation is $-3/\sqrt{11}\approx -0.90$.
Show solution
Writing both estimators as weighted sums turns every covariance into one sum of weight products.
Sign check: with $\bar x>0$ every fitted line passes near $(\bar x,\bar y)$, so a steeper line must start lower at $x=0$, which is a negative covariance.
The pivot of the sampling picture is the point of means; centring the input ($\bar x=0$) removes the covariance altogether.
3§03.5 — the noise level's own estimate
Under i.i.d. Gaussian noise the log-likelihood also depends on $\sigma$. Maximizing it over $\sigma$ as well as $\beta$ gives a second estimate of $\sigma^2$, to compare with the RSE.
Find
(a) With $\beta=\hat\beta_{\mathrm{RSS}}$ fixed, find the $\sigma^2$ that maximizes $l$.
(b) Evaluate it for the café fit and compare with $\mathrm{RSE}^2=\mathrm{RSS}/(n-2)$.
Second derivative at that point: $\frac{n}{\sigma^2}-\frac{3\,\mathrm{RSS}}{\sigma^4}=\frac{n}{\sigma^2}-\frac{3n}{\sigma^2}<0$, so it is a maximum.
The MLE of the noise level runs small because residuals are smaller than errors; the course's intervals use the RSE.
4§03.6 — a filter trained on counts
A spam filter's first training set has $20$ e-mails, $7$ of them spam (coded $1$) and $13$ not (coded $0$). With no inputs yet, it fits $y=\beta_0+\varepsilon$ by least squares and flags an e-mail when $\hat y>0.5$.
Find(a) What is $\hat\beta_0$, and what does the filter do with every new e-mail?
Given
$20$ labels: $7$ ones and $13$ zeros
model $y=\beta_0+\varepsilon$
flag if $\hat y>0.5$
Hint 1/4
With only an intercept, least squares looks for the single number closest to all twenty labels.
Hint 2/4
Minimizing $\sum_i(y_i-\beta_0)^2$ gives $\hat\beta_0=\bar y$; for 0/1 labels that is the fraction of ones, the rate MLE $N_1/N$.
Hint 3/4
$7$ ones among $20$ labels: $\bar y=7/20$.
Hint 4/4
$\hat\beta_0=0.35\le 0.5$, so no e-mail is ever flagged.
Show solution
The one-parameter least squares problem is the warm-up minimization of this section, so its answer is the mean.
centres $\mu_j$ and scale $s$ are chosen before fitting
Check yourself
Close the page and write from memory: the RSS; the formulas for $\hat\beta_1$ and $\hat\beta_0$; the two variances and the 95% interval; the normal equations with their rank condition; what $X^Te=0$ says; the two reasons for squares; the $0.5$ decision rule; and when a polynomial fit is unique. Then compare with the formula card.
Fit a line from a table of pairs and check it with $\sum e_i=0$ and $\sum x_ie_i=0$?
c-simple-ls
Compute the RSE and a 95% interval for the slope and the intercept, and say which one is wider and why?
c-accuracy
Derive $X^TX\beta=X^Ty$ with the course's derivative rules and solve a $2\times 2$ or diagonal case?
c-matrix-ls
Explain why $e$ is perpendicular to every column of $X$, and use that to test a proposed fit?
c-projection
State what Gauss-Markov assumes, and show that Gaussian noise turns the MLE into least squares?
c-why-squares
Find a decision boundary from a fit on 0/1 labels and name its two weaknesses?
c-classification
Build a design matrix with polynomial, interaction or Gaussian columns and decide whether the fit is unique?
c-basis
Glossary (29 terms)
linear regressiondoğrusal regresyon
A model that predicts the response by a linear combination of the coefficients, $\hat y=\hat\beta_0+\sum_j\hat\beta_jx_j$.
simple linear regressionbasit doğrusal regresyon
Linear regression with one input: $Y\approx\beta_0+\beta_1X$.
multiple linear regressionçoklu doğrusal regresyon
Linear regression with several inputs, written $y\approx X\beta$.
bağımlı değişken
The quantity being predicted, Y.
predictoraçıklayıcı değişken
An input variable used to predict the response.
interceptsabit terim
The coefficient $\beta_0$, the prediction when every input is zero; the lecture also calls it the bias.
slopeeğim
The change in the prediction per unit change of an input, $\beta_1$ in simple regression.
residualartık
The observed response minus the fitted one, $e_i=y_i-\hat y_i$.
residual sum of squaresartık kareler toplamı
$\mathrm{RSS}=\sum_ie_i^2$, the score that least squares minimizes.
least squaresen küçük kareler
The method that chooses the coefficients minimizing the residual sum of squares.
sıradan en küçük kareler
Least squares with every observation weighted equally, giving $\hat\beta=(X^TX)^{-1}X^Ty$.
fitted value
The model's prediction at an observed input, $\hat y_i=\hat\beta^Tx_i$.
merkezleme
Subtracting the mean from a variable so that its values sum to zero.
residual standard error
$\mathrm{RSE}=\sqrt{\mathrm{RSS}/(n-2)}$, the estimate of the noise standard deviation in simple regression.
unbiased estimatoryansız kestirici
An estimator whose mean over repeated samples equals the true value, $E[\hat\beta]=\beta^{\mathrm{true}}$.
design matrixtasarım matrisi
The matrix $X$ with one row per observation and one column per coefficient.
full column rank
No column of $X$ is a linear combination of the others; it makes $X^TX$ invertible.
normal equationsnormal denklemler
The system $X^TX\beta=X^Ty$ that the least squares coefficients solve.
Jacobi matrisi
The matrix of partial derivatives $\partial h_i/\partial g_j$ of a vector function $h$ of a vector $g$.
hat matrixşapka matrisi
$H=X(X^TX)^{-1}X^T$, which maps $y$ to $\hat y$; the lecture calls it the projection matrix.
orthogonal projectiondik izdüşüm
The closest point of a subspace to a given vector, reached along a direction perpendicular to the subspace.
column spacesütun uzayı
All vectors of the form $X\beta$: the combinations of the columns of $X$.
Gauss-Markov theoremGauss Markov teoremi
Under zero-mean, equal-variance, uncorrelated noise, least squares has the smallest variance among linear unbiased estimators.
BLUE
Best linear unbiased estimator; under the Gauss-Markov assumptions it is the least squares estimator.
decision boundarykarar sınırı
The set of inputs where a classifier switches class; for a regression fit on 0/1 labels, $\hat\beta^Tx=0.5$.
etkileşim terimi
A product column such as $x_1x_2$ that lets the effect of one input depend on another.
polynomial regressionpolinom regresyonu
Linear regression on the columns $1,x,x^2,\dots,x^p$.
Vandermonde matrixVandermonde matrisi
The design matrix of polynomial regression, with rows $[1, \allowbreak x_i, \allowbreak x_i^2, \allowbreak \dots, \allowbreak x_i^p]$.
basis functiontaban fonksiyonu
A fixed function $\phi_j(x)$ of the inputs used as a column of the design matrix.
What comes next
§04 · Measuring performance: bias-variance and cross-validation
Here every fit was judged on the same points it was fitted to, and the RSS only went down as columns were added. Next comes the question that number cannot answer: how well does a fitted model do on data it has never seen?
Sources
textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer, 2003 The textbook named in the syllabus. This week's syllabus line gives no section numbers, so none are cited here.
course materialEEE 485/585 chapter 3 lecture slides and lecture notes, Fall 2026 Topic order and notation follow them: beta true, RSS, beta hat RSS, the row convention for matrix derivatives, the projection matrix H and the 0.5 decision rule. All wording, data, examples and exercises here are original.
course materialEEE 485 syllabus page on STARS, printed 21 September 2026 Source of the weekly line, the assessment weights quoted in the card and the conditions for sitting the final exam.
textbookOther books the syllabus recommends: G. James et al., An Introduction to Statistical Learning (2013); K. P. Murphy, Machine Learning: A Probabilistic Perspective (2012); C. M. Bishop, Pattern Recognition and Machine Learning (2011) Recommended, not required. The lecture slides borrow several of their regression figures from James et al.