7 concepts24 worked examples30 exercises4 exam-level7 figures
What are you here for?
09 PCA, ICA and blind source separation: directions of largest variance, unmixing and the ICA likelihood
Start with this
One question before you read anything. Getting it wrong is the point: it shows you what this section is for.
§09.1 — two least squares lines through the same points
Five students' centered quiz scores $(x_1,x_2)$ are $(-3,-4)$, $(-1,-3)$, $(-1,2)$, $(1,3)$ and $(4,2)$. Regressing quiz 2 on quiz 1 through the origin gives the line $x_2=0.857\,x_1$.
Find(a) Now regress quiz 1 on quiz 2 and draw that line in the same picture, with quiz 2 on the vertical axis. What is its slope in that picture?
Solve for $x_2$ to put quiz 2 on the vertical axis.
Answer $$\boxed{\text{slope }1.75\ne0.857}$$
Check
If the two fits gave one line, the product of the two regression slopes, $\frac{24}{28}\cdot\frac{24}{42}=0.490$, would equal $1$; it equals the squared correlation instead.
Least squares is not symmetric in its two variables; the summary line this section builds will be.
Five students took two quizzes, and a committee wants one number per student that spreads them out as much as possible. With weights of length 1, quiz 2 alone keeps $8.4$ of the $14$ units of variance and equal weights keep $11.8$. The weights $0.6$ and $0.8$ keep $12$, and no other pair of length 1 keeps more.
By the end you can find that weighting for any number of features, say what share of the variance $k$ components keep, and write and compare the likelihood that pulls mixed signals apart.
In 60 seconds
PCA describes centered data by the eigenvectors of $\Sigma=\frac1nX^TX$ with the largest eigenvalues, each keeping variance $\lambda_m$; ICA models recordings as $x=As$ with independent, non-Gaussian sources and estimates $W=A^{-1}$ by maximum likelihood, up to order, scale and sign.
comparing candidate unmixing matrices, or deriving the estimate $\hat W$
Three most common mistakes
Running PCA on uncentered data, or on raw features in arbitrary units: the components then follow the means or the units, not the structure.
Taking the eigenvector of the wrong eigenvalue, or leaving it unnormalized: $u_1$ belongs to $\lambda_1$ and has length 1.
Dropping $\log\lvert\det W\rvert$ from the ICA log-likelihood, or adding it once instead of $T$ times.
Two course documents give different weights:
Chapter 1 slides, undergraduate line: midterm 25, final 25, four quizzes 20, two-phase project 30.
STARS syllabus page printed on 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. It is the later document; confirm which split applies.
The STARS weekly list splits this chapter: blind signal separation is its seventh item, and with feature selection its ninth; the lecture slides make PCA, ICA and blind source separation chapter 9.
How much time do you have?
10 minutes
The eigenvector recipe with a worked 2 by 2 case, the algorithm with scores and , and the mixing model with its three ambiguities.
The 60-second card · The best direction is the top eigenvector of Σ · The PCA algorithm · Blind source separation · Formula card
45 minutes
Every result once with a worked example, then PCA by hand from a full solution down to a bare problem.
The 60-second card · Projecting onto a direction · The best direction is the top eigenvector of Σ · The PCA algorithm · How many components · Blind source separation · Why the sources must not be Gaussian · ICA by maximum likelihood · Scaffolding comes off · Formula card
full read
Adds the derivations, the look-alike pairs, a full exam-style question and mixed practice where you pick the method yourself.
The opening pages · Recall first · Projecting onto a direction · The best direction is the top eigenvector of Σ · The PCA algorithm · How many components · Blind source separation · Why the sources must not be Gaussian · ICA by maximum likelihood · Method boxes · Look-alike pairs · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
Compute the variance of the scores along a unit vector as $u^T\Sigma u$, and find the best direction for two features by turning an angle.
Derive with a that principal components are eigenvectors of $\Sigma$, compute them for 2 by 2 and block 3 by 3 matrices, and obtain the next component by .
Apply the PCA algorithm: center, standardize when units differ, compute scores and reconstructions, and read the loadings.
Choose the number of components from the proportion of variance explained and the scree plot, and relate the discarded eigenvalues to the reconstruction error.
Model blind source separation as $x=As$, unmix with $W=A^{-1}$, and tell an acceptable output (order, scale, sign) from a failed one.
Explain why Gaussian sources cannot be unmixed, by showing that $A$ and $AQ$ give the same distribution of $x$.
Evaluate the density $p_X(x)=\prod_jp_{S_j}(w_j^Tx)\lvert\det W\rvert$ and the ICA log-likelihood $l(W)$, and compare candidate unmixing matrices.
Syllabus coverage
covered
PCA — covered
centering and scaling the columns
projections and the variance of the scores
the first component by a Lagrange multiplier
the top eigenvectors of the sample covariance matrix
the algorithm, scores, loadings and reconstruction
standardization
the proportion of variance explained and the choice of k
Spread over four blocks, in the lecture's order. The deflation argument for the second component, set as an exercise in the lecture notes, is worked out in the eigenvector block.
covered
ICA — covered
the
the order, scale and sign ambiguities
the need for non-Gaussian sources
the density of a linear transformation
the likelihood and log-likelihood of the unmixing matrix
its maximum likelihood estimate
The ambiguities sit in the two blocks before the likelihood.
covered
blind source separation — covered
The : T observations of p sensors, each a weighted sum of p unobserved sources with unknown coefficients, and the goal of estimating the source signals. The lecture lists image denoising, medical signal processing, brain-computer interfaces and time series analysis as uses.
The applications are named, not developed.
covered
Blind signal separation — covered
The official weekly line's name for the same problem.
It is the seventh item of the STARS weekly list; the lecture teaches it inside chapter 9.
covered
Feature extraction — covered
as new features: k numbers per point in place of p, used on their own or as a pre-processing step before a supervised method.
The ninth item of the STARS weekly list pairs it with feature selection; in this chapter the lecture's feature extraction method is PCA.
deferred
feature selection — deferred
Choosing a subset of the original features.
The lecturer teaches it as a chapter of its own, 'Feature selection', and a later section of these notes covers it. PCA builds new features out of all the old ones; it selects none of them.
covered
Unsupervised learning — covered
Training data without labels, used on their own or as a pre-processing step before a supervised method.
The chapter's opening slide; the first block starts from it.
off syllabus
Reconstruction error and the discarded eigenvalues — off syllabus
$\frac1n\sum_i\lVert x_i-\hat x_i\rVert^2=\lambda_{k+1}+\dots+\lambda_p$, so keeping the most variance and making the smallest squared reconstruction error pick the same components.
Further reading: the lecture asks whether a point can be recovered exactly from its scores; this identity says how much is lost when it cannot. Used in one worked example, one practice question and the check of the exam example.
off syllabus
Gradient of the ICA log-likelihood — off syllabus
$\nabla_Wl=\sum_tg(Wx(t))\,x(t)^T+T(W^{-1})^T$, the direction in which an implementation climbs $l(W)$.
Further reading: the lecture stops at $\hat W=\arg\max l(W)$; the gradient is what an implementation climbs.
off syllabus
At most one Gaussian source — off syllabus
The rotation argument needs every source Gaussian; separation stays possible when a single source is Gaussian.
Further reading, stated without proof; no exercise depends on it.
Recall first
Covariance matrix and weighted sums
For a random vector $X$ with mean $\mu$, $\Sigma=E[(X-\mu)(X-\mu)^T]$ is symmetric and positive semidefinite, and $\operatorname{Var}(a^TX)=a^T\Sigma a$. For $n$ centered data points the sample version is $\frac1n\sum_ix_ix_i^T$.
The variance of every projection in this section is this quadratic form.
Eigenvalues and eigenvectors of a symmetric matrix
$\Sigma u=\lambda u$ with $u\ne0$. A real symmetric $p\times p$ matrix has $p$ real eigenvalues and eigenvectors $u_1,\dots,u_p$, and $\Sigma=\sum_j\lambda_ju_ju_j^T$. For $2\times2$, $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$; in general the is the sum of the eigenvalues.
Principal components are these eigenvectors, and the eigenvalues are the variances they keep.
Lagrange multipliers
To optimize $f(u)$ subject to $g(u)=0$, look for stationary points of $L(u,\lambda)=f(u)-\lambda g(u)$. For symmetric $\Sigma$, $\nabla_u(u^T\Sigma u)=2\Sigma u$ and $\nabla_u(u^Tu)=2u$.
The lecture derives the first component this way.
Least squares through the origin
For centered data, the line $y=bx$ with $b=\frac{\sum_ix_iy_i}{\sum_ix_i^2}$ minimizes the vertical misses $\sum_i(y_i-bx_i)^2$; with several predictors, $Z^TZ\beta=Z^Ty$.
The first failure of the section compares it with the first principal component, and one practice question regresses on scores.
Gaussian vectors and independence
A Gaussian vector is fixed by its mean and covariance, and if $s$ is Gaussian with mean $\mu$ and covariance $C$, then $As$ is Gaussian with mean $A\mu$ and covariance $ACA^T$. Independent variables have a joint density equal to the product of their densities; jointly Gaussian variables that are uncorrelated are independent.
The argument against Gaussian sources and the ICA likelihood both rest on these facts.
and orthogonal matrices
$\det(AB)=\det A\det B$ and $\det(A^{-1})=1/\det A$; $\lvert\det A\rvert$ is the factor by which $A$ scales areas and volumes. An orthogonal $Q$ has $Q^TQ=QQ^T=I$ and $\lvert\det Q\rvert=1$. $\begin{bmatrix}a&b\\c&d\end{bmatrix}^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix}$.
Unmixing, the density of a mixture and the rotation argument all use them.
Maximum likelihood
$\hat\theta=\arg\max_\theta\log p(D\mid\theta)$; for independent observations the log-likelihood is a sum over them, and taking the log does not move the maximizer.
ICA estimates the unmixing matrix this way.
Try it yourself first (2 questions)
1§09.1 — the variance of a rescaled sum
Five students' two quiz deviations $x_1$ and $x_2$ add up to a sum $x_1+x_2$ whose variance is $23.6$.
Find(a) What is the variance of the average $\frac12(x_1+x_2)$?
Given
$\operatorname{Var}(x_1+x_2)=23.6$
$\operatorname{Var}(cY)=c^2\operatorname{Var}(Y)$
Hint 1/4
The average is the sum multiplied by a constant; ask how a constant factor acts on a variance.
The other eigenvalue must be $\operatorname{tr}\Sigma-12=2$; indeed $\Sigma(4,-3)^T=(8,-6)=2\,(4,-3)^T$.
This section will show that the eigenvector with the largest eigenvalue is the most informative direction in the data.
Notation
symbol
reads as
means
watch out
$x_i=[x_{i1},\dots,x_{ip}]^T,\ \ X$
point i, the data matrix
$X$ is $n\times p$ with rows $x_i^T$, centered unless stated otherwise
In the ICA blocks $X$ and $S$ are random vectors instead, as in the lecture.
$\bar x_j,\ \ \sigma_j$
x bar j, sigma j
the mean of column $j$; its standard deviation, with $1/n$
Divide by $\sigma_j$ only after centering.
$u,\ \ \lVert u\rVert=1$
a unit vector
a direction in $\mathbb R^p$
$u$ and $-u$ give the same line and the same variance.
$\Sigma=\tfrac1nX^TX$
Sigma
the sample covariance matrix of the centered data
$1/n$ as in the lecture; $1/(n-1)$ scales every eigenvalue and leaves the eigenvectors unchanged.
$\lambda_m,\ \ u_m$
lambda m, u m
the $m$-th largest eigenvalue and its unit eigenvector, the $m$-th principal component
Here $\lambda$ is an eigenvalue and the Lagrange multiplier, not a penalty weight.
$z_{im}=x_i^Tu_m,\ \ z_i$
the score of point i on component m
$z_i=[z_{i1},\dots,z_{ik}]^T$ holds point $i$'s $k$ scores
A score is a number; the projection $(x_i^Tu_m)\,u_m$ is a vector.
$\hat x_i$
x hat i
the reconstruction $z_{i1}u_1+\dots+z_{ik}u_k$
Add the means back to return to raw units.
$\mathrm{PVE}(m)$
proportion of variance explained
$\lambda_m/\sum_j\lambda_j$
A number between 0 and 1.
$Y=X-Xu_1u_1^T$
the deflated data
what is left after the first component is removed
$\frac1nY^TY=\Sigma-\lambda_1u_1u_1^T$.
$s(t),\ \ x(t),\ \ t=1,\dots,T$
sources and recordings at time t
the hidden and the observed vectors
Independent over t in the ICA model.
$A=[a_{ij}]$
the mixing matrix
$p\times p$, unknown and invertible; $a_{ij}$ is the weight of source $j$ in sensor $i$
Square: as many sensors as sources.
$W=A^{-1},\ \ w_j^T$
the unmixing matrix, its j-th row
$\hat s_j=w_j^Tx$
Recovered only up to order, scale and sign.
$P,\ \ D$
a , a diagonal matrix
one 1 in each row and column; nonzero diagonal entries
$D$ with entries $\pm1$ gives the sign flips.
$Q$
an
$QQ^T=Q^TQ=I$; in 2D a rotation
$AQ$ mixes Gaussian sources exactly like $A$.
$p_S,\ \ p_{S_j},\ \ p_X$
densities
of the source vector, of source j, of the recordings
The source densities are assumed known.
$L(W),\ \ l(W),\ \ \hat W$
likelihood, log-likelihood, estimate
$\prod_tp_X(x(t))$, its natural log, and its maximizer
Maximized, not minimized.
Conventions used here
Centered data.
$X$ is centered before any PCA formula is applied, and $\Sigma=\frac1nX^TX$ uses $1/n$, as in the lecture.
With 1/(n − 1) every eigenvalue scales by the same factor; directions and PVE do not change.
Unit length and sign.
Principal components are unit vectors. Since $u$ and $-u$ are equally valid, this page takes the first nonzero entry positive.
Two correct answers can differ by a sign; the convention makes answers comparable.
Order of the components.
Eigenvalues are sorted, $\lambda_1\ge\lambda_2\ge\dots\ge\lambda_p$, and component $m$ belongs to $\lambda_m$.
First means largest variance, not first feature.
What λ means here.
$\lambda$ is the Lagrange multiplier and, at the solution, the eigenvalue; the two coincide.
In earlier sections the same letter was a penalty weight.
Standardize when units differ.
If the features are measured in different or arbitrary units, divide each centered column by its $\sigma_j$ first; if they share a unit, center only.
This is the lecture's rule, and the answer otherwise depends on the units.
PVE as a proportion.
$\mathrm{PVE}$ is reported as a number between 0 and 1, and targets such as 'at least 0.85' use $\ge$.
Meeting a target exactly counts.
Square mixing.
ICA here has as many sensors as sources, $A$ is $p\times p$ and invertible, and $W=A^{-1}$.
This is the lecture's model.
An acceptable ICA output.
An output counts as correct when it equals $DPs$: each output one source, in any order, times any nonzero constant.
Order, scale and sign cannot be recovered from the recordings.
Logarithms.
$\log$ is the natural logarithm, in every likelihood on this page.
Another base would only rescale the log-likelihood.
Rounding.
Intermediate values keep at least four significant figures; answers are given to three decimals and angles to one decimal of a degree.
Eigenvalues and angles computed from rounded inputs move in the last digit.
9.1Projecting onto a direction: the spread one number keeps
Turns each point into one number, its score $x^Tu$ along a unit vector $u$, and measures the spread those scores keep: $u^T\Sigma u$.
Every model so far learned from labelled pairs $(x_i,y_i)$; now the training data are the $x_i$ alone, and the first job is to describe them with fewer numbers.
Solvable with what we have
Compute each quiz's variance and the covariance: $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$, total variance $5.6+8.4=14$.
Compute the variance of any fixed weighting, for example the average: $\tfrac14(5.6+2\cdot4.8+8.4)=5.9$.
Fit a least squares line that predicts quiz 2 from quiz 1: slope $24/28=0.857$.
Not solvable yet
Choose the weights that spread the five students out the most when neither quiz is a response.
Say how much of the total 14 a single number keeps, and how much it loses.
Do the same for four features, or four hundred.
Use least squares anyway. Quiz 2 on quiz 1 gives slope $\textcolor{#8250df}{0.857}$; quiz 1 on quiz 2, drawn in the same picture, gives slope $\textcolor{#8250df}{1.75}$. One data set, two answers. The direction built in this block has slope $\textcolor{#1f6feb}{4/3}$, and its scores keep 12 of the 14 units of variance.
Why it fails
Least squares counts misses along the response axis only, so its answer depends on which quiz we call the response. Here there is no response: the spread has to be measured in a way that treats both quizzes alike.
DefinitionDefinition 9.1: Score, projection and variance along a direction
Conditions
the $n$ points $x_1,\dots,x_n\in\mathbb R^p$ are centered: each feature has mean $0$
$u\in\mathbb R^p$ is a unit vector, $\lVert u\rVert=1$, that is $u^Tu=1$
$\Sigma=\frac1n\sum_{i=1}^nx_ix_i^T=\frac1nX^TX$ is the sample covariance matrix, with $1/n$ as in the lecture
Drop each point perpendicularly onto the line through the origin in direction $u$; the signed distance from the origin to the foot is the score. Centered data give scores with mean zero, so their mean square is their variance, and it equals $u^T\Sigma u$. The first principal component is the unit direction that makes this number largest.
Proof
The foot is at $x^Tu$. The closest point $cu$ on the line minimizes $\lVert x-cu\rVert^2=\lVert x\rVert^2-2c\,x^Tu+c^2$; the derivative $-2x^Tu+2c$ vanishes at $c=x^Tu$.
Scores have mean zero.$\frac1n\sum_ix_i^Tu=\big(\frac1n\sum_ix_i\big)^Tu=0^Tu=0$.
Their mean square is a quadratic form.$(x_i^Tu)^2=u^Tx_i\,x_i^Tu$, so $\frac1n\sum_i(x_i^Tu)^2=u^T\big(\frac1n\sum_ix_ix_i^T\big)u=u^T\Sigma u$.
Why the unit length. For $a=cu$ the same formula gives $c^2\,u^T\Sigma u$: without $\lVert u\rVert=1$ the variance grows without bound and has no maximum.
The five students after centering, as open circles. The blue line is the direction $\textcolor{#1f6feb}{u=(0.6,0.8)}$ found in the first worked example. Each dashed orange piece drops a point onto the line; the blue feet sit at scores $-5,-3,1,3,4$, whose mean square is $12$. The orange pieces have lengths $0,1,2,1,2$ and mean square $\textcolor{#d1690a}{2}$, and $12+2=14$ is the total variance.
Looks like this, but is not
The plain sum $x_1+x_2$, weights $a=(1,1)$, spreads the students with variance $a^T\Sigma a=23.6$, more than the total $14$ of the two quizzes.
$a$ is not a unit vector: $\lVert a\rVert=\sqrt2$, and doubling any direction multiplies $a^T\Sigma a$ by four. As a unit vector, $(1,1)/\sqrt2$ keeps $11.8$, below the best $12$.
angle θ
unit vector u
variance of the scores
$0^\circ$
$(1,\ 0)$
$5.6$
$30^\circ$
$(0.866,\ 0.5)$
$10.457$
$53.1^\circ$
$(0.6,\ 0.8)$
$12$
$90^\circ$
$(0,\ 1)$
$8.4$
$143.1^\circ$
$(-0.8,\ 0.6)$
$2$
The spread rises to $12$ at $53.1^\circ$ and falls to $2$ exactly $90^\circ$ later; the two axes keep only the single-quiz variances $5.6$ and $8.4$.
The best single weighting for the five students, found by turning a unit vector
Five students' quiz scores, out of 20, are $(9,10)$, $(11,11)$, $(11,16)$, $(13,17)$, $(16,16)$. Find the unit vector $u$ whose scores $x_i^Tu$ have the largest variance, and that variance.
FindThe best $u=(\cos\theta,\sin\theta)$ and the variance $u^T\Sigma u$ it keeps.
Given
raw scores $(9,10),\ \allowbreak (11,11),\ \allowbreak (11,16),\ \allowbreak (13,17),\ \allowbreak (16,16)$ with means $(12,14)$
With two features every unit vector is $(\cos\theta,\sin\theta)$, so the problem becomes maximizing a function of one angle; no new theory is needed yet.
$\tan\theta=4/3$ is the 3-4-5 triangle, so the cosine is $3/5$ and the sine $4/5$.
Answer $$\boxed{u=\textcolor{#1f6feb}{(0.6,\ 0.8)},\qquad u^T\Sigma u=12\ \text{ of the total }14}$$
Check
Compute the scores directly: $0.6x_1+0.8x_2$ gives $-5,-3,1,3,4$, with mean $0$ and mean square $\frac{25+9+1+9+16}5=12$. The smallest value, $7-5=2$, sits at $\theta=143.13^\circ$, perpendicular to $u$.
That is the opening's weighting, $0.6\times$ quiz 1 $+\ 0.8\times$ quiz 2, keeping 12 of the 14. With more features the unit vectors form a sphere, and the next block replaces the angle by an eigenvector equation.
How much spread the average, quiz 2 alone and the two least squares lines keep
For the same five students compare four candidate directions: the average $(1,1)/\sqrt2$, quiz 2 alone $(0,1)$, and the two least squares lines, of slopes $6/7$ and $7/4$.
Find$u^T\Sigma u$ for each candidate, against the best value $12$.
Every candidate stays below $12$ and above $2$, the smallest possible value found in the first example; the two least squares lines straddle the best angle, $40.6^\circ<53.1^\circ<60.3^\circ$.
Directions near the best one keep almost as much. What sets the best one apart is that the data alone define it, with no response chosen.
Checkpoint
§09.1 — the spread along a swapped direction
A classmate says that $u=(0.8,0.6)$ must keep the same variance as $u=(0.6,0.8)$, since it uses the same two numbers.
Find(a) Compute $u^T\Sigma u$ for $u=(0.8,0.6)$.
Given
$\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$ for the five students
$u=(0.8,0.6)$, a unit vector
Hint 1/4
The variance along a unit vector is a quadratic form in its entries; write out all four terms.
Here $u=(0.8,0.6)$, $\Sigma_{11}=5.6$, $\Sigma_{12}=4.8$, $\Sigma_{22}=8.4$: $5.6\cdot0.64+2\cdot4.8\cdot0.48+8.4\cdot0.36$.
Hint 4/4
The three terms add to $11.216$, below the $12$ of $(0.6,0.8)$.
Show solution
Term by term is the fastest route for a 2 by 2 matrix.
Three terms
$$5.6\,(0.8)^2+2\,(4.8)(0.8)(0.6)+8.4\,(0.6)^2$$
Squares on the diagonal terms, the cross term counted twice.
$$=3.584+4.608+3.024=11.216$$
These three terms are the whole quadratic form, since the cross term was already doubled.
Answer $$\boxed{u^T\Sigma u=11.216}$$
Check
The angle formula agrees: $u=(0.8,0.6)$ is $\theta=36.87^\circ$, with $\cos2\theta=0.28$ and $\sin2\theta=0.96$, so $7-1.4\cdot0.28+4.8\cdot0.96=11.216$.
Which feature gets the larger weight matters as much as the weights themselves.
⚠ Using weights that are not a unit vector
The plain sum, or an eigenvector left unscaled, looks like a direction.
wrong$$\operatorname{Var}(x_1+x_2)=23.6\ \text{ as the spread along the diagonal}$$
Step the direction round the five students. The blue feet are the scores and the dashed orange pieces are what one number throws away; the two mean squares always add to $14$, and the kept part peaks at $\textcolor{#1f6feb}{12}$ at $53.1^\circ$.
At the edges
θ = 53.1° 12
The largest spread, along $u=(0.6,0.8)$.
θ = 143.1° 2
The smallest spread, perpendicular to the best direction; this is what one number loses.
9.2The best direction is the top eigenvector of Σ
Solves $\max u^T\Sigma u$ subject to $u^Tu=1$: the answer is the eigenvector of $\Sigma$ with the largest eigenvalue, and the maximum is that eigenvalue.
Turning an angle worked because $p=2$; with four features the unit vectors form a sphere, and we need an equation that singles out the best one.
TheoremTheorem 9.2: Principal components are eigenvectors of Σ
Conditions
$\Sigma$ is the $p\times p$ sample covariance matrix of centered data: symmetric and positive semidefinite
its eigenvalues are sorted, $\lambda_1\ge\lambda_2\ge\dots\ge\lambda_p\ge0$, with orthonormal eigenvectors $u_1,\dots,u_p$
if $\lambda_1=\lambda_2$ the first direction is not unique: every unit vector in the span of $u_1$ and $u_2$ is optimal
$$\boxed{\begin{aligned}&\max_{u^Tu=1}u^T\Sigma u:\qquad L(u,\lambda)=u^T\Sigma u-\lambda\,(u^Tu-1)\\ &\nabla_uL=0\ \Rightarrow\ \Sigma u=\lambda u,\qquad u^T\Sigma u=\lambda\\ &\text{maximum }\textcolor{#1f6feb}{\lambda_1}\text{ at }\textcolor{#1f6feb}{u_1};\qquad\text{the first }k\text{ components are }u_1,\dots,u_k\end{aligned}}$$
Setting the gradient of the Lagrangian to zero says the best $u$ must be an eigenvector of $\Sigma$, and at an eigenvector the variance equals its eigenvalue. So the winner is the eigenvector with the largest eigenvalue. The next component, the best direction perpendicular to it, is the eigenvector of the second largest eigenvalue, and so on.
Proof
The constraint.$\partial L/\partial\lambda=-(u^Tu-1)=0$ gives back $u^Tu=1$.
The stationary points. For symmetric $\Sigma$, $\nabla_u(u^T\Sigma u)=2\Sigma u$ and $\nabla_u(u^Tu)=2u$, so $\nabla_uL=2\Sigma u-2\lambda u=0$: $\Sigma u=\lambda u$.
Their values. At a unit eigenvector $u^T\Sigma u=u^T(\lambda u)=\lambda\,u^Tu=\lambda$, so the largest eigenvalue gives the largest candidate.
A global maximum, not just a stationary point. Expand any unit $u=\sum_jc_ju_j$ with $\sum_jc_j^2=1$. Then $u^T\Sigma u=\sum_j\lambda_jc_j^2\le\lambda_1\sum_jc_j^2=\lambda_1$, with equality at $u=u_1$.
The later components. Over unit vectors perpendicular to $u_1,\dots,u_{m-1}$ the first $m-1$ coefficients are $0$, and the same bound gives $\lambda_m$, reached at $u_m$.
The variance $u^T\Sigma u$ of the five students' scores as $u=(\cos\theta,\sin\theta)$ turns through half a circle. It peaks at $\textcolor{#1f6feb}{\lambda_1=12}$ at $53.1^\circ$, the direction $u_1=(0.6,0.8)$, and bottoms out at $\textcolor{#d1690a}{\lambda_2=2}$ at $143.1^\circ$, the line of $u_2=(0.8,-0.6)$. The two extremes are the eigenvalues, reached $90^\circ$ apart, and the curve oscillates around $7=\operatorname{tr}\Sigma/2$.
Looks like this, but is not
Every eigenvector satisfies the Lagrange condition $\Sigma u=\lambda u$, so for the five students $u=(0.8,-0.6)$ is also a solution.
The condition finds every stationary point, and $(0.8,-0.6)$ is the minimum: its variance is $2$. Compare the eigenvalues before choosing.
Eigenvalues and eigenvectors of the five students' covariance matrix
Find the eigenvalues and unit eigenvectors of $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$ and name the first principal component.
Multiply out: $\Sigma u_1=(5.6\cdot0.6+4.8\cdot0.8,\ 4.8\cdot0.6+8.4\cdot0.8)=(7.2,\ 9.6)=12\,(0.6,0.8)$. The angle $\arctan(4/3)=53.1^\circ$ matches the turning-angle example.
Trace and determinant give both eigenvalues of any 2 by 2 covariance matrix in one line, and the two always add up to the total variance.
A 3 by 3 covariance where the largest single variance does not win
Three sensors have covariance matrix $\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$. Sensor 3 has the largest variance. Is it the first principal component?
FindAll eigenvalues and eigenvectors, and the first component.
The eigenvalues add to $5+4+1=10=\operatorname{tr}\Sigma=3+3+4$, and $\Sigma\,(1, \allowbreak 1, \allowbreak 0)^T=(5, \allowbreak 5, \allowbreak 0)^T$ confirms the first one.
Correlated features pool their variance, so a component can beat every single feature: read the eigenvalues, not the diagonal.
The second component by deflation, the route of the lecture notes
Remove the first component from the five students, $Y=X-Xu_1u_1^T$, and find the best direction for what is left.
Find$\frac1nY^TY$, its top eigenvector, and that eigenvector's variance.
Row by row, $Y$ has rows $x_i-z_{i1}u_1$: $(0,0)$, $(0.8,-0.6)$, $(-1.6,1.2)$, $(-0.8,0.6)$, $(1.6,-1.2)$, with mean square length $\frac{0+1+4+1+4}5=2=\lambda_2$.
For two features this is a detour, since $u_2$ is just $u_1$ turned by $90^\circ$; its value is that it works the same way for any $p$, one component at a time.
Checkpoint
§09.2 — the first component of a diagonal covariance
Two uncorrelated features have variances $4$ and $9$, so their covariance matrix is diagonal.
Find(a) Give the first principal component and the variance its scores keep.
9.3The PCA algorithm: prepare the columns, score, reconstruct
Turns raw data into $k$ new features per point: center, standardize if the units differ, eigendecompose $\Sigma$, and take $k$ dot products.
We know which directions to use; this block is the full recipe the lecture writes as an algorithm, starting from raw columns.
MethodMethod 9.3: The PCA algorithm
Conditions
Center every column: $x_{ij}\leftarrow x_{ij}-\bar x_j$, where $\bar x_j=\frac1n\sum_ix_{ij}$.
If the features are in different units, also divide by $\sigma_j$, where $\sigma_j^2=\frac1n\sum_ix_{ij}^2$ after centering; $\Sigma$ then has ones on its diagonal and is the .
Compute the covariance matrix of the prepared data, sort its eigenvectors by eigenvalue, and describe each point by $k$ scores, one dot product per kept component. Going back, $\hat x_i$ adds the components with the scores as weights. It equals $x_i$ exactly when $k=p$, or when $x_i$ already lies in the span of $u_1,\dots,u_k$.
Proof
Why $\hat x_i$ is the natural way back. The eigenvectors $u_1,\dots,u_p$ form an orthonormal basis, so $x_i=\sum_{m=1}^p(x_i^Tu_m)\,u_m=\sum_{m=1}^pz_{im}u_m$.
Keeping the first $k$ terms gives $\hat x_i$; the dropped part $x_i-\hat x_i=\sum_{m>k}z_{im}u_m$ is perpendicular to every kept component.
Matrix form. With $U_k=[u_1,\dots,u_k]$, a $p\times k$ matrix, the scores are $Z=XU_k$ ($n\times k$) and the reconstructions are $\hat X=ZU_k^T$.
Why standardize. Multiplying feature $j$ by $c$ multiplies the off-diagonal entries of row and column $j$ of $\Sigma$ by $c$ and the diagonal entry by $c^2$, which turns the eigenvectors. After standardizing, every feature has variance $1$ whatever its unit.
The first component's angle for the five students when quiz 2 is multiplied by $c$. In points, $c=1$, it is $\textcolor{#1f6feb}{53.1^\circ}$; with quiz 2 in percent, $c=5$, it swings to $\textcolor{#1f6feb}{83.4^\circ}$, almost pure quiz 2; with quiz 1 in percent, the same as $c=0.2$, it drops to $\textcolor{#1f6feb}{10.0^\circ}$. Standardized first, it is $\textcolor{#d1690a}{45^\circ}$ for every $c$.
Looks like this, but is not
Changing the unit of a feature is harmless in least squares, where its coefficient simply rescales, so PCA should not care about units either.
PCA maximizes variance, and a change of unit multiplies that feature's variance by the square of the factor. Reporting quiz 2 in percent multiplies its variance by $25$ and swings the first component from $53.1^\circ$ to $83.4^\circ$.
Scores and one-component reconstructions for the five students
Run the algorithm with $k=1$ on the five students and rebuild each student's two quiz scores from one number.
Find$z_{i1}$ and $\hat x_i$ in points, for all five students.
For the third student $12+0.6\cdot1=12.6$ and $14+0.8\cdot1=14.8$.
Answer $$\boxed{z_{\cdot1}=(-5,-3,1,3,4);\qquad \hat x_3=(12.6,\ 14.8)\ \text{ for the raw }(11,16)}$$
Check
The miss $x_3-\hat x_3=(11,16)-(12.6,14.8)=(-1.6,1.2)$ is perpendicular to $u_1$: $-1.6\cdot0.6+1.2\cdot0.8=0$. The first student is rebuilt exactly, because that point lies on the line.
One score per student keeps the order and the spread along the main axis; what it loses is always perpendicular to the kept direction.
Reading loadings: four exam parts, two components
A course has four exam parts: calculus, linear algebra, reading and writing. After standardizing, their covariance matrix $\Sigma$, which is now a correlation matrix, has the eigen-pairs below. Interpret the first two components and summarize a student whose standardized scores are $x=(1.2,\ 0.8,\ -0.4,\ 0)$.
FindA reading of $u_1$ and $u_2$, the student's two scores, and $\hat x$ with $k=2$.
The miss $x-\hat x=(0.2,-0.2,-0.2,0.2)$ lies along $(1,-1,-1,1)$, the direction of $u_4$, so it is perpendicular to both kept components: $0.2-0.2-0.2+0.2=0$.
Name a component by the pattern of its loadings, then read a score as how much of that pattern a point has.
The same quizzes in other units: why the lecture standardizes
Quiz 2 is now reported in percent, points times 5, while quiz 1 stays in points. Find the first component, then redo PCA on standardized data.
Find$u_1$ for the percent data; $u_1$, $\lambda_1$ and $\lambda_2$ for standardized data.
Given
$\Sigma$ in points: $\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$
quiz 2 in percent: $x_2'=5x_2$
Solution
A change of units acts on $\Sigma$ directly, so there is no need to go back to the raw scores.
The eigenvalues add up to the traces, $212.78+2.82=215.6$ and $1.7+0.3=2$. Standardized, the answer is the same whether quiz 2 is in points or in percent, because $r$ does not change.
Use the raw covariance only when every feature is in the same unit and its size means something; otherwise standardize first.
Checkpoint
§09.3 — scoring and rebuilding a new student
A sixth student, not used to fit the PCA, scored $13$ on quiz 1 and $15$ on quiz 2. The fitted model has means $(12,14)$ and first component $u_1=(0.6,0.8)$.
Find(a) Find the student's score $z_1$ and the one-component reconstruction $\hat x$ in points.
Given
means $\bar x=(12,14)$
$u_1=(0.6,0.8)$
new raw scores $(13,15)$
Hint 1/4
Put the new student into the coordinates the model was fitted in, then go one step along $u_1$ and back.
Hint 2/4
$z_1=(x-\bar x)^Tu_1$ and $\hat x=\bar x+z_1u_1$.
Hint 3/4
Here $x-\bar x=(13,15)-(12,14)=(1,1)$ and $u_1=(0.6,0.8)$.
Hint 4/4
$z_1=1.4$ and $\hat x=(12+0.84,\ 14+1.12)=(12.84,\ 15.12)$.
Show solution
A new point is centered with the means of the data the model was fitted on, not with its own values.
The miss $(13,15)-(12.84,15.12)=(0.16,-0.12)$ is perpendicular to $u_1$: $0.096-0.096=0$.
Store the training means with the eigenvectors: every new point needs them twice, once going in and once coming back.
⚠ Forgetting to add the means back
The algorithm writes $\hat x_i$ for centered data, and the last step back to raw units is easy to skip.
wrong$$\hat x=z_1u_1=(0.84,\ 1.12)$$
right$$\hat x=\bar x+z_1u_1=(12.84,\ 15.12)$$
⚠ Running PCA on the raw covariance of features in arbitrary units
Standardizing looks like an optional extra step.
wrong$$\text{quiz 2 in percent, raw }\Sigma:\ \ u_1=(0.115,\ 0.993)$$
right$$\text{standardized first: }u_1=\tfrac1{\sqrt2}(1,1)\ \text{ in any units}$$
9.4How many components: the proportion of variance explained
Measures what the first $k$ components keep, $\frac{\lambda_1+\dots+\lambda_k}{\lambda_1+\dots+\lambda_p}$, so $k$ can be read off a scree plot or a target.
The algorithm left $k$ open; the eigenvalues already computed say exactly what each extra component buys.
DefinitionDefinition 9.4: Proportion of variance explained (PVE)
Conditions
$X$ is centered, and standardized if the features were
$\lambda_1\ge\dots\ge\lambda_p$ are the eigenvalues of $\Sigma=\frac1nX^TX$ and $z_{im}=x_i^Tu_m$ are the scores
The total variance is the sum of the feature variances, and it is also the sum of the eigenvalues. Component $m$ carries $\lambda_m$ of it, so its share is $\lambda_m$ over the total, and the first $k$ components carry the sum of their shares. One minus that sum is what the discarded components hold.
Proof
Numerator.$\frac1n\sum_iz_{im}^2=u_m^T\Sigma u_m=\lambda_m$, by Definition 9.1 and $\Sigma u_m=\lambda_mu_m$.
Denominator.$\sum_j\frac1n\sum_ix_{ij}^2=\sum_j\Sigma_{jj}=\operatorname{tr}\Sigma$, and the trace of a symmetric matrix is the sum of its eigenvalues. The factors $\frac1n$ cancel in the ratio.
What the rest costs (further reading). With $k$ components the average squared miss is $$\frac1n\sum_i\lVert x_i-\hat x_i\rVert^2=\frac1n\sum_i\sum_{m>k}z_{im}^2=\lambda_{k+1}+\dots+\lambda_p,$$ because the $u_m$ are orthonormal. So $1-\mathrm{PVE}(\text{first }k)$ is the reconstruction error as a share of the total.
Scree plot for the four exam parts. The bars are $\textcolor{#1f6feb}{\mathrm{PVE}(m)}=0.60,\ \allowbreak 0.25,\ \allowbreak 0.10,\ \allowbreak 0.05$; the line is the cumulative $\textcolor{#d1690a}{\mathrm{PVE}(\text{first }k)}=0.60,\ \allowbreak 0.85,\ \allowbreak 0.95,\ \allowbreak 1.00$. The drop flattens after the second bar, and a target of $0.9$ needs three components.
Looks like this, but is not
A first component with $\mathrm{PVE}(1)=0.987$ looks like a great compression, and four components with $0.25$ each look like a useless PCA.
PVE only compares variances. The $0.987$ came from quiz 2 recorded in percent, a unit artefact. Equal eigenvalues say that no direction is special, which is a finding about the data, not a failure.
component m
eigenvalue
PVE(m)
PVE(first m)
$1$
$2.4$
$0.60$
$0.60$
$2$
$1.0$
$0.25$
$0.85$
$3$
$0.4$
$0.10$
$0.95$
$4$
$0.2$
$0.05$
$1.00$
Two components keep $0.85$ of the variance; the third adds $0.10$ and the fourth only $0.05$.
PVE, the scree plot and the choice of k for the four exam parts
The standardized exam data have eigenvalues $2.4,\ \allowbreak 1.0,\ \allowbreak 0.4,\ \allowbreak 0.2$. Compute every PVE and the cumulative PVE, and choose $k$ by the elbow and by a target of $0.9$.
Find$\mathrm{PVE}(m)$ for $m=1,\dots,4$, the cumulative values, and $k$.
Given$\lambda=(2.4,\ 1.0,\ 0.4,\ 0.2)$ from the $4\times4$ correlation matrix
Solution
With standardized data every feature has variance $1$, so the total is simply $p=4$; check that the eigenvalues add up to it before dividing.
Shares
$$\sum_j\lambda_j=2.4+1.0+0.4+0.2=4=p$$
A correlation matrix has ones on its diagonal, so its trace is $p$.
With standardized data the eigenvalues look like shares, but they add up to $p$, not to 1.
wrong$$\mathrm{PVE}(1)=\lambda_1=2.4$$
right$$\mathrm{PVE}(1)=\frac{2.4}{4}=0.6$$
9.5Blind source separation: the model x = As and what it cannot recover
Models sensors as unknown mixtures of independent sources, $x=As$, and recovers $s=Wx$ with $W=A^{-1}$, up to order, scale and sign.
PCA kept variance; the second half of the chapter looks for hidden signals that were added together, and asks what the recordings can reveal about them.
DefinitionDefinition 9.5: The blind source separation model
Conditions
$p$ sources and $p$ sensors; the mixing matrix $A$ is $p\times p$, invertible and unknown
the sources are independent of each other, and $s(1),\dots,s(T)$ are independent over time
only $x(1),\dots,x(T)$ are observed; the lecture drops the time index when one time step is enough
$$\boxed{\begin{aligned}x(t)&=A\,s(t),\qquad x_i(t)=\sum_{j=1}^pa_{ij}\,s_j(t)\\ W&=A^{-1},\qquad \textcolor{#1f6feb}{\hat s(t)}=W\,x(t)\\ As&=(AP^TD^{-1})(DPs)\quad\text{for every permutation }P\text{ and invertible diagonal }D\end{aligned}}$$
Each sensor records a weighted sum of all sources, with weights nobody knows. If $A$ were known we would simply invert it. Since only $x$ is seen, reordering the sources and rescaling each one, with the matching change absorbed by $A$, explains the recordings equally well: order, scale and sign cannot be recovered.
Proof
A permutation matrix is orthogonal. $P$ has exactly one $1$ in each row and each column, so $P^TP=I$.
Same recordings.$(AP^TD^{-1})(DPs)=AP^T(D^{-1}D)Ps=AP^TPs=As$.
Still a valid model. $DPs$ lists the same sources in another order, each multiplied by a nonzero constant, so its entries are still independent, and $AP^TD^{-1}$ is still invertible.
Sign is the special case of a $D$ with entries $\pm1$.
Two hidden sources, a $\textcolor{#1a7f37}{\text{sawtooth }s_1}$ and a $\textcolor{#1a7f37}{\text{sine }s_2}$, reach two microphones through $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$. A correct ICA output is $\textcolor{#1f6feb}{\hat s_1=-s_2}$, $\textcolor{#1f6feb}{\hat s_2=2s_1}$: both shapes are back, in another order, one of them flipped and the other twice as tall.
Looks like this, but is not
An unmixing matrix whose output is $\hat s=(-s_2,\ 2s_1)$ has failed: it returned the wrong signal first, upside down and at the wrong volume.
That output is $DPs$ with a swap $P$ and $D=\operatorname{diag}(-1,2)$, which Definition 9.5 says the recordings cannot rule out. For telling voices apart, their order and loudness do not matter.
Two microphones, two speakers: mixing and unmixing with a known A
Two speakers take turns; at four instants their signals are $s(1)=(2,0)$, $s(2)=(0,-1)$, $s(3)=(1,0)$ and $s(4)=(0,2)$. The microphones mix them with $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$. Find the recordings, then recover the sources.
Find$x(t)=As(t)$ for $t=1,\dots,4$, the unmixing matrix $W$, and $Wx(t)$.
With $A$ known this is linear algebra, one product to mix and one 2 by 2 inverse to unmix; the point is to see what a blind method must reproduce without $A$.
Recordings of one active source line up along a column of A; that is the kind of structure a blind method can exploit.
Which unmixing matrices are equally good answers?
The true unmixing matrix is $W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$. Which of the four candidates below could an ICA method return with equal justification?
FindFor each candidate, its output $W_\ast x=W_\ast As$ in terms of $s_1$ and $s_2$.
A Gaussian vector is fixed by its mean and covariance. Mixing with $A$ or with $AQ$ gives the same mean and the same covariance, so the same distribution, and no amount of data separates the two. The unmixing matrices $W$ and $Q^TW$ look equally good, yet $Q^TW$ returns rotated mixtures of the sources.
Proof
Gaussian in, Gaussian out. A linear transformation of a Gaussian vector is Gaussian, so $x$ is Gaussian with $E[x]=A\,E[s]=0$.
Its covariance.$E[xx^T]=E[As\,s^TA^T]=A\,E[ss^T]\,A^T=AIA^T=AA^T$.
Rotate the mixing matrix.$E[x'x'^T]=AQ\,I\,Q^TA^T=A(QQ^T)A^T=AA^T$: same mean, same covariance, same Gaussian.
Where non-Gaussian sources escape. For them equal means and covariances do not force equal distributions, as the uniform square in the figure shows.
Left: sources uniform on $[-1,1]$ fill a square, and mixing with $\textcolor{#1f6feb}{A}$ or with the rotated $\textcolor{#d1690a}{AQ}$, $Q=\begin{bmatrix}0.6&-0.8\\0.8&0.6\end{bmatrix}$, gives two different parallelograms, although both clouds have covariance $\frac13AA^T$. Right: standard Gaussian sources give one ellipse, the same for $A$ and for $AQ$.
Looks like this, but is not
If one of two sources is Gaussian, the same argument should rule out ICA, since that source can be rotated freely.
The argument needs every source Gaussian. With one Gaussian and one uniform source, $x=As$ and $x'=AQs$ have different distributions; the standard result, further reading beyond the lecture, allows at most one Gaussian source.
Two different mixing matrices that Gaussian data cannot tell apart
Take $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$ and the rotation $Q=\begin{bmatrix}0.6&-0.8\\0.8&0.6\end{bmatrix}$. Show that $A$ and $A'=AQ$ give Gaussian recordings with the same covariance, and find what the unmixing matrix of $A'$ returns.
Find$A'$, $AA^T$, $A'A'^T$, and $A'^{-1}x$ in terms of $s$.
Both entries lie in [−1, 1], so the point is inside the blue parallelogram.
Under AQ
$$W'x=(0.48+0.6,\ -1.44+1.2)=(1.08,\ -0.24)$$
The first entry exceeds 1: no source values produce this recording.
Answer $$\boxed{x=(1.2,1.2)\ \text{is possible under }A\text{ and impossible under }AQ}$$
Check
Both clouds still share the covariance $\frac13AA^T$, because a uniform source on $[-1,1]$ has variance $\frac13$; the difference lives entirely in the shape.
Enough recordings near the corners reveal the mixing matrix; Gaussian recordings have no corners to reveal.
Checkpoint
§09.6 — separating two Gaussian speakers
Two independent speakers are modelled as standard Gaussian signals and mixed by an unknown invertible $A$. An engineer plans to record for an hour instead of a minute.
Find(a) True or false: with enough recordings, ICA recovers the two speakers up to order, scale and sign.
Given$s\sim N(0,I)$, $x=As$, $A$ unknown and invertible
Hint 1/4
Ask what the distribution of $x$ depends on, and whether different mixing matrices can share it.
Hint 2/4
For Gaussian sources $x\sim N(0,AA^T)$, and $(AQ)(AQ)^T=AA^T$ for every orthogonal $Q$.
Hint 3/4
Here $s\sim N(0,I)$ and $A$ is unknown, so the recordings reveal only $AA^T$.
Hint 4/4
A rotation of $A$ stays hidden however long we record: false.
Show solution
One second mixing matrix with the same distribution is enough to refute a claim about any amount of data.
The distribution of x
$$x\sim N(0,\ AA^T)$$
A linear transformation of a Gaussian, with covariance $AIA^T$.
A second matrix with the same distribution
$$(AQ)(AQ)^T=AQQ^TA^T=AA^T$$
Any orthogonal $Q$ that is not a signed permutation gives a genuinely different unmixing.
Answer $$\boxed{\text{False}}$$
Check
The rotation example is a concrete pair: $A$ and $AQ$, with the 3-4-5 rotation $Q$, share $\begin{bmatrix}1.25&1\\1&1.6\end{bmatrix}$.
is a property of the model, not of the sample size; more data cannot repair it.
⚠ Treating uncorrelated outputs as separated sources
PCA makes its outputs uncorrelated, and that looks like separation.
wrong$$\operatorname{Cov}(\hat s)=I\ \Rightarrow\ \hat s=s\ \text{up to order and sign}$$
right$$\operatorname{Cov}(Q^Ts)=I\ \text{for every orthogonal }Q$$
⚠ Blaming a single Gaussian source
The theorem is remembered as 'any Gaussian source breaks ICA'.
To get the density of a recording, unmix it, evaluate each source density at its unmixed value, multiply, and multiply once more by $\lvert\det W\rvert$, the correction for how much $A$ stretches areas. The log-likelihood adds this up over the $T$ recordings; the estimate is the $W$ that makes the recordings most probable.
Proof
One dimension. If $X=aS$ with $a>0$, then $P(X\le x)=P(S\le x/a)$, and differentiating gives $p_X(x)=p_S(x/a)\cdot\frac1a$. For $a<0$ the inequality flips and the factor becomes $\frac1{\lvert a\rvert}$.
Many dimensions. $A$ maps a small box of volume $V$ around $s$ to a parallelogram of volume $\lvert\det A\rvert\,V$ around $x=As$, carrying the same probability, so the density is divided by $\lvert\det A\rvert$.
In terms of W.$\lvert\det A\rvert^{-1}=\lvert\det A^{-1}\rvert=\lvert\det W\rvert$, and independence splits $p_S(Wx)$ into $\prod_jp_{S_j}(w_j^Tx)$, since the $j$-th entry of $Wx$ is $w_j^Tx$.
Over time. Independent recordings multiply, $L(W)=\prod_tp_X(x(t))$, and the log turns the product into the sum $l(W)$; the term $\log\lvert\det W\rvert$ appears once per recording, $T$ times in all.
Climbing it (further reading). With $g_j=(\log p_{S_j})'$ applied entry by entry, $\nabla_Wl=\sum_tg\big(Wx(t)\big)\,x(t)^T+T\,(W^{-1})^T$. For a $2\times2$ matrix the last term comes from $\frac{\partial}{\partial w_{11}}\log\lvert w_{11}w_{22}-w_{12}w_{21}\rvert=\frac{w_{22}}{\det W}$ and its three analogues.
Source values and recordings drawn in one plane. The $\textcolor{#1a7f37}{\text{unit square}}$ of source values, probability $1$ spread evenly, is mapped by $A=\begin{bmatrix}2&1\\0.5&1.5\end{bmatrix}$ to a $\textcolor{#1f6feb}{\text{parallelogram}}$ of area $\lvert\det A\rvert=2.5$, so the same probability over $2.5$ times the area means density $1/2.5=0.4$. To evaluate $p_X$ at $x=(2,1.5)$, unmix it: $Wx=(0.6,0.8)$ lies in the square, so $p_X(x)=1\cdot0.4$.
Looks like this, but is not
Since $s=Wx$, the density of a recording should simply be $p_S(Wx)$.
That forgets the stretching. With $X=2S$ and $S$ uniform on $[0,1]$, $p_S(x/2)=1$ on $[0,2]$ would integrate to $2$; the factor $\lvert\det W\rvert=\frac12$ brings the total back to $1$.
The density of a stretched and of a mixed variable
(a) $S$ is uniform on $[0,1]$ and $X=2S$; find $p_X$. (b) $s$ is uniform on the unit square and $x=As$ with $A=\begin{bmatrix}2&1\\0.5&1.5\end{bmatrix}$; find $p_X$ at $x=(2,1.5)$ and at $x=(0.5,1.5)$.
Find$p_X$ in (a), and the two density values in (b).
Given
$S\sim U[0,1]$ and $X=2S$
$s$ uniform on $[0,1]^2$, $A=\begin{bmatrix}2&1\\0.5&1.5\end{bmatrix}$
Solution
Both parts follow the theorem's recipe: unmix, evaluate the source density, multiply by $\lvert\det W\rvert$.
Outside the square: no source values produce this recording.
Answer $$\boxed{\text{(a) }p_X=\tfrac12\text{ on }[0,2];\qquad \text{(b) }p_X(2,1.5)=0.4,\quad p_X(0.5,1.5)=0}$$
Check
Total probability: in (a) $\frac12\cdot2=1$; in (b) the parallelogram has area $2.5$ and density $0.4$, and $2.5\cdot0.4=1$.
Every density evaluation in ICA follows this recipe: unmix, look up the source densities, multiply by $\lvert\det W\rvert$.
The log-likelihood of four recordings under three candidate unmixing matrices
The four recordings of the two speakers are $x(1)=(2,0.8)$, $x(2)=(-0.5,-1.2)$, $x(3)=(1,0.4)$ and $x(4)=(1,2.4)$. Model each source with the Laplace density $p(s)=\frac12e^{-\lvert s\rvert}$ and compare $l$ for the true $W$, for $I$, and for the rotated $W'=Q^TW$.
Find$l(W)$, $l(I)$, $l(W')$ and the best of the three.
$W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$ and $W'=\begin{bmatrix}0.4&0.5\\-1.2&1\end{bmatrix}$; all three candidates have $\lvert\det\rvert=1$
$p(s)=\frac12e^{-\lvert s\rvert}$ for both sources
Solution
With Laplace sources $\log p(s)=-\log2-\lvert s\rvert$, so $l$ reduces to a sum of absolute values of the unmixed recordings plus the determinant term.
Swap the Laplace density for a standard Gaussian: $l$ then depends on squares, $\sum_t\lVert Q^Ts(t)\rVert^2=\sum_t\lVert s(t)\rVert^2=10$, and $W$ and $W'$ tie at $-12.352$, as the Gaussian theorem predicts.
Four recordings and three candidates: twelve unmixings by hand. A computer does the same for every W along a gradient.
A peaked, heavy-tailed source model rewards unmixing matrices whose outputs are often near zero, and that is what lets the likelihood choose a rotation.
The scale the likelihood picks for one source
One Laplace source and one sensor: $x_t=a\,s_t$ with $a$ unknown, and four recordings $3,\ -1,\ 0.5,\ -1.5$. Find the maximum likelihood unmixing weight $\hat w$ and the recovered sources.
Find$\hat w$ and $\hat s_t=\hat w\,x_t$.
Given
$x=(3,\ -1,\ 0.5,\ -1.5)$
$p(s)=\frac12e^{-\lvert s\rvert}$; $W=w$ is a nonzero number
Solution
With $p=1$ the determinant is $w$ itself, so $l$ is a function of one number and calculus finds its maximum.
$l(\tfrac23)=-8.394>l(1)=-8.773$, and the recovered values have mean absolute value $\frac{2+0.667+0.333+1}4=1$, the mean absolute value of a Laplace source.
Fixing the source density fixes the scale, but $l(-w)=l(w)$: the sign stays free, exactly as the list of ambiguities says.
Checkpoint
§09.7 — the density of a flipped and stretched variable
A sensor records $X=-3S$, where $S$ is uniform on $[0,1]$.
Find(a) Find $p_X(x)$ and where it is positive.
Given
$S\sim U[0,1]$
$X=-3S$
Hint 1/4
Unmix: express $S$ through $X$, then correct for the stretching.
Hint 2/4
$p_X(x)=p_S(wx)\,\lvert w\rvert$ with $w=a^{-1}$ when $X=aS$.
Hint 3/4
Here $a=-3$, so $w=-\frac13$, and $p_S=1$ on $[0,1]$: $p_X(x)=1\cdot\frac13$ wherever $-x/3\in[0,1]$.
Hint 4/4
$p_X=\frac13$ on $[-3,0]$ and $0$ elsewhere.
Show solution
The one-dimensional case of the theorem needs one unmixing and one absolute value.
A question gives a few two-dimensional points or a 2 by 2 covariance matrix and asks for components, scores or PVE.
Center
Subtract each column's mean from that column.
Covariance
$\Sigma=\frac1nX^TX$: two sums of squares and one sum of cross-products, each divided by $n$.
Eigenvalues
Solve $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$ and sort the roots.
Eigenvector
From $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$, then divide by the length; $u_2$ is $u_1$ turned by $90^\circ$.
Scores and PVE
$z_{i1}=x_i^Tu_1$ on the centered points; $\mathrm{PVE}(1)=\lambda_1/(\lambda_1+\lambda_2)$.
Where it goes wrong
Writing $\lambda_1-\Sigma_{11}$ in step 4: a sign slip that tilts the direction.
Skipping the division by the length, which inflates every score.
Scoring the raw points after fitting on the centered ones.
Evaluating the ICA log-likelihood of a candidate W
A question gives recordings, source densities and one or more candidate unmixing matrices.
Unmix
Compute $Wx(t)$ for every recording.
Source log-densities
Add $\log p_{S_j}$ of every entry; for Laplace sources this is $-\log2-\lvert\cdot\rvert$.
Determinant
Add $T\log\lvert\det W\rvert$, once per recording.
Compare
The candidate with the larger total is the more likely unmixing.
Where it goes wrong
Dropping $T\log\lvert\det W\rvert$, which lets $W$ shrink toward $0$ and win.
Adding $\log\lvert\det W\rvert$ once instead of $T$ times.
Reading a tie between W and a row-swapped or sign-flipped W as a failure: when the sources share one symmetric density, those changes never alter the log-likelihood.
Checking whether an ICA output is acceptable
A question gives the true mixing matrix, or the true sources, and a returned unmixing matrix or returned signals.
Multiply
Form $M=W_\ast A$, or write each output as $Ms$.
Count
Each row and each column of $M$ must hold exactly one nonzero entry.
Read
Then $M=DP$: the permutation says which source each output is, and $D$ gives its scale and sign.
Where it goes wrong
Calling a swapped, flipped or rescaled output a failure.
Accepting an output whose row of $M$ has two nonzero entries: that output is still a mixture.
Least squares: quiz 2 predicted from quiz 1
Fit $x_2=bx_1$ to the five centered students by least squares and measure the misses vertically.
The first component's line, slope $4/3$, has mean squared vertical miss $\frac{50}9=5.556$, larger, as it must be: least squares is the best at vertical misses.
PCA: the line that treats both quizzes alike
Take the first component of the same students and measure the misses perpendicular to its line.
FindThe slope of the line and the mean squared perpendicular miss.
The least squares direction $(7,6)/\sqrt{85}$ keeps $11.529$, so its perpendicular misses average $14-11.529=2.471>2$.
Same five points, two lines: least squares measures misses vertically and gets slope $0.857$; the first component measures them perpendicular to the line and gets slope $1.333$. Swapping the roles of the quizzes changes the first answer, never the second.
How to tell them apart
A response to predict means least squares; a summary of unlabelled features that treats them alike means PCA.
PCA on the recordings of two speakers who take turns
Speakers take turns, $s=(2,0),\ \allowbreak (-2,0),\ \allowbreak (0,1),\ \allowbreak (0,-1)$, and $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$ mixes them. Run PCA on the four recordings.
FindThe two principal directions.
Givenrecordings $(2,0.8),\ \allowbreak (-2,-0.8),\ \allowbreak (0.5,1.2),\ \allowbreak (-0.5,-1.2)$, with mean $(0,0)$
Solution
The recordings are already centered, so $\Sigma$ comes straight from them.
Trace $3.165$ and determinant $2.21-1.21=1$; directions from $(\Sigma-\lambda I)u=0$.
Answer $$\boxed{u_1\text{ at }31.9^\circ,\qquad u_2\text{ at }121.9^\circ\ \ \text{(perpendicular)}}$$
Check
The speakers' own directions are the columns of $A$, at $\arctan0.4=21.8^\circ$ and $\arctan2.4=67.4^\circ$; PCA's directions match neither.
ICA versus PCA on two mixed speakers
Two speakers who take turns are mixed into four centered recordings. Compare the Laplace log-likelihood of the true unmixing matrix $W$ with that of PCA's rotation $U^T$, whose rows are the principal directions $u_1^T$ and $u_2^T$.
FindWhich unmixing the Laplace likelihood prefers.
source density $p(s)=\tfrac12e^{-\lvert s\rvert}$, so $l(W)=\sum_t\sum_j\log p(w_j^Tx(t))+4\log\lvert\det W\rvert$
$W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$
$U^T$ rows $u_1\approx(0.849,\,0.528)$ and $u_2\approx(-0.528,\,0.849)$, at $31.9^\circ$ and $121.9^\circ$
both matrices have $\lvert\det\rvert=1$
Solution
With equal determinants, $l$ is decided by the sum of the absolute unmixed values.
The true unmixing
$$\sum_t\sum_j\lvert w_j^Tx(t)\rvert=2+2+1+1=6$$
W returns the sources, and each has one zero entry.
PCA's rotation
$$\sum_t\sum_j\lvert u_j^Tx(t)\rvert=8.62$$
Each recording is spread over both principal directions.
Answer $$\boxed{l(W)-l(U^T)=8.62-6=2.62>0}$$
Check
The columns of $A$ are not perpendicular, $1\cdot0.5+0.4\cdot1.2=0.98\ne0$, so no rotation such as $U^T$ can undo $A$.
On the same four recordings PCA returns perpendicular directions at $31.9^\circ$ and $121.9^\circ$, while the speakers sit along the columns of $A$ at $21.8^\circ$ and $67.4^\circ$; only $W=A^{-1}$ turns each recording back into one active speaker.
How to tell them apart
PCA looks for uncorrelated directions of large variance and always returns perpendicular ones; ICA looks for independent sources, whose directions need not be perpendicular.
Scaffolding comes off
The common skeleton
Center each column by subtracting its mean.
Form $\Sigma=\frac1nX^TX$.
Solve $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$ and sort the roots.
Get $u_1$ from $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$ and normalize it.
Score the centered points, $z_{i1}=x_i^Tu_1$, and report $\mathrm{PVE}(1)=\lambda_1/(\lambda_1+\lambda_2)$.
1 · fully worked
Full PCA of four points by hand
Four points are $(2,1)$, $(4,3)$, $(6,5)$ and $(8,3)$. Find the first principal component, the scores and $\mathrm{PVE}(1)$.
Find$u_1$, the scores $z_{i1}$ and $\mathrm{PVE}(1)$.
The scores sum to $0$, and their mean square is $\frac{64+4+16+36}{5\cdot4}=\frac{120}{20}=6=\lambda_1$.
The same five moves solve every two-feature PCA question; only the arithmetic changes.
2 · you write the reasoning
Easier, and this time you write the reasons. The covariance matrix is given directly: $\Sigma=\begin{bmatrix}3&1\\1&3\end{bmatrix}$. Find $u_1$, $\mathrm{PVE}(1)$ and the score of the centered point $(2,0)$. For each line, write why it is allowed.
The first row of $(\Sigma-4I)u=0$ forces $u_b=u_a$; dividing by $\sqrt2$ makes it a unit vector.
$\mathrm{PVE}(1)=\tfrac4{4+2}=0.667$
reasoning
The first eigenvalue over the sum of both.
$z=(2,0)^Tu_1=\tfrac2{\sqrt2}=1.414$
reasoning
The point is already centered, so its score is one dot product with $u_1$.
3 · find the buried error
Harder, with two errors buried in the solution. Four points are $(5,3)$, $(9,3)$, $(11,7)$ and $(15,7)$. A student finds the first component, the score of the first point and $\mathrm{PVE}(1)$. Which two steps are wrong?
Step 1. Means $(10,5)$; centered points $(-5,-2)$, $(-1,-2)$, $(1,2)$, $(5,2)$.
Follow the five moves of the skeleton: center, covariance, eigenvalues, eigenvector, then scores and PVE.
Hint 2/4
$\Sigma=\frac1nX^TX$ on centered data; $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$; $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$.
Hint 3/4
Here the readings $(1,3), \allowbreak (3,5), \allowbreak (5,9), \allowbreak (7,7)$ have means $(4,6)$, so the centered points are $(-3,-3), \allowbreak (-1,-1), \allowbreak (1,3), \allowbreak (3,1)$ and $\Sigma=\frac14\begin{bmatrix}20&16\\16&20\end{bmatrix}$.
The scores sum to $0$ and their mean square is $\frac{36+4+16+16}{2\cdot4}=\frac{72}8=9=\lambda_1$.
When the two diagonal entries of a 2 by 2 covariance are equal, the components are always the diagonals (1, 1) and (1, −1), whatever the off-diagonal entry.
Full exam-style question
Three sensors, four readings: PCA by hand and an ICA follow-upexam format
Three sensors in the same unit give four readings: $(13,21,7)$, $(11,23,3)$, $(9,17,3)$ and $(7,19,7)$.
(a) Center the data and compute $\Sigma=\frac14X^TX$.
(b) Find all eigenvalues and unit eigenvectors.
(c) Give $\mathrm{PVE}(1)$ and the smallest $k$ with $\mathrm{PVE}(\text{first }k)\ge0.8$.
(d) For the first reading, give the scores $z_1$, $z_2$ and the reconstruction with $k=2$ in the original units.
(e) Suppose the three signals are mixtures of independent non-Gaussian sources $s_1,s_2,s_3$, and ICA returns $\hat s=(-2s_3,\ s_1,\ 0.5s_2)$. Has it failed?
Find$\Sigma$; every $\lambda_m$ and $u_m$; the PVE and $k$; $z_1$, $z_2$ and $\hat x$ for the first reading; a verdict on the ICA output.
Each output is one source times a nonzero constant, so it has not failed.
Answer $$\boxed{\lambda=(8,4,2);\quad k=2;\quad \hat x_1=(12,\ 22,\ 7);\quad \text{ICA has not failed}}$$
Check
The reconstruction misses the first reading $(13,21,7)$ by $(1,-1,0)=\sqrt2\,u_3$, of squared length $2=\lambda_3$; every reading here misses by the same amount, so the average is $\lambda_3$, as it must be.
Check Σ for zero blocks first; and judge an ICA output by whether each output is a single rescaled source, never by its order or size.
Practice
A · concept 4 questions
1§09.1 — is the first component a regression line?
A classmate fits a least squares line of quiz 2 on quiz 1 to the five students and calls it the first principal component.
Find(a) True or false: the first principal component of two features is the least squares line of the second feature on the first.
Least squares minimizes $\sum_i(x_{i2}-bx_{i1})^2$, the vertical misses; the first component maximizes $u^T\Sigma u$, which minimizes the perpendicular misses.
Hint 3/4
Here the least squares slope is $24/28=0.857$, and $u_1=(0.6,0.8)$ has slope $0.8/0.6$.
Hint 4/4
$0.857\ne1.333$: false.
Show solution
One pair of numbers that differ refutes the claim.
A variance scales with the square of the factor, a covariance with the factor itself.
A concrete case
$$\text{quiz 2}\times5:\qquad u_1\ \text{turns from }53.1^\circ\text{ to }83.4^\circ$$
Computed in the algorithm block.
Answer $$\boxed{\text{False}}$$
Check
After standardizing, both versions have the same correlation matrix and so the same components.
Units are part of the input to PCA; standardize when they are arbitrary.
3§09.6 — uncorrelated scores, independent scores?
The scores on different principal components always have zero sample covariance.
Find(a) True or false: because the scores on the first two components are uncorrelated, they are independent.
Given
$z_{im}=x_i^Tu_m$
$\frac1n\sum_iz_{i1}z_{i2}=u_1^T\Sigma u_2=0$
Hint 1/4
Separate two claims: zero covariance, and a joint distribution that factorizes.
Hint 2/4
$\frac1n\sum_iz_{i1}z_{i2}=u_1^T\Sigma u_2=\lambda_2u_1^Tu_2=0$, but independence needs $p(z_1,z_2)=p(z_1)\,p(z_2)$.
Hint 3/4
Here, for points spread evenly round the unit circle, $\Sigma=\frac12I$: any two perpendicular directions give uncorrelated scores, yet $z_1^2+z_2^2=1$ ties them.
Hint 4/4
Uncorrelated but not independent: false.
Show solution
A general proof for 'uncorrelated' and one counterexample for 'independent'.
Each has one nonzero entry per row and column: a $DP$.
Answer $$\boxed{(s_1+s_2,\ s_2)}$$
Check
Independence check: $s_1+s_2$ and $s_2$ have covariance $\operatorname{Var}(s_2)\ne0$, so these outputs are not even uncorrelated.
Judge an unmixing by whether each output is a single source, never by the order, sign or size of the outputs.
B · computation 7 questions
1§09.1 — variance along the all-ones direction in three dimensions
Three sensors have covariance $\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$. An engineer summarizes each reading by the plain average of the three sensors, turned into a unit vector.
Find
(a) Compute $u^T\Sigma u$.
(b) Compare it with the best possible value, $\lambda_1$.
Run PCA twice, once on $\Sigma$ and once on the correlation matrix, and compare.
Hint 2/4
Standardizing turns $\Sigma_{jk}$ into $\frac{\Sigma_{jk}}{\sigma_j\sigma_k}$; a 2 by 2 correlation matrix has eigenvalues $1\pm r$ and eigenvectors $\frac1{\sqrt2}(1,\pm1)$.
Hint 3/4
Here $\operatorname{tr}\Sigma=13$, $\det\Sigma=36-9=27$, $\sigma_1=2$, $\sigma_2=3$ and $r=\frac{3}{2\cdot3}$.
The raw eigenvalues add to $13=4+9$; the standardized ones add to $2$, the number of features.
The raw answer leans toward the feature with the larger variance; standardizing removes the lean and gives each feature one unit of variance.
5§09.4 — one component for the four exam parts
A student's standardized scores on calculus, linear algebra, reading and writing are $x=(0.6,\ 1.0,\ -0.2,\ 0.2)$. The first component is $u_1=\frac12(1,1,1,1)$, with $\lambda_1=2.4$ out of a total of $4$.
Find
(a) Find the score $z_1$ and the reconstruction $\hat x$ with $k=1$.
(b) Find the squared reconstruction error, and check it against the other three scores.
Go one step along $u_1$ and back, then measure what is left.
Hint 2/4
$z_1=x^Tu_1$, $\hat x=z_1u_1$, and $\lVert x-\hat x\rVert^2=z_2^2+z_3^2+z_4^2$.
Hint 3/4
Here $x=(0.6,1.0,-0.2,0.2)$ and every entry of $u_1$ is $\frac12$; the other directions are $\frac12(1,1,-1,-1)$, $\frac12(1,-1,1,-1)$ and $\frac12(1,-1,-1,1)$.
Hint 4/4
$z_1=0.8$, $\hat x=(0.4,0.4,0.4,0.4)$, and the error is $0.8$.
Show solution
Computing the error directly and through the dropped scores checks the arithmetic.
Mix the second one back: $A(-2,0.5)^T=(-2+1,\ -1+1)=(-1,0)=x(2)$.
With A known this is routine; ICA has to find the same W from the recordings alone.
7§09.7 — the density of one recording
Two Laplace sources, each with density $p(s)=\frac12e^{-\lvert s\rvert}$, are mixed, and a candidate unmixing matrix is $W=\begin{bmatrix}1&1\\-1&1\end{bmatrix}$. One recording is $x=(1,1)$.
Find
(a) Compute $p_X(x)$ under this $W$.
(b) Give the term this recording contributes to $l(W)$.
Given
$W=\begin{bmatrix}1&1\\-1&1\end{bmatrix}$
$x=(1,1)$
$p(s)=\frac12e^{-\lvert s\rvert}$ for each source
Hint 1/4
Unmix the recording, evaluate both source densities, and correct for the stretching.
Term by term in the log-likelihood: $(-\log2-2)+(-\log2-0)+\log2=-\log2-2$, the same number.
Each recording contributes one such term, and $l(W)$ is their sum over time.
C · exam level 4 questions
1§09.4 — choosing k for six sensor channels
A PCA on six standardized sensor channels returns eigenvalues $2.7,\ \allowbreak 1.5,\ \allowbreak 0.9,\ \allowbreak 0.5,\ \allowbreak 0.3,\ \allowbreak 0.1$. The team wants the fewest components that keep at least $0.85$ of the variance.
Components add their eigenvalues, so the shares accumulate.
Answer $$\boxed{k=3}$$
Check
The discarded eigenvalues $0.5+0.3+0.1=0.9$ are $\frac{0.9}6=0.15$ of the total, and $1-0.15=0.85$.
Meeting a target exactly counts: read 'at least' as $\ge$.
2§09.7 — two candidate unmixing matrices
Three recordings $x(1)=(0.5,0)$, $x(2)=(0,-0.5)$ and $x(3)=(0.25,0.25)$ come from two Laplace sources with density $p(s)=\frac12e^{-\lvert s\rvert}$. The candidates are $W_1=I$ and $W_2=2I$.
Find(a) Which candidate has the larger log-likelihood?
Compute both log-likelihoods in full, the determinant term included.
Hint 2/4
$l(W)=-2T\log2-\sum_t\sum_j\lvert w_j^Tx(t)\rvert+T\log\lvert\det W\rvert$ with $T=3$.
Hint 3/4
Here $\sum_t\sum_j\lvert x_j(t)\rvert=0.5+0.5+0.5=1.5$; $W_2$ doubles it; $\det W_1=1$ and $\det W_2=4$.
Hint 4/4
$l(W_1)=-5.659$ and $l(W_2)=-3.000$: $W_2$ wins.
Show solution
Both candidates share the constant, so compute it once and add the two parts that differ.
Common part
$$-2T\log2=-6\log2=-4.159$$
Each of the $2T=6$ unmixed values contributes $-\log2$.
W₁
$$l(W_1)=-4.159-1.5+3\log1=-5.659$$
Absolute sum $0.5+0.5+(0.25+0.25)=1.5$.
W₂
$$l(W_2)=-4.159-3+3\log4=-4.159-3+4.159=-3.000$$
Doubling $W$ doubles the absolute sum, and $\det(2I)=4$.
Answer $$\boxed{W_2}$$
Check
Along $W=cI$, $l(c)=-4.159-1.5c+6\log c$ peaks at $c=\frac6{1.5}=4$, so $l$ rises from $c=1$ through $c=2$, consistent with $W_2$ beating $W_1$.
Without the determinant term, shrinking $W$ toward $0$ would always look better; the term keeps the recovered sources at the scale of the assumed density.
3§09.6 — which sources can be separated
An engineer can model two independent sources in four ways before they are mixed by an unknown invertible $A$.
Find(a) In which case can ICA recover the sources up to order, scale and sign?
Given
$x=As$ with $A$ unknown and invertible
the two sources are independent
Hint 1/4
Ask whether a genuinely different mixing matrix produces the same distribution of $x$.
Hint 2/4
Gaussian sources: $A$ and $AQ$ give the same distribution for every orthogonal $Q$, and the argument needs every source Gaussian.
Hint 3/4
Here the four options are uniform with Laplace; Gaussian with different variances; Gaussian with a long recording; Gaussian after standardizing.
Hint 4/4
Only the uniform and Laplace pair has no Gaussian source.
Show solution
The rotation argument either applies to a case or it does not; checking it case by case settles the question.
Gaussian cases
$$s=Ds',\ \ s'\sim N(0,I):\qquad x=(AD)s'\ \text{and}\ (ADQ)s'\ \text{agree in distribution}$$
Different variances only rescale; a longer recording and standardizing leave the model unchanged.
Non-Gaussian case
$$\text{uniform and Laplace}:\qquad AQs\ne As\ \text{in distribution for }Q\ne DP$$
Corners and heavy tails turn with Q and show in the data.
Answer $$\boxed{\text{uniform and Laplace}}$$
Check
The uniform example gives a concrete test: the recording $(1.2,1.2)$ is possible under $A$ and impossible under $AQ$.
Whether ICA can work is decided by the source distributions, not by preprocessing or by how much data there is.
4§09.2 — three equally correlated features
Three features have equal variances $2$ and equal covariances $1$.
Find
(a) Show that $u_1=\frac1{\sqrt3}(1,1,1)$ is an eigenvector and give $\lambda_1$.
(b) Find $\lambda_2$ and $\lambda_3$ from the trace, and say whether the second component is unique.
Test $(1,1,1)$ first, then use what the trace says about the other two eigenvalues.
Hint 2/4
$\operatorname{tr}\Sigma=\lambda_1+\lambda_2+\lambda_3$, and any vector perpendicular to $(1,1,1)$ can be tested with $\Sigma v=\lambda v$.
Hint 3/4
Here each row of $\Sigma=\begin{bmatrix}2&1&1\\1&2&1\\1&1&2\end{bmatrix}$ sums to $4$, the trace is $6$, and $v=(1,-1,0)$ is perpendicular to $(1,1,1)$.
Hint 4/4
$\lambda_1=4$; $\Sigma v=v$, so $\lambda_2=\lambda_3=1$ and the second component is not unique; $\mathrm{PVE}(1)=\frac46=0.667$.
Show solution
Equal row sums give one eigenvector at once, and the trace gives the sum of the rest.
$\Sigma=I+\mathbf 1\mathbf 1^T$ with $\mathbf 1=(1,1,1)^T$, so on vectors perpendicular to $\mathbf 1$ it acts as the identity: eigenvalue $1$ on a whole plane.
Tied eigenvalues mean a tied direction: software returns one choice, and a small change in the data may return a rotated one.
D · interleaved 4 questions
1§09.1 — five students and a final exam
The five students' first-component scores are $z_1=(-5, \allowbreak -3, \allowbreak 1, \allowbreak 3, \allowbreak 4)$ and their second-component scores $z_2=(0, \allowbreak 1, \allowbreak -2, \allowbreak -1, \allowbreak 2)$. Their centered final exam marks are $y=(-10, \allowbreak -5, \allowbreak 0, \allowbreak 5, \allowbreak 10)$.
Find
(a) Fit $y=\beta_1z_1$ by least squares.
(b) Fit $y=\beta_1z_1+\beta_2z_2$. Does $\hat\beta_1$ change?
Regressing on principal component scores decouples the coefficients, because the scores are uncorrelated.
2§09.3 — preparing data before cross-validation
A team standardizes 500 samples, runs PCA on all of them and keeps 10 components. It then runs 5-fold cross-validation of a regression that uses the 10 scores as inputs.
Find(a) What, if anything, is wrong with this procedure?
$$\text{fold }f:\quad \bar x^{(-f)},\ \sigma^{(-f)},\ U^{(-f)}\ \text{from the other four folds only}$$
The held-out fold is then transformed with quantities it did not help to estimate.
Answer $$\boxed{\text{fit the standardization and the PCA inside each fold}}$$
Check
Apply the rule a future sample must obey: a new sample never contributes to the means, the spreads or the components it is transformed with, so a held-out sample must not either.
Unsupervised steps are still fitted steps, and anything fitted belongs inside the cross-validation loop.
3§09.7 — one sensor, one peaked source, a climb
One source with density $p(s)=\frac12e^{-\lvert s\rvert}$ is recorded by one sensor as $x_t=a\,s_t$. The recordings are $3,\ -1,\ 0.5,\ -1.5$, and we estimate $w=1/a>0$.
Find
(a) Find $\hat w$ in closed form.
(b) Starting from $w=1$, take two gradient ascent steps $w\leftarrow w+\eta\,l'(w)$.
Given
$x=(3,\ -1,\ 0.5,\ -1.5)$
$l(w)=-4\log2-w\sum_t\lvert x_t\rvert+4\log w$ for $w>0$
learning rate $\eta=0.1$
Hint 1/4
One parameter: set the derivative to zero for (a), and follow it uphill for (b).
Hint 2/4
$l'(w)=-\sum_t\lvert x_t\rvert+\frac4w$.
Hint 3/4
Here $\sum_t\lvert x_t\rvert=3+1+0.5+1.5=6$, so $l'(w)=-6+\frac4w$, with $\eta=0.1$ and $w=1$ to start.
Hint 4/4
$\hat w=\frac23$; the steps give $1\to0.8\to0.7$.
Show solution
The closed form gives the target; the two steps show how an iterative method approaches it.
The exponent splits into a $z_1$ term and a $z_2$ term.
Otherwise
$$z_1^2+z_2^2=1\ \text{on the unit circle}:\quad \operatorname{Cov}=0,\ \text{yet dependent}$$
A nonlinear tie that covariance cannot see.
Answer $$\boxed{\text{when the scores are jointly Gaussian}}$$
Check
The two speakers who take turns give PCA scores that are uncorrelated, yet only the true unmixing matrix returns outputs that are zero whenever a speaker is silent.
PCA decorrelates; ICA goes further and looks for independence, which needs non-Gaussian sources to be visible.
Mistake ledger (16 entries)
⚠ Using weights that are not a unit vector
The plain sum, or an eigenvector left unscaled, looks like a direction.
wrong$$\operatorname{Var}(x_1+x_2)=23.6\ \text{ as the spread along the diagonal}$$
Blind source separation that assumes independent, non-Gaussian sources and estimates the unmixing matrix, here by maximum likelihood.
permutation matrixpermütasyon matrisi
A square matrix with exactly one 1 in each row and each column and zeros elsewhere; it reorders the entries of a vector and is orthogonal.
orthogonal matrixortogonal matris
A square matrix $Q$ with $Q^TQ=QQ^T=I$; it preserves lengths and angles, and in two dimensions it is a rotation or a reflection.
identifiabilitytanımlanabilirlik
The property that different parameter values give different distributions of the data, so that enough data can tell them apart.
determinantdeterminant
A number attached to a square matrix; its absolute value is the factor by which the matrix scales areas and volumes, and $\det(A^{-1})=1/\det A$.
What comes next
§10 · Clustering: mixture of Gaussians and EM
Here unlabelled data were summarized by directions and pulled apart into hidden sources. Next they are split into groups: each point comes from one of a few Gaussian clusters, and the unknown group of every point has to be estimated together with the clusters themselves.
Sources
textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The weekly line names no chapter of the book, so no section numbers are cited here.
course materialEEE 485 lecture slides and lecture notes, chapter 9: PCA, ICA, blind source separation Scope, order of topics and notation (the centered data matrix X, the covariance matrix Σ with 1/n, the unit vector u, the multiplier λ, the components u_m and eigenvalues λ_m, the scores z_i, the reconstructions, PVE, the deflated data Y, x(t), s(t), a_ij, A, W and its rows w_j, P, Q, p_S, p_X, L(W) and l(W)) follow these materials; all explanations, data, figures and exercises here are original.
course materialEEE 485 syllabus page on STARS, Fall 2026-27, printed 21 September 2026 Assessment weights, the weekly topic list and the recommended books.
textbookG. James, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning, Springer, 2013 Recommended in the syllabus; the lecture's arrest-rate example and its PCA figures come from it.
textbookK. P. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press, 2012 Recommended in the syllabus; the lecture's ICA example figure comes from it.
standard resultEigen-decomposition of symmetric matrices, Lagrange multipliers and the change of variables for densities Standard results used in the derivations. Every number and figure on this page was computed for it.