← back to EEE 485
Week 9119 min full read
7 concepts24 worked examples30 exercises4 exam-level7 figures
What are you here for?

09 PCA, ICA and blind source separation: directions of largest variance, unmixing and the ICA likelihood

Start with this

One question before you read anything. Getting it wrong is the point: it shows you what this section is for.

§09.1 — two least squares lines through the same points

Five students' centered quiz scores $(x_1,x_2)$ are $(-3,-4)$, $(-1,-3)$, $(-1,2)$, $(1,3)$ and $(4,2)$. Regressing quiz 2 on quiz 1 through the origin gives the line $x_2=0.857\,x_1$.

Find(a) Now regress quiz 1 on quiz 2 and draw that line in the same picture, with quiz 2 on the vertical axis. What is its slope in that picture?
Given
  • $\sum x_1^2=28$, $\sum x_2^2=42$, $\sum x_1x_2=24$

  • least squares through the origin: slope $=\frac{\sum(\text{input})(\text{response})}{\sum(\text{input})^2}$

Hint 1/4

Decide which quiz plays the response in the second fit, then rewrite that line with quiz 2 as a function of quiz 1.

Hint 2/4

Quiz 1 on quiz 2: $x_1=b\,x_2$ with $b=\frac{\sum x_1x_2}{\sum x_2^2}$; in the picture the slope is $1/b$.

Hint 3/4

Here $\sum x_1x_2=24$ and $\sum x_2^2=42$, so $b=\frac{24}{42}$.

Hint 4/4

$b=0.571$ and the slope in the picture is $\frac{42}{24}=1.75$, not $0.857$.

Show solution

Only the roles change, so the same one-coefficient formula applies with the roles swapped.

Fit with quiz 1 as the response

$$b=\frac{\sum x_1x_2}{\sum x_2^2}=\frac{24}{42}=0.571$$

The input is now quiz 2, so its sum of squares goes in the denominator.

Redraw

$$x_1=0.571\,x_2\ \iff\ x_2=\frac{42}{24}\,x_1=1.75\,x_1$$

Solve for $x_2$ to put quiz 2 on the vertical axis.

Answer $$\boxed{\text{slope }1.75\ne0.857}$$
Check

If the two fits gave one line, the product of the two regression slopes, $\frac{24}{28}\cdot\frac{24}{42}=0.490$, would equal $1$; it equals the squared correlation instead.

Least squares is not symmetric in its two variables; the summary line this section builds will be.

Five students took two quizzes, and a committee wants one number per student that spreads them out as much as possible. With weights of length 1, quiz 2 alone keeps $8.4$ of the $14$ units of variance and equal weights keep $11.8$. The weights $0.6$ and $0.8$ keep $12$, and no other pair of length 1 keeps more.

By the end you can find that weighting for any number of features, say what share of the variance $k$ components keep, and write and compare the likelihood that pulls mixed signals apart.

In 60 seconds

PCA describes centered data by the eigenvectors of $\Sigma=\frac1nX^TX$ with the largest eigenvalues, each keeping variance $\lambda_m$; ICA models recordings as $x=As$ with independent, non-Gaussian sources and estimates $W=A^{-1}$ by maximum likelihood, up to order, scale and sign.

$$\Sigma u_m=\lambda_mu_m,\qquad z_{im}=x_i^Tu_m,\qquad \tfrac1n\textstyle\sum_iz_{im}^2=\lambda_m$$

finding directions, scores and the variance each direction keeps

Proportion of variance explained
$$\mathrm{PVE}(\text{first }k)=\frac{\lambda_1+\dots+\lambda_k}{\lambda_1+\dots+\lambda_p}$$

choosing $k$, or saying what $k$ components keep

Mixing model
$$x=As,\qquad W=A^{-1},\qquad \hat s=Wx\ \ \text{up to order, scale and sign}$$

unmixing recordings, or judging whether an output is acceptable

ICA log-likelihood
$$l(W)=\sum_{t=1}^T\Big(\sum_{j=1}^p\log p_{S_j}\big(w_j^Tx(t)\big)+\log\lvert\det W\rvert\Big)$$

comparing candidate unmixing matrices, or deriving the estimate $\hat W$

Three most common mistakes
  1. Running PCA on uncentered data, or on raw features in arbitrary units: the components then follow the means or the units, not the structure.

  2. Taking the eigenvector of the wrong eigenvalue, or leaving it unnormalized: $u_1$ belongs to $\lambda_1$ and has length 1.

  3. Dropping $\log\lvert\det W\rvert$ from the ICA log-likelihood, or adding it once instead of $T$ times.

Two course documents give different weights:

  • Chapter 1 slides, undergraduate line: midterm 25, final 25, four quizzes 20, two-phase project 30.
  • STARS syllabus page printed on 21 September 2026: midterm 30, final 30, problem sets and quizzes 20, project 20. It is the later document; confirm which split applies.
  • The STARS weekly list splits this chapter: blind signal separation is its seventh item, and with feature selection its ninth; the lecture slides make PCA, ICA and blind source separation chapter 9.
How much time do you have?
10 minutes

The eigenvector recipe with a worked 2 by 2 case, the algorithm with scores and , and the mixing model with its three ambiguities.

The 60-second card · The best direction is the top eigenvector of Σ · The PCA algorithm · Blind source separation · Formula card
45 minutes

Every result once with a worked example, then PCA by hand from a full solution down to a bare problem.

The 60-second card · Projecting onto a direction · The best direction is the top eigenvector of Σ · The PCA algorithm · How many components · Blind source separation · Why the sources must not be Gaussian · ICA by maximum likelihood · Scaffolding comes off · Formula card
full read

Adds the derivations, the look-alike pairs, a full exam-style question and mixed practice where you pick the method yourself.

The opening pages · Recall first · Projecting onto a direction · The best direction is the top eigenvector of Σ · The PCA algorithm · How many components · Blind source separation · Why the sources must not be Gaussian · ICA by maximum likelihood · Method boxes · Look-alike pairs · Scaffolding comes off · Full exam-style question · Practice set · Check yourself
By the end of this section
  1. Compute the variance of the scores along a unit vector as $u^T\Sigma u$, and find the best direction for two features by turning an angle.

  2. Derive with a that principal components are eigenvectors of $\Sigma$, compute them for 2 by 2 and block 3 by 3 matrices, and obtain the next component by .

  3. Apply the PCA algorithm: center, standardize when units differ, compute scores and reconstructions, and read the loadings.

  4. Choose the number of components from the proportion of variance explained and the scree plot, and relate the discarded eigenvalues to the reconstruction error.

  5. Model blind source separation as $x=As$, unmix with $W=A^{-1}$, and tell an acceptable output (order, scale, sign) from a failed one.

  6. Explain why Gaussian sources cannot be unmixed, by showing that $A$ and $AQ$ give the same distribution of $x$.

  7. Evaluate the density $p_X(x)=\prod_jp_{S_j}(w_j^Tx)\lvert\det W\rvert$ and the ICA log-likelihood $l(W)$, and compare candidate unmixing matrices.

Syllabus coverage

PCA — covered

  • centering and scaling the columns
  • projections and the variance of the scores
  • the first component by a Lagrange multiplier
  • the top eigenvectors of the sample covariance matrix
  • the algorithm, scores, loadings and reconstruction
  • standardization
  • the proportion of variance explained and the choice of k

Spread over four blocks, in the lecture's order. The deflation argument for the second component, set as an exercise in the lecture notes, is worked out in the eigenvector block.

ICA — covered

  • the
  • the order, scale and sign ambiguities
  • the need for non-Gaussian sources
  • the density of a linear transformation
  • the likelihood and log-likelihood of the unmixing matrix
  • its maximum likelihood estimate

The ambiguities sit in the two blocks before the likelihood.

blind source separation — covered

The : T observations of p sensors, each a weighted sum of p unobserved sources with unknown coefficients, and the goal of estimating the source signals. The lecture lists image denoising, medical signal processing, brain-computer interfaces and time series analysis as uses.

The applications are named, not developed.

Blind signal separation — covered

The official weekly line's name for the same problem.

It is the seventh item of the STARS weekly list; the lecture teaches it inside chapter 9.

Feature extraction — covered

as new features: k numbers per point in place of p, used on their own or as a pre-processing step before a supervised method.

The ninth item of the STARS weekly list pairs it with feature selection; in this chapter the lecture's feature extraction method is PCA.

feature selection — deferred

Choosing a subset of the original features.

The lecturer teaches it as a chapter of its own, 'Feature selection', and a later section of these notes covers it. PCA builds new features out of all the old ones; it selects none of them.

Unsupervised learning — covered

Training data without labels, used on their own or as a pre-processing step before a supervised method.

The chapter's opening slide; the first block starts from it.

Reconstruction error and the discarded eigenvalues — off syllabus

$\frac1n\sum_i\lVert x_i-\hat x_i\rVert^2=\lambda_{k+1}+\dots+\lambda_p$, so keeping the most variance and making the smallest squared reconstruction error pick the same components.

Further reading: the lecture asks whether a point can be recovered exactly from its scores; this identity says how much is lost when it cannot. Used in one worked example, one practice question and the check of the exam example.

Gradient of the ICA log-likelihood — off syllabus

$\nabla_Wl=\sum_tg(Wx(t))\,x(t)^T+T(W^{-1})^T$, the direction in which an implementation climbs $l(W)$.

Further reading: the lecture stops at $\hat W=\arg\max l(W)$; the gradient is what an implementation climbs.

At most one Gaussian source — off syllabus

The rotation argument needs every source Gaussian; separation stays possible when a single source is Gaussian.

Further reading, stated without proof; no exercise depends on it.

Recall first
Covariance matrix and weighted sums

For a random vector $X$ with mean $\mu$, $\Sigma=E[(X-\mu)(X-\mu)^T]$ is symmetric and positive semidefinite, and $\operatorname{Var}(a^TX)=a^T\Sigma a$. For $n$ centered data points the sample version is $\frac1n\sum_ix_ix_i^T$.

The variance of every projection in this section is this quadratic form.

Eigenvalues and eigenvectors of a symmetric matrix

$\Sigma u=\lambda u$ with $u\ne0$. A real symmetric $p\times p$ matrix has $p$ real eigenvalues and eigenvectors $u_1,\dots,u_p$, and $\Sigma=\sum_j\lambda_ju_ju_j^T$. For $2\times2$, $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$; in general the is the sum of the eigenvalues.

Principal components are these eigenvectors, and the eigenvalues are the variances they keep.

Lagrange multipliers

To optimize $f(u)$ subject to $g(u)=0$, look for stationary points of $L(u,\lambda)=f(u)-\lambda g(u)$. For symmetric $\Sigma$, $\nabla_u(u^T\Sigma u)=2\Sigma u$ and $\nabla_u(u^Tu)=2u$.

The lecture derives the first component this way.

Least squares through the origin

For centered data, the line $y=bx$ with $b=\frac{\sum_ix_iy_i}{\sum_ix_i^2}$ minimizes the vertical misses $\sum_i(y_i-bx_i)^2$; with several predictors, $Z^TZ\beta=Z^Ty$.

The first failure of the section compares it with the first principal component, and one practice question regresses on scores.

Gaussian vectors and independence

A Gaussian vector is fixed by its mean and covariance, and if $s$ is Gaussian with mean $\mu$ and covariance $C$, then $As$ is Gaussian with mean $A\mu$ and covariance $ACA^T$. Independent variables have a joint density equal to the product of their densities; jointly Gaussian variables that are uncorrelated are independent.

The argument against Gaussian sources and the ICA likelihood both rest on these facts.

and orthogonal matrices

$\det(AB)=\det A\det B$ and $\det(A^{-1})=1/\det A$; $\lvert\det A\rvert$ is the factor by which $A$ scales areas and volumes. An orthogonal $Q$ has $Q^TQ=QQ^T=I$ and $\lvert\det Q\rvert=1$. $\begin{bmatrix}a&b\\c&d\end{bmatrix}^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix}$.

Unmixing, the density of a mixture and the rotation argument all use them.

Maximum likelihood

$\hat\theta=\arg\max_\theta\log p(D\mid\theta)$; for independent observations the log-likelihood is a sum over them, and taking the log does not move the maximizer.

ICA estimates the unmixing matrix this way.

Try it yourself first (2 questions)
1§09.1 — the variance of a rescaled sum

Five students' two quiz deviations $x_1$ and $x_2$ add up to a sum $x_1+x_2$ whose variance is $23.6$.

Find(a) What is the variance of the average $\frac12(x_1+x_2)$?
Given
  • $\operatorname{Var}(x_1+x_2)=23.6$

  • $\operatorname{Var}(cY)=c^2\operatorname{Var}(Y)$

Hint 1/4

The average is the sum multiplied by a constant; ask how a constant factor acts on a variance.

Hint 2/4

$\operatorname{Var}(cY)=c^2\operatorname{Var}(Y)$.

Hint 3/4

Here $Y=x_1+x_2$ with variance $23.6$, and $c=\frac12$.

Hint 4/4

$\frac14\cdot23.6=5.9$.

Show solution

A constant factor comes out of a variance squared, so no data are needed.

Pull out the constant

$$\operatorname{Var}\big(\tfrac12(x_1+x_2)\big)=\tfrac14\cdot23.6=5.9$$

$c^2=\frac14$ for $c=\frac12$.

Answer $$\boxed{5.9}$$
Check

Directly from the covariance matrix: $\frac14(5.6+2\cdot4.8+8.4)=\frac{23.6}4=5.9$.

Weights can inflate or shrink a variance at will, so comparing directions only makes sense with weights of a fixed length.

2§09.2 — recognizing an eigenvector

The five students' covariance matrix is $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$.

Find(a) Which vector is an eigenvector of $\Sigma$, and with which eigenvalue?
Given
  • $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$

  • $v$ is an eigenvector when $\Sigma v=\lambda v$ for some number $\lambda$

Hint 1/4

Multiply each candidate by the matrix and see whether the result points the same way.

Hint 2/4

$v$ is an eigenvector when $\Sigma v=\lambda v$; then $\lambda$ is the common ratio of the entries.

Hint 3/4

Here $\Sigma(3,4)^T=(5.6\cdot3+4.8\cdot4,\ 4.8\cdot3+8.4\cdot4)$ and $\Sigma(4,3)^T=(5.6\cdot4+4.8\cdot3,\ 4.8\cdot4+8.4\cdot3)$.

Hint 4/4

$\Sigma(3,4)^T=(36,48)=12\,(3,4)^T$: eigenvector $(3,4)$ with eigenvalue $12$.

Show solution

Multiplying is faster than solving the characteristic polynomial when candidates are given.

Multiply

$$\Sigma(3,4)^T=(16.8+19.2,\ 14.4+33.6)=(36,\ 48)=12\,(3,4)^T$$

Both entries are 12 times the input.

$$\Sigma(4,3)^T=(36.8,\ 44.4),\qquad \frac{36.8}4\ne\frac{44.4}3$$

The ratios differ, so the direction changed.

Answer $$\boxed{v=(3,4),\ \ \lambda=12}$$
Check

The other eigenvalue must be $\operatorname{tr}\Sigma-12=2$; indeed $\Sigma(4,-3)^T=(8,-6)=2\,(4,-3)^T$.

This section will show that the eigenvector with the largest eigenvalue is the most informative direction in the data.

Notation
symbolreads asmeanswatch out
$x_i=[x_{i1},\dots,x_{ip}]^T,\ \ X$

point i, the data matrix

$X$ is $n\times p$ with rows $x_i^T$, centered unless stated otherwise

In the ICA blocks $X$ and $S$ are random vectors instead, as in the lecture.

$\bar x_j,\ \ \sigma_j$

x bar j, sigma j

the mean of column $j$; its standard deviation, with $1/n$

Divide by $\sigma_j$ only after centering.

$u,\ \ \lVert u\rVert=1$

a unit vector

a direction in $\mathbb R^p$

$u$ and $-u$ give the same line and the same variance.

$\Sigma=\tfrac1nX^TX$

Sigma

the sample covariance matrix of the centered data

$1/n$ as in the lecture; $1/(n-1)$ scales every eigenvalue and leaves the eigenvectors unchanged.

$\lambda_m,\ \ u_m$

lambda m, u m

the $m$-th largest eigenvalue and its unit eigenvector, the $m$-th principal component

Here $\lambda$ is an eigenvalue and the Lagrange multiplier, not a penalty weight.

$z_{im}=x_i^Tu_m,\ \ z_i$

the score of point i on component m

$z_i=[z_{i1},\dots,z_{ik}]^T$ holds point $i$'s $k$ scores

A score is a number; the projection $(x_i^Tu_m)\,u_m$ is a vector.

$\hat x_i$

x hat i

the reconstruction $z_{i1}u_1+\dots+z_{ik}u_k$

Add the means back to return to raw units.

$\mathrm{PVE}(m)$

proportion of variance explained

$\lambda_m/\sum_j\lambda_j$

A number between 0 and 1.

$Y=X-Xu_1u_1^T$

the deflated data

what is left after the first component is removed

$\frac1nY^TY=\Sigma-\lambda_1u_1u_1^T$.

$s(t),\ \ x(t),\ \ t=1,\dots,T$

sources and recordings at time t

the hidden and the observed vectors

Independent over t in the ICA model.

$A=[a_{ij}]$

the mixing matrix

$p\times p$, unknown and invertible; $a_{ij}$ is the weight of source $j$ in sensor $i$

Square: as many sensors as sources.

$W=A^{-1},\ \ w_j^T$

the unmixing matrix, its j-th row

$\hat s_j=w_j^Tx$

Recovered only up to order, scale and sign.

$P,\ \ D$

a , a diagonal matrix

one 1 in each row and column; nonzero diagonal entries

$D$ with entries $\pm1$ gives the sign flips.

$Q$

an

$QQ^T=Q^TQ=I$; in 2D a rotation

$AQ$ mixes Gaussian sources exactly like $A$.

$p_S,\ \ p_{S_j},\ \ p_X$

densities

of the source vector, of source j, of the recordings

The source densities are assumed known.

$L(W),\ \ l(W),\ \ \hat W$

likelihood, log-likelihood, estimate

$\prod_tp_X(x(t))$, its natural log, and its maximizer

Maximized, not minimized.

Conventions used here
Centered data.

$X$ is centered before any PCA formula is applied, and $\Sigma=\frac1nX^TX$ uses $1/n$, as in the lecture.

With 1/(n − 1) every eigenvalue scales by the same factor; directions and PVE do not change.

Unit length and sign.

Principal components are unit vectors. Since $u$ and $-u$ are equally valid, this page takes the first nonzero entry positive.

Two correct answers can differ by a sign; the convention makes answers comparable.

Order of the components.

Eigenvalues are sorted, $\lambda_1\ge\lambda_2\ge\dots\ge\lambda_p$, and component $m$ belongs to $\lambda_m$.

First means largest variance, not first feature.

What λ means here.

$\lambda$ is the Lagrange multiplier and, at the solution, the eigenvalue; the two coincide.

In earlier sections the same letter was a penalty weight.

Standardize when units differ.

If the features are measured in different or arbitrary units, divide each centered column by its $\sigma_j$ first; if they share a unit, center only.

This is the lecture's rule, and the answer otherwise depends on the units.

PVE as a proportion.

$\mathrm{PVE}$ is reported as a number between 0 and 1, and targets such as 'at least 0.85' use $\ge$.

Meeting a target exactly counts.

Square mixing.

ICA here has as many sensors as sources, $A$ is $p\times p$ and invertible, and $W=A^{-1}$.

This is the lecture's model.

An acceptable ICA output.

An output counts as correct when it equals $DPs$: each output one source, in any order, times any nonzero constant.

Order, scale and sign cannot be recovered from the recordings.

Logarithms.

$\log$ is the natural logarithm, in every likelihood on this page.

Another base would only rescale the log-likelihood.

Rounding.

Intermediate values keep at least four significant figures; answers are given to three decimals and angles to one decimal of a degree.

Eigenvalues and angles computed from rounded inputs move in the last digit.

9.1Projecting onto a direction: the spread one number keeps

Turns each point into one number, its score $x^Tu$ along a unit vector $u$, and measures the spread those scores keep: $u^T\Sigma u$.

Every model so far learned from labelled pairs $(x_i,y_i)$; now the training data are the $x_i$ alone, and the first job is to describe them with fewer numbers.

Solvable with what we have
  • Compute each quiz's variance and the covariance: $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$, total variance $5.6+8.4=14$.

  • Compute the variance of any fixed weighting, for example the average: $\tfrac14(5.6+2\cdot4.8+8.4)=5.9$.

  • Fit a least squares line that predicts quiz 2 from quiz 1: slope $24/28=0.857$.

Not solvable yet
  • Choose the weights that spread the five students out the most when neither quiz is a response.

  • Say how much of the total 14 a single number keeps, and how much it loses.

  • Do the same for four features, or four hundred.

Use least squares anyway. Quiz 2 on quiz 1 gives slope $\textcolor{#8250df}{0.857}$; quiz 1 on quiz 2, drawn in the same picture, gives slope $\textcolor{#8250df}{1.75}$. One data set, two answers. The direction built in this block has slope $\textcolor{#1f6feb}{4/3}$, and its scores keep 12 of the 14 units of variance.

Why it fails

Least squares counts misses along the response axis only, so its answer depends on which quiz we call the response. Here there is no response: the spread has to be measured in a way that treats both quizzes alike.

DefinitionDefinition 9.1: Score, projection and variance along a direction
Conditions
  • the $n$ points $x_1,\dots,x_n\in\mathbb R^p$ are centered: each feature has mean $0$

  • $u\in\mathbb R^p$ is a unit vector, $\lVert u\rVert=1$, that is $u^Tu=1$

  • $\Sigma=\frac1n\sum_{i=1}^nx_ix_i^T=\frac1nX^TX$ is the sample covariance matrix, with $1/n$ as in the lecture

$$\boxed{\begin{aligned}z_i&=x_i^T\textcolor{#1f6feb}{u}\ \ \text{(score)},\qquad (x_i^T\textcolor{#1f6feb}{u})\,\textcolor{#1f6feb}{u}\ \ \text{(projection)}\\ \frac1n\sum_{i=1}^nz_i^2&=\textcolor{#1f6feb}{u}^T\Sigma\,\textcolor{#1f6feb}{u}\\ \text{first principal component:}&\quad \max_{u^Tu=1}\ u^T\Sigma u\end{aligned}}$$

Drop each point perpendicularly onto the line through the origin in direction $u$; the signed distance from the origin to the foot is the score. Centered data give scores with mean zero, so their mean square is their variance, and it equals $u^T\Sigma u$. The first principal component is the unit direction that makes this number largest.

Proof

The foot is at $x^Tu$. The closest point $cu$ on the line minimizes $\lVert x-cu\rVert^2=\lVert x\rVert^2-2c\,x^Tu+c^2$; the derivative $-2x^Tu+2c$ vanishes at $c=x^Tu$.

Scores have mean zero. $\frac1n\sum_ix_i^Tu=\big(\frac1n\sum_ix_i\big)^Tu=0^Tu=0$.

Their mean square is a quadratic form. $(x_i^Tu)^2=u^Tx_i\,x_i^Tu$, so $\frac1n\sum_i(x_i^Tu)^2=u^T\big(\frac1n\sum_ix_ix_i^T\big)u=u^T\Sigma u$.

Why the unit length. For $a=cu$ the same formula gives $c^2\,u^T\Sigma u$: without $\lVert u\rVert=1$ the variance grows without bound and has no maximum.

Looks like this, but is not

The plain sum $x_1+x_2$, weights $a=(1,1)$, spreads the students with variance $a^T\Sigma a=23.6$, more than the total $14$ of the two quizzes.

$a$ is not a unit vector: $\lVert a\rVert=\sqrt2$, and doubling any direction multiplies $a^T\Sigma a$ by four. As a unit vector, $(1,1)/\sqrt2$ keeps $11.8$, below the best $12$.

angle θunit vector uvariance of the scores

$0^\circ$

$(1,\ 0)$

$5.6$

$30^\circ$

$(0.866,\ 0.5)$

$10.457$

$53.1^\circ$

$(0.6,\ 0.8)$

$12$

$90^\circ$

$(0,\ 1)$

$8.4$

$143.1^\circ$

$(-0.8,\ 0.6)$

$2$

The spread rises to $12$ at $53.1^\circ$ and falls to $2$ exactly $90^\circ$ later; the two axes keep only the single-quiz variances $5.6$ and $8.4$.

The best single weighting for the five students, found by turning a unit vector

Five students' quiz scores, out of 20, are $(9,10)$, $(11,11)$, $(11,16)$, $(13,17)$, $(16,16)$. Find the unit vector $u$ whose scores $x_i^Tu$ have the largest variance, and that variance.

FindThe best $u=(\cos\theta,\sin\theta)$ and the variance $u^T\Sigma u$ it keeps.
Given
  • raw scores $(9,10),\ \allowbreak (11,11),\ \allowbreak (11,16),\ \allowbreak (13,17),\ \allowbreak (16,16)$ with means $(12,14)$

  • centered: $(-3,-4),\ \allowbreak (-1,-3),\ \allowbreak (-1,2),\ \allowbreak (1,3),\ \allowbreak (4,2)$

Solution

With two features every unit vector is $(\cos\theta,\sin\theta)$, so the problem becomes maximizing a function of one angle; no new theory is needed yet.

Build Σ from the centered scores

$$\sum_ix_{i1}^2=28,\qquad \sum_ix_{i2}^2=42,\qquad \sum_ix_{i1}x_{i2}=24$$

Each sum runs over the five centered students, for example $9+1+1+1+16=28$.

$$\Sigma=\frac15\begin{bmatrix}28&24\\24&42\end{bmatrix}=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$$

Divide by $n=5$, the lecture's $\frac1nX^TX$.

Write the variance as a function of the angle

$$u^T\Sigma u=5.6\cos^2\theta+9.6\sin\theta\cos\theta+8.4\sin^2\theta$$

The cross term appears twice, $2\cdot4.8=9.6$, because $\Sigma_{12}=\Sigma_{21}$.

$$=7-1.4\cos2\theta+4.8\sin2\theta$$

Double angles: $\cos^2\theta=\frac{1+\cos2\theta}2$, $\sin^2\theta=\frac{1-\cos2\theta}2$, $2\sin\theta\cos\theta=\sin2\theta$.

Maximize over the angle

$$-1.4\cos2\theta+4.8\sin2\theta=5\cos(2\theta-106.26^\circ)$$

Amplitude $\sqrt{1.4^2+4.8^2}=\sqrt{25}=5$; the phase $\varphi$ has $\cos\varphi=-0.28$ and $\sin\varphi=0.96$.

$$2\theta=106.26^\circ\ \Rightarrow\ \theta=53.13^\circ,\qquad \tan\theta=\tfrac43$$

The cosine is largest, equal to 1, where its argument is 0.

$$u=(\cos53.13^\circ,\ \sin53.13^\circ)=(0.6,\ 0.8),\qquad u^T\Sigma u=7+5=12$$

$\tan\theta=4/3$ is the 3-4-5 triangle, so the cosine is $3/5$ and the sine $4/5$.

Answer $$\boxed{u=\textcolor{#1f6feb}{(0.6,\ 0.8)},\qquad u^T\Sigma u=12\ \text{ of the total }14}$$
Check

Compute the scores directly: $0.6x_1+0.8x_2$ gives $-5,-3,1,3,4$, with mean $0$ and mean square $\frac{25+9+1+9+16}5=12$. The smallest value, $7-5=2$, sits at $\theta=143.13^\circ$, perpendicular to $u$.

That is the opening's weighting, $0.6\times$ quiz 1 $+\ 0.8\times$ quiz 2, keeping 12 of the 14. With more features the unit vectors form a sphere, and the next block replaces the angle by an eigenvector equation.

How much spread the average, quiz 2 alone and the two least squares lines keep

For the same five students compare four candidate directions: the average $(1,1)/\sqrt2$, quiz 2 alone $(0,1)$, and the two least squares lines, of slopes $6/7$ and $7/4$.

Find$u^T\Sigma u$ for each candidate, against the best value $12$.
Given
  • $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$

  • least squares slopes $24/28=6/7$ (quiz 2 on quiz 1) and $42/24=7/4$ (quiz 1 on quiz 2, redrawn)

Solution

Each candidate is a fixed direction, so one formula, $u^T\Sigma u$ with $u$ scaled to length 1, answers all four.

The average and quiz 2 alone

$$u=\tfrac1{\sqrt2}(1,1):\quad \tfrac12(5.6+2\cdot4.8+8.4)=11.8$$

Both entries are $1/\sqrt2$, so every term of $u^T\Sigma u$ carries a factor $\frac12$.

$$u=(0,1):\quad u^T\Sigma u=\Sigma_{22}=8.4$$

A coordinate axis keeps exactly that feature's variance.

The two least squares lines

$$u=\tfrac{(7,6)}{\sqrt{85}}:\quad \tfrac{5.6\cdot49+9.6\cdot42+8.4\cdot36}{85}=\tfrac{980}{85}=11.529$$

A line of slope $6/7$ has direction $(7,6)$; divide by its length $\sqrt{85}$.

$$u=\tfrac{(4,7)}{\sqrt{65}}:\quad \tfrac{5.6\cdot16+9.6\cdot28+8.4\cdot49}{65}=\tfrac{770}{65}=11.846$$

Slope $7/4$ means direction $(4,7)$, of length $\sqrt{65}$.

Answer $$\boxed{11.8,\quad 8.4,\quad 11.529,\quad 11.846\ \ <\ \ 12}$$
Check

Every candidate stays below $12$ and above $2$, the smallest possible value found in the first example; the two least squares lines straddle the best angle, $40.6^\circ<53.1^\circ<60.3^\circ$.

Directions near the best one keep almost as much. What sets the best one apart is that the data alone define it, with no response chosen.

Checkpoint
§09.1 — the spread along a swapped direction

A classmate says that $u=(0.8,0.6)$ must keep the same variance as $u=(0.6,0.8)$, since it uses the same two numbers.

Find(a) Compute $u^T\Sigma u$ for $u=(0.8,0.6)$.
Given
  • $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$ for the five students

  • $u=(0.8,0.6)$, a unit vector

Hint 1/4

The variance along a unit vector is a quadratic form in its entries; write out all four terms.

Hint 2/4

$u^T\Sigma u=\Sigma_{11}u_1^2+2\Sigma_{12}u_1u_2+\Sigma_{22}u_2^2$.

Hint 3/4

Here $u=(0.8,0.6)$, $\Sigma_{11}=5.6$, $\Sigma_{12}=4.8$, $\Sigma_{22}=8.4$: $5.6\cdot0.64+2\cdot4.8\cdot0.48+8.4\cdot0.36$.

Hint 4/4

The three terms add to $11.216$, below the $12$ of $(0.6,0.8)$.

Show solution

Term by term is the fastest route for a 2 by 2 matrix.

Three terms

$$5.6\,(0.8)^2+2\,(4.8)(0.8)(0.6)+8.4\,(0.6)^2$$

Squares on the diagonal terms, the cross term counted twice.

$$=3.584+4.608+3.024=11.216$$

These three terms are the whole quadratic form, since the cross term was already doubled.

Answer $$\boxed{u^T\Sigma u=11.216}$$
Check

The angle formula agrees: $u=(0.8,0.6)$ is $\theta=36.87^\circ$, with $\cos2\theta=0.28$ and $\sin2\theta=0.96$, so $7-1.4\cdot0.28+4.8\cdot0.96=11.216$.

Which feature gets the larger weight matters as much as the weights themselves.

⚠ Using weights that are not a unit vector

The plain sum, or an eigenvector left unscaled, looks like a direction.

wrong$$\operatorname{Var}(x_1+x_2)=23.6\ \text{ as the spread along the diagonal}$$
right$$u=\tfrac1{\sqrt2}(1,1):\quad u^T\Sigma u=11.8$$
⚠ Projecting raw, uncentered scores

The formula $u^T\Sigma u$ is quoted without the centering it assumes.

wrong$$\tfrac15\sum_i(\tilde x_i^Tu_1)^2=350.56\ \ \text{(raw points }\tilde x_i)$$
right$$\tfrac15\sum_i\big((\tilde x_i-\bar x)^Tu_1\big)^2=12$$
−4−224−4−224x₁ (quiz 1)x₂ (quiz 2)θ = 0°u = (1, 0)scores along the line:−3, −1, −1, 1, 4mean square 5.6dashed pieces thrown away:mean square 8.4total 14 = kept + thrown away

Step the direction round the five students. The blue feet are the scores and the dashed orange pieces are what one number throws away; the two mean squares always add to $14$, and the kept part peaks at $\textcolor{#1f6feb}{12}$ at $53.1^\circ$.

At the edges
θ = 53.1° 12

The largest spread, along $u=(0.6,0.8)$.

θ = 143.1° 2

The smallest spread, perpendicular to the best direction; this is what one number loses.

9.2The best direction is the top eigenvector of Σ

Solves $\max u^T\Sigma u$ subject to $u^Tu=1$: the answer is the eigenvector of $\Sigma$ with the largest eigenvalue, and the maximum is that eigenvalue.

Turning an angle worked because $p=2$; with four features the unit vectors form a sphere, and we need an equation that singles out the best one.

TheoremTheorem 9.2: Principal components are eigenvectors of Σ
Conditions
  • $\Sigma$ is the $p\times p$ sample covariance matrix of centered data: symmetric and positive semidefinite

  • its eigenvalues are sorted, $\lambda_1\ge\lambda_2\ge\dots\ge\lambda_p\ge0$, with orthonormal eigenvectors $u_1,\dots,u_p$

  • if $\lambda_1=\lambda_2$ the first direction is not unique: every unit vector in the span of $u_1$ and $u_2$ is optimal

$$\boxed{\begin{aligned}&\max_{u^Tu=1}u^T\Sigma u:\qquad L(u,\lambda)=u^T\Sigma u-\lambda\,(u^Tu-1)\\ &\nabla_uL=0\ \Rightarrow\ \Sigma u=\lambda u,\qquad u^T\Sigma u=\lambda\\ &\text{maximum }\textcolor{#1f6feb}{\lambda_1}\text{ at }\textcolor{#1f6feb}{u_1};\qquad\text{the first }k\text{ components are }u_1,\dots,u_k\end{aligned}}$$

Setting the gradient of the Lagrangian to zero says the best $u$ must be an eigenvector of $\Sigma$, and at an eigenvector the variance equals its eigenvalue. So the winner is the eigenvector with the largest eigenvalue. The next component, the best direction perpendicular to it, is the eigenvector of the second largest eigenvalue, and so on.

Proof

The constraint. $\partial L/\partial\lambda=-(u^Tu-1)=0$ gives back $u^Tu=1$.

The stationary points. For symmetric $\Sigma$, $\nabla_u(u^T\Sigma u)=2\Sigma u$ and $\nabla_u(u^Tu)=2u$, so $\nabla_uL=2\Sigma u-2\lambda u=0$: $\Sigma u=\lambda u$.

Their values. At a unit eigenvector $u^T\Sigma u=u^T(\lambda u)=\lambda\,u^Tu=\lambda$, so the largest eigenvalue gives the largest candidate.

A global maximum, not just a stationary point. Expand any unit $u=\sum_jc_ju_j$ with $\sum_jc_j^2=1$. Then $u^T\Sigma u=\sum_j\lambda_jc_j^2\le\lambda_1\sum_jc_j^2=\lambda_1$, with equality at $u=u_1$.

The later components. Over unit vectors perpendicular to $u_1,\dots,u_{m-1}$ the first $m-1$ coefficients are $0$, and the same bound gives $\lambda_m$, reached at $u_m$.

Looks like this, but is not

Every eigenvector satisfies the Lagrange condition $\Sigma u=\lambda u$, so for the five students $u=(0.8,-0.6)$ is also a solution.

The condition finds every stationary point, and $(0.8,-0.6)$ is the minimum: its variance is $2$. Compare the eigenvalues before choosing.

Eigenvalues and eigenvectors of the five students' covariance matrix

Find the eigenvalues and unit eigenvectors of $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$ and name the first principal component.

Find$\lambda_1\ge\lambda_2$, $u_1$ and $u_2$.
Given$\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$
Solution

For a 2 by 2 matrix the characteristic polynomial is a quadratic in $\lambda$, and each eigenvector comes from one row of $(\Sigma-\lambda I)u=0$.

Eigenvalues

$$\det(\Sigma-\lambda I)=\lambda^2-(\operatorname{tr}\Sigma)\,\lambda+\det\Sigma=\lambda^2-14\lambda+24$$

Trace $5.6+8.4=14$; determinant $5.6\cdot8.4-4.8^2=47.04-23.04=24$.

$$\lambda=\frac{14\pm\sqrt{196-96}}2=\frac{14\pm10}2:\qquad \lambda_1=12,\ \ \lambda_2=2$$

Quadratic formula, then sort in decreasing order.

First eigenvector

$$(5.6-12)\,u_a+4.8\,u_b=0\ \Rightarrow\ u_b=\tfrac{6.4}{4.8}\,u_a=\tfrac43\,u_a$$

The first row of $(\Sigma-\lambda_1I)u=0$; the second row is a multiple of it.

$$u_1=\tfrac15(3,4)=(0.6,\ 0.8)$$

Normalize: $\lVert(3,4)\rVert=5$. The first entry is taken positive.

Second eigenvector

$$u_2=(0.8,\ -0.6),\qquad u_1^Tu_2=0.48-0.48=0$$

Eigenvectors of a symmetric matrix for different eigenvalues are perpendicular, so in two dimensions $u_2$ is $u_1$ turned by $90^\circ$.

Answer $$\boxed{\lambda_1=12,\ u_1=(0.6,\ 0.8);\qquad \lambda_2=2,\ u_2=(0.8,\ -0.6)}$$
Check

Multiply out: $\Sigma u_1=(5.6\cdot0.6+4.8\cdot0.8,\ 4.8\cdot0.6+8.4\cdot0.8)=(7.2,\ 9.6)=12\,(0.6,0.8)$. The angle $\arctan(4/3)=53.1^\circ$ matches the turning-angle example.

Trace and determinant give both eigenvalues of any 2 by 2 covariance matrix in one line, and the two always add up to the total variance.

A 3 by 3 covariance where the largest single variance does not win

Three sensors have covariance matrix $\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$. Sensor 3 has the largest variance. Is it the first principal component?

FindAll eigenvalues and eigenvectors, and the first component.
Given$\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$
Solution

Sensor 3 is uncorrelated with the other two, so $\Sigma$ splits into a 2 by 2 block and a 1 by 1 block that can be solved separately.

The block of sensors 1 and 2

$$\begin{bmatrix}3&2\\2&3\end{bmatrix}:\quad \lambda^2-6\lambda+5=0,\quad \lambda=5,\ 1$$

Trace $6$, determinant $9-4=5$.

$$\lambda=5:\ \tfrac1{\sqrt2}(1,1,0);\qquad \lambda=1:\ \tfrac1{\sqrt2}(1,-1,0)$$

Equal diagonal entries always give the directions $(1,1)$ and $(1,-1)$; a $0$ fills the place of sensor 3.

The block of sensor 3

$$\lambda=4:\quad (0,0,1)$$

Its row and column are zero off the diagonal, so sensor 3 alone is an eigenvector.

Sort

$$\lambda_1=5>\lambda_2=4>\lambda_3=1$$

The combined direction of the correlated pair beats sensor 3.

Answer $$\boxed{u_1=\tfrac1{\sqrt2}(1,1,0),\ \lambda_1=5;\quad u_2=(0,0,1),\ \lambda_2=4;\quad u_3=\tfrac1{\sqrt2}(1,-1,0),\ \lambda_3=1}$$
Check

The eigenvalues add to $5+4+1=10=\operatorname{tr}\Sigma=3+3+4$, and $\Sigma\,(1, \allowbreak 1, \allowbreak 0)^T=(5, \allowbreak 5, \allowbreak 0)^T$ confirms the first one.

Correlated features pool their variance, so a component can beat every single feature: read the eigenvalues, not the diagonal.

The second component by deflation, the route of the lecture notes

Remove the first component from the five students, $Y=X-Xu_1u_1^T$, and find the best direction for what is left.

Find$\frac1nY^TY$, its top eigenvector, and that eigenvector's variance.
Given
  • centered rows $(-3,-4),\ \allowbreak (-1,-3),\ \allowbreak (-1,2),\ \allowbreak (1,3),\ \allowbreak (4,2)$

  • $u_1=(0.6,0.8)$, $\lambda_1=12$, $\Sigma=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$

Solution

Working with $\frac1nY^TY=\Sigma-\lambda_1u_1u_1^T$ is shorter than building $Y$ row by row; the rows then serve as the check.

The deflated covariance

$$\tfrac1nY^TY=\tfrac1n(X-Xu_1u_1^T)^T(X-Xu_1u_1^T)=\Sigma-\lambda_1u_1u_1^T$$

Expand the product; $\frac1nX^TXu_1=\Sigma u_1=\lambda_1u_1$ and $u_1^Tu_1=1$ collapse the three cross terms.

$$=\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}-12\begin{bmatrix}0.36&0.48\\0.48&0.64\end{bmatrix}=\begin{bmatrix}1.28&-0.96\\-0.96&0.72\end{bmatrix}$$

$u_1u_1^T$ has entries $0.6^2$, $0.6\cdot0.8$ and $0.8^2$.

Its best direction

$$\begin{bmatrix}1.28&-0.96\\-0.96&0.72\end{bmatrix}=2\begin{bmatrix}0.64&-0.48\\-0.48&0.36\end{bmatrix}=2\,u_2u_2^T$$

Every entry is twice the matching entry of $u_2u_2^T$ with $u_2=(0.8,-0.6)$.

$$\tfrac1nY^TY\,u_2=2\,u_2,\qquad \tfrac1nY^TY\,u_1=0$$

The rank-one matrix $2u_2u_2^T$ sends $u_2$ to $2u_2$ and every vector perpendicular to $u_2$ to $0$.

Answer $$\boxed{\text{maximizer }u_2=(0.8,\ -0.6),\ \text{variance }\lambda_2=2;\ \ u_1\text{ now keeps }0}$$
Check

Row by row, $Y$ has rows $x_i-z_{i1}u_1$: $(0,0)$, $(0.8,-0.6)$, $(-1.6,1.2)$, $(-0.8,0.6)$, $(1.6,-1.2)$, with mean square length $\frac{0+1+4+1+4}5=2=\lambda_2$.

For two features this is a detour, since $u_2$ is just $u_1$ turned by $90^\circ$; its value is that it works the same way for any $p$, one component at a time.

Checkpoint
§09.2 — the first component of a diagonal covariance

Two uncorrelated features have variances $4$ and $9$, so their covariance matrix is diagonal.

Find(a) Give the first principal component and the variance its scores keep.
Given$\Sigma=\begin{bmatrix}4&0\\0&9\end{bmatrix}$
Hint 1/4

Decide which direction keeps more spread: a single feature or a direction in between.

Hint 2/4

The first component is the eigenvector of the largest eigenvalue, and its scores have variance equal to that eigenvalue.

Hint 3/4

Here $\Sigma=\begin{bmatrix}4&0\\0&9\end{bmatrix}$: $\Sigma(1,0)^T=(4,0)^T$ and $\Sigma(0,1)^T=(0,9)^T$.

Hint 4/4

The eigenvalues are $4$ and $9$; the first component is $(0,1)$, with variance $9$.

Show solution

A diagonal matrix shows its eigen-pairs without any computation.

Read off the eigen-pairs

$$\Sigma(1,0)^T=4\,(1,0)^T,\qquad \Sigma(0,1)^T=9\,(0,1)^T$$

A diagonal matrix scales each axis by its diagonal entry.

$$\lambda_1=9,\ u_1=(0,1);\qquad \lambda_2=4,\ u_2=(1,0)$$

Sort by eigenvalue, not by feature number.

Answer $$\boxed{u_1=(0,1),\ \ \lambda_1=9}$$
Check

Any other unit vector $(\cos\theta,\sin\theta)$ keeps $4\cos^2\theta+9\sin^2\theta\le9$, with equality only at $\theta=90^\circ$.

With uncorrelated features PCA only sorts them by variance; it rotates the axes only when features are correlated.

⚠ Taking the eigenvector of the smallest eigenvalue

The eigenvector found first, or printed first by some software, is assumed to be the best one.

wrong$$u_1=(0.8,-0.6)\ \Rightarrow\ \text{variance }2$$
right$$u_1=(0.6,0.8)\ \Rightarrow\ \text{variance }12$$
⚠ Leaving the eigenvector unnormalized

$(3,4)$ solves $(\Sigma-12I)u=0$ just as well, and the scaling step is skipped.

wrong$$z_{i1}=3x_{i1}+4x_{i2}:\ \text{scores five times too large}$$
right$$u_1=\tfrac15(3,4):\quad z_{i1}=0.6x_{i1}+0.8x_{i2}$$
⚠ Flipping the sign inside Σ − λI

$\lambda-\Sigma_{11}$ and $\Sigma_{11}-\lambda$ look alike in a line copied fast.

wrong$$(12-5.6)\,u_a+4.8\,u_b=0\ \Rightarrow\ u\propto(3,-4)$$
right$$(5.6-12)\,u_a+4.8\,u_b=0\ \Rightarrow\ u\propto(3,4)$$

9.3The PCA algorithm: prepare the columns, score, reconstruct

Turns raw data into $k$ new features per point: center, standardize if the units differ, eigendecompose $\Sigma$, and take $k$ dot products.

We know which directions to use; this block is the full recipe the lecture writes as an algorithm, starting from raw columns.

MethodMethod 9.3: The PCA algorithm
Conditions
  • Center every column: $x_{ij}\leftarrow x_{ij}-\bar x_j$, where $\bar x_j=\frac1n\sum_ix_{ij}$.

  • If the features are in different units, also divide by $\sigma_j$, where $\sigma_j^2=\frac1n\sum_ix_{ij}^2$ after centering; $\Sigma$ then has ones on its diagonal and is the .

  • Choose $k\le p$; the next block shows how.

$$\boxed{\begin{aligned}&1.\ \ \Sigma=\tfrac1nX^TX\qquad 2.\ \ \Sigma u_m=\lambda_mu_m,\ \ \lambda_1\ge\lambda_2\ge\dots\ge\lambda_p\\ &3.\ \ z_i=\big[\,x_i^Tu_1,\ \dots,\ x_i^Tu_k\,\big]^T\qquad 4.\ \ \hat x_i=z_{i1}u_1+z_{i2}u_2+\dots+z_{ik}u_k\end{aligned}}$$

Compute the covariance matrix of the prepared data, sort its eigenvectors by eigenvalue, and describe each point by $k$ scores, one dot product per kept component. Going back, $\hat x_i$ adds the components with the scores as weights. It equals $x_i$ exactly when $k=p$, or when $x_i$ already lies in the span of $u_1,\dots,u_k$.

Proof

Why $\hat x_i$ is the natural way back. The eigenvectors $u_1,\dots,u_p$ form an orthonormal basis, so $x_i=\sum_{m=1}^p(x_i^Tu_m)\,u_m=\sum_{m=1}^pz_{im}u_m$.

Keeping the first $k$ terms gives $\hat x_i$; the dropped part $x_i-\hat x_i=\sum_{m>k}z_{im}u_m$ is perpendicular to every kept component.

Matrix form. With $U_k=[u_1,\dots,u_k]$, a $p\times k$ matrix, the scores are $Z=XU_k$ ($n\times k$) and the reconstructions are $\hat X=ZU_k^T$.

Why standardize. Multiplying feature $j$ by $c$ multiplies the off-diagonal entries of row and column $j$ of $\Sigma$ by $c$ and the diagonal entry by $c^2$, which turns the eigenvectors. After standardizing, every feature has variance $1$ whatever its unit.

Looks like this, but is not

Changing the unit of a feature is harmless in least squares, where its coefficient simply rescales, so PCA should not care about units either.

PCA maximizes variance, and a change of unit multiplies that feature's variance by the square of the factor. Reporting quiz 2 in percent multiplies its variance by $25$ and swings the first component from $53.1^\circ$ to $83.4^\circ$.

Scores and one-component reconstructions for the five students

Run the algorithm with $k=1$ on the five students and rebuild each student's two quiz scores from one number.

Find$z_{i1}$ and $\hat x_i$ in points, for all five students.
Given
  • raw scores $(9,10),\ \allowbreak (11,11),\ \allowbreak (11,16),\ \allowbreak (13,17),\ \allowbreak (16,16)$

  • $u_1=(0.6,0.8)$ from the eigenvector block

Solution

Both quizzes are scored in points out of 20, so we only center; the lecture drops the division by $\sigma_j$ when the features share a unit.

Center and score

$$x_i-\bar x:\quad (-3,-4),\ (-1,-3),\ (-1,2),\ (1,3),\ (4,2)$$

Subtract the means $(12,14)$, since $u_1$ was found for centered data.

$$z_{i1}=0.6\,x_{i1}+0.8\,x_{i2}:\quad -5,\ -3,\ 1,\ 3,\ 4$$

A score is the centered point's coordinate along $u_1$.

Rebuild in points

$$\hat x_i=\bar x+z_{i1}u_1=(12,14)+z_{i1}\,(0.6,\ 0.8)$$

The algorithm's $\hat x_i$ lives in centered units; adding the means returns to points.

$$(9,10),\ (10.2,11.6),\ (12.6,14.8),\ (13.8,16.4),\ (14.4,17.2)$$

For the third student $12+0.6\cdot1=12.6$ and $14+0.8\cdot1=14.8$.

Answer $$\boxed{z_{\cdot1}=(-5,-3,1,3,4);\qquad \hat x_3=(12.6,\ 14.8)\ \text{ for the raw }(11,16)}$$
Check

The miss $x_3-\hat x_3=(11,16)-(12.6,14.8)=(-1.6,1.2)$ is perpendicular to $u_1$: $-1.6\cdot0.6+1.2\cdot0.8=0$. The first student is rebuilt exactly, because that point lies on the line.

One score per student keeps the order and the spread along the main axis; what it loses is always perpendicular to the kept direction.

Reading loadings: four exam parts, two components

A course has four exam parts: calculus, linear algebra, reading and writing. After standardizing, their covariance matrix $\Sigma$, which is now a correlation matrix, has the eigen-pairs below. Interpret the first two components and summarize a student whose standardized scores are $x=(1.2,\ 0.8,\ -0.4,\ 0)$.

FindA reading of $u_1$ and $u_2$, the student's two scores, and $\hat x$ with $k=2$.
Given
  • $\Sigma=\begin{bmatrix}1&0.7&0.4&0.3\\0.7&1&0.3&0.4\\0.4&0.3&1&0.7\\0.3&0.4&0.7&1\end{bmatrix}$

  • $u_1=\tfrac12(1,1,1,1)$, $\lambda_1=2.4$; $u_2=\tfrac12(1,1,-1,-1)$, $\lambda_2=1.0$

  • $x=(1.2,\ 0.8,\ -0.4,\ 0)$ for one student

Solution

Loadings are read by their signs and sizes; the scores are two dot products, and the reconstruction adds the two components back.

Check and read the loadings

$$\Sigma u_1=\tfrac12(2.4,2.4,2.4,2.4)=2.4\,u_1,\qquad \Sigma u_2=1.0\,u_2$$

Every row of $\Sigma$ sums to $2.4$; row 1 against $u_2$ gives $\frac12(1+0.7-0.4-0.3)=0.5=1.0\cdot\frac12$.

$$u_1:\ \text{four equal loadings};\qquad u_2:\ \text{calculus, algebra }+,\ \ \text{reading, writing }-$$

Equal loadings measure the overall level; opposite signs measure a contrast between the two groups.

Two scores

$$z_1=\tfrac12(1.2+0.8-0.4+0)=0.8,\qquad z_2=\tfrac12(1.2+0.8+0.4-0)=1.2$$

Dot products with $u_1$ and $u_2$; the minus signs of $u_2$ turn $-0.4$ into $+0.4$.

Reconstruct with k = 2

$$\hat x=0.8\cdot\tfrac12(1,1,1,1)+1.2\cdot\tfrac12(1,1,-1,-1)=(1.0,\ 1.0,\ -0.2,\ -0.2)$$

Add the two components with the scores as weights.

Answer $$\boxed{z=(0.8,\ 1.2),\qquad \hat x=(1.0,\ 1.0,\ -0.2,\ -0.2)}$$
Check

The miss $x-\hat x=(0.2,-0.2,-0.2,0.2)$ lies along $(1,-1,-1,1)$, the direction of $u_4$, so it is perpendicular to both kept components: $0.2-0.2-0.2+0.2=0$.

Name a component by the pattern of its loadings, then read a score as how much of that pattern a point has.

The same quizzes in other units: why the lecture standardizes

Quiz 2 is now reported in percent, points times 5, while quiz 1 stays in points. Find the first component, then redo PCA on standardized data.

Find$u_1$ for the percent data; $u_1$, $\lambda_1$ and $\lambda_2$ for standardized data.
Given
  • $\Sigma$ in points: $\begin{bmatrix}5.6&4.8\\4.8&8.4\end{bmatrix}$

  • quiz 2 in percent: $x_2'=5x_2$

Solution

A change of units acts on $\Sigma$ directly, so there is no need to go back to the raw scores.

Covariance in the new units

$$\Sigma'=\begin{bmatrix}5.6&5\cdot4.8\\5\cdot4.8&25\cdot8.4\end{bmatrix}=\begin{bmatrix}5.6&24\\24&210\end{bmatrix}$$

Scaling $x_2$ by $5$ scales its covariance by $5$ and its variance by $25$.

$$\lambda^2-215.6\,\lambda+600=0:\qquad \lambda_1=212.78,\ \ \lambda_2=2.82$$

Trace $215.6$; determinant $5.6\cdot210-24^2=1176-576=600$.

$$u_1\propto(24,\ 212.78-5.6)=(24,\ 207.18):\qquad u_1=(0.115,\ 0.993),\ \ 83.4^\circ$$

From the first row, $(5.6-\lambda_1)u_a+24u_b=0$, so $u\propto(24,\lambda_1-5.6)$; then normalize.

Standardized data

$$r=\frac{4.8}{\sqrt{5.6\cdot8.4}}=\frac{4.8}{6.859}=0.700,\qquad \Sigma_{\text{std}}\approx\begin{bmatrix}1&0.7\\0.7&1\end{bmatrix}$$

Dividing each quiz by its standard deviation leaves the correlation off the diagonal, in any units.

$$\lambda=1\pm r:\quad \lambda_1=1.700,\ u_1=\tfrac1{\sqrt2}(1,1);\qquad \lambda_2=0.300,\ u_2=\tfrac1{\sqrt2}(1,-1)$$

Every 2 by 2 correlation matrix with $r>0$ has these eigenvectors; only the eigenvalues depend on $r$.

Answer $$\boxed{\text{percent: }u_1=(0.115,\ 0.993);\qquad \text{standardized: }u_1=\tfrac1{\sqrt2}(1,1),\ \ \lambda=1.700,\ 0.300}$$
Check

The eigenvalues add up to the traces, $212.78+2.82=215.6$ and $1.7+0.3=2$. Standardized, the answer is the same whether quiz 2 is in points or in percent, because $r$ does not change.

Use the raw covariance only when every feature is in the same unit and its size means something; otherwise standardize first.

Checkpoint
§09.3 — scoring and rebuilding a new student

A sixth student, not used to fit the PCA, scored $13$ on quiz 1 and $15$ on quiz 2. The fitted model has means $(12,14)$ and first component $u_1=(0.6,0.8)$.

Find(a) Find the student's score $z_1$ and the one-component reconstruction $\hat x$ in points.
Given
  • means $\bar x=(12,14)$

  • $u_1=(0.6,0.8)$

  • new raw scores $(13,15)$

Hint 1/4

Put the new student into the coordinates the model was fitted in, then go one step along $u_1$ and back.

Hint 2/4

$z_1=(x-\bar x)^Tu_1$ and $\hat x=\bar x+z_1u_1$.

Hint 3/4

Here $x-\bar x=(13,15)-(12,14)=(1,1)$ and $u_1=(0.6,0.8)$.

Hint 4/4

$z_1=1.4$ and $\hat x=(12+0.84,\ 14+1.12)=(12.84,\ 15.12)$.

Show solution

A new point is centered with the means of the data the model was fitted on, not with its own values.

Score

$$z_1=(1,1)^T(0.6,\ 0.8)=1.4$$

Center with $(12,14)$, then take one dot product.

Rebuild

$$\hat x=(12,14)+1.4\,(0.6,\ 0.8)=(12.84,\ 15.12)$$

Go $1.4$ along $u_1$ from the mean point.

Answer $$\boxed{z_1=1.4,\qquad \hat x=(12.84,\ 15.12)}$$
Check

The miss $(13,15)-(12.84,15.12)=(0.16,-0.12)$ is perpendicular to $u_1$: $0.096-0.096=0$.

Store the training means with the eigenvectors: every new point needs them twice, once going in and once coming back.

⚠ Forgetting to add the means back

The algorithm writes $\hat x_i$ for centered data, and the last step back to raw units is easy to skip.

wrong$$\hat x=z_1u_1=(0.84,\ 1.12)$$
right$$\hat x=\bar x+z_1u_1=(12.84,\ 15.12)$$
⚠ Running PCA on the raw covariance of features in arbitrary units

Standardizing looks like an optional extra step.

wrong$$\text{quiz 2 in percent, raw }\Sigma:\ \ u_1=(0.115,\ 0.993)$$
right$$\text{standardized first: }u_1=\tfrac1{\sqrt2}(1,1)\ \text{ in any units}$$

9.4How many components: the proportion of variance explained

Measures what the first $k$ components keep, $\frac{\lambda_1+\dots+\lambda_k}{\lambda_1+\dots+\lambda_p}$, so $k$ can be read off a scree plot or a target.

The algorithm left $k$ open; the eigenvalues already computed say exactly what each extra component buys.

DefinitionDefinition 9.4: Proportion of variance explained (PVE)
Conditions
  • $X$ is centered, and standardized if the features were

  • $\lambda_1\ge\dots\ge\lambda_p$ are the eigenvalues of $\Sigma=\frac1nX^TX$ and $z_{im}=x_i^Tu_m$ are the scores

$$\boxed{\begin{aligned}\text{total variance}&=\sum_{j=1}^p\frac1n\sum_{i=1}^nx_{ij}^2=\operatorname{tr}\Sigma=\sum_{j=1}^p\lambda_j\\ \mathrm{PVE}(m)&=\frac{\sum_{i=1}^nz_{im}^2}{\sum_{j=1}^p\sum_{i=1}^nx_{ij}^2}=\frac{\textcolor{#1f6feb}{\lambda_m}}{\sum_{j=1}^p\lambda_j}\\ \mathrm{PVE}(\text{first }k)&=\sum_{m=1}^k\mathrm{PVE}(m)\end{aligned}}$$

The total variance is the sum of the feature variances, and it is also the sum of the eigenvalues. Component $m$ carries $\lambda_m$ of it, so its share is $\lambda_m$ over the total, and the first $k$ components carry the sum of their shares. One minus that sum is what the discarded components hold.

Proof

Numerator. $\frac1n\sum_iz_{im}^2=u_m^T\Sigma u_m=\lambda_m$, by Definition 9.1 and $\Sigma u_m=\lambda_mu_m$.

Denominator. $\sum_j\frac1n\sum_ix_{ij}^2=\sum_j\Sigma_{jj}=\operatorname{tr}\Sigma$, and the trace of a symmetric matrix is the sum of its eigenvalues. The factors $\frac1n$ cancel in the ratio.

What the rest costs (further reading). With $k$ components the average squared miss is $$\frac1n\sum_i\lVert x_i-\hat x_i\rVert^2=\frac1n\sum_i\sum_{m>k}z_{im}^2=\lambda_{k+1}+\dots+\lambda_p,$$ because the $u_m$ are orthonormal. So $1-\mathrm{PVE}(\text{first }k)$ is the reconstruction error as a share of the total.

Looks like this, but is not

A first component with $\mathrm{PVE}(1)=0.987$ looks like a great compression, and four components with $0.25$ each look like a useless PCA.

PVE only compares variances. The $0.987$ came from quiz 2 recorded in percent, a unit artefact. Equal eigenvalues say that no direction is special, which is a finding about the data, not a failure.

component meigenvaluePVE(m)PVE(first m)

$1$

$2.4$

$0.60$

$0.60$

$2$

$1.0$

$0.25$

$0.85$

$3$

$0.4$

$0.10$

$0.95$

$4$

$0.2$

$0.05$

$1.00$

Two components keep $0.85$ of the variance; the third adds $0.10$ and the fourth only $0.05$.

PVE, the scree plot and the choice of k for the four exam parts

The standardized exam data have eigenvalues $2.4,\ \allowbreak 1.0,\ \allowbreak 0.4,\ \allowbreak 0.2$. Compute every PVE and the cumulative PVE, and choose $k$ by the elbow and by a target of $0.9$.

Find$\mathrm{PVE}(m)$ for $m=1,\dots,4$, the cumulative values, and $k$.
Given$\lambda=(2.4,\ 1.0,\ 0.4,\ 0.2)$ from the $4\times4$ correlation matrix
Solution

With standardized data every feature has variance $1$, so the total is simply $p=4$; check that the eigenvalues add up to it before dividing.

Shares

$$\sum_j\lambda_j=2.4+1.0+0.4+0.2=4=p$$

A correlation matrix has ones on its diagonal, so its trace is $p$.

$$\mathrm{PVE}=\tfrac{2.4}4,\ \tfrac{1.0}4,\ \tfrac{0.4}4,\ \tfrac{0.2}4=0.60,\ 0.25,\ 0.10,\ 0.05$$

Each component keeps its eigenvalue out of the total.

Cumulative shares and the choice

$$\mathrm{PVE}(\text{first }k)=0.60,\ 0.85,\ 0.95,\ 1.00$$

Components add their variances, so their shares add too.

$$\text{elbow: }k=2;\qquad \text{target }0.9:\ k=3$$

The bars drop by $0.35$ from the first to the second, then by only $0.15$ and $0.05$; and $0.85<0.9\le0.95$.

Answer $$\boxed{\mathrm{PVE}=(0.60,\ 0.25,\ 0.10,\ 0.05);\quad k=2\ \text{by the elbow},\ \ k=3\ \text{for }0.9}$$
Check

Out of every 100 units of variance the first two components keep 85 and the last two hold 15, and $\frac{0.4+0.2}4=0.15$: the two views agree.

Say which rule chose $k$; the elbow and a fixed target can disagree, as they do here.

PVE the lecture's way, straight from the data

Compute $\mathrm{PVE}(1)$ for the five students from the scores and the data, without eigenvalues.

Find$\mathrm{PVE}(1)$ from sums of squares.
Given
  • centered $x_i$: $(-3,-4),\ \allowbreak (-1,-3),\ \allowbreak (-1,2),\ \allowbreak (1,3),\ \allowbreak (4,2)$

  • first-component scores $z_{i1}=-5,\ \allowbreak -3,\ \allowbreak 1,\ \allowbreak 3,\ \allowbreak 4$

Solution

The lecture defines PVE by sums of squares; using them directly also checks the eigenvalue shortcut.

Two sums of squares

$$\sum_iz_{i1}^2=25+9+1+9+16=60$$

The numerator of the lecture's PVE: the scores' sum of squares.

$$\sum_j\sum_ix_{ij}^2=28+42=70$$

Quiz 1 contributes $28$ and quiz 2 contributes $42$.

Divide

$$\mathrm{PVE}(1)=\tfrac{60}{70}=0.857$$

The factor $\frac1n$ would cancel between numerator and denominator, so it is left out.

Answer $$\boxed{\mathrm{PVE}(1)=\tfrac{60}{70}=\tfrac67=0.857}$$
Check

The eigenvalue route gives $\frac{\lambda_1}{\lambda_1+\lambda_2}=\frac{12}{12+2}=\frac67$, the same number.

Both formulas are the same ratio: use sums of squares when you have the scores, eigenvalues when you have $\Sigma$.

What the discarded component costs: the reconstruction error (further reading)

For the five students with $k=1$, compute the average squared distance between $x_i$ and $\hat x_i$ and compare it with $\lambda_2$.

Find$\frac15\sum_i\lVert x_i-\hat x_i\rVert^2$.
Given
  • second-component scores $z_{i2}=0,\ \allowbreak 1,\ \allowbreak -2,\ \allowbreak -1,\ \allowbreak 2$

  • $\lambda_2=2$

Solution

The miss $x_i-\hat x_i=z_{i2}u_2$ has length $\lvert z_{i2}\rvert$, so the second scores give the error directly.

Squared misses

$$\lVert x_i-\hat x_i\rVert^2=z_{i2}^2:\quad 0,\ 1,\ 4,\ 1,\ 4$$

$u_2$ has length 1, so the miss vector's squared length is the square of its score.

Average

$$\tfrac15(0+1+4+1+4)=2=\lambda_2$$

The mean square of the dropped scores.

Answer $$\boxed{\tfrac15\sum_i\lVert x_i-\hat x_i\rVert^2=2=\lambda_2}$$
Check

Kept plus lost equals total, $12+2=14$, and $\frac2{14}=1-\mathrm{PVE}(1)=1-0.857$.

Keeping the most variance and making the smallest average squared reconstruction error pick the same components.

Checkpoint
§09.4 — the share kept by two components

A PCA on four features measured in the same unit returns eigenvalues $5,\ 3,\ 1,\ 1$.

Find(a) What is $\mathrm{PVE}(\text{first }2)$?
Given$\lambda=(5,\ 3,\ 1,\ 1)$
Hint 1/4

PVE compares what the kept components hold with the total; identify both before dividing.

Hint 2/4

$\mathrm{PVE}(\text{first }k)=\frac{\lambda_1+\dots+\lambda_k}{\lambda_1+\dots+\lambda_p}$.

Hint 3/4

Here the eigenvalues are $5,3,1,1$: the kept sum is $5+3$ and the total is $5+3+1+1$.

Hint 4/4

$\frac{8}{10}=0.8$.

Show solution

The eigenvalues are given, so no data are needed: one ratio.

Kept over total

$$\frac{5+3}{5+3+1+1}=\frac{8}{10}=0.8$$

Numerator: the two largest eigenvalues. Denominator: all four.

Answer $$\boxed{0.8}$$
Check

The discarded share is $\frac{1+1}{10}=0.2$, and $0.8+0.2=1$.

The denominator is always the whole trace, never only the kept part.

⚠ Dividing by the kept eigenvalues only

The word 'proportion' invites dividing by what is on the page.

wrong$$\mathrm{PVE}(\text{first }2)=\frac{5+3}{5+3}=1$$
right$$\mathrm{PVE}(\text{first }2)=\frac{5+3}{5+3+1+1}=0.8$$
⚠ Reading an eigenvalue as a proportion

With standardized data the eigenvalues look like shares, but they add up to $p$, not to 1.

wrong$$\mathrm{PVE}(1)=\lambda_1=2.4$$
right$$\mathrm{PVE}(1)=\frac{2.4}{4}=0.6$$

9.5Blind source separation: the model x = As and what it cannot recover

Models sensors as unknown mixtures of independent sources, $x=As$, and recovers $s=Wx$ with $W=A^{-1}$, up to order, scale and sign.

PCA kept variance; the second half of the chapter looks for hidden signals that were added together, and asks what the recordings can reveal about them.

DefinitionDefinition 9.5: The blind source separation model
Conditions
  • $p$ sources and $p$ sensors; the mixing matrix $A$ is $p\times p$, invertible and unknown

  • the sources are independent of each other, and $s(1),\dots,s(T)$ are independent over time

  • only $x(1),\dots,x(T)$ are observed; the lecture drops the time index when one time step is enough

$$\boxed{\begin{aligned}x(t)&=A\,s(t),\qquad x_i(t)=\sum_{j=1}^pa_{ij}\,s_j(t)\\ W&=A^{-1},\qquad \textcolor{#1f6feb}{\hat s(t)}=W\,x(t)\\ As&=(AP^TD^{-1})(DPs)\quad\text{for every permutation }P\text{ and invertible diagonal }D\end{aligned}}$$

Each sensor records a weighted sum of all sources, with weights nobody knows. If $A$ were known we would simply invert it. Since only $x$ is seen, reordering the sources and rescaling each one, with the matching change absorbed by $A$, explains the recordings equally well: order, scale and sign cannot be recovered.

Proof

A permutation matrix is orthogonal. $P$ has exactly one $1$ in each row and each column, so $P^TP=I$.

Same recordings. $(AP^TD^{-1})(DPs)=AP^T(D^{-1}D)Ps=AP^TPs=As$.

Still a valid model. $DPs$ lists the same sources in another order, each multiplied by a nonzero constant, so its entries are still independent, and $AP^TD^{-1}$ is still invertible.

Sign is the special case of a $D$ with entries $\pm1$.

Looks like this, but is not

An unmixing matrix whose output is $\hat s=(-s_2,\ 2s_1)$ has failed: it returned the wrong signal first, upside down and at the wrong volume.

That output is $DPs$ with a swap $P$ and $D=\operatorname{diag}(-1,2)$, which Definition 9.5 says the recordings cannot rule out. For telling voices apart, their order and loudness do not matter.

Two microphones, two speakers: mixing and unmixing with a known A

Two speakers take turns; at four instants their signals are $s(1)=(2,0)$, $s(2)=(0,-1)$, $s(3)=(1,0)$ and $s(4)=(0,2)$. The microphones mix them with $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$. Find the recordings, then recover the sources.

Find$x(t)=As(t)$ for $t=1,\dots,4$, the unmixing matrix $W$, and $Wx(t)$.
Given
  • $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$

  • $s(1)=(2,0),\ \allowbreak s(2)=(0,-1),\ \allowbreak s(3)=(1,0),\ \allowbreak s(4)=(0,2)$

Solution

With $A$ known this is linear algebra, one product to mix and one 2 by 2 inverse to unmix; the point is to see what a blind method must reproduce without $A$.

Mix

$$x(1)=(2,\ 0.8),\ \ x(2)=(-0.5,\ -1.2),\ \ x(3)=(1,\ 0.4),\ \ x(4)=(1,\ 2.4)$$

When only speaker $j$ talks, $x$ is $s_j$ times column $j$ of $A$, that is $(1,0.4)$ or $(0.5,1.2)$.

Unmix

$$\det A=1\cdot1.2-0.5\cdot0.4=1,\qquad W=A^{-1}=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$$

The 2 by 2 inverse: swap the diagonal, negate the off-diagonal, divide by the determinant.

$$Wx(1)=(1.2\cdot2-0.5\cdot0.8,\ -0.4\cdot2+1\cdot0.8)=(2,\ 0)$$

Row by row; the other three recordings come back the same way.

Answer $$\boxed{W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix},\qquad Wx(t)=s(t)\ \text{for all four }t}$$
Check

$AW=\begin{bmatrix}1.2-0.2&-0.5+0.5\\0.48-0.48&-0.2+1.2\end{bmatrix}=I$.

Recordings of one active source line up along a column of A; that is the kind of structure a blind method can exploit.

Which unmixing matrices are equally good answers?

The true unmixing matrix is $W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$. Which of the four candidates below could an ICA method return with equal justification?

FindFor each candidate, its output $W_\ast x=W_\ast As$ in terms of $s_1$ and $s_2$.
Given
  • $W_a=\begin{bmatrix}-0.4&1\\1.2&-0.5\end{bmatrix}$, $W_b=\begin{bmatrix}2.4&-1\\-0.4&1\end{bmatrix}$

  • $W_c=\begin{bmatrix}1.2&-0.5\\0.4&-1\end{bmatrix}$, $W_d=\begin{bmatrix}1.2&-0.5\\0.8&0.5\end{bmatrix}$

  • $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$

Solution

An output is acceptable exactly when $W_\ast A$ has one nonzero entry in each row and each column, that is when it equals $DP$.

Swap, scale and sign

$$W_aA=\begin{bmatrix}0&1\\1&0\end{bmatrix}:\ (s_2,\ s_1);\qquad W_bA=\begin{bmatrix}2&0\\0&1\end{bmatrix}:\ (2s_1,\ s_2)$$

$W_a$ lists the rows of $W$ in the other order; $W_b$ doubles the first row.

$$W_cA=\begin{bmatrix}1&0\\0&-1\end{bmatrix}:\ (s_1,\ -s_2)$$

$W_c$ negates the second row.

A mixture in disguise

$$W_dA=\begin{bmatrix}1&0\\1&1\end{bmatrix}:\ (s_1,\ s_1+s_2)$$

Its second row is $w_1^T+w_2^T$, so the second output still carries both speakers.

Answer $$\boxed{W_a,\ W_b,\ W_c\ \text{acceptable};\quad W_d\ \text{not: its second output is }s_1+s_2}$$
Check

Second row of $W_d$ times $A$: $(0.8\cdot1+0.5\cdot0.4,\ 0.8\cdot0.5+0.5\cdot1.2)=(1,\ 1)$, two nonzero entries in one row.

In a simulation, test a candidate by multiplying it with A: a scaled permutation passes, and any row with two nonzero entries is still mixing.

Checkpoint
§09.5 — unmixing one recording

Two sources are mixed by $A=\begin{bmatrix}2&1\\1&1\end{bmatrix}$, and one recording reads $x=(5,3)$.

Find(a) Find the source values $s$ that produced this recording.
Given
  • $A=\begin{bmatrix}2&1\\1&1\end{bmatrix}$

  • $x=(5,3)$

Hint 1/4

The recording is a mixture; undo the mixing matrix.

Hint 2/4

$s=A^{-1}x$, and $\begin{bmatrix}a&b\\c&d\end{bmatrix}^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix}$.

Hint 3/4

Here $ad-bc=2-1=1$, so $A^{-1}=\begin{bmatrix}1&-1\\-1&2\end{bmatrix}$, applied to $x=(5,3)$.

Hint 4/4

$s=(5-3,\ -5+6)=(2,\ 1)$.

Show solution

A 2 by 2 inverse is the shortest route when A is known.

Invert and apply

$$A^{-1}=\begin{bmatrix}1&-1\\-1&2\end{bmatrix}$$

Determinant $2\cdot1-1\cdot1=1$.

$$s=A^{-1}x=(5-3,\ -5+6)=(2,\ 1)$$

Each source value is one row of $A^{-1}$ times the recording.

Answer $$\boxed{s=(2,\ 1)}$$
Check

Mix back: $As=(2\cdot2+1,\ 2+1)=(5,\ 3)=x$.

When A is known, unmixing is one inverse; the whole difficulty of ICA is that A is not known.

⚠ Multiplying by A to unmix

A and W sit side by side in the model and are easy to swap.

wrong$$\hat s=Ax$$
right$$\hat s=A^{-1}x=Wx$$
⚠ Expecting the first output to be the first source

The index in $\hat s_1$ suggests a match with $s_1$.

wrong$$\hat s_1=s_1$$
right$$\hat s_1=\alpha\,s_j\ \ \text{for some source }j\text{ and some }\alpha\ne0$$

9.6Why the sources must not be Gaussian

Shows that Gaussian sources leave the mixing matrix free up to a rotation: $A$ and $AQ$ give the same distribution of $x$.

Order, scale and sign are harmless ambiguities; this block shows the fourth one on the lecture's list, which is not harmless at all.

TheoremTheorem 9.6: Gaussian sources cannot be unmixed
Conditions
  • $s\sim N(0,I)$: the $p$ sources are independent standard Gaussians

  • $Q$ is orthogonal, $QQ^T=Q^TQ=I$; in two dimensions a rotation is an example

  • $A$ is invertible and $A'=AQ$

$$\boxed{\begin{aligned}x=\textcolor{#1f6feb}{A}s&\ \sim\ N\big(0,\ AA^T\big)\\ x'=(\textcolor{#d1690a}{AQ})s&\ \sim\ N\big(0,\ AQQ^TA^T\big)=N\big(0,\ AA^T\big)\end{aligned}}$$

A Gaussian vector is fixed by its mean and covariance. Mixing with $A$ or with $AQ$ gives the same mean and the same covariance, so the same distribution, and no amount of data separates the two. The unmixing matrices $W$ and $Q^TW$ look equally good, yet $Q^TW$ returns rotated mixtures of the sources.

Proof

Gaussian in, Gaussian out. A linear transformation of a Gaussian vector is Gaussian, so $x$ is Gaussian with $E[x]=A\,E[s]=0$.

Its covariance. $E[xx^T]=E[As\,s^TA^T]=A\,E[ss^T]\,A^T=AIA^T=AA^T$.

Rotate the mixing matrix. $E[x'x'^T]=AQ\,I\,Q^TA^T=A(QQ^T)A^T=AA^T$: same mean, same covariance, same Gaussian.

Where non-Gaussian sources escape. For them equal means and covariances do not force equal distributions, as the uniform square in the figure shows.

Looks like this, but is not

If one of two sources is Gaussian, the same argument should rule out ICA, since that source can be rotated freely.

The argument needs every source Gaussian. With one Gaussian and one uniform source, $x=As$ and $x'=AQs$ have different distributions; the standard result, further reading beyond the lecture, allows at most one Gaussian source.

Two different mixing matrices that Gaussian data cannot tell apart

Take $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$ and the rotation $Q=\begin{bmatrix}0.6&-0.8\\0.8&0.6\end{bmatrix}$. Show that $A$ and $A'=AQ$ give Gaussian recordings with the same covariance, and find what the unmixing matrix of $A'$ returns.

Find$A'$, $AA^T$, $A'A'^T$, and $A'^{-1}x$ in terms of $s$.
Given
  • $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$, $Q=\begin{bmatrix}0.6&-0.8\\0.8&0.6\end{bmatrix}$

  • $s\sim N(0,I)$

Solution

A Gaussian distribution is pinned by its covariance, so comparing the two covariance matrices settles the question.

The rotated mixing matrix

$$A'=AQ=\begin{bmatrix}0.6+0.4&-0.8+0.3\\0.24+0.96&-0.32+0.72\end{bmatrix}=\begin{bmatrix}1&-0.5\\1.2&0.4\end{bmatrix}$$

Each entry is a row of $A$ times a column of $Q$.

$$Q^TQ=\begin{bmatrix}0.36+0.64&-0.48+0.48\\-0.48+0.48&0.64+0.36\end{bmatrix}=I$$

$Q$ is orthogonal, so the theorem applies.

Same covariance

$$AA^T=\begin{bmatrix}1+0.25&0.4+0.6\\0.4+0.6&0.16+1.44\end{bmatrix}=\begin{bmatrix}1.25&1\\1&1.6\end{bmatrix}$$

The rows of $A$ dotted with each other.

$$A'A'^T=\begin{bmatrix}1+0.25&1.2-0.2\\1.2-0.2&1.44+0.16\end{bmatrix}=\begin{bmatrix}1.25&1\\1&1.6\end{bmatrix}$$

The same dot products for the rows of $A'$.

What unmixing with the wrong matrix returns

$$A'^{-1}x=Q^TA^{-1}As=Q^Ts=(0.6s_1+0.8s_2,\ \ -0.8s_1+0.6s_2)$$

$(AQ)^{-1}=Q^{-1}A^{-1}=Q^TA^{-1}$.

Answer $$\boxed{AA^T=A'A'^T=\begin{bmatrix}1.25&1\\1&1.6\end{bmatrix},\qquad A'^{-1}x=Q^Ts\ \text{(still mixtures)}}$$
Check

Both matrices have determinant $1$: $1\cdot1.2-0.5\cdot0.4=1$ and $1\cdot0.4-(-0.5)\cdot1.2=1$, as they must, since $\det Q=0.36+0.64=1$.

Matching second moments leaves a rotation free; only features beyond the covariance, such as corners or heavy tails, can pin it down.

Uniform sources: the same two mixing matrices give different data

Now the sources are independent and uniform on $[-1,1]$. Is the recording $x=(1.2,\ 1.2)$ possible under $A$? Under $A'=AQ$?

FindThe source values that would produce $x=(1.2,1.2)$ under each mixing matrix.
Given
  • $W=A^{-1}=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$

  • $W'=A'^{-1}=\begin{bmatrix}0.4&0.5\\-1.2&1\end{bmatrix}$

  • sources independent and uniform on $[-1,1]$

Solution

A recording is possible exactly when unmixing it lands inside the square $[-1,1]^2$ of possible source values.

Under A

$$Wx=(1.44-0.6,\ -0.48+1.2)=(0.84,\ 0.72)\in[-1,1]^2$$

Both entries lie in [−1, 1], so the point is inside the blue parallelogram.

Under AQ

$$W'x=(0.48+0.6,\ -1.44+1.2)=(1.08,\ -0.24)$$

The first entry exceeds 1: no source values produce this recording.

Answer $$\boxed{x=(1.2,1.2)\ \text{is possible under }A\text{ and impossible under }AQ}$$
Check

Both clouds still share the covariance $\frac13AA^T$, because a uniform source on $[-1,1]$ has variance $\frac13$; the difference lives entirely in the shape.

Enough recordings near the corners reveal the mixing matrix; Gaussian recordings have no corners to reveal.

Checkpoint
§09.6 — separating two Gaussian speakers

Two independent speakers are modelled as standard Gaussian signals and mixed by an unknown invertible $A$. An engineer plans to record for an hour instead of a minute.

Find(a) True or false: with enough recordings, ICA recovers the two speakers up to order, scale and sign.
Given$s\sim N(0,I)$, $x=As$, $A$ unknown and invertible
Hint 1/4

Ask what the distribution of $x$ depends on, and whether different mixing matrices can share it.

Hint 2/4

For Gaussian sources $x\sim N(0,AA^T)$, and $(AQ)(AQ)^T=AA^T$ for every orthogonal $Q$.

Hint 3/4

Here $s\sim N(0,I)$ and $A$ is unknown, so the recordings reveal only $AA^T$.

Hint 4/4

A rotation of $A$ stays hidden however long we record: false.

Show solution

One second mixing matrix with the same distribution is enough to refute a claim about any amount of data.

The distribution of x

$$x\sim N(0,\ AA^T)$$

A linear transformation of a Gaussian, with covariance $AIA^T$.

A second matrix with the same distribution

$$(AQ)(AQ)^T=AQQ^TA^T=AA^T$$

Any orthogonal $Q$ that is not a signed permutation gives a genuinely different unmixing.

Answer $$\boxed{\text{False}}$$
Check

The rotation example is a concrete pair: $A$ and $AQ$, with the 3-4-5 rotation $Q$, share $\begin{bmatrix}1.25&1\\1&1.6\end{bmatrix}$.

is a property of the model, not of the sample size; more data cannot repair it.

⚠ Treating uncorrelated outputs as separated sources

PCA makes its outputs uncorrelated, and that looks like separation.

wrong$$\operatorname{Cov}(\hat s)=I\ \Rightarrow\ \hat s=s\ \text{up to order and sign}$$
right$$\operatorname{Cov}(Q^Ts)=I\ \text{for every orthogonal }Q$$
⚠ Blaming a single Gaussian source

The theorem is remembered as 'any Gaussian source breaks ICA'.

wrong$$\text{one Gaussian source}\ \Rightarrow\ \text{no separation}$$
right$$\text{every source Gaussian}\ \Rightarrow\ \text{no separation}$$

9.7ICA by maximum likelihood: the density of x = As and the log-likelihood of W

Writes each recording's density as $\prod_jp_{S_j}(w_j^Tx)\,\lvert\det W\rvert$, adds the logs over time, and picks the $W$ that maximizes the sum.

We know what can be recovered; to actually estimate $W$ the lecture writes the probability of the recordings as a function of $W$ and maximizes it.

TheoremTheorem 9.7: The density of a mixture and the ICA log-likelihood
Conditions
  • $x=As$ with $A$ invertible; $W=A^{-1}$ has rows $w_1^T,\dots,w_p^T$

  • the sources are independent with known densities $p_{S_1},\dots,p_{S_p}$, and $x(1),\dots,x(T)$ are independent

  • $\lvert\det W\rvert$ is the absolute value of the determinant, and $\log$ is the natural logarithm

$$\boxed{\begin{aligned}p_X(x)&=p_S(A^{-1}x)\,\lvert\det A\rvert^{-1}=p_S(Wx)\,\lvert\det W\rvert=\prod_{j=1}^pp_{S_j}(w_j^Tx)\,\lvert\det W\rvert\\ l(W)&=\sum_{t=1}^T\Big(\sum_{j=1}^p\log p_{S_j}\big(w_j^Tx(t)\big)+\log\lvert\det W\rvert\Big),\qquad \hat W=\arg\max_W\,l(W)\end{aligned}}$$

To get the density of a recording, unmix it, evaluate each source density at its unmixed value, multiply, and multiply once more by $\lvert\det W\rvert$, the correction for how much $A$ stretches areas. The log-likelihood adds this up over the $T$ recordings; the estimate is the $W$ that makes the recordings most probable.

Proof

One dimension. If $X=aS$ with $a>0$, then $P(X\le x)=P(S\le x/a)$, and differentiating gives $p_X(x)=p_S(x/a)\cdot\frac1a$. For $a<0$ the inequality flips and the factor becomes $\frac1{\lvert a\rvert}$.

Many dimensions. $A$ maps a small box of volume $V$ around $s$ to a parallelogram of volume $\lvert\det A\rvert\,V$ around $x=As$, carrying the same probability, so the density is divided by $\lvert\det A\rvert$.

In terms of W. $\lvert\det A\rvert^{-1}=\lvert\det A^{-1}\rvert=\lvert\det W\rvert$, and independence splits $p_S(Wx)$ into $\prod_jp_{S_j}(w_j^Tx)$, since the $j$-th entry of $Wx$ is $w_j^Tx$.

Over time. Independent recordings multiply, $L(W)=\prod_tp_X(x(t))$, and the log turns the product into the sum $l(W)$; the term $\log\lvert\det W\rvert$ appears once per recording, $T$ times in all.

Climbing it (further reading). With $g_j=(\log p_{S_j})'$ applied entry by entry, $\nabla_Wl=\sum_tg\big(Wx(t)\big)\,x(t)^T+T\,(W^{-1})^T$. For a $2\times2$ matrix the last term comes from $\frac{\partial}{\partial w_{11}}\log\lvert w_{11}w_{22}-w_{12}w_{21}\rvert=\frac{w_{22}}{\det W}$ and its three analogues.

Looks like this, but is not

Since $s=Wx$, the density of a recording should simply be $p_S(Wx)$.

That forgets the stretching. With $X=2S$ and $S$ uniform on $[0,1]$, $p_S(x/2)=1$ on $[0,2]$ would integrate to $2$; the factor $\lvert\det W\rvert=\frac12$ brings the total back to $1$.

The density of a stretched and of a mixed variable

(a) $S$ is uniform on $[0,1]$ and $X=2S$; find $p_X$. (b) $s$ is uniform on the unit square and $x=As$ with $A=\begin{bmatrix}2&1\\0.5&1.5\end{bmatrix}$; find $p_X$ at $x=(2,1.5)$ and at $x=(0.5,1.5)$.

Find$p_X$ in (a), and the two density values in (b).
Given
  • $S\sim U[0,1]$ and $X=2S$

  • $s$ uniform on $[0,1]^2$, $A=\begin{bmatrix}2&1\\0.5&1.5\end{bmatrix}$

Solution

Both parts follow the theorem's recipe: unmix, evaluate the source density, multiply by $\lvert\det W\rvert$.

(a) One dimension

$$p_X(x)=p_S\big(\tfrac x2\big)\cdot\tfrac12=\tfrac12\ \ \text{for }0\le x\le2,\ \ \text{and }0\text{ otherwise}$$

$W=\frac12$, and $p_S=1$ wherever $x/2\in[0,1]$.

(b) The unmixing matrix

$$\det A=2\cdot1.5-1\cdot0.5=2.5,\qquad W=\tfrac1{2.5}\begin{bmatrix}1.5&-1\\-0.5&2\end{bmatrix}$$

The 2 by 2 inverse; $\lvert\det W\rvert=\frac1{2.5}=0.4$.

(b) Two points

$$W(2,\ 1.5)=\tfrac1{2.5}(3-1.5,\ -1+3)=(0.6,\ 0.8):\qquad p_X=1\cdot0.4=0.4$$

Inside the square, where $p_S=1$.

$$W(0.5,\ 1.5)=\tfrac1{2.5}(0.75-1.5,\ -0.25+3)=(-0.3,\ 1.1):\qquad p_X=0$$

Outside the square: no source values produce this recording.

Answer $$\boxed{\text{(a) }p_X=\tfrac12\text{ on }[0,2];\qquad \text{(b) }p_X(2,1.5)=0.4,\quad p_X(0.5,1.5)=0}$$
Check

Total probability: in (a) $\frac12\cdot2=1$; in (b) the parallelogram has area $2.5$ and density $0.4$, and $2.5\cdot0.4=1$.

Every density evaluation in ICA follows this recipe: unmix, look up the source densities, multiply by $\lvert\det W\rvert$.

The log-likelihood of four recordings under three candidate unmixing matrices

The four recordings of the two speakers are $x(1)=(2,0.8)$, $x(2)=(-0.5,-1.2)$, $x(3)=(1,0.4)$ and $x(4)=(1,2.4)$. Model each source with the Laplace density $p(s)=\frac12e^{-\lvert s\rvert}$ and compare $l$ for the true $W$, for $I$, and for the rotated $W'=Q^TW$.

Find$l(W)$, $l(I)$, $l(W')$ and the best of the three.
Given
  • $x(t)$: $(2,0.8)$, $(-0.5,-1.2)$, $(1,0.4)$, $(1,2.4)$

  • $W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$ and $W'=\begin{bmatrix}0.4&0.5\\-1.2&1\end{bmatrix}$; all three candidates have $\lvert\det\rvert=1$

  • $p(s)=\frac12e^{-\lvert s\rvert}$ for both sources

Solution

With Laplace sources $\log p(s)=-\log2-\lvert s\rvert$, so $l$ reduces to a sum of absolute values of the unmixed recordings plus the determinant term.

Simplify the log-likelihood

$$l(W)=-2T\log2-\sum_{t=1}^T\big(\lvert w_1^Tx(t)\rvert+\lvert w_2^Tx(t)\rvert\big)+T\log\lvert\det W\rvert$$

Two sources and $T$ recordings give $2T$ copies of $-\log2$.

$$T=4:\qquad -8\log2=-5.545,\qquad \log\lvert\det\rvert=0$$

All three candidates have $\lvert\det\rvert=1$, so only the absolute values differ.

Sum the absolute unmixed values

$$W:\ (2,0),\ (0,-1),\ (1,0),\ (0,2)\ \Rightarrow\ 2+1+1+2=6$$

The true W returns the sources.

$$I:\ \ 2.8+1.7+1.4+3.4=9.3$$

$\lvert x_1\rvert+\lvert x_2\rvert$ for each recording.

$$W':\ (1.2,-1.6),\ (-0.8,-0.6),\ (0.6,-0.8),\ (1.6,1.2)\ \Rightarrow\ 8.4$$

$W'x(t)=Q^Ts(t)$ smears each single active speaker over both outputs.

Compare

$$l(W)=-11.545,\qquad l(I)=-14.845,\qquad l(W')=-13.945$$

$-5.545$ minus each sum.

Answer $$\boxed{l(W)=-11.545\ >\ l(W')=-13.945\ >\ l(I)=-14.845}$$
Check

Swap the Laplace density for a standard Gaussian: $l$ then depends on squares, $\sum_t\lVert Q^Ts(t)\rVert^2=\sum_t\lVert s(t)\rVert^2=10$, and $W$ and $W'$ tie at $-12.352$, as the Gaussian theorem predicts.

Four recordings and three candidates: twelve unmixings by hand. A computer does the same for every W along a gradient.

A peaked, heavy-tailed source model rewards unmixing matrices whose outputs are often near zero, and that is what lets the likelihood choose a rotation.

The scale the likelihood picks for one source

One Laplace source and one sensor: $x_t=a\,s_t$ with $a$ unknown, and four recordings $3,\ -1,\ 0.5,\ -1.5$. Find the maximum likelihood unmixing weight $\hat w$ and the recovered sources.

Find$\hat w$ and $\hat s_t=\hat w\,x_t$.
Given
  • $x=(3,\ -1,\ 0.5,\ -1.5)$

  • $p(s)=\frac12e^{-\lvert s\rvert}$; $W=w$ is a nonzero number

Solution

With $p=1$ the determinant is $w$ itself, so $l$ is a function of one number and calculus finds its maximum.

Write l(w)

$$l(w)=\sum_{t=1}^4\big(-\log2-\lvert wx_t\rvert\big)+4\log\lvert w\rvert=-4\log2-6\lvert w\rvert+4\log\lvert w\rvert$$

$\sum_t\lvert x_t\rvert=3+1+0.5+1.5=6$.

Maximize

$$w>0:\qquad l'(w)=-6+\tfrac4w=0\ \Rightarrow\ \hat w=\tfrac46=\tfrac23$$

The second derivative $-4/w^2$ is negative, so this is a maximum.

$$\hat s=\tfrac23\,(3,\ -1,\ 0.5,\ -1.5)=(2,\ -0.667,\ 0.333,\ -1)$$

Multiplying by $\hat w$ undoes the unknown gain $a$ as far as the data allow.

Answer $$\boxed{\hat w=\pm\tfrac23,\qquad \hat s=\pm(2,\ -0.667,\ 0.333,\ -1)}$$
Check

$l(\tfrac23)=-8.394>l(1)=-8.773$, and the recovered values have mean absolute value $\frac{2+0.667+0.333+1}4=1$, the mean absolute value of a Laplace source.

Fixing the source density fixes the scale, but $l(-w)=l(w)$: the sign stays free, exactly as the list of ambiguities says.

Checkpoint
§09.7 — the density of a flipped and stretched variable

A sensor records $X=-3S$, where $S$ is uniform on $[0,1]$.

Find(a) Find $p_X(x)$ and where it is positive.
Given
  • $S\sim U[0,1]$

  • $X=-3S$

Hint 1/4

Unmix: express $S$ through $X$, then correct for the stretching.

Hint 2/4

$p_X(x)=p_S(wx)\,\lvert w\rvert$ with $w=a^{-1}$ when $X=aS$.

Hint 3/4

Here $a=-3$, so $w=-\frac13$, and $p_S=1$ on $[0,1]$: $p_X(x)=1\cdot\frac13$ wherever $-x/3\in[0,1]$.

Hint 4/4

$p_X=\frac13$ on $[-3,0]$ and $0$ elsewhere.

Show solution

The one-dimensional case of the theorem needs one unmixing and one absolute value.

Unmix and correct

$$p_X(x)=p_S\big(-\tfrac x3\big)\cdot\big\lvert-\tfrac13\big\rvert$$

Theorem 9.7 with $p=1$ and $W=-\frac13$.

$$-\tfrac x3\in[0,1]\iff -3\le x\le0$$

Where the source density equals 1.

Answer $$\boxed{p_X=\tfrac13\ \text{on }[-3,0]}$$
Check

Total probability: $\frac13\cdot3=1$.

The absolute value in $\lvert\det W\rvert$ is what handles a flip.

⚠ Dropping the determinant

$s=Wx$ makes $p_S(Wx)$ look complete.

wrong$$p_X(x)=p_S(Wx)$$
right$$p_X(x)=p_S(Wx)\,\lvert\det W\rvert$$
⚠ Multiplying by det A instead of dividing

The factor is remembered, its direction is not.

wrong$$p_X(x)=p_S(Wx)\,\lvert\det A\rvert$$
right$$p_X(x)=p_S(Wx)\,/\,\lvert\det A\rvert$$
⚠ Adding the determinant term once

$\log\lvert\det W\rvert$ does not depend on $t$, so it looks like a constant outside the sum.

wrong$$l(W)=\sum_t\sum_j\log p_{S_j}(w_j^Tx(t))+\log\lvert\det W\rvert$$
right$$l(W)=\sum_t\sum_j\log p_{S_j}(w_j^Tx(t))+T\log\lvert\det W\rvert$$
PCA by hand for two features

A question gives a few two-dimensional points or a 2 by 2 covariance matrix and asks for components, scores or PVE.

  1. Center

    Subtract each column's mean from that column.

  2. Covariance

    $\Sigma=\frac1nX^TX$: two sums of squares and one sum of cross-products, each divided by $n$.

  3. Eigenvalues

    Solve $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$ and sort the roots.

  4. Eigenvector

    From $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$, then divide by the length; $u_2$ is $u_1$ turned by $90^\circ$.

  5. Scores and PVE

    $z_{i1}=x_i^Tu_1$ on the centered points; $\mathrm{PVE}(1)=\lambda_1/(\lambda_1+\lambda_2)$.

Where it goes wrong
  • Writing $\lambda_1-\Sigma_{11}$ in step 4: a sign slip that tilts the direction.

  • Skipping the division by the length, which inflates every score.

  • Scoring the raw points after fitting on the centered ones.

Evaluating the ICA log-likelihood of a candidate W

A question gives recordings, source densities and one or more candidate unmixing matrices.

  1. Unmix

    Compute $Wx(t)$ for every recording.

  2. Source log-densities

    Add $\log p_{S_j}$ of every entry; for Laplace sources this is $-\log2-\lvert\cdot\rvert$.

  3. Determinant

    Add $T\log\lvert\det W\rvert$, once per recording.

  4. Compare

    The candidate with the larger total is the more likely unmixing.

Where it goes wrong
  • Dropping $T\log\lvert\det W\rvert$, which lets $W$ shrink toward $0$ and win.

  • Adding $\log\lvert\det W\rvert$ once instead of $T$ times.

  • Reading a tie between W and a row-swapped or sign-flipped W as a failure: when the sources share one symmetric density, those changes never alter the log-likelihood.

Checking whether an ICA output is acceptable

A question gives the true mixing matrix, or the true sources, and a returned unmixing matrix or returned signals.

  1. Multiply

    Form $M=W_\ast A$, or write each output as $Ms$.

  2. Count

    Each row and each column of $M$ must hold exactly one nonzero entry.

  3. Read

    Then $M=DP$: the permutation says which source each output is, and $D$ gives its scale and sign.

Where it goes wrong
  • Calling a swapped, flipped or rescaled output a failure.

  • Accepting an output whose row of $M$ has two nonzero entries: that output is still a mixture.

Least squares: quiz 2 predicted from quiz 1

Fit $x_2=bx_1$ to the five centered students by least squares and measure the misses vertically.

Find$b$ and the mean squared vertical miss.
Given$\sum x_1^2=28$, $\sum x_1x_2=24$, $\sum x_2^2=42$
Solution

Least squares through the origin has a closed form, and its misses are measured along the quiz 2 axis.

Slope

$$b=\frac{\sum x_1x_2}{\sum x_1^2}=\frac{24}{28}=0.857$$

The normal equation for one coefficient.

Vertical misses

$$\tfrac15\sum_i(x_{i2}-bx_{i1})^2=\tfrac15\Big(42-\tfrac{24^2}{28}\Big)=\tfrac{30}7=4.286$$

The residual sum of squares of a one-coefficient fit is $\sum y^2-(\sum xy)^2/\sum x^2$.

Answer $$\boxed{b=0.857,\qquad \text{mean squared vertical miss}=4.286}$$
Check

The first component's line, slope $4/3$, has mean squared vertical miss $\frac{50}9=5.556$, larger, as it must be: least squares is the best at vertical misses.

PCA: the line that treats both quizzes alike

Take the first component of the same students and measure the misses perpendicular to its line.

FindThe slope of the line and the mean squared perpendicular miss.
Given$u_1=(0.6,0.8)$, $\lambda_1=12$, $\lambda_2=2$
Solution

Each point's perpendicular miss is its second score, so no new computation is needed.

Slope

$$\frac{0.8}{0.6}=\frac43=1.333$$

Rise over run of $u_1$.

Perpendicular misses

$$\tfrac15\sum_iz_{i2}^2=\lambda_2=2$$

The miss vector is $z_{i2}u_2$.

Answer $$\boxed{\text{slope }1.333,\qquad \text{mean squared perpendicular miss}=2}$$
Check

The least squares direction $(7,6)/\sqrt{85}$ keeps $11.529$, so its perpendicular misses average $14-11.529=2.471>2$.

Same five points, two lines: least squares measures misses vertically and gets slope $0.857$; the first component measures them perpendicular to the line and gets slope $1.333$. Swapping the roles of the quizzes changes the first answer, never the second.

How to tell them apart

A response to predict means least squares; a summary of unlabelled features that treats them alike means PCA.

PCA on the recordings of two speakers who take turns

Speakers take turns, $s=(2,0),\ \allowbreak (-2,0),\ \allowbreak (0,1),\ \allowbreak (0,-1)$, and $A=\begin{bmatrix}1&0.5\\0.4&1.2\end{bmatrix}$ mixes them. Run PCA on the four recordings.

FindThe two principal directions.
Givenrecordings $(2,0.8),\ \allowbreak (-2,-0.8),\ \allowbreak (0.5,1.2),\ \allowbreak (-0.5,-1.2)$, with mean $(0,0)$
Solution

The recordings are already centered, so $\Sigma$ comes straight from them.

Covariance

$$\Sigma=\tfrac14\begin{bmatrix}8.5&4.4\\4.4&4.16\end{bmatrix}=\begin{bmatrix}2.125&1.1\\1.1&1.04\end{bmatrix}$$

Sums of squares and of cross-products over the four recordings.

Eigenvectors

$$\lambda_1=2.809,\ \ \lambda_2=0.356;\qquad u_1\ \text{at }31.9^\circ,\ \ u_2\ \text{at }121.9^\circ$$

Trace $3.165$ and determinant $2.21-1.21=1$; directions from $(\Sigma-\lambda I)u=0$.

Answer $$\boxed{u_1\text{ at }31.9^\circ,\qquad u_2\text{ at }121.9^\circ\ \ \text{(perpendicular)}}$$
Check

The speakers' own directions are the columns of $A$, at $\arctan0.4=21.8^\circ$ and $\arctan2.4=67.4^\circ$; PCA's directions match neither.

ICA versus PCA on two mixed speakers

Two speakers who take turns are mixed into four centered recordings. Compare the Laplace log-likelihood of the true unmixing matrix $W$ with that of PCA's rotation $U^T$, whose rows are the principal directions $u_1^T$ and $u_2^T$.

FindWhich unmixing the Laplace likelihood prefers.
Given
  • recordings $x(t)=(2,0.8),\ \allowbreak (-2,-0.8),\ \allowbreak (0.5,1.2),\ \allowbreak (-0.5,-1.2)$

  • source density $p(s)=\tfrac12e^{-\lvert s\rvert}$, so $l(W)=\sum_t\sum_j\log p(w_j^Tx(t))+4\log\lvert\det W\rvert$

  • $W=\begin{bmatrix}1.2&-0.5\\-0.4&1\end{bmatrix}$

  • $U^T$ rows $u_1\approx(0.849,\,0.528)$ and $u_2\approx(-0.528,\,0.849)$, at $31.9^\circ$ and $121.9^\circ$

  • both matrices have $\lvert\det\rvert=1$

Solution

With equal determinants, $l$ is decided by the sum of the absolute unmixed values.

The true unmixing

$$\sum_t\sum_j\lvert w_j^Tx(t)\rvert=2+2+1+1=6$$

W returns the sources, and each has one zero entry.

PCA's rotation

$$\sum_t\sum_j\lvert u_j^Tx(t)\rvert=8.62$$

Each recording is spread over both principal directions.

Answer $$\boxed{l(W)-l(U^T)=8.62-6=2.62>0}$$
Check

The columns of $A$ are not perpendicular, $1\cdot0.5+0.4\cdot1.2=0.98\ne0$, so no rotation such as $U^T$ can undo $A$.

On the same four recordings PCA returns perpendicular directions at $31.9^\circ$ and $121.9^\circ$, while the speakers sit along the columns of $A$ at $21.8^\circ$ and $67.4^\circ$; only $W=A^{-1}$ turns each recording back into one active speaker.

How to tell them apart

PCA looks for uncorrelated directions of large variance and always returns perpendicular ones; ICA looks for independent sources, whose directions need not be perpendicular.

Scaffolding comes off
The common skeleton
  1. Center each column by subtracting its mean.

  2. Form $\Sigma=\frac1nX^TX$.

  3. Solve $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$ and sort the roots.

  4. Get $u_1$ from $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$ and normalize it.

  5. Score the centered points, $z_{i1}=x_i^Tu_1$, and report $\mathrm{PVE}(1)=\lambda_1/(\lambda_1+\lambda_2)$.

1 · fully worked

Full PCA of four points by hand

Four points are $(2,1)$, $(4,3)$, $(6,5)$ and $(8,3)$. Find the first principal component, the scores and $\mathrm{PVE}(1)$.

Find$u_1$, the scores $z_{i1}$ and $\mathrm{PVE}(1)$.
Givenraw points $(2,1),\ \allowbreak (4,3),\ \allowbreak (6,5),\ \allowbreak (8,3)$
Solution

Two features, so the 2 by 2 route through trace and determinant is quicker than any general method.

Center

$$\bar x=(5,3):\quad (-3,-2),\ (-1,0),\ (1,2),\ (3,0)$$

Column means $\frac{20}4$ and $\frac{12}4$.

Covariance

$$\Sigma=\tfrac14\begin{bmatrix}20&8\\8&8\end{bmatrix}=\begin{bmatrix}5&2\\2&2\end{bmatrix}$$

$9+1+1+9=20$, $4+0+4+0=8$ and $6+0+2+0=8$.

Eigenvalues

$$\lambda^2-7\lambda+6=0:\quad \lambda_1=6,\ \lambda_2=1$$

Trace $7$, determinant $10-4=6$.

Eigenvector

$$(5-6)\,u_a+2\,u_b=0\ \Rightarrow\ u_b=\tfrac12u_a:\quad u_1=\tfrac1{\sqrt5}(2,1)=(0.894,\ 0.447)$$

First row of $(\Sigma-6I)u=0$, then divide by $\sqrt5$.

Scores and PVE

$$z_{i1}=\tfrac{2x_{i1}+x_{i2}}{\sqrt5}=\tfrac{-8,\ -2,\ 4,\ 6}{\sqrt5}=-3.578,\ -0.894,\ 1.789,\ 2.683$$

Scores use the centered points, because $u_1$ was fitted to them.

$$\mathrm{PVE}(1)=\tfrac6{6+1}=0.857$$

PVE divides the kept eigenvalue by the sum of all of them.

Answer $$\boxed{u_1=\tfrac1{\sqrt5}(2,1),\quad z_{\cdot1}=(-3.578,\ -0.894,\ 1.789,\ 2.683),\quad \mathrm{PVE}(1)=0.857}$$
Check

The scores sum to $0$, and their mean square is $\frac{64+4+16+36}{5\cdot4}=\frac{120}{20}=6=\lambda_1$.

The same five moves solve every two-feature PCA question; only the arithmetic changes.

2 · you write the reasoning

Easier, and this time you write the reasons. The covariance matrix is given directly: $\Sigma=\begin{bmatrix}3&1\\1&3\end{bmatrix}$. Find $u_1$, $\mathrm{PVE}(1)$ and the score of the centered point $(2,0)$. For each line, write why it is allowed.

  1. $\lambda^2-6\lambda+8=0\ \Rightarrow\ \lambda_1=4,\ \ \lambda_2=2$

    reasoning

    Trace $3+3=6$ and determinant $9-1=8$ give the characteristic polynomial, whose roots are $4$ and $2$.

  2. $(3-4)\,u_a+u_b=0\ \Rightarrow\ u_1=\tfrac1{\sqrt2}(1,1)$

    reasoning

    The first row of $(\Sigma-4I)u=0$ forces $u_b=u_a$; dividing by $\sqrt2$ makes it a unit vector.

  3. $\mathrm{PVE}(1)=\tfrac4{4+2}=0.667$

    reasoning

    The first eigenvalue over the sum of both.

  4. $z=(2,0)^Tu_1=\tfrac2{\sqrt2}=1.414$

    reasoning

    The point is already centered, so its score is one dot product with $u_1$.

3 · find the buried error

Harder, with two errors buried in the solution. Four points are $(5,3)$, $(9,3)$, $(11,7)$ and $(15,7)$. A student finds the first component, the score of the first point and $\mathrm{PVE}(1)$. Which two steps are wrong?

  1. Step 1. Means $(10,5)$; centered points $(-5,-2)$, $(-1,-2)$, $(1,2)$, $(5,2)$.

  2. Step 2. $\Sigma=\frac14\begin{bmatrix}52&24\\24&16\end{bmatrix}=\begin{bmatrix}13&6\\6&4\end{bmatrix}$.

  3. Step 3. $\lambda^2-17\lambda+16=0$, so $\lambda_1=16$ and $\lambda_2=1$.

  4. Step 4. From $(16-13)u_a+6u_b=0$: $u_b=-\frac12u_a$, so $u_1=\frac1{\sqrt5}(2,-1)=(0.894,\ -0.447)$.

  5. Step 5. Score of the first point: $z_{11}=(5,3)^Tu_1=\frac{10-3}{\sqrt5}=3.130$.

  6. Step 6. $\mathrm{PVE}(1)=\frac{16}{16+1}=0.941$.

the two buried errors (2)
⚠ step 4

The row of $(\Sigma-\lambda_1I)u=0$ is $(13-16)u_a+6u_b=0$, not $(16-13)u_a+6u_b=0$. The sign slip turns $u_b=\frac12u_a$ into $u_b=-\frac12u_a$.

Both $\lambda-\Sigma_{11}$ and $\Sigma_{11}-\lambda$ appear in textbooks, and a switch is easy to miss in a copied line.

right

$u_1=\frac1{\sqrt5}(2,1)=(0.894,\ 0.447)$; check: $\Sigma(2,1)^T=(32,16)^T=16\,(2,1)^T$.

⚠ step 5

The score uses the raw point $(5,3)$; the model was fitted to centered data, in which the first point is $(-5,-2)$.

Step 1 computed the centered points, but the raw ones are still at the top of the page.

right

$z_{11}=(-5,-2)^T\frac1{\sqrt5}(2,1)=\frac{-12}{\sqrt5}=-5.367$.

4 · the bare problem
§09.2 — four readings, the first component, scores and PVE

Four readings of two gauges in the same unit are $(1,3)$, $(3,5)$, $(5,9)$ and $(7,7)$.

Find
  1. (a) Find $\Sigma$ and its eigenvalues.

  2. (b) Find $u_1$ and the four scores.

  3. (c) Find $\mathrm{PVE}(1)$.

Given$(1,3),\ \allowbreak (3,5),\ \allowbreak (5,9),\ \allowbreak (7,7)$
Hint 1/4

Follow the five moves of the skeleton: center, covariance, eigenvalues, eigenvector, then scores and PVE.

Hint 2/4

$\Sigma=\frac1nX^TX$ on centered data; $\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$; $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$.

Hint 3/4

Here the readings $(1,3), \allowbreak (3,5), \allowbreak (5,9), \allowbreak (7,7)$ have means $(4,6)$, so the centered points are $(-3,-3), \allowbreak (-1,-1), \allowbreak (1,3), \allowbreak (3,1)$ and $\Sigma=\frac14\begin{bmatrix}20&16\\16&20\end{bmatrix}$.

Hint 4/4

$\lambda=9,\ 1$; $u_1=\frac1{\sqrt2}(1,1)$; scores $-4.243,\ \allowbreak -1.414,\ \allowbreak 2.828,\ \allowbreak 2.828$; $\mathrm{PVE}(1)=0.9$.

Show solution

Same unit for both gauges, so centering is the only preparation, and the 2 by 2 route applies.

Center

$$\bar x=(4,6):\quad (-3,-3),\ (-1,-1),\ (1,3),\ (3,1)$$

Column means $\frac{16}4$ and $\frac{24}4$.

Covariance

$$\Sigma=\tfrac14\begin{bmatrix}20&16\\16&20\end{bmatrix}=\begin{bmatrix}5&4\\4&5\end{bmatrix}$$

$9+1+1+9=20$, $9+1+9+1=20$ and $9+1+3+3=16$.

Eigenvalues

$$\lambda^2-10\lambda+9=0:\quad \lambda_1=9,\ \lambda_2=1$$

Trace $10$, determinant $25-16=9$.

Eigenvector

$$(5-9)\,u_a+4\,u_b=0\ \Rightarrow\ u_b=u_a:\quad u_1=\tfrac1{\sqrt2}(1,1)$$

First row of $(\Sigma-9I)u=0$, then normalize.

Scores and PVE

$$z_{i1}=\tfrac{x_{i1}+x_{i2}}{\sqrt2}=\tfrac{-6,\ -2,\ 4,\ 4}{\sqrt2}=-4.243,\ -1.414,\ 2.828,\ 2.828$$

The centered readings, not the raw ones, since $\Sigma$ came from them.

$$\mathrm{PVE}(1)=\tfrac{9}{9+1}=0.9$$

PVE divides the kept eigenvalue by the sum of all of them.

Answer $$\boxed{u_1=\tfrac1{\sqrt2}(1,1),\quad z_{\cdot1}=(-4.243,\ -1.414,\ 2.828,\ 2.828),\quad \mathrm{PVE}(1)=0.9}$$
Check

The scores sum to $0$ and their mean square is $\frac{36+4+16+16}{2\cdot4}=\frac{72}8=9=\lambda_1$.

When the two diagonal entries of a 2 by 2 covariance are equal, the components are always the diagonals (1, 1) and (1, −1), whatever the off-diagonal entry.

Full exam-style question

Three sensors, four readings: PCA by hand and an ICA follow-upexam format

Three sensors in the same unit give four readings: $(13,21,7)$, $(11,23,3)$, $(9,17,3)$ and $(7,19,7)$.

  • (a) Center the data and compute $\Sigma=\frac14X^TX$.
  • (b) Find all eigenvalues and unit eigenvectors.
  • (c) Give $\mathrm{PVE}(1)$ and the smallest $k$ with $\mathrm{PVE}(\text{first }k)\ge0.8$.
  • (d) For the first reading, give the scores $z_1$, $z_2$ and the reconstruction with $k=2$ in the original units.
  • (e) Suppose the three signals are mixtures of independent non-Gaussian sources $s_1,s_2,s_3$, and ICA returns $\hat s=(-2s_3,\ s_1,\ 0.5s_2)$. Has it failed?
Find$\Sigma$; every $\lambda_m$ and $u_m$; the PVE and $k$; $z_1$, $z_2$ and $\hat x$ for the first reading; a verdict on the ICA output.
Given
  • readings $(13,21,7),\ \allowbreak (11,23,3),\ \allowbreak (9,17,3),\ \allowbreak (7,19,7)$

  • ICA output $\hat s=(-2s_3,\ s_1,\ 0.5s_2)$

Solution

Look for zeros in $\Sigma$ before solving a cubic: an uncorrelated sensor splits the problem into a 2 by 2 block and a 1 by 1 block.

(a) Center and covariance

$$\bar x=(10,20,5):\quad (3,1,2),\ (1,3,-2),\ (-1,-3,-2),\ (-3,-1,2)$$

Column means $\frac{40}4$, $\frac{80}4$ and $\frac{20}4$.

$$\Sigma=\tfrac14\begin{bmatrix}20&12&0\\12&20&0\\0&0&16\end{bmatrix}=\begin{bmatrix}5&3&0\\3&5&0\\0&0&4\end{bmatrix}$$

For example $\sum x_1x_2=3+3+3+3=12$ and $\sum x_1x_3=6-2+2-6=0$.

(b) Eigen-pairs

$$\begin{bmatrix}5&3\\3&5\end{bmatrix}:\quad \lambda=8,\ \tfrac1{\sqrt2}(1,1);\qquad \lambda=2,\ \tfrac1{\sqrt2}(1,-1)$$

Equal diagonal entries: eigenvalues $5\pm3$ along the two diagonals.

$$\lambda_1=8,\ u_1=\tfrac1{\sqrt2}(1,1,0);\quad \lambda_2=4,\ u_2=(0,0,1);\quad \lambda_3=2,\ u_3=\tfrac1{\sqrt2}(1,-1,0)$$

Sensor 3 alone has eigenvalue $4$; sort all three.

(c) PVE and k

$$\mathrm{PVE}(1)=\tfrac8{14}=0.571,\qquad \mathrm{PVE}(\text{first }2)=\tfrac{12}{14}=0.857\ \Rightarrow\ k=2$$

The total is $8+4+2=14=5+5+4$; one component falls short of $0.8$, two pass it.

(d) The first reading

$$z_1=\tfrac{3+1}{\sqrt2}=2.828,\qquad z_2=2$$

Dot products of the centered $(3,1,2)$ with $u_1$ and $u_2$.

$$\hat x=2.828\cdot\tfrac1{\sqrt2}(1,1,0)+2\,(0,0,1)=(2,2,2)\ \Rightarrow\ (12,\ 22,\ 7)$$

Add the means $(10,20,5)$ to return to the original units.

(e) The ICA output

$$\hat s=DPs,\qquad Ps=(s_3,\ s_1,\ s_2),\qquad D=\operatorname{diag}(-2,\ 1,\ 0.5)$$

Each output is one source times a nonzero constant, so it has not failed.

Answer $$\boxed{\lambda=(8,4,2);\quad k=2;\quad \hat x_1=(12,\ 22,\ 7);\quad \text{ICA has not failed}}$$
Check

The reconstruction misses the first reading $(13,21,7)$ by $(1,-1,0)=\sqrt2\,u_3$, of squared length $2=\lambda_3$; every reading here misses by the same amount, so the average is $\lambda_3$, as it must be.

Check Σ for zero blocks first; and judge an ICA output by whether each output is a single rescaled source, never by its order or size.

Practice

A · concept 4 questions
1§09.1 — is the first component a regression line?

A classmate fits a least squares line of quiz 2 on quiz 1 to the five students and calls it the first principal component.

Find(a) True or false: the first principal component of two features is the least squares line of the second feature on the first.
Givencentered students $(-3,-4),\ \allowbreak (-1,-3),\ \allowbreak (-1,2),\ \allowbreak (1,3),\ \allowbreak (4,2)$
Hint 1/4

Ask what each line is chosen to do.

Hint 2/4

Least squares minimizes $\sum_i(x_{i2}-bx_{i1})^2$, the vertical misses; the first component maximizes $u^T\Sigma u$, which minimizes the perpendicular misses.

Hint 3/4

Here the least squares slope is $24/28=0.857$, and $u_1=(0.6,0.8)$ has slope $0.8/0.6$.

Hint 4/4

$0.857\ne1.333$: false.

Show solution

One pair of numbers that differ refutes the claim.

Two slopes

$$b_{\text{LS}}=\tfrac{24}{28}=0.857,\qquad b_{\text{PC}}=\tfrac{0.8}{0.6}=1.333$$

The slope of each line through the origin.

Answer $$\boxed{\text{False}}$$
Check

Swap the roles: least squares of quiz 1 on quiz 2 gives slope $1.75$ in the same picture, while the first component stays where it is.

Regression needs a response and is not symmetric; PCA has no response and is.

2§09.3 — does a change of units move the components?

A data set records height in metres and mass in kilograms. A student re-records height in centimetres and then runs PCA on the raw covariance matrix.

Find(a) True or false: the principal component directions stay the same.
Given
  • height multiplied by $100$

  • PCA on the covariance matrix, without standardizing

Hint 1/4

Track what the change of unit does to each entry of $\Sigma$.

Hint 2/4

Multiplying feature $j$ by $c$ multiplies $\Sigma_{jj}$ by $c^2$ and the other entries of row and column $j$ by $c$.

Hint 3/4

Here $c=100$ for height: its variance grows by a factor $10^4$, its covariance with mass by $100$.

Hint 4/4

The first component turns toward height: false.

Show solution

The five-quiz example already shows the size of the effect, so one general rule and one number are enough.

What rescaling does to Σ

$$x_j\to c\,x_j:\qquad \Sigma_{jj}\to c^2\Sigma_{jj},\qquad \Sigma_{jk}\to c\,\Sigma_{jk}$$

A variance scales with the square of the factor, a covariance with the factor itself.

A concrete case

$$\text{quiz 2}\times5:\qquad u_1\ \text{turns from }53.1^\circ\text{ to }83.4^\circ$$

Computed in the algorithm block.

Answer $$\boxed{\text{False}}$$
Check

After standardizing, both versions have the same correlation matrix and so the same components.

Units are part of the input to PCA; standardize when they are arbitrary.

3§09.6 — uncorrelated scores, independent scores?

The scores on different principal components always have zero sample covariance.

Find(a) True or false: because the scores on the first two components are uncorrelated, they are independent.
Given
  • $z_{im}=x_i^Tu_m$

  • $\frac1n\sum_iz_{i1}z_{i2}=u_1^T\Sigma u_2=0$

Hint 1/4

Separate two claims: zero covariance, and a joint distribution that factorizes.

Hint 2/4

$\frac1n\sum_iz_{i1}z_{i2}=u_1^T\Sigma u_2=\lambda_2u_1^Tu_2=0$, but independence needs $p(z_1,z_2)=p(z_1)\,p(z_2)$.

Hint 3/4

Here, for points spread evenly round the unit circle, $\Sigma=\frac12I$: any two perpendicular directions give uncorrelated scores, yet $z_1^2+z_2^2=1$ ties them.

Hint 4/4

Uncorrelated but not independent: false.

Show solution

A general proof for 'uncorrelated' and one counterexample for 'independent'.

Uncorrelated

$$\tfrac1n\sum_iz_{i1}z_{i2}=u_1^T\Sigma u_2=\lambda_2\,u_1^Tu_2=0$$

$\Sigma u_2=\lambda_2u_2$, and the eigenvectors are perpendicular.

Not independent

$$z_1^2+z_2^2=1\ \text{ for points on the unit circle}$$

Knowing $z_1=1$ forces $z_2=0$, so the conditional distribution of $z_2$ changes.

Answer $$\boxed{\text{False}}$$
Check

The five students' scores give $\frac15\sum_iz_{i1}z_{i2}=\frac15(0-3-2-3+8)=0$: uncorrelated, as promised.

PCA stops at uncorrelated; ICA is the step from uncorrelated to independent.

4§09.5 — spotting a failed unmixing

Two independent sources $s_1$ and $s_2$ are mixed, and four different methods then try to unmix them.

Find(a) Which output shows that the unmixing failed?
Given
  • true sources $s_1$ and $s_2$

  • each method returns two output signals

Hint 1/4

An acceptable output lists each source once, possibly reordered, rescaled or flipped.

Hint 2/4

Acceptable means $\hat s=DPs$: exactly one source in each output, times a nonzero constant.

Hint 3/4

Check each option: $(s_1+s_2,\ s_2)$, $(-s_2,\ 3s_1)$, $(2s_1,\ -s_2)$, $(s_2,\ s_1)$.

Hint 4/4

Only $(s_1+s_2,\ s_2)$ has an output that contains two sources.

Show solution

An acceptable output has $M=DP$, one nonzero entry per row and column, so writing each $M$ settles it.

Write each output as M s

$$(s_1+s_2,\ s_2)=\begin{bmatrix}1&1\\0&1\end{bmatrix}s$$

Two nonzero entries in the first row: a mixture.

$$(-s_2,3s_1),\ (2s_1,-s_2),\ (s_2,s_1):\quad M=\begin{bmatrix}0&-1\\3&0\end{bmatrix},\ \begin{bmatrix}2&0\\0&-1\end{bmatrix},\ \begin{bmatrix}0&1\\1&0\end{bmatrix}$$

Each has one nonzero entry per row and column: a $DP$.

Answer $$\boxed{(s_1+s_2,\ s_2)}$$
Check

Independence check: $s_1+s_2$ and $s_2$ have covariance $\operatorname{Var}(s_2)\ne0$, so these outputs are not even uncorrelated.

Judge an unmixing by whether each output is a single source, never by the order, sign or size of the outputs.

B · computation 7 questions
1§09.1 — variance along the all-ones direction in three dimensions

Three sensors have covariance $\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$. An engineer summarizes each reading by the plain average of the three sensors, turned into a unit vector.

Find
  1. (a) Compute $u^T\Sigma u$.

  2. (b) Compare it with the best possible value, $\lambda_1$.

Given
  • $\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$

  • $u=\frac1{\sqrt3}(1,1,1)$

Hint 1/4

The variance kept by a unit direction is one quadratic form; the best possible value is the largest eigenvalue.

Hint 2/4

$u^T\Sigma u=\sum_j\sum_k\Sigma_{jk}u_ju_k$; when every $u_j=\frac1{\sqrt3}$ this is $\frac13$ times the sum of all entries of $\Sigma$.

Hint 3/4

Here the entries of $\Sigma=\begin{bmatrix}3&2&0\\2&3&0\\0&0&4\end{bmatrix}$ add to $3+2+2+3+4=14$, and its eigenvalues are $5,4,1$.

Hint 4/4

$u^T\Sigma u=\frac{14}3=4.667$, below $\lambda_1=5$.

Show solution

With equal weights the quadratic form is a plain sum of the entries of $\Sigma$.

Quadratic form

$$u^T\Sigma u=\tfrac13\sum_{j,k}\Sigma_{jk}=\tfrac13(3+2+0+2+3+0+0+0+4)=\tfrac{14}3=4.667$$

Every product $u_ju_k$ equals $\frac13$.

Best value

$$\lambda_1=5\ \text{ at }\ u_1=\tfrac1{\sqrt2}(1,1,0)$$

From the block structure: eigenvalues $3\pm2$ and $4$.

Answer $$\boxed{u^T\Sigma u=4.667<\lambda_1=5}$$
Check

Expand $u$ in the eigenvectors: $c_1^2=\frac23$, $c_2^2=\frac13$, $c_3^2=0$, so $u^T\Sigma u=5\cdot\frac23+4\cdot\frac13=\frac{14}3$.

An evenly weighted average is a direction like any other; the eigenvalues say how far it is from the best one.

2§09.2 — eigen-decomposition of a 2 by 2 covariance

Two correlated measurements of the same process have covariance matrix $\Sigma=\begin{bmatrix}13&6\\6&4\end{bmatrix}$.

Find
  1. (a) Find $\lambda_1$ and $\lambda_2$.

  2. (b) Find $u_1$ and $u_2$ as unit vectors.

  3. (c) Give $\mathrm{PVE}(1)$.

Given$\Sigma=\begin{bmatrix}13&6\\6&4\end{bmatrix}$
Hint 1/4

Two eigen-pairs of a 2 by 2 matrix: eigenvalues first, then one eigenvector, then the perpendicular one.

Hint 2/4

$\lambda^2-(\operatorname{tr}\Sigma)\lambda+\det\Sigma=0$ and $(\Sigma_{11}-\lambda_1)u_a+\Sigma_{12}u_b=0$.

Hint 3/4

Here $\operatorname{tr}\Sigma=13+4=17$, $\det\Sigma=52-36=16$, and the first row of $\Sigma-\lambda_1I$ is $(13-\lambda_1,\ 6)$.

Hint 4/4

$\lambda=16,\ 1$; $u_1=\frac1{\sqrt5}(2,1)$ and $u_2=\frac1{\sqrt5}(1,-2)$; $\mathrm{PVE}(1)=0.941$.

Show solution

The characteristic polynomial factors here, which is faster than the quadratic formula.

Eigenvalues

$$\lambda^2-17\lambda+16=(\lambda-16)(\lambda-1)=0$$

Trace $17$, determinant $16$: two numbers that add to $17$ and multiply to $16$.

Eigenvectors

$$(13-16)\,u_a+6\,u_b=0\ \Rightarrow\ u_b=\tfrac12u_a:\qquad u_1=\tfrac1{\sqrt5}(2,1)$$

A component must have length 1, and $\lVert(2,1)\rVert=\sqrt5$.

$$u_2=\tfrac1{\sqrt5}(1,-2)$$

Perpendicular to $u_1$, first entry positive.

PVE

$$\mathrm{PVE}(1)=\tfrac{16}{16+1}=0.941$$

PVE divides the kept eigenvalue by the sum of all of them.

Answer $$\boxed{\lambda=(16,\ 1),\quad u_1=\tfrac1{\sqrt5}(2,1),\quad \mathrm{PVE}(1)=0.941}$$
Check

$\Sigma(2, \allowbreak 1)^T=(26+6,\ \allowbreak 12+4)=(32, \allowbreak 16)=16\,(2, \allowbreak 1)$ and $\Sigma(1,-2)^T=(13-12,\ 6-8)=(1,-2)$.

When the discriminant is a perfect square the eigenvalues factor; otherwise the quadratic formula gives them to three decimals.

3§09.2 — deflating a covariance matrix

For $\Sigma=\begin{bmatrix}13&6\\6&4\end{bmatrix}$ the first component is $u_1=\frac1{\sqrt5}(2,1)$, with $\lambda_1=16$.

Find
  1. (a) Compute $\Sigma-\lambda_1u_1u_1^T$.

  2. (b) Find its top eigenvector and eigenvalue, and compare them with $u_2$ and $\lambda_2$.

Given
  • $\Sigma=\begin{bmatrix}13&6\\6&4\end{bmatrix}$

  • $u_1=\frac1{\sqrt5}(2,1)$ and $\lambda_1=16$

Hint 1/4

Remove what the first component explains, then look for the best direction in what is left.

Hint 2/4

The deflated covariance is $\frac1nY^TY=\Sigma-\lambda_1u_1u_1^T$, with $Y=X-Xu_1u_1^T$.

Hint 3/4

Here $u_1u_1^T=\frac15\begin{bmatrix}4&2\\2&1\end{bmatrix}$ and $\lambda_1=16$, subtracted from $\Sigma=\begin{bmatrix}13&6\\6&4\end{bmatrix}$.

Hint 4/4

$\begin{bmatrix}0.2&-0.4\\-0.4&0.8\end{bmatrix}=1\cdot u_2u_2^T$: top eigenvector $u_2=\frac1{\sqrt5}(1,-2)$ with eigenvalue $1=\lambda_2$.

Show solution

Subtract entry by entry, then recognize the remainder as a single outer product.

Subtract the first component

$$16\cdot\tfrac15\begin{bmatrix}4&2\\2&1\end{bmatrix}=\begin{bmatrix}12.8&6.4\\6.4&3.2\end{bmatrix}$$

$\lambda_1u_1u_1^T$ entry by entry.

$$\Sigma-\lambda_1u_1u_1^T=\begin{bmatrix}0.2&-0.4\\-0.4&0.8\end{bmatrix}$$

The remainder is what the first component leaves unexplained.

Read the remainder

$$\begin{bmatrix}0.2&-0.4\\-0.4&0.8\end{bmatrix}=\tfrac15\begin{bmatrix}1&-2\\-2&4\end{bmatrix}=u_2u_2^T$$

$u_2=\frac1{\sqrt5}(1,-2)$ gives exactly these entries, so its eigenvalue is $1$.

Answer $$\boxed{\text{top eigen-pair of the remainder: }\big(1,\ \tfrac1{\sqrt5}(1,-2)\big)=(\lambda_2,\ u_2)}$$
Check

Trace check: $0.2+0.8=1=17-16$, the variance left after the first component.

Deflation peels the components off one at a time; the top eigen-pair of each remainder is the next component.

4§09.3 — standardizing before PCA

Two features in different units have covariance matrix $\Sigma=\begin{bmatrix}4&3\\3&9\end{bmatrix}$.

Find
  1. (a) Find the first component of the raw data as an angle from feature 1, and its PVE.

  2. (b) Standardize: give the correlation matrix, its first component and its PVE.

Given$\Sigma=\begin{bmatrix}4&3\\3&9\end{bmatrix}$
Hint 1/4

Run PCA twice, once on $\Sigma$ and once on the correlation matrix, and compare.

Hint 2/4

Standardizing turns $\Sigma_{jk}$ into $\frac{\Sigma_{jk}}{\sigma_j\sigma_k}$; a 2 by 2 correlation matrix has eigenvalues $1\pm r$ and eigenvectors $\frac1{\sqrt2}(1,\pm1)$.

Hint 3/4

Here $\operatorname{tr}\Sigma=13$, $\det\Sigma=36-9=27$, $\sigma_1=2$, $\sigma_2=3$ and $r=\frac{3}{2\cdot3}$.

Hint 4/4

Raw: $\lambda_1=10.405$, $u_1$ at $64.9^\circ$, PVE $0.800$. Standardized: $r=0.5$, $u_1=\frac1{\sqrt2}(1,1)$, PVE $0.75$.

Show solution

The raw case needs the quadratic formula; the standardized case has a closed form in $r$.

Raw data

$$\lambda=\tfrac{13\pm\sqrt{169-108}}2=\tfrac{13\pm7.810}2=10.405,\ 2.595$$

Trace $13$ and determinant $27$.

$$u_1\propto(3,\ 10.405-4)=(3,\ 6.405):\qquad \arctan\tfrac{6.405}3=64.9^\circ$$

First row of $(\Sigma-\lambda_1I)u=0$.

$$\mathrm{PVE}(1)=\tfrac{10.405}{13}=0.800$$

PVE divides the kept eigenvalue by the sum of all of them.

Standardized data

$$r=\tfrac{3}{2\cdot3}=0.5,\qquad \begin{bmatrix}1&0.5\\0.5&1\end{bmatrix}$$

Divide the covariance by $\sigma_1\sigma_2=6$.

$$\lambda=1\pm0.5=1.5,\ 0.5;\qquad u_1=\tfrac1{\sqrt2}(1,1);\qquad \mathrm{PVE}(1)=\tfrac{1.5}2=0.75$$

The 2 by 2 correlation pattern.

Answer $$\boxed{\text{raw: }64.9^\circ,\ \ 0.800;\qquad \text{standardized: }45^\circ,\ \ 0.75}$$
Check

The raw eigenvalues add to $13=4+9$; the standardized ones add to $2$, the number of features.

The raw answer leans toward the feature with the larger variance; standardizing removes the lean and gives each feature one unit of variance.

5§09.4 — one component for the four exam parts

A student's standardized scores on calculus, linear algebra, reading and writing are $x=(0.6,\ 1.0,\ -0.2,\ 0.2)$. The first component is $u_1=\frac12(1,1,1,1)$, with $\lambda_1=2.4$ out of a total of $4$.

Find
  1. (a) Find the score $z_1$ and the reconstruction $\hat x$ with $k=1$.

  2. (b) Find the squared reconstruction error, and check it against the other three scores.

Given
  • $x=(0.6,\ 1.0,\ -0.2,\ 0.2)$

  • $u_1=\frac12(1,1,1,1)$

  • $u_2=\frac12(1,1,-1,-1)$, $u_3=\frac12(1,-1,1,-1)$, $u_4=\frac12(1,-1,-1,1)$

Hint 1/4

Go one step along $u_1$ and back, then measure what is left.

Hint 2/4

$z_1=x^Tu_1$, $\hat x=z_1u_1$, and $\lVert x-\hat x\rVert^2=z_2^2+z_3^2+z_4^2$.

Hint 3/4

Here $x=(0.6,1.0,-0.2,0.2)$ and every entry of $u_1$ is $\frac12$; the other directions are $\frac12(1,1,-1,-1)$, $\frac12(1,-1,1,-1)$ and $\frac12(1,-1,-1,1)$.

Hint 4/4

$z_1=0.8$, $\hat x=(0.4,0.4,0.4,0.4)$, and the error is $0.8$.

Show solution

Computing the error directly and through the dropped scores checks the arithmetic.

Score and reconstruction

$$z_1=\tfrac12(0.6+1.0-0.2+0.2)=0.8,\qquad \hat x=0.8\cdot\tfrac12(1,1,1,1)=(0.4,\ 0.4,\ 0.4,\ 0.4)$$

One dot product and one step back.

The error two ways

$$x-\hat x=(0.2,\ 0.6,\ -0.6,\ -0.2):\qquad 0.04+0.36+0.36+0.04=0.8$$

The squared length of the miss vector, entry by entry.

$$z_2=0.8,\ \ z_3=-0.4,\ \ z_4=0:\qquad 0.64+0.16+0=0.8$$

Through the dropped scores, since the $u_m$ are orthonormal.

Answer $$\boxed{z_1=0.8,\qquad \lVert x-\hat x\rVert^2=0.8}$$
Check

The two routes agree, and $z_1^2+0.8=0.64+0.8=1.44=\lVert x\rVert^2=0.36+1+0.04+0.04$.

One student's error can be above or below the average $\lambda_2+\lambda_3+\lambda_4=1.6$; the eigenvalues describe the average over all students.

6§09.5 — unmixing three recordings

Two sources are mixed by $A=\begin{bmatrix}1&2\\0.5&2\end{bmatrix}$. Three recordings are $x(1)=(3,\ 2.5)$, $x(2)=(-1,\ 0)$ and $x(3)=(2,\ 1.5)$.

Find
  1. (a) Find $W=A^{-1}$.

  2. (b) Recover $s(1)$, $s(2)$ and $s(3)$.

Given
  • $A=\begin{bmatrix}1&2\\0.5&2\end{bmatrix}$

  • $x(1)=(3,2.5)$, $x(2)=(-1,0)$, $x(3)=(2,1.5)$

Hint 1/4

Undo the mixing: one inverse, then three matrix-vector products.

Hint 2/4

$W=A^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix}$ and $s(t)=Wx(t)$.

Hint 3/4

Here $a=1$, $b=2$, $c=0.5$, $d=2$, so $ad-bc=2-1=1$; the recordings are $(3,2.5)$, $(-1,0)$ and $(2,1.5)$.

Hint 4/4

$W=\begin{bmatrix}2&-2\\-0.5&1\end{bmatrix}$; $s=(1,1)$, $(-2,0.5)$ and $(1,0.5)$.

Show solution

The determinant is 1, so the inverse needs no division.

Inverse

$$W=\begin{bmatrix}2&-2\\-0.5&1\end{bmatrix}$$

Determinant $1\cdot2-2\cdot0.5=1$.

Unmix

$$Wx(1)=(6-5,\ -1.5+2.5)=(1,\ 1)$$

Each source value is one row of $W$ times the recording.

$$Wx(2)=(-2,\ 0.5),\qquad Wx(3)=(4-3,\ -1+1.5)=(1,\ 0.5)$$

The same rows serve every recording, since $W$ does not depend on $t$.

Answer $$\boxed{s(1)=(1,1),\quad s(2)=(-2,0.5),\quad s(3)=(1,0.5)}$$
Check

Mix the second one back: $A(-2,0.5)^T=(-2+1,\ -1+1)=(-1,0)=x(2)$.

With A known this is routine; ICA has to find the same W from the recordings alone.

7§09.7 — the density of one recording

Two Laplace sources, each with density $p(s)=\frac12e^{-\lvert s\rvert}$, are mixed, and a candidate unmixing matrix is $W=\begin{bmatrix}1&1\\-1&1\end{bmatrix}$. One recording is $x=(1,1)$.

Find
  1. (a) Compute $p_X(x)$ under this $W$.

  2. (b) Give the term this recording contributes to $l(W)$.

Given
  • $W=\begin{bmatrix}1&1\\-1&1\end{bmatrix}$

  • $x=(1,1)$

  • $p(s)=\frac12e^{-\lvert s\rvert}$ for each source

Hint 1/4

Unmix the recording, evaluate both source densities, and correct for the stretching.

Hint 2/4

$p_X(x)=p_{S_1}(w_1^Tx)\,p_{S_2}(w_2^Tx)\,\lvert\det W\rvert$.

Hint 3/4

Here $w_1^Tx=1+1$, $w_2^Tx=-1+1$, and $\det W=1\cdot1-1\cdot(-1)$.

Hint 4/4

$Wx=(2,0)$ and $\det W=2$: $p_X=\frac12e^{-2}\cdot\frac12\cdot2=0.0677$, and $\log p_X=-2.693$.

Show solution

Theorem 9.7 applied once: one unmixing, two density values, one determinant.

Unmix

$$Wx=(1+1,\ -1+1)=(2,\ 0),\qquad \det W=1+1=2$$

Each unmixed value is one row of $W$ times $x$; a 2 by 2 determinant is $ad-bc$.

Densities

$$p_X(x)=\tfrac12e^{-2}\cdot\tfrac12e^{0}\cdot2=\tfrac12e^{-2}=0.0677$$

Both Laplace densities, times $\lvert\det W\rvert=2$.

$$\log p_X(x)=-\log2-2=-2.693$$

The term this recording adds to $l(W)$.

Answer $$\boxed{p_X(x)=0.0677,\qquad \log p_X(x)=-2.693}$$
Check

Term by term in the log-likelihood: $(-\log2-2)+(-\log2-0)+\log2=-\log2-2$, the same number.

Each recording contributes one such term, and $l(W)$ is their sum over time.

C · exam level 4 questions
1§09.4 — choosing k for six sensor channels

A PCA on six standardized sensor channels returns eigenvalues $2.7,\ \allowbreak 1.5,\ \allowbreak 0.9,\ \allowbreak 0.5,\ \allowbreak 0.3,\ \allowbreak 0.1$. The team wants the fewest components that keep at least $0.85$ of the variance.

Find(a) Which $k$ do they choose?
Given
  • $\lambda=(2.7,\ \allowbreak 1.5,\ \allowbreak 0.9,\ \allowbreak 0.5,\ \allowbreak 0.3,\ \allowbreak 0.1)$

  • target $\mathrm{PVE}(\text{first }k)\ge0.85$

Hint 1/4

Build the running share and stop at the first $k$ that meets the target.

Hint 2/4

$\mathrm{PVE}(\text{first }k)=\frac{\lambda_1+\dots+\lambda_k}{\sum_j\lambda_j}$.

Hint 3/4

Here the total is $2.7+1.5+0.9+0.5+0.3+0.1=6$ and the running sums are $2.7$, $4.2$, $5.1$, $5.6$.

Hint 4/4

The shares are $0.45,\ \allowbreak 0.70,\ \allowbreak 0.85,\ \allowbreak 0.933$, so $k=3$.

Show solution

The shares only need running sums and one division each.

Total

$$\sum_j\lambda_j=6=p$$

With standardized data the trace equals the number of channels.

Running shares

$$\tfrac{2.7}6,\ \tfrac{4.2}6,\ \tfrac{5.1}6,\ \tfrac{5.6}6=0.45,\ 0.70,\ 0.85,\ 0.933$$

Components add their eigenvalues, so the shares accumulate.

Answer $$\boxed{k=3}$$
Check

The discarded eigenvalues $0.5+0.3+0.1=0.9$ are $\frac{0.9}6=0.15$ of the total, and $1-0.15=0.85$.

Meeting a target exactly counts: read 'at least' as $\ge$.

2§09.7 — two candidate unmixing matrices

Three recordings $x(1)=(0.5,0)$, $x(2)=(0,-0.5)$ and $x(3)=(0.25,0.25)$ come from two Laplace sources with density $p(s)=\frac12e^{-\lvert s\rvert}$. The candidates are $W_1=I$ and $W_2=2I$.

Find(a) Which candidate has the larger log-likelihood?
Given
  • $x(1)=(0.5,0)$, $x(2)=(0,-0.5)$, $x(3)=(0.25,0.25)$

  • $W_1=I$ and $W_2=2I$

  • $l(W)=-2T\log2-\sum_t\sum_j\lvert w_j^Tx(t)\rvert+T\log\lvert\det W\rvert$

Hint 1/4

Compute both log-likelihoods in full, the determinant term included.

Hint 2/4

$l(W)=-2T\log2-\sum_t\sum_j\lvert w_j^Tx(t)\rvert+T\log\lvert\det W\rvert$ with $T=3$.

Hint 3/4

Here $\sum_t\sum_j\lvert x_j(t)\rvert=0.5+0.5+0.5=1.5$; $W_2$ doubles it; $\det W_1=1$ and $\det W_2=4$.

Hint 4/4

$l(W_1)=-5.659$ and $l(W_2)=-3.000$: $W_2$ wins.

Show solution

Both candidates share the constant, so compute it once and add the two parts that differ.

Common part

$$-2T\log2=-6\log2=-4.159$$

Each of the $2T=6$ unmixed values contributes $-\log2$.

W₁

$$l(W_1)=-4.159-1.5+3\log1=-5.659$$

Absolute sum $0.5+0.5+(0.25+0.25)=1.5$.

W₂

$$l(W_2)=-4.159-3+3\log4=-4.159-3+4.159=-3.000$$

Doubling $W$ doubles the absolute sum, and $\det(2I)=4$.

Answer $$\boxed{W_2}$$
Check

Along $W=cI$, $l(c)=-4.159-1.5c+6\log c$ peaks at $c=\frac6{1.5}=4$, so $l$ rises from $c=1$ through $c=2$, consistent with $W_2$ beating $W_1$.

Without the determinant term, shrinking $W$ toward $0$ would always look better; the term keeps the recovered sources at the scale of the assumed density.

3§09.6 — which sources can be separated

An engineer can model two independent sources in four ways before they are mixed by an unknown invertible $A$.

Find(a) In which case can ICA recover the sources up to order, scale and sign?
Given
  • $x=As$ with $A$ unknown and invertible

  • the two sources are independent

Hint 1/4

Ask whether a genuinely different mixing matrix produces the same distribution of $x$.

Hint 2/4

Gaussian sources: $A$ and $AQ$ give the same distribution for every orthogonal $Q$, and the argument needs every source Gaussian.

Hint 3/4

Here the four options are uniform with Laplace; Gaussian with different variances; Gaussian with a long recording; Gaussian after standardizing.

Hint 4/4

Only the uniform and Laplace pair has no Gaussian source.

Show solution

The rotation argument either applies to a case or it does not; checking it case by case settles the question.

Gaussian cases

$$s=Ds',\ \ s'\sim N(0,I):\qquad x=(AD)s'\ \text{and}\ (ADQ)s'\ \text{agree in distribution}$$

Different variances only rescale; a longer recording and standardizing leave the model unchanged.

Non-Gaussian case

$$\text{uniform and Laplace}:\qquad AQs\ne As\ \text{in distribution for }Q\ne DP$$

Corners and heavy tails turn with Q and show in the data.

Answer $$\boxed{\text{uniform and Laplace}}$$
Check

The uniform example gives a concrete test: the recording $(1.2,1.2)$ is possible under $A$ and impossible under $AQ$.

Whether ICA can work is decided by the source distributions, not by preprocessing or by how much data there is.

4§09.2 — three equally correlated features

Three features have equal variances $2$ and equal covariances $1$.

Find
  1. (a) Show that $u_1=\frac1{\sqrt3}(1,1,1)$ is an eigenvector and give $\lambda_1$.

  2. (b) Find $\lambda_2$ and $\lambda_3$ from the trace, and say whether the second component is unique.

  3. (c) Give $\mathrm{PVE}(1)$.

Given$\Sigma=\begin{bmatrix}2&1&1\\1&2&1\\1&1&2\end{bmatrix}$
Hint 1/4

Test $(1,1,1)$ first, then use what the trace says about the other two eigenvalues.

Hint 2/4

$\operatorname{tr}\Sigma=\lambda_1+\lambda_2+\lambda_3$, and any vector perpendicular to $(1,1,1)$ can be tested with $\Sigma v=\lambda v$.

Hint 3/4

Here each row of $\Sigma=\begin{bmatrix}2&1&1\\1&2&1\\1&1&2\end{bmatrix}$ sums to $4$, the trace is $6$, and $v=(1,-1,0)$ is perpendicular to $(1,1,1)$.

Hint 4/4

$\lambda_1=4$; $\Sigma v=v$, so $\lambda_2=\lambda_3=1$ and the second component is not unique; $\mathrm{PVE}(1)=\frac46=0.667$.

Show solution

Equal row sums give one eigenvector at once, and the trace gives the sum of the rest.

First eigen-pair

$$\Sigma(1,1,1)^T=(4,4,4)^T=4\,(1,1,1)^T$$

Each row sums to $2+1+1=4$.

The rest

$$\lambda_2+\lambda_3=\operatorname{tr}\Sigma-\lambda_1=6-4=2$$

The trace is the sum of the eigenvalues.

$$\Sigma(1,-1,0)^T=(1,-1,0)^T,\qquad \Sigma(1,0,-1)^T=(1,0,-1)^T$$

Two independent perpendicular vectors, both with eigenvalue $1$, so $\lambda_2=\lambda_3=1$.

PVE

$$\mathrm{PVE}(1)=\tfrac46=0.667$$

PVE divides the kept eigenvalue by the sum of all of them.

Answer $$\boxed{\lambda=(4,\ 1,\ 1),\quad u_2\ \text{not unique},\quad \mathrm{PVE}(1)=0.667}$$
Check

$\Sigma=I+\mathbf 1\mathbf 1^T$ with $\mathbf 1=(1,1,1)^T$, so on vectors perpendicular to $\mathbf 1$ it acts as the identity: eigenvalue $1$ on a whole plane.

Tied eigenvalues mean a tied direction: software returns one choice, and a small change in the data may return a rotated one.

D · interleaved 4 questions
1§09.1 — five students and a final exam

The five students' first-component scores are $z_1=(-5, \allowbreak -3, \allowbreak 1, \allowbreak 3, \allowbreak 4)$ and their second-component scores $z_2=(0, \allowbreak 1, \allowbreak -2, \allowbreak -1, \allowbreak 2)$. Their centered final exam marks are $y=(-10, \allowbreak -5, \allowbreak 0, \allowbreak 5, \allowbreak 10)$.

Find
  1. (a) Fit $y=\beta_1z_1$ by least squares.

  2. (b) Fit $y=\beta_1z_1+\beta_2z_2$. Does $\hat\beta_1$ change?

Given
  • $z_1=(-5, \allowbreak -3, \allowbreak 1, \allowbreak 3, \allowbreak 4)$ and $z_2=(0, \allowbreak 1, \allowbreak -2, \allowbreak -1, \allowbreak 2)$

  • $y=(-10, \allowbreak -5, \allowbreak 0, \allowbreak 5, \allowbreak 10)$

  • all three centered

Hint 1/4

Treat the scores as two new predictors and apply least squares.

Hint 2/4

Through the origin, $\hat\beta=\frac{\sum_iz_iy_i}{\sum_iz_i^2}$; with two predictors the normal equations are $Z^TZ\beta=Z^Ty$.

Hint 3/4

Here $\sum z_1y=50+15+0+15+40$, $\sum z_1^2=60$, $\sum z_2y=0-5+0-5+20$, $\sum z_2^2=10$ and $\sum z_1z_2=0$.

Hint 4/4

$\hat\beta_1=2$ in both fits and $\hat\beta_2=1$; the second fit is exact.

Show solution

The least squares formulas apply unchanged; the scores just happen to be uncorrelated.

One predictor

$$\hat\beta_1=\tfrac{\sum z_{i1}y_i}{\sum z_{i1}^2}=\tfrac{120}{60}=2$$

The one-coefficient least squares formula.

Two predictors

$$Z^TZ=\begin{bmatrix}60&0\\0&10\end{bmatrix},\qquad Z^Ty=\begin{bmatrix}120\\10\end{bmatrix}$$

Scores on different components are uncorrelated, so $Z^TZ$ is diagonal.

$$\hat\beta=(2,\ 1)$$

Each equation now involves one coefficient only.

Answer $$\boxed{\hat\beta_1=2\ \text{in both fits},\qquad \hat\beta_2=1}$$
Check

$2z_1+z_2=(-10,\ \allowbreak -6+1,\ \allowbreak 2-2,\ \allowbreak 6-1,\ \allowbreak 8+2)=(-10, \allowbreak -5, \allowbreak 0, \allowbreak 5, \allowbreak 10)=y$: zero residuals.

Regressing on principal component scores decouples the coefficients, because the scores are uncorrelated.

2§09.3 — preparing data before cross-validation

A team standardizes 500 samples, runs PCA on all of them and keeps 10 components. It then runs 5-fold cross-validation of a regression that uses the 10 scores as inputs.

Find(a) What, if anything, is wrong with this procedure?
Given
  • $n=500$ samples, $p=40$ features, $k=10$ components

  • 5-fold cross-validation of a regression on the scores

Hint 1/4

Ask which samples each fitted quantity has seen.

Hint 2/4

Cross-validation is honest only if everything fitted, the preprocessing included, is fitted on the training folds alone.

Hint 3/4

Here the means, the standard deviations and the 10 eigenvectors were computed from all 500 samples, each held-out fold included.

Hint 4/4

The preprocessing saw the held-out folds; it has to be fitted inside each fold.

Show solution

List every fitted quantity, then ask which samples were used to fit it.

What was fitted

$$\bar x,\ \ \sigma_j,\ \ u_1,\dots,u_{10}\ \ \text{from all }500\text{ samples}$$

Each of these is estimated from data.

What cross-validation needs

$$\text{fold }f:\quad \bar x^{(-f)},\ \sigma^{(-f)},\ U^{(-f)}\ \text{from the other four folds only}$$

The held-out fold is then transformed with quantities it did not help to estimate.

Answer $$\boxed{\text{fit the standardization and the PCA inside each fold}}$$
Check

Apply the rule a future sample must obey: a new sample never contributes to the means, the spreads or the components it is transformed with, so a held-out sample must not either.

Unsupervised steps are still fitted steps, and anything fitted belongs inside the cross-validation loop.

3§09.7 — one sensor, one peaked source, a climb

One source with density $p(s)=\frac12e^{-\lvert s\rvert}$ is recorded by one sensor as $x_t=a\,s_t$. The recordings are $3,\ -1,\ 0.5,\ -1.5$, and we estimate $w=1/a>0$.

Find
  1. (a) Find $\hat w$ in closed form.

  2. (b) Starting from $w=1$, take two gradient ascent steps $w\leftarrow w+\eta\,l'(w)$.

Given
  • $x=(3,\ -1,\ 0.5,\ -1.5)$

  • $l(w)=-4\log2-w\sum_t\lvert x_t\rvert+4\log w$ for $w>0$

  • learning rate $\eta=0.1$

Hint 1/4

One parameter: set the derivative to zero for (a), and follow it uphill for (b).

Hint 2/4

$l'(w)=-\sum_t\lvert x_t\rvert+\frac4w$.

Hint 3/4

Here $\sum_t\lvert x_t\rvert=3+1+0.5+1.5=6$, so $l'(w)=-6+\frac4w$, with $\eta=0.1$ and $w=1$ to start.

Hint 4/4

$\hat w=\frac23$; the steps give $1\to0.8\to0.7$.

Show solution

The closed form gives the target; the two steps show how an iterative method approaches it.

Closed form

$$l'(w)=-6+\tfrac4w=0\ \Rightarrow\ \hat w=\tfrac23$$

The same calculus as any one-parameter maximum likelihood estimate.

Two ascent steps

$$w_1=1+0.1\,(-6+4)=0.8$$

$l'(1)=-2$: the likelihood falls as $w$ grows, so the step goes down.

$$w_2=0.8+0.1\,(-6+5)=0.7$$

$l'(0.8)=-1$: a smaller step, closer to $\frac23$.

Answer $$\boxed{\hat w=0.667,\qquad w:\ 1\to0.8\to0.7}$$
Check

A third step gives $0.7+0.1(-6+5.714)=0.671$, closing in on $0.667$ from above.

In $p$ dimensions the same climb uses the gradient of $l(W)$ with respect to the whole matrix.

4§09.6 — independent or merely uncorrelated?

PCA scores of any data set are uncorrelated. A student concludes that PCA already solves blind source separation.

Find(a) When are uncorrelated scores also independent?
Given$\frac1n\sum_iz_{i1}z_{i2}=0$ for PCA scores
Hint 1/4

Recall the one family of distributions in which uncorrelated and independent coincide.

Hint 2/4

Jointly Gaussian and uncorrelated implies independent; in general, uncorrelated only rules out a linear relation.

Hint 3/4

Here PCA gives $\frac1n\sum_iz_{i1}z_{i2}=0$, and the question is about the joint distribution of $(z_1,z_2)$.

Hint 4/4

Only for jointly Gaussian scores, and in that case the Gaussian theorem says ICA could not do better anyway.

Show solution

The Gaussian density's exponent is the one place where zero covariance splits the joint density.

Gaussian case

$$(z_1,z_2)\ \text{Gaussian},\ \operatorname{Cov}=\operatorname{diag}(\lambda_1,\lambda_2)\ \Rightarrow\ p(z_1,z_2)=p(z_1)\,p(z_2)$$

The exponent splits into a $z_1$ term and a $z_2$ term.

Otherwise

$$z_1^2+z_2^2=1\ \text{on the unit circle}:\quad \operatorname{Cov}=0,\ \text{yet dependent}$$

A nonlinear tie that covariance cannot see.

Answer $$\boxed{\text{when the scores are jointly Gaussian}}$$
Check

The two speakers who take turns give PCA scores that are uncorrelated, yet only the true unmixing matrix returns outputs that are zero whenever a speaker is silent.

PCA decorrelates; ICA goes further and looks for independence, which needs non-Gaussian sources to be visible.

Mistake ledger (16 entries)
⚠ Using weights that are not a unit vector

The plain sum, or an eigenvector left unscaled, looks like a direction.

wrong$$\operatorname{Var}(x_1+x_2)=23.6\ \text{ as the spread along the diagonal}$$
right$$u=\tfrac1{\sqrt2}(1,1):\quad u^T\Sigma u=11.8$$
⚠ Projecting raw, uncentered scores

The formula $u^T\Sigma u$ is quoted without the centering it assumes.

wrong$$\tfrac15\sum_i(\tilde x_i^Tu_1)^2=350.56\ \ \text{(raw points }\tilde x_i)$$
right$$\tfrac15\sum_i\big((\tilde x_i-\bar x)^Tu_1\big)^2=12$$
⚠ Taking the eigenvector of the smallest eigenvalue

The eigenvector found first, or printed first by some software, is assumed to be the best one.

wrong$$u_1=(0.8,-0.6)\ \Rightarrow\ \text{variance }2$$
right$$u_1=(0.6,0.8)\ \Rightarrow\ \text{variance }12$$
⚠ Leaving the eigenvector unnormalized

$(3,4)$ solves $(\Sigma-12I)u=0$ just as well, and the scaling step is skipped.

wrong$$z_{i1}=3x_{i1}+4x_{i2}:\ \text{scores five times too large}$$
right$$u_1=\tfrac15(3,4):\quad z_{i1}=0.6x_{i1}+0.8x_{i2}$$
⚠ Flipping the sign inside Σ − λI

$\lambda-\Sigma_{11}$ and $\Sigma_{11}-\lambda$ look alike in a line copied fast.

wrong$$(12-5.6)\,u_a+4.8\,u_b=0\ \Rightarrow\ u\propto(3,-4)$$
right$$(5.6-12)\,u_a+4.8\,u_b=0\ \Rightarrow\ u\propto(3,4)$$
⚠ Forgetting to add the means back

The algorithm writes $\hat x_i$ for centered data, and the last step back to raw units is easy to skip.

wrong$$\hat x=z_1u_1=(0.84,\ 1.12)$$
right$$\hat x=\bar x+z_1u_1=(12.84,\ 15.12)$$
⚠ Running PCA on the raw covariance of features in arbitrary units

Standardizing looks like an optional extra step.

wrong$$\text{quiz 2 in percent, raw }\Sigma:\ \ u_1=(0.115,\ 0.993)$$
right$$\text{standardized first: }u_1=\tfrac1{\sqrt2}(1,1)\ \text{ in any units}$$
⚠ Dividing by the kept eigenvalues only

The word 'proportion' invites dividing by what is on the page.

wrong$$\mathrm{PVE}(\text{first }2)=\frac{5+3}{5+3}=1$$
right$$\mathrm{PVE}(\text{first }2)=\frac{5+3}{5+3+1+1}=0.8$$
⚠ Reading an eigenvalue as a proportion

With standardized data the eigenvalues look like shares, but they add up to $p$, not to 1.

wrong$$\mathrm{PVE}(1)=\lambda_1=2.4$$
right$$\mathrm{PVE}(1)=\frac{2.4}{4}=0.6$$
⚠ Multiplying by A to unmix

A and W sit side by side in the model and are easy to swap.

wrong$$\hat s=Ax$$
right$$\hat s=A^{-1}x=Wx$$
⚠ Expecting the first output to be the first source

The index in $\hat s_1$ suggests a match with $s_1$.

wrong$$\hat s_1=s_1$$
right$$\hat s_1=\alpha\,s_j\ \ \text{for some source }j\text{ and some }\alpha\ne0$$
⚠ Treating uncorrelated outputs as separated sources

PCA makes its outputs uncorrelated, and that looks like separation.

wrong$$\operatorname{Cov}(\hat s)=I\ \Rightarrow\ \hat s=s\ \text{up to order and sign}$$
right$$\operatorname{Cov}(Q^Ts)=I\ \text{for every orthogonal }Q$$
⚠ Blaming a single Gaussian source

The theorem is remembered as 'any Gaussian source breaks ICA'.

wrong$$\text{one Gaussian source}\ \Rightarrow\ \text{no separation}$$
right$$\text{every source Gaussian}\ \Rightarrow\ \text{no separation}$$
⚠ Dropping the determinant

$s=Wx$ makes $p_S(Wx)$ look complete.

wrong$$p_X(x)=p_S(Wx)$$
right$$p_X(x)=p_S(Wx)\,\lvert\det W\rvert$$
⚠ Multiplying by det A instead of dividing

The factor is remembered, its direction is not.

wrong$$p_X(x)=p_S(Wx)\,\lvert\det A\rvert$$
right$$p_X(x)=p_S(Wx)\,/\,\lvert\det A\rvert$$
⚠ Adding the determinant term once

$\log\lvert\det W\rvert$ does not depend on $t$, so it looks like a constant outside the sum.

wrong$$l(W)=\sum_t\sum_j\log p_{S_j}(w_j^Tx(t))+\log\lvert\det W\rvert$$
right$$l(W)=\sum_t\sum_j\log p_{S_j}(w_j^Tx(t))+T\log\lvert\det W\rvert$$
Formula card
Score and variance along a direction
$$z_i=x_i^Tu,\qquad \tfrac1n\textstyle\sum_iz_i^2=u^T\Sigma u,\qquad \lVert u\rVert=1$$

centered data; u a unit vector

Principal components
$$\max_{u^Tu=1}u^T\Sigma u:\ \ \Sigma u=\lambda u,\ \ \text{maximum }\lambda_1\text{ at }u_1$$

eigenvalues sorted; eigenvectors of unit length

Deflation
$$\tfrac1nY^TY=\Sigma-\lambda_1u_1u_1^T,\qquad Y=X-Xu_1u_1^T$$

the top eigen-pair of the remainder is the next component

PCA algorithm
$$\Sigma=\tfrac1nX^TX,\quad z_i=[x_i^Tu_1,\dots,x_i^Tu_k]^T,\quad \hat x_i=\textstyle\sum_{m=1}^kz_{im}u_m$$

center first; standardize when units differ; add the means back for raw units

Proportion of variance explained
$$\mathrm{PVE}(m)=\frac{\lambda_m}{\sum_j\lambda_j},\qquad \mathrm{PVE}(\text{first }k)=\sum_{m=1}^k\mathrm{PVE}(m)$$

$\sum_j\lambda_j=\operatorname{tr}\Sigma$, which is $p$ for standardized data

Reconstruction error (further reading)
$$\tfrac1n\textstyle\sum_i\lVert x_i-\hat x_i\rVert^2=\lambda_{k+1}+\dots+\lambda_p$$

k components kept

Mixing model and ambiguities
$$x=As,\quad W=A^{-1},\quad As=(AP^TD^{-1})(DPs)$$

p sources, p sensors, A invertible, sources independent

Gaussian sources cannot be unmixed
$$x=As\sim N(0,AA^T),\qquad (AQ)s\sim N(0,AA^T)$$

s ~ N(0, I); Q orthogonal

Density of a mixture and the ICA log-likelihood
$$p_X(x)=\textstyle\prod_jp_{S_j}(w_j^Tx)\,\lvert\det W\rvert,\qquad l(W)=\sum_t\Big(\sum_j\log p_{S_j}(w_j^Tx(t))+\log\lvert\det W\rvert\Big)$$

independent sources with known densities; recordings independent over time

Check yourself

Close the page and write down from memory:

  • the variance of the scores along a unit vector;
  • the Lagrangian of the first component and the equation it leads to;
  • the four steps of the PCA algorithm and when to standardize;
  • the formula for PVE;
  • the mixing model, the unmixing matrix and the three harmless ambiguities;
  • why Gaussian sources fail;
  • the density of $x=As$ and the ICA log-likelihood.

Then reopen the page and compare; whatever is missing is your reread list.

  • Compute $u^T\Sigma u$ for a given unit vector and find the best angle for two features?

    c-projection

  • Derive $\Sigma u=\lambda u$ from the Lagrangian, eigen-decompose a 2 by 2 matrix, and deflate it?

    c-eigen

  • Center and, when needed, standardize data, then compute scores, reconstructions and a reading of the loadings?

    c-pca-algorithm

  • Turn eigenvalues into PVE, pick $k$ by a target or an elbow, and state the reconstruction error?

    c-choose-k

  • Unmix recordings with a known mixing matrix and say whether a returned output is acceptable?

    c-bss

  • Show that $A$ and $AQ$ give the same distribution for Gaussian sources and a different one for uniform sources?

    c-non-gaussian

  • Evaluate $p_X(x)$ and $l(W)$ for a candidate $W$, determinant term included?

    c-ica-mle

Glossary (24 terms)
principal component analysistemel bileşenler analizi

An unsupervised method that describes centered data by a few directions of largest variance, the leading eigenvectors of the sample covariance matrix.

principal componenttemel bileşen

A unit eigenvector $u_m$ of the sample covariance matrix; the first one is the unit direction along which the data vary most.

principal component scorebileşen skoru

The number $z_{im}=x_i^Tu_m$: where point $i$ lands along component $m$.

The entries of a principal component $u_m$, read as the weights that the component puts on the original features.

eigenvectorözvektör

A nonzero vector $u$ that a matrix only rescales: $\Sigma u=\lambda u$.

eigenvalueözdeğer

The factor $\lambda$ in $\Sigma u=\lambda u$; for a covariance matrix, the variance of the scores along the eigenvector.

Lagrange multiplierLagrange çarpanı

The extra variable $\lambda$ in $L(u,\lambda)=f(u)-\lambda g(u)$, whose stationary points solve 'optimize $f$ subject to $g=0$'.

orthonormalortonormal

Of a set of vectors: each has length 1 and every two are perpendicular.

traceiz

The sum of the diagonal entries of a square matrix; for a symmetric matrix it equals the sum of the eigenvalues.

correlation matrixkorelasyon matrisi

The covariance matrix of standardized features: ones on the diagonal and correlations off it.

reconstruction

The point $\hat x_i=\sum_{m\le k}z_{im}u_m$ rebuilt from $k$ scores; it equals $x_i$ when $k=p$.

proportion of variance explainedaçıklanan varyans oranı

The share $\lambda_m/\sum_j\lambda_j$ of the total variance carried by component $m$; summed over the first $k$ components it measures what they keep.

scree plot

A plot of the proportion of variance explained by each component in order, used to pick the number of components at the point where the drop flattens.

deflation

Removing a found component from the data, $Y=X-Xu_1u_1^T$, so that the best direction of what is left is the next component.

feature extractionöznitelik çıkarımı

Building new features out of the original ones, for example principal component scores, instead of choosing among the originals.

blind source separationkör kaynak ayrıştırma

Recovering unobserved source signals from recordings that are unknown mixtures of them, $x=As$ with $A$ unknown.

cocktail party problemkokteyl parti problemi

The standard picture of blind source separation: several microphones in a room each record a mixture of several speakers.

mixing matrixkarıştırma matrisi

The unknown invertible matrix $A$ in $x=As$; entry $a_{ij}$ is the weight of source $j$ in sensor $i$.

unmixing matrixayrıştırma matrisi

The matrix $W=A^{-1}$ that turns recordings back into sources, $s=Wx$.

independent component analysisbağımsız bileşen analizi

Blind source separation that assumes independent, non-Gaussian sources and estimates the unmixing matrix, here by maximum likelihood.

permutation matrixpermütasyon matrisi

A square matrix with exactly one 1 in each row and each column and zeros elsewhere; it reorders the entries of a vector and is orthogonal.

orthogonal matrixortogonal matris

A square matrix $Q$ with $Q^TQ=QQ^T=I$; it preserves lengths and angles, and in two dimensions it is a rotation or a reflection.

identifiabilitytanımlanabilirlik

The property that different parameter values give different distributions of the data, so that enough data can tell them apart.

determinantdeterminant

A number attached to a square matrix; its absolute value is the factor by which the matrix scales areas and volumes, and $\det(A^{-1})=1/\det A$.

What comes next
§10 · Clustering: mixture of Gaussians and EM

Here unlabelled data were summarized by directions and pulled apart into hidden sources. Next they are split into groups: each point comes from one of a few Gaussian clusters, and the unknown group of every point has to be estimated together with the clusters themselves.

Sources
  • textbookT. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer (course textbook) The weekly line names no chapter of the book, so no section numbers are cited here.
  • course materialEEE 485 lecture slides and lecture notes, chapter 9: PCA, ICA, blind source separation Scope, order of topics and notation (the centered data matrix X, the covariance matrix Σ with 1/n, the unit vector u, the multiplier λ, the components u_m and eigenvalues λ_m, the scores z_i, the reconstructions, PVE, the deflated data Y, x(t), s(t), a_ij, A, W and its rows w_j, P, Q, p_S, p_X, L(W) and l(W)) follow these materials; all explanations, data, figures and exercises here are original.
  • course materialEEE 485 syllabus page on STARS, Fall 2026-27, printed 21 September 2026 Assessment weights, the weekly topic list and the recommended books.
  • textbookG. James, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning, Springer, 2013 Recommended in the syllabus; the lecture's arrest-rate example and its PCA figures come from it.
  • textbookK. P. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press, 2012 Recommended in the syllabus; the lecture's ICA example figure comes from it.
  • standard resultEigen-decomposition of symmetric matrices, Lagrange multipliers and the change of variables for densities Standard results used in the derivations. Every number and figure on this page was computed for it.

Spotted something missing or wrong? tell us · share your own notes or an old exam.

Last updated .