Sulba
000 / 100

Eigenvectors, the SVD and PCA

When two inputs rise and fall together, one number can say almost as much as both. Learn the covariance matrix, which gives the spread of the data along any direction, eigenvectors and eigenvalues, and principal component analysis (PCA), which keeps the directions with the most spread. Then see how libraries compute it with the singular value decomposition (SVD).

Lesson 8 stepsPractice 5 problemsExercise Python, 8 functionsQuiz 5 questions

Warm-up

One question before the lesson. Choose an answer and check it.

21 students’ maths and physics marks rise and fall together. You may keep only one number per student. Which keeps the most of what the two marks say?

Show the answer

C: Each student’s position along the direction in which the marks spread most. The marks spread most along one tilted direction, where both rise together. Positions along it keep 93.5% of the spread of the two marks. The maths mark alone keeps 62.2%, and the difference between the marks only 8.2%. This lesson finds that direction.

Step 1 Two tests, one idea

21 students’ marks on two tests, maths and physics,each out of 100. They rise and fall together.Could one number per student say almost as much?First, centre the marks: take away each test’s mean,66 for maths and 61 for physics.Principal component analysis, PCA, finds that number.406080100406080100mathsphysicsmean 66mean 61
01/08

21 students sat two tests, maths and physics, each marked out of 100. Students who do well in one test tend to do well in the other, so the two marks overlap: much of what one says, the other says too.

Principal component analysis, PCA for short, finds new directions to describe data, in order of how much the data spreads along them. The first few directions can then stand in for all of them, so each student is described by fewer numbers. The question here is whether one number per student can say almost as much as the two marks.

PCA starts from centred data: each test’s mean taken away from its marks, so that the students sit round (0, 0). The means:

maths: (66 + 72 + 78 + 51 + 68 + 91 + 64 + 76 + 64 + 52 + 59 + 51 + 74 + 46 + 53 + 73 + 81 + 74 + 70 + 66 + 57) ÷ 21 = 1386 ÷ 21 = 66

physics: (60 + 65 + 74 + 51 + 63 + 75 + 59 + 63 + 57 + 53 + 64 + 52 + 77 + 47 + 55 + 63 + 69 + 70 + 62 + 59 + 43) ÷ 21 = 1281 ÷ 21 = 61

Each student’s centred marks are a = maths − 66 and b = physics − 61. A student above both means has a and b both positive, and on the board it sits above and to the right of the point where the dashed lines cross:

StudentMathsPhysicsa = maths − 66b = physics − 61
1666066 − 66 = 060 − 61 = −1
2726572 − 66 = 665 − 61 = 4
3787478 − 66 = 1274 − 61 = 13
4515151 − 66 = −1551 − 61 = −10
5686368 − 66 = 263 − 61 = 2
6917591 − 66 = 2575 − 61 = 14
7645964 − 66 = −259 − 61 = −2
8766376 − 66 = 1063 − 61 = 2
9645764 − 66 = −257 − 61 = −4
10525352 − 66 = −1453 − 61 = −8
11596459 − 66 = −764 − 61 = 3
12515251 − 66 = −1552 − 61 = −9
13747774 − 66 = 877 − 61 = 16
14464746 − 66 = −2047 − 61 = −14
15535553 − 66 = −1355 − 61 = −6
16736373 − 66 = 763 − 61 = 2
17816981 − 66 = 1569 − 61 = 8
18747074 − 66 = 870 − 61 = 9
19706270 − 66 = 462 − 61 = 1
20665966 − 66 = 059 − 61 = −2
21574357 − 66 = −943 − 61 = −18

Try it yourself

Turn a direction through the students’ marks and watch the spread along it change, with every student’s position worked out. Find the direction with the most spread. Then mark physics out of 10, or standardise both tests, and see the first component move.

40506070809010030405060708090mathsphysics0°45°90°135°180°0100200spread along u

Violet: the line through the means along u, and each student joined to its foot on it. Below: the spread along u at every angle.

Spread along u 134

u = (cos 0°, sin 0°) = (1, 0)

Each position is u₁ × a + u₂ × b, and the spread is their squares added up and divided by n − 1.

Show the calculation
StudentabPosition u₁a + u₂bSquared
10−11 × 0 + 0 × (−1) = 00² = 0
2641 × 6 + 0 × 4 = 66² = 36
312131 × 12 + 0 × 13 = 1212² = 144
4−15−101 × (−15) + 0 × (−10) = −15(−15)² = 225
5221 × 2 + 0 × 2 = 22² = 4
625141 × 25 + 0 × 14 = 2525² = 625
7−2−21 × (−2) + 0 × (−2) = −2(−2)² = 4
81021 × 10 + 0 × 2 = 1010² = 100
9−2−41 × (−2) + 0 × (−4) = −2(−2)² = 4
10−14−81 × (−14) + 0 × (−8) = −14(−14)² = 196
11−731 × (−7) + 0 × 3 = −7(−7)² = 49
12−15−91 × (−15) + 0 × (−9) = −15(−15)² = 225
138161 × 8 + 0 × 16 = 88² = 64
14−20−141 × (−20) + 0 × (−14) = −20(−20)² = 400
15−13−61 × (−13) + 0 × (−6) = −13(−13)² = 169
16721 × 7 + 0 × 2 = 77² = 49
171581 × 15 + 0 × 8 = 1515² = 225
18891 × 8 + 0 × 9 = 88² = 64
19411 × 4 + 0 × 1 = 44² = 16
200−21 × 0 + 0 × (−2) = 00² = 0
21−9−181 × (−9) + 0 × (−18) = −9(−9)² = 81
Total2680
Spread2680 ÷ 20 = 134

Turn u and watch the spread. Where is it largest, and where is it smallest?

Practice problems

Work each problem out on paper, then type your answer and press Check. Every problem has hints and a full solution.

Score: 0 of 10 points

  1. Problem 1

    1 point

    What is the position of the centred point (4, 2) along the direction u = (0.6, 0.8)?

    Hint 1

    The position is the dot product of u with the point: multiply the matching parts, then add.

    Solution

    position = 0.6 × 4 + 0.8 × 2 = 2.4 + 1.6 = 4

  2. Problem 2

    2 points

    Multiply v = (1, 1) by the matrix with rows (2, 1) and (1, 2). The result is a multiple of v, so v is an eigenvector. What is its eigenvalue?

    Hint 1

    Each part of the result is a row of the matrix times v: 2 × 1 + 1 × 1 for the first.

    Hint 2

    The eigenvalue is how many times v the result is.

    Solution

    first part: 2 × 1 + 1 × 1 = 3

    second part: 1 × 1 + 2 × 1 = 3

    (3, 3) = 3 × (1, 1), so λ = 3

  3. Problem 3

    2 points

    A covariance matrix has rows (4, 2) and (2, 1). What is its larger eigenvalue?

    Hint 1

    λ = (Σ₁₁ + Σ₂₂) ÷ 2 ± √(((Σ₁₁ − Σ₂₂) ÷ 2)² + Σ₁₂²), with Σ₁₁ = 4, Σ₁₂ = 2 and Σ₂₂ = 1.

    Hint 2

    (Σ₁₁ + Σ₂₂) ÷ 2 = 2.5 and (Σ₁₁ − Σ₂₂) ÷ 2 = 1.5.

    Solution

    (Σ₁₁ + Σ₂₂) ÷ 2 = (4 + 1) ÷ 2 = 2.5

    (Σ₁₁ − Σ₂₂) ÷ 2 = (4 − 1) ÷ 2 = 1.5

    √(1.5² + 2²) = √(2.25 + 4) = √6.25 = 2.5

    λ₁ = 2.5 + 2.5 = 5

  4. Problem 4

    2 points

    Three principal components have eigenvalues 6, 3 and 1. What share of the spread do the first two hold together?

    Hint 1

    The total spread is the sum of all the eigenvalues.

    Hint 2

    Divide the first two eigenvalues’ sum by the total.

    Solution

    total = 6 + 3 + 1 = 10

    share = (6 + 3) ÷ 10 = 0.9

  5. Problem 5

    3 points

    A covariance matrix has rows (4, 2) and (2, 3). What is the spread along u = (0.6, 0.8)?

    Hint 1

    The spread along u is uᵀΣu = u₁² × Σ₁₁ + 2 × u₁ × u₂ × Σ₁₂ + u₂² × Σ₂₂.

    Hint 2

    Here u₁² = 0.36, u₁ × u₂ = 0.48 and u₂² = 0.64.

    Solution

    u₁² × Σ₁₁ = 0.36 × 4 = 1.44

    2 × u₁ × u₂ × Σ₁₂ = 0.96 × 2 = 1.92

    u₂² × Σ₂₂ = 0.64 × 3 = 1.92

    uᵀΣu = 1.44 + 1.92 + 1.92 = 5.28

Programming exercise

Write PCA for two inputs in plain Python: the covariance matrix, its eigenvalues and eigenvectors, the scores and the rebuilt points, and power iteration. Save pca.py and test_pca.py in the same folder, fill in each function in pca.py, and run the tests:

python test_pca.py

"""Eigenvectors, the SVD and PCA: programming exercise. Write principal component analysis for data with two inputs in plainPython: centre the data, build its covariance matrix, find the matrix'seigenvalues and eigenvectors, score each point along a direction and rebuildit from its score, and find the first component a second way, by poweriteration. Run the tests from this folder:     python test_pca.py A point is a list of two numbers, such as [maths mark, physics mark]. A2 x 2 matrix is a list of two rows, such as [[134, 90], [90, 81.5]]. Adirection is a list of two numbers whose squares add up to 1, such as[0.8, 0.6].""" import math  # noqa: F401  (you will need math.sqrt)  def mean_centre(points):    """Return (means, centred): the mean of each input, and every point with the means taken away."""    raise NotImplementedError  def covariance(centred):    """Return the covariance matrix of centred points, dividing by n - 1.     It is [[s00, s01], [s01, s11]], where s00 adds up p[0] * p[0] over the    points, s01 adds up p[0] * p[1] and s11 adds up p[1] * p[1], each divided    by n - 1.    """    raise NotImplementedError  def variance_along(cov, u):    """Return the spread along the direction u: u[0]**2 * cov[0][0] + 2 * u[0] * u[1] * cov[0][1] + u[1]**2 * cov[1][1]."""    raise NotImplementedError  def eigen2(cov):    """Return ((l1, l2), (v1, v2)): a 2 x 2 covariance matrix's eigenvalues, largest first, and their eigenvectors.     With a = cov[0][0], b = cov[0][1] and d = cov[1][1], the eigenvalues are    (a + d) / 2 plus or minus sqrt(((a - d) / 2)**2 + b**2). The eigenvector    for an eigenvalue l is [b, l - a] divided by its length. If its first part    is negative, multiply both parts by -1. You may assume b is not 0.    """    raise NotImplementedError  def project(centred, v):    """Return each centred point's score along the direction v: v[0] * p[0] + v[1] * p[1]."""    raise NotImplementedError  def rebuild(means, scores, v):    """Return each point rebuilt from its score alone: [means[0] + score * v[0], means[1] + score * v[1]]."""    raise NotImplementedError  def explained(values):    """Return each eigenvalue's share of their total."""    raise NotImplementedError  def power_iteration(cov, start, steps):    """Return the direction power iteration reaches.     Start from start divided by its length. Then, steps times, multiply the    direction by cov and divide the result by its length.    """    raise NotImplementedError 

Stuck? Write mean_centre and covariance first. eigen2 uses the formula from step 4, with the eigenvector (Σ₁₂, λ − Σ₁₁) divided by its length. project, rebuild and explained take a line or two each. The Solution tab has one way to write each function.

In practice: PCA in real projects

Scale the inputs first. PCA follows the inputs with the biggest numbers, so inputs in different units, such as kilograms and centimetres, should be standardised first, as in Lesson 6. scikit-learn’s StandardScaler does this.

scikit-learn. PCA(n_components=k) centres the data itself, but does not scale it. It finds the components with the SVD or, when there are many more rows than inputs, from the covariance matrix. components_ holds the components, one per row, and explained_variance_ratio_ each one’s share of the spread. PCA(n_components=0.95, svd_solver="full") keeps as many components as it takes to hold 95% of the spread. Fit it on the training data only, then apply the same components to the validation and test data.

Signs. If v is an eigenvector, so is −v: the spread along both is the same. Two libraries, or two versions of one, may give the same component with opposite signs, so compare components up to their sign.

Uses. PCA squeezes many inputs into a few before training another model, which can make it faster and less likely to overfit. It plots data with many inputs in two dimensions, and it removes noise by rebuilding from the first few components. It does not look at the labels, so the direction with the most spread need not be the one that best separates the classes.

Test your knowledge

  1. 01What does PCA do?Show answer

    It finds new directions for the data, the principal components, in order of how much the data spreads along each. They are the eigenvectors of the covariance matrix, and each eigenvalue is the spread along its component. Keeping the first few components describes each point with fewer numbers and keeps most of the spread.

  2. 02What are an eigenvector and an eigenvalue?Show answer

    An eigenvector of a matrix is a direction the matrix does not turn: multiplying by the matrix only stretches it. The number it is stretched by is its eigenvalue, so Av = λv. A symmetric matrix, such as a covariance matrix, has eigenvectors at right angles to each other.

  3. 03Why standardise the inputs before PCA?Show answer

    PCA finds the directions with the most spread, and spread depends on units. An input measured in big numbers spreads more and takes over the first component, whether or not it says more. Standardising gives every input a standard deviation of 1, so each counts the same.

  4. 04How do you choose how many components to keep?Show answer

    Add up the components’ shares of the spread, largest first, until they reach a target such as 90% or 95%. Or look for an elbow, where the next component adds little. If the components feed another model, keep the number that gives the best validation score.

  5. 05How is PCA related to the SVD?Show answer

    Write the centred data as X = USVᵀ. The columns of V are the principal components, each eigenvalue is a singular value squared divided by n − 1, and US holds the scores. Libraries mostly compute PCA this way, because building the covariance matrix multiplies the data by itself, which loses precision when some directions have far less spread than others.

Exit ticket

One last question on the main idea of the lesson.

Physics is re-marked out of 10 instead of 100, and nothing else changes. What happens to the first principal component?

Show the answer

B: It swings almost onto maths, because maths now has much bigger numbers. Out of 10, physics’ variance falls 100 times, from 81.5 to 0.815. The first component becomes (0.9977, 0.0671), almost all maths, and holds 99.84% of the spread. PCA follows the bigger numbers, not the more useful test. Standardise both tests first and it is (0.7071, 0.7071).