Math Core

Lesson 10.4 · Data and Statistics

Correlation and causation

Do students who sleep more get better grades? Do cities with more trees have cooler summers? Questions like these are about the relationship between two numerical variables. In this lesson you'll measure how strong a linear relationship is with a single number, the correlation coefficient, and learn why a strong relationship still doesn't prove that one variable causes the other.

Describing a scatter plot

A scatter plot shows two-variable data as points (x,y)(x, y). The variable on the horizontal axis is often called the explanatory variable and the one on the vertical axis the response variable. When you describe the pattern, discuss three things:

  • Direction. In a positive association, yy tends to increase as xx increases. In a negative association, yy tends to decrease as xx increases.
  • Form. Do the points follow a straight line (linear) or a curve (nonlinear)?
  • Strength. How tightly do the points cluster around that line or curve?
x-axis: hours spent studying. y-axis: test score. Strong, positive, linear.Open in grapher →

The points above rise from left to right and hug an imaginary line closely: a strong, positive, linear association.

The correlation coefficient

Words like "strong" and "weak" are a matter of opinion. The correlation coefficient replaces them with a number.

Definition

Correlation coefficient

The correlation coefficient rr is a number from −1-1 to 11 that measures the direction and strength of a linear association between two numerical variables.

  • The sign of rr gives the direction: positive or negative.
  • The size of rr (how close it is to 11 or −1-1) gives the strength.
  • r=1r = 1 or r=−1r = -1 means every point lies exactly on a line. r=0r = 0 means there is no linear association at all.

The correlation coefficient is found with a calculator or spreadsheet. These rough guidelines help you put it into words:

value of rrdescription
0.8≤∣r∣≤10.8 \le \lvert r \rvert \le 1strong
0.5≤∣r∣<0.80.5 \le \lvert r \rvert < 0.8moderate
0.2≤∣r∣<0.50.2 \le \lvert r \rvert < 0.5weak
∣r∣<0.2\lvert r \rvert < 0.2little or no linear association

The study-hours data above have r≈0.99r \approx 0.99. The two plots below show a moderate negative association and no association.

x-axis: minutes waiting in line. y-axis: customer rating out of 10. Moderate negative, r ≈ −0.66.Open in grapher →
x-axis: day of the month. y-axis: number of emails received. No linear association, r ≈ −0.01.Open in grapher →

Worked example: Interpreting r

For 4040 used cars of the same model, the correlation between the number of miles driven and the selling price is r=−0.87r = -0.87. Interpret this value.

The sign is negative, so cars with more miles tend to sell for less. Since ∣−0.87∣=0.87\lvert -0.87 \rvert = 0.87 is at least 0.80.8, the association is strong. In context: there is a strong, negative, linear association between mileage and price for these cars.

Common mistake

A negative rr is not a weaker rr. The number r=−0.9r = -0.9 describes a stronger relationship than r=0.6r = 0.6, because −0.9-0.9 is closer to −1-1 than 0.60.6 is to 11. The sign only tells you the direction. Compare strengths using ∣r∣\lvert r \rvert.

Where r comes from

You won't need to compute rr by hand, but you can see why its sign works. Draw a vertical line at the mean xˉ\bar{x} and a horizontal line at the mean yˉ\bar{y}. They cut the scatter plot into four regions.

  • A point up and to the right of the center has x−xˉ>0x - \bar{x} > 0 and y−yˉ>0y - \bar{y} > 0, so the product (x−xˉ)(y−yˉ)(x - \bar{x})(y - \bar{y}) is positive. A point down and to the left has two negative factors, so its product is also positive.
  • Points in the other two regions have one positive and one negative factor, so their products are negative.

The correlation coefficient is built by adding up these products (after rescaling so that the units cancel). When most points sit in the lower-left and upper-right regions, the sum is positive and so is rr. When most sit in the upper-left and lower-right regions, rr is negative. When the points are spread evenly among all four, the products cancel and rr is near 00.

This explains several useful facts:

  • rr has no units, and it doesn't change if you change units (inches to centimeters, dollars to euros).
  • Switching which variable is xx and which is yy doesn't change rr.
  • The sign of rr always matches the sign of the slope of the line of best fit. But rr is not the slope: a nearly flat line can still have rr close to 11 if the points hug it tightly.

What r can't tell you

r only measures linear relationships

A correlation near 00 means there is no linear association. It does not mean there is no relationship. Always look at the scatter plot before trusting rr.

Worked example: A perfect relationship with r = 0

The points (−3,9)(-3, 9), (−2,4)(-2, 4), (−1,1)(-1, 1), (0,0)(0, 0), (1,1)(1, 1), (2,4)(2, 4), (3,9)(3, 9) lie exactly on the parabola y=x2y = x^2. Their correlation coefficient is exactly 00. Why?

The points lie on y = x², but r = 0.Open in grapher →

The left half of the pattern goes down and the right half goes up. The negative trend on the left exactly cancels the positive trend on the right, so there is no overall linear trend. The relationship is perfect, but it's a curve, and rr can't see curves.

A single outlier can also change rr dramatically.

Worked example: One point changes everything

These eight points have r≈0.98r \approx 0.98:

(1,2), (2,3), (3,3), (4,5), (5,6), (6,6), (7,8), (8,9)(1, 2), \ (2, 3), \ (3, 3), \ (4, 5), \ (5, 6), \ (6, 6), \ (7, 8), \ (8, 9)

Adding a ninth point, (9,1)(9, 1), drops the correlation to r≈0.42r \approx 0.42.

The point (9, 1) turns a strong correlation into a weak one.Open in grapher →

The eight original points still follow a clear line. One point far from that line is enough to make the association look weak. When you see an outlier, report rr both with and without it, and try to find out why that point is different.

Correlation is not causation

Even a strong correlation does not show that changes in one variable cause changes in the other.

Definition

Lurking variable

A lurking variable is a variable that is not part of the study but affects both of the variables being studied. It can create a correlation between two variables that have no direct effect on each other.

Worked example: Spot the lurking variable

Across the days of a year, a town finds a strong positive correlation between the number of people at the public pool and the number of air conditioners sold. Does going to the pool make people buy air conditioners?

No. Both variables go up on hot days. Temperature is a lurking variable that drives both. Closing the pool would not reduce air conditioner sales.

A correlation can arise in several ways: xx might cause yy; yy might cause xx; a lurking variable might drive both; or the pattern might be a coincidence, especially in a small data set. To show that one variable actually causes a change in another, researchers run a randomized experiment: they randomly assign subjects to groups, change only the variable being tested, and compare the results. Random assignment balances out lurking variables between the groups.

Tip

When you read a headline like "People who do X live longer," ask two questions. Was this an experiment or just an observation? And what else might be different about the people who do X?

Practice

Practice 1

Which correlation coefficient shows the strongest linear association?

Practice 2

Which value is the best estimate of the correlation coefficient for this scatter plot?

x-axis: age of a car in years. y-axis: resale value as a percent of the original price.Open in grapher →
Practice 3

For a group of hikers, the correlation between the altitude they reach and the air temperature they measure there is r=−0.8r = -0.8. Is the association positive or negative?

Type your answer

Practice 4

The correlation between the heights of 3030 plants, measured in inches, and their numbers of leaves is r=0.74r = 0.74. The heights are converted to centimeters (11 inch =2.54= 2.54 cm). What is the new correlation coefficient?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A scatter plot of the height of a thrown ball against time forms a clear upside-down U shape. Its correlation coefficient is r=0.03r = 0.03. Which conclusion is correct?

Practice 6

In a study of towns, the number of churches is strongly positively correlated with the number of restaurants. Which is the most likely explanation?

Practice 7

A data set of 1212 points has r=0.35r = 0.35. One point lies far below the others. When that point is removed, which new value of rr is most plausible if the other 1111 points lie close to a rising line?

Practice 8

A researcher finds that students who eat breakfast have higher test scores. Which study design would give the strongest evidence that eating breakfast causes higher scores?