Math Core

Lesson 5.6 · Linear Functions

Scatter plots and lines of fit

Real data is messy. Hours of study and test scores are related, but no formula predicts a score exactly. A scatter plot shows the pattern in two-variable data, and a line of fit turns that pattern into a linear model you can use to make predictions.

Scatter plots

A scatter plot graphs paired data as points (x,y)(x, y), one point for each person, day or object. Here are seven students' study hours and test scores.

hours studied, xx11223344556677
score, yy6262666671717373808084848686
Study hours (x) and test scores (y) for seven students. The points drift upward from left to right.Open in grapher →

The points don't lie on one line, but they clearly trend upward: more study time goes with higher scores.

Correlation

Definition

Correlation

Correlation describes how two variables move together.

  • Positive correlation: as xx increases, yy tends to increase. The points trend upward.
  • Negative correlation: as xx increases, yy tends to decrease. The points trend downward.
  • No correlation: there is no clear upward or downward trend.

A correlation is strong when the points hug a line closely and weak when they are widely scattered around it. Statisticians measure this with the correlation coefficient rr, a number from −1-1 to 11 that a calculator can compute. Values near 11 mean a strong positive correlation, values near −1-1 a strong negative correlation, and values near 00 little or no linear correlation.

The value of a used car (in thousands of dollars) against its age in years: a strong negative correlation.Open in grapher →

Lines of fit

A line of fit (or trend line) is a line that follows the general direction of the data, with roughly as many points above it as below it. It doesn't have to pass through any data point. You can draw one by eye and then write its equation.

Writing a line of fit

  1. Draw a line that follows the trend, with points balanced on both sides.
  2. Choose two points on your line (they may or may not be data points) that are far apart.
  3. Find the slope and write the equation, just as for any two points.

Worked example: A line of fit for study time

Write an equation for a line of fit for the study data. Then interpret the slope and yy-intercept.

The line through (1,62)(1, 62) and (7,86)(7, 86) follows the trend well. Its slope is

m=86−627−1=246=4.m = \frac{86 - 62}{7 - 1} = \frac{24}{6} = 4.

Using (1,62)(1, 62): y−62=4(x−1)y - 62 = 4(x - 1), so y=4x+58y = 4x + 58.

The line of fit y = 4x + 58. Some points are above it and some below.Open in grapher →

Slope: each extra hour of study goes with about 44 more points on the test. yy-intercept: a student who studies 00 hours would be predicted to score about 5858.

Different people will draw slightly different lines, and that's fine, as long as each one follows the trend. A calculator can find the single best line using a method called linear regression; you'll see this in statistics.

Common mistake

A line of fit describes a trend, not a guarantee. The model predicts a score of 4(4)+58=744(4) + 58 = 74 for 44 hours, but the actual student scored 7373. Predictions are estimates, so use words like "about" or "predicted."

Making predictions

Substitute a value into the equation of the line of fit to predict the other variable.

  • Predicting inside the range of the data (here, between 11 and 77 hours) is called interpolation. It's usually reliable.
  • Predicting outside that range is called extrapolation. It's riskier, because the trend may not continue. The model predicts that 1212 hours of study gives 4(12)+58=1064(12) + 58 = 106, which is impossible on a 100100-point test.

The difference between an actual value and the predicted value is called the residual: residual == actual −- predicted. For the student who studied 44 hours, the residual is 73−74=−173 - 74 = -1. A negative residual means the point lies below the line.

Worked example: Predicting a car's value

For the used-car data, a line of fit passes through (1,22)(1, 22) and (5,12)(5, 12). Write its equation and predict the value of a car that is 3.53.5 years old.

m=12−225−1=−104=−2.5.m = \frac{12 - 22}{5 - 1} = \frac{-10}{4} = -2.5.

Using (1,22)(1, 22): y−22=−2.5(x−1)y - 22 = -2.5(x - 1), so y=−2.5x+24.5y = -2.5x + 24.5.

For x=3.5x = 3.5: y=−2.5(3.5)+24.5=−8.75+24.5=15.75y = -2.5(3.5) + 24.5 = -8.75 + 24.5 = 15.75. The predicted value is about $15,750. The slope says the car loses about $2,500 of value per year.

Correlation is not causation

A strong correlation does not prove that one variable causes the other. Ice cream sales and sunburns are positively correlated, but ice cream doesn't cause sunburns: hot, sunny weather drives both. Before claiming cause and effect, look for a hidden third variable like this, or for a controlled experiment.

Tip

To check a line of fit, count points above and below it. If nearly all of them are on one side, adjust the line.

Practice

Practice 1

What kind of correlation does this scatter plot show?

(1, 9)(2, 8.5)(3, 7)(4, 6.5)(5, 5)(6, 4.2)(7, 3)Open in grapher →
Practice 2

Which pair of variables would most likely have a negative correlation?

Practice 3

A line of fit for a plant's height yy (in cm) after xx weeks is y=1.5x+20y = 1.5x + 20. Predict the plant's height after 88 weeks.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

In the model y=1.5x+20y = 1.5x + 20 from the previous problem, what does the slope 1.51.5 mean?

Practice 5

A line of fit passes through the points (2,30)(2, 30) and (10,70)(10, 70). Write its equation in slope-intercept form.

Enter an expression, e.g. 3x^2 - 2x + 1

Practice 6

The line of fit y=−3.2x+88y = -3.2x + 88 models the number of tickets still available, yy, for a concert xx days after sales open. Predict the number of tickets left after 1515 days.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A line of fit is y=0.8x+12y = 0.8x + 12. One data point is (40,47)(40, 47). What is the residual for this point?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

In a city, the number of ice cream cones sold each day has a strong positive correlation with the number of people at the beach. Which conclusion is best supported?