Lesson 5.6 · Linear Functions
Scatter plots and lines of fit
Real data is messy. Hours of study and test scores are related, but no formula predicts a score exactly. A scatter plot shows the pattern in two-variable data, and a line of fit turns that pattern into a linear model you can use to make predictions.
Scatter plots
A scatter plot graphs paired data as points , one point for each person, day or object. Here are seven students' study hours and test scores.
| hours studied, | |||||||
|---|---|---|---|---|---|---|---|
| score, |
The points don't lie on one line, but they clearly trend upward: more study time goes with higher scores.
Correlation
Definition
Correlation
Correlation describes how two variables move together.
- Positive correlation: as increases, tends to increase. The points trend upward.
- Negative correlation: as increases, tends to decrease. The points trend downward.
- No correlation: there is no clear upward or downward trend.
A correlation is strong when the points hug a line closely and weak when they are widely scattered around it. Statisticians measure this with the correlation coefficient , a number from to that a calculator can compute. Values near mean a strong positive correlation, values near a strong negative correlation, and values near little or no linear correlation.
Lines of fit
A line of fit (or trend line) is a line that follows the general direction of the data, with roughly as many points above it as below it. It doesn't have to pass through any data point. You can draw one by eye and then write its equation.
Writing a line of fit
- Draw a line that follows the trend, with points balanced on both sides.
- Choose two points on your line (they may or may not be data points) that are far apart.
- Find the slope and write the equation, just as for any two points.
Worked example: A line of fit for study time
Write an equation for a line of fit for the study data. Then interpret the slope and -intercept.
The line through and follows the trend well. Its slope is
Using : , so .
Slope: each extra hour of study goes with about more points on the test. -intercept: a student who studies hours would be predicted to score about .
Different people will draw slightly different lines, and that's fine, as long as each one follows the trend. A calculator can find the single best line using a method called linear regression; you'll see this in statistics.
Common mistake
A line of fit describes a trend, not a guarantee. The model predicts a score of for hours, but the actual student scored . Predictions are estimates, so use words like "about" or "predicted."
Making predictions
Substitute a value into the equation of the line of fit to predict the other variable.
- Predicting inside the range of the data (here, between and hours) is called interpolation. It's usually reliable.
- Predicting outside that range is called extrapolation. It's riskier, because the trend may not continue. The model predicts that hours of study gives , which is impossible on a -point test.
The difference between an actual value and the predicted value is called the residual: residual actual predicted. For the student who studied hours, the residual is . A negative residual means the point lies below the line.
Worked example: Predicting a car's value
For the used-car data, a line of fit passes through and . Write its equation and predict the value of a car that is years old.
Using : , so .
For : . The predicted value is about $15,750. The slope says the car loses about $2,500 of value per year.
Correlation is not causation
A strong correlation does not prove that one variable causes the other. Ice cream sales and sunburns are positively correlated, but ice cream doesn't cause sunburns: hot, sunny weather drives both. Before claiming cause and effect, look for a hidden third variable like this, or for a controlled experiment.
Tip
To check a line of fit, count points above and below it. If nearly all of them are on one side, adjust the line.
Practice
What kind of correlation does this scatter plot show?
Which pair of variables would most likely have a negative correlation?
A line of fit for a plant's height (in cm) after weeks is . Predict the plant's height after weeks.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In the model from the previous problem, what does the slope mean?
A line of fit passes through the points and . Write its equation in slope-intercept form.
Enter an expression, e.g. 3x^2 - 2x + 1
The line of fit models the number of tickets still available, , for a concert days after sales open. Predict the number of tickets left after days.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A line of fit is . One data point is . What is the residual for this point?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In a city, the number of ice cream cones sold each day has a strong positive correlation with the number of people at the beach. Which conclusion is best supported?