Lesson 2.5 · Exploring Two-Variable Data
Residuals
A regression line almost never passes through every point. The leftover vertical distances, the residuals, are not just errors to shrug off: they tell you how far off a typical prediction is, whether a line was the right model in the first place, and which points have an outsized effect on the results.
What a residual is
Definition
Residual
A residual is the difference between an actual value of the response and the value predicted by the regression line:
A positive residual means the point lies above the line: the model underestimated the actual value. A negative residual means the point lies below the line and the model overestimated.
Recall the study-hours data and its LSRL, .
| Hours, | ||||||
|---|---|---|---|---|---|---|
| Score, | ||||||
| Predicted, | ||||||
| Residual, |
The residuals add to . That's no accident: the residuals from a least-squares line always sum to , so their mean is as well.
Worked example: Interpreting a residual
For the -hour quiz, interpret the residual in context.
The residual is . Her actual score was points lower than the score predicted by the least-squares regression line for a student who studied hours.
Common mistake
Subtract in the right order: actual minus predicted, . Reversing it flips every sign, so points above the line would look like they are below it. A memory aid: "AP" stands for Actual minus Predicted.
Residual plots
A residual plot graphs each residual against its -value (or against ). It takes the regression line and "flattens" it into the horizontal line residual , which magnifies any departures from the linear pattern.
Reading a residual plot
- No leftover pattern (random scatter above and below ): a linear model is appropriate.
- A curved pattern: the relationship is nonlinear, so a linear model is not appropriate.
- A fan shape (residuals spread out as increases): predictions are more precise for some -values than for others.
With only six points, the study-hours plot shows no clear curve, though the -hour quiz stands out below the others.
Worked example: A high r doesn't mean a line fits
A biologist measures a plant's height (cm) at weeks through : . The LSRL is with . Is a linear model appropriate?
The residuals are .
The residual plot has a clear U shape: the line overestimates in the middle weeks and underestimates at both ends. Even though , a linear model is not appropriate. The growth is curved, and the last lesson of this unit shows how to handle it.
Here's what a healthy residual plot looks like. A courier regressed the delivery cost (dollars) on the distance of deliveries (km), getting .
The residuals bounce above and below with no curve and no fan, so the linear model is appropriate.
The standard deviation of the residuals
Definition
Standard deviation of the residuals
The standard deviation of the residuals, , measures the typical size of a prediction error:
Interpretation: "The actual [] is typically about [units] away from the value predicted by the LSRL."
The divisor is rather than because the line used two estimated quantities, a slope and an intercept. Regression output labels this value S.
For the study-hours data, the squared residuals are , which add to . Then . The actual quiz scores are typically about points away from the scores predicted by the LSRL with hours studied.
Use together with : tells you what fraction of the variation the model accounts for, and tells you, in the units of , how big a typical miss is.
Outliers, high leverage and influential points
Unusual points affect a regression in different ways, and AP Statistics distinguishes three kinds.
- An outlier in regression is a point with a large residual: it lies far above or below the line.
- A high-leverage point has an -value much farther from than the other points. Such points have the potential to pull the line strongly.
- An influential point is one whose removal would substantially change the slope, the intercept or the correlation. Outliers and high-leverage points are often, but not always, influential.
Worked example: Which points are influential?
Eight points, , , , , , , , , have the LSRL with . Classify each added point.
(a) Add . The new LSRL is , and rises to .
(b) Add . The new LSRL is , and falls to .
(c) Add . The new LSRL is , and falls to .
Solutions:
(a) is far from the other -values, so the point has high leverage. But it lies almost exactly on the line's extension, so the slope and intercept barely change: it is not influential on the line. (It even strengthens .)
(b) Also high leverage, and it falls far below the pattern. It pulls the slope from down to : it is influential. Notice that its own residual, about , is no bigger than the residuals of several ordinary points, because the line bent toward it. Influential points can hide themselves in a residual plot.
(c) This point has , right at , so it has low leverage. It lies far above the line, so it is an outlier with a residual of about . The slope doesn't change at all, but the intercept rises and drops sharply. It is influential on the intercept and on , but not on the slope.
Tip
To test whether a point is influential, fit the line with and without it and compare. Never delete a point just because it is unusual. Investigate it: it may be a recording error, or it may be the most interesting observation in the data.
Practice
In the previous lesson, the LSRL for fuel efficiency (mpg) on weight (thousands of pounds) was . A car weighing pounds gets mpg. Find its residual, to two decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A regression line is . For the point with , the residual is . What is the actual -value?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A least-squares line is fit to five points. Four of the residuals are , , and . What is the fifth residual?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
The five residuals from the previous problem are , , , and . Compute the standard deviation of the residuals , to two decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Here is a residual plot from a linear regression of a population's size on the year. What does it suggest?
For the delivery-cost regression, the output shows . Which is the best interpretation?
A line fit to countries' data predicts life expectancy from health spending. One country has health spending far larger than all the others, and removing it changes the slope from to . Which term best describes this country's point?
A regression predicts the selling price of houses from their size. One house has a residual of +$38,000. Which statement is correct?