Math Core

Lesson 2.5 · Exploring Two-Variable Data

Residuals

A regression line almost never passes through every point. The leftover vertical distances, the residuals, are not just errors to shrug off: they tell you how far off a typical prediction is, whether a line was the right model in the first place, and which points have an outsized effect on the results.

What a residual is

Definition

Residual

A residual is the difference between an actual value of the response and the value predicted by the regression line:

residual=y−y^=actual−predicted.\text{residual} = y - \hat{y} = \text{actual} - \text{predicted}.

A positive residual means the point lies above the line: the model underestimated the actual value. A negative residual means the point lies below the line and the model overestimated.

Recall the study-hours data and its LSRL, y^=55.2+4.8x\hat{y} = 55.2 + 4.8x.

Hours, xx112233445566
Score, yy606066667070717180808585
Predicted, y^\hat{y}60.060.064.864.869.669.674.474.479.279.284.084.0
Residual, y−y^y - \hat{y}001.21.20.40.4−3.4-3.40.80.81.01.0

The residuals add to 0+1.2+0.4−3.4+0.8+1.0=00 + 1.2 + 0.4 - 3.4 + 0.8 + 1.0 = 0. That's no accident: the residuals from a least-squares line always sum to 00, so their mean is 00 as well.

Worked example: Interpreting a residual

For the 44-hour quiz, interpret the residual in context.

The residual is 71−74.4=−3.471 - 74.4 = -3.4. Her actual score was 3.43.4 points lower than the score predicted by the least-squares regression line for a student who studied 44 hours.

Common mistake

Subtract in the right order: actual minus predicted, y−y^y - \hat{y}. Reversing it flips every sign, so points above the line would look like they are below it. A memory aid: "AP" stands for Actual minus Predicted.

Residual plots

A residual plot graphs each residual against its xx-value (or against y^\hat{y}). It takes the regression line and "flattens" it into the horizontal line residual =0= 0, which magnifies any departures from the linear pattern.

Residual plot for the study-hours data. x-axis: hours studied. y-axis: residual (points).Open in grapher →

Reading a residual plot

  • No leftover pattern (random scatter above and below 00): a linear model is appropriate.
  • A curved pattern: the relationship is nonlinear, so a linear model is not appropriate.
  • A fan shape (residuals spread out as xx increases): predictions are more precise for some xx-values than for others.

With only six points, the study-hours plot shows no clear curve, though the 44-hour quiz stands out below the others.

Worked example: A high r doesn't mean a line fits

A biologist measures a plant's height yy (cm) at weeks x=1x = 1 through 88: 2,3,5,8,12,17,23,302, 3, 5, 8, 12, 17, 23, 30. The LSRL is y^=−5.5+4x\hat{y} = -5.5 + 4x with r≈0.97r \approx 0.97. Is a linear model appropriate?

x-axis: week. y-axis: plant height (cm). The line looks close, but the points bend.Open in grapher →

The residuals are 3.5,0.5,−1.5,−2.5,−2.5,−1.5,0.5,3.53.5, 0.5, -1.5, -2.5, -2.5, -1.5, 0.5, 3.5.

Residual plot for the plant data: a clear U shape.Open in grapher →

The residual plot has a clear U shape: the line overestimates in the middle weeks and underestimates at both ends. Even though r≈0.97r \approx 0.97, a linear model is not appropriate. The growth is curved, and the last lesson of this unit shows how to handle it.

Here's what a healthy residual plot looks like. A courier regressed the delivery cost (dollars) on the distance of 88 deliveries (km), getting y^=0.70+0.392x\hat{y} = 0.70 + 0.392x.

Residual plot for delivery cost. x-axis: distance (km). y-axis: residual (dollars). No leftover pattern.Open in grapher →

The residuals bounce above and below 00 with no curve and no fan, so the linear model is appropriate.

The standard deviation of the residuals

Definition

Standard deviation of the residuals

The standard deviation of the residuals, ss, measures the typical size of a prediction error:

s=∑(yi−y^i)2n−2=∑residuals2n−2.s = \sqrt{\frac{\sum (y_i - \hat{y}_i)^2}{n - 2}} = \sqrt{\frac{\sum \text{residuals}^2}{n - 2}}.

Interpretation: "The actual [yy] is typically about ss [units] away from the value predicted by the LSRL."

The divisor is n−2n - 2 rather than n−1n - 1 because the line used two estimated quantities, a slope and an intercept. Regression output labels this value S.

For the study-hours data, the squared residuals are 0,1.44,0.16,11.56,0.64,10, 1.44, 0.16, 11.56, 0.64, 1, which add to 14.814.8. Then s=14.8/4=3.7≈1.92s = \sqrt{14.8 / 4} = \sqrt{3.7} \approx 1.92. The actual quiz scores are typically about 1.921.92 points away from the scores predicted by the LSRL with hours studied.

Use ss together with r2r^2: r2r^2 tells you what fraction of the variation the model accounts for, and ss tells you, in the units of yy, how big a typical miss is.

Outliers, high leverage and influential points

Unusual points affect a regression in different ways, and AP Statistics distinguishes three kinds.

  • An outlier in regression is a point with a large residual: it lies far above or below the line.
  • A high-leverage point has an xx-value much farther from xˉ\bar{x} than the other points. Such points have the potential to pull the line strongly.
  • An influential point is one whose removal would substantially change the slope, the intercept or the correlation. Outliers and high-leverage points are often, but not always, influential.

Worked example: Which points are influential?

Eight points, (1,3)(1, 3), (2,4)(2, 4), (3,6)(3, 6), (4,6)(4, 6), (5,8)(5, 8), (6,9)(6, 9), (7,10)(7, 10), (8,12)(8, 12), have the LSRL y^=1.68+1.24x\hat{y} = 1.68 + 1.24x with r=0.99r = 0.99. Classify each added point.

(a) Add (20,27)(20, 27). The new LSRL is y^=1.55+1.27x\hat{y} = 1.55 + 1.27x, and rr rises to 0.9990.999.

(b) Add (20,10)(20, 10). The new LSRL is y^=5.37+0.35x\hat{y} = 5.37 + 0.35x, and rr falls to 0.660.66.

The point (20, 10) drags the regression line down toward itself.Open in grapher →

(c) Add (4.5,15)(4.5, 15). The new LSRL is y^=2.54+1.24x\hat{y} = 2.54 + 1.24x, and rr falls to 0.740.74.

Solutions:

(a) x=20x = 20 is far from the other xx-values, so the point has high leverage. But it lies almost exactly on the line's extension, so the slope and intercept barely change: it is not influential on the line. (It even strengthens rr.)

(b) Also high leverage, and it falls far below the pattern. It pulls the slope from 1.241.24 down to 0.350.35: it is influential. Notice that its own residual, about −2.4-2.4, is no bigger than the residuals of several ordinary points, because the line bent toward it. Influential points can hide themselves in a residual plot.

(c) This point has x=4.5x = 4.5, right at xˉ\bar{x}, so it has low leverage. It lies far above the line, so it is an outlier with a residual of about 6.96.9. The slope doesn't change at all, but the intercept rises and rr drops sharply. It is influential on the intercept and on rr, but not on the slope.

Tip

To test whether a point is influential, fit the line with and without it and compare. Never delete a point just because it is unusual. Investigate it: it may be a recording error, or it may be the most interesting observation in the data.

Practice

Practice 1

In the previous lesson, the LSRL for fuel efficiency (mpg) on weight (thousands of pounds) was mpg^=52.802−7.468(weight)\widehat{\text{mpg}} = 52.802 - 7.468(\text{weight}). A car weighing 3,0003{,}000 pounds gets 3030 mpg. Find its residual, to two decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

A regression line is y^=12+0.8x\hat{y} = 12 + 0.8x. For the point with x=15x = 15, the residual is −2.5-2.5. What is the actual yy-value?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A least-squares line is fit to five points. Four of the residuals are 0.40.4, 0.70.7, −2.0-2.0 and 0.30.3. What is the fifth residual?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

The five residuals from the previous problem are 0.40.4, 0.70.7, −2.0-2.0, 0.30.3 and 0.60.6. Compute the standard deviation of the residuals ss, to two decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

Here is a residual plot from a linear regression of a population's size on the year. What does it suggest?

x-axis: year. y-axis: residual (thousands).Open in grapher →
Practice 6

For the delivery-cost regression, the output shows S=0.72S = 0.72. Which is the best interpretation?

Practice 7

A line fit to 1010 countries' data predicts life expectancy from health spending. One country has health spending far larger than all the others, and removing it changes the slope from 0.90.9 to 2.42.4. Which term best describes this country's point?

Practice 8

A regression predicts the selling price of houses from their size. One house has a residual of +$38,000. Which statement is correct?