Math Core

Lesson 2.4 · Exploring Two-Variable Data

Least-squares regression

Correlation tells you how strong a linear relationship is. A regression line goes further: it gives you an equation you can use to predict the response from the explanatory variable. Of all the lines you could draw through a scatterplot, statistics picks one specific line, the least-squares regression line, and this lesson explains how it is chosen, how to find it, and how to interpret it.

Regression lines and predicted values

A regression line summarizes a linear relationship with an equation of the form

y^=a+bx\hat{y} = a + bx

where y^\hat{y} ("y-hat") is the predicted value of the response for a given xx, bb is the slope, and aa is the y-intercept. AP Statistics writes the intercept first, and always uses y^\hat{y} rather than yy, because the line gives predictions, not actual data values.

A student tracked the hours she studied for six quizzes and her scores.

Hours studied, xx112233445566
Quiz score, yy606066667070717180808585

Technology gives the regression line y^=55.2+4.8x\hat{y} = 55.2 + 4.8x.

x-axis: hours studied. y-axis: quiz score. The line is ŷ = 55.2 + 4.8x.Open in grapher →

For x=4x = 4 hours the line predicts y^=55.2+4.8(4)=74.4\hat{y} = 55.2 + 4.8(4) = 74.4 points. Her actual score was 7171, so the prediction was off by 71−74.4=−3.471 - 74.4 = -3.4 points. That difference, actual minus predicted, is called a residual. You'll study residuals in detail in the next lesson; for now, a residual is the vertical distance from a point to the line.

The least-squares criterion

Every line through the data leaves some residuals. Some are positive (points above the line) and some are negative (points below). To choose the "best" line, we square each residual, which makes them all positive and penalizes big misses heavily, and add them up.

Definition

Least-squares regression line

The least-squares regression line (LSRL) is the line that makes the sum of the squared residuals as small as possible. Its slope and intercept are

b=r sysxanda=yˉ−bxˉ.b = r \, \frac{s_y}{s_x} \qquad \text{and} \qquad a = \bar{y} - b\bar{x}.

Two consequences are worth remembering:

  • Since a=yˉ−bxˉa = \bar{y} - b\bar{x}, rearranging gives yˉ=a+bxˉ\bar{y} = a + b\bar{x}. The LSRL always passes through the point (xˉ,yˉ)(\bar{x}, \bar{y}).
  • The slope and rr always have the same sign, because sxs_x and sys_y are positive.

Worked example: The LSRL from summary statistics

For a class of students, the number of minutes spent reviewing notes before a quiz has mean xˉ=20\bar{x} = 20 and standard deviation sx=4s_x = 4. Quiz scores have mean yˉ=75\bar{y} = 75 and standard deviation sy=10s_y = 10. The correlation is r=0.6r = 0.6. Find the equation of the LSRL.

b=r sysx=0.6⋅104=1.5b = r\,\frac{s_y}{s_x} = 0.6 \cdot \frac{10}{4} = 1.5a=yˉ−bxˉ=75−1.5(20)=45a = \bar{y} - b\bar{x} = 75 - 1.5(20) = 45

The LSRL is score^=45+1.5(minutes)\widehat{\text{score}} = 45 + 1.5(\text{minutes}). Check: at x=20x = 20 the line gives 45+30=75=yˉ45 + 30 = 75 = \bar{y}, as it must.

Interpreting slope and intercept

Interpretations must use context and must describe predicted values.

AP-style interpretations

  • Slope: "For each additional 1 [unit of xx], the predicted [yy] increases (or decreases) by bb [units of yy]."
  • y-intercept: "When [xx] is 00, the predicted [yy] is aa [units]." Only interpret it if x=0x = 0 makes sense and is near the data.

For the review-minutes model: for each additional minute of reviewing notes, the predicted quiz score increases by 1.51.5 points. The intercept says a student who reviews for 00 minutes has a predicted score of 4545. Is that trustworthy? Minutes had a mean of 2020 and a standard deviation of 44, so 00 minutes is 55 standard deviations below the mean, far outside the data. Using the line there is extrapolation.

Common mistake

Extrapolation means using a regression line to predict for xx-values outside the range of the data. The linear pattern may not continue, so extrapolated predictions are often badly wrong. The study-hours line predicts that 1010 hours of studying gives a score of 55.2+4.8(10)=103.255.2 + 4.8(10) = 103.2, which is impossible on a 100100-point quiz.

The coefficient of determination

How good is the line at predicting? The coefficient of determination r2r^2 answers this.

Definition

Coefficient of determination

r2r^2 is the proportion (or percent) of the variation in the response variable that is accounted for by the least-squares regression line with the explanatory variable. It is literally the square of the correlation.

For the study-hours data, r≈0.982r \approx 0.982, so r2≈0.965r^2 \approx 0.965. Interpretation: about 96.5%96.5\% of the variation in quiz scores is accounted for by the least-squares regression line with hours studied. The remaining 3.5%3.5\% is due to other factors.

If you know r2r^2 and want rr, take the square root and give it the sign of the slope. An r2r^2 of 0.640.64 with a negative slope means r=−0.8r = -0.8.

Reading computer output

AP questions often give regression output instead of an equation. An engineer regressed the highway fuel efficiency (miles per gallon) of 88 cars on their weight (thousands of pounds):

PredictorCoefSE CoefTP
Constant52.80252.8022.4662.46621.4121.410.0000.000
Weight−7.468-7.4680.7410.741−10.08-10.080.0000.000

S=1.61311R-Sq=94.4%S = 1.61311 \qquad \text{R-Sq} = 94.4\%

The Coef column holds the intercept (next to "Constant") and the slope (next to the explanatory variable's name). The other columns are for inference, which comes much later in the course.

Worked example: Using regression output

Use the output above. (a) Write the LSRL. (b) Interpret the slope. (c) Predict the fuel efficiency of a car weighing 3,5003{,}500 pounds. (d) Find rr.

(a) mpg^=52.802−7.468(weight)\widehat{\text{mpg}} = 52.802 - 7.468(\text{weight}), with weight in thousands of pounds.

(b) For each additional 1,0001{,}000 pounds of weight, the predicted fuel efficiency decreases by about 7.477.47 miles per gallon.

(c) Weight is in thousands, so use x=3.5x = 3.5: y^=52.802−7.468(3.5)=52.802−26.138≈26.7\hat{y} = 52.802 - 7.468(3.5) = 52.802 - 26.138 \approx 26.7 mpg.

(d) r2=0.944r^2 = 0.944 and the slope is negative, so r=−0.944≈−0.972r = -\sqrt{0.944} \approx -0.972.

Regression toward the mean

The formula b=r sysxb = r\,\dfrac{s_y}{s_x} says something surprising. If xx is one standard deviation above xˉ\bar{x}, the predicted yy is only rr standard deviations above yˉ\bar{y}. Since ∣r∣≤1\lvert r \rvert \le 1, predictions are always closer to the mean (in standard units) than the xx-values they come from. Tall parents tend to have tall children, but on average not quite as tall. This is regression toward the mean, and it is where the word "regression" comes from.

Tip

The LSRL of yy on xx is not the same as the LSRL of xx on yy. The first minimizes vertical distances; the second would minimize horizontal ones. Always regress the response on the explanatory variable, and don't solve the equation backward to predict xx from yy.

Practice

Practice 1

Using the study-hours model y^=55.2+4.8x\hat{y} = 55.2 + 4.8x, predict the quiz score for a student who studies 3.53.5 hours.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

For a group of adult men, heights have xˉ=64\bar{x} = 64 inches and sx=3s_x = 3 inches, weights have yˉ=150\bar{y} = 150 pounds and sy=20s_y = 20 pounds, and r=0.45r = 0.45. What is the slope of the least-squares regression line for predicting weight from height?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

Using the same summary statistics (xˉ=64\bar{x} = 64, yˉ=150\bar{y} = 150, slope b=3b = 3), find the yy-intercept of the least-squares regression line.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

For the car output in the lesson, which is the correct interpretation of the slope?

Practice 5

A regression of the number of ice cream cones sold (yy) on the daily high temperature in °F (xx) produces this output. Which is the equation of the least-squares regression line?

PredictorCoefSE CoefTP
Constant−118.4-118.422.722.7−5.22-5.220.0000.000
Temperature4.734.730.310.3115.2615.260.0000.000
Practice 6

A least-squares regression of a car's resale value on its age has a negative slope and r2=0.81r^2 = 0.81. What is the correlation rr?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

For the study-hours data, r2≈0.965r^2 \approx 0.965. Which is the correct interpretation?

Practice 8

In the review-minutes example, r=0.6r = 0.6. A student reviewed for a number of minutes that is 22 standard deviations above the mean. Her predicted quiz score is how many standard deviations above the mean score?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.