Math Core

Lesson 9.1 · Inference for Slopes

Confidence intervals for slope

In Unit 2 you fit a least-squares line to a set of data and interpreted its slope. But that slope came from one sample. A different random sample would give a different line and a different slope. In this lesson you'll build a confidence interval that estimates the slope of the true relationship in the whole population, using the same "statistic plus or minus margin of error" idea you used for proportions and means.

The population regression line

Imagine every individual in a population, not just the ones in your sample. For each value of xx, the responses yy have a distribution with some mean μy\mu_y. If those means fall on a straight line, that line is the population regression line:

μy=α+βx\mu_y = \alpha + \beta x

The slope β\beta (beta) and intercept α\alpha (alpha) are parameters. You never know them exactly. Your sample gives the least-squares line y^=a+bx\hat{y} = a + bx, and its slope bb is a statistic that estimates β\beta.

Definition

Population slope and sample slope

The population slope β\beta is the change in the mean response μy\mu_y for each one-unit increase in xx, for the whole population. The sample slope bb is the slope of the least-squares line computed from a sample. It is a point estimate of β\beta.

How much does the sample slope vary?

If you took many random samples of the same size and computed bb each time, the values of bb would form a sampling distribution. When the conditions below are met, that distribution is approximately Normal, centered at β\beta (so bb is an unbiased estimator), with standard deviation

σb=σσxn,\sigma_b = \frac{\sigma}{\sigma_x\sqrt{n}},

where σ\sigma is the standard deviation of the responses around the population line. Since σ\sigma is unknown, you estimate it with ss, the standard deviation of the residuals. That gives the standard error of the slope:

SEb=ssxn−1SE_b = \frac{s}{s_x\sqrt{n-1}}

You will almost never compute SEbSE_b by hand. On the AP exam it is given in computer output, in the row for the explanatory variable, under a heading like "SE Coef."

The formula still tells you what makes a slope estimate precise. SEbSE_b is smaller when the points hug the line (small ss), when the xx values are spread out (large sxs_x), and when the sample is large (large nn).

Conditions: LINER

Before using any inference procedure for slope, check five conditions. The acronym LINER helps you remember them.

Conditions for regression inference (LINER)

  • Linear: The true relationship between xx and yy is linear. Check that the scatterplot looks linear and the residual plot shows no curved pattern.
  • Independent: Individual observations are independent. When sampling without replacement, check the 10% condition: n≤10%n \le 10\% of the population.
  • Normal: For any fixed xx, the responses vary Normally around the population line. Check that a dotplot, histogram or Normal probability plot of the residuals shows no strong skew or outliers.
  • Equal SD: The standard deviation of yy is the same for every value of xx. Check that the residual plot has roughly equal vertical spread everywhere, with no "fan" or "funnel" shape.
  • Random: The data come from a random sample or a randomized experiment.

Reading computer output

Here is the kind of output you'll see. A coffee cart owner takes a random sample of 1010 days and records the daily high temperature (xx, in °F) and the number of iced drinks sold (yy).

x-axis: daily high temperature (°F). y-axis: iced drinks sold.Open in grapher →
PredictorCoefSE CoefTP
Constant-26.014114.7730-1.760.116
Temperature2.33130.194411.990.000

S = 5.2974, R-Sq = 94.7%

  • The least-squares line is drinks^=−26.0141+2.3313(temperature)\widehat{\text{drinks}} = -26.0141 + 2.3313(\text{temperature}). The slope b=2.3313b = 2.3313 is in the Temperature row, Coef column.
  • SEb=0.1944SE_b = 0.1944 is in the same row, SE Coef column. Don't grab 14.773014.7730 from the Constant row by mistake.
  • s=5.2974s = 5.2974 is the standard deviation of the residuals: actual sales typically differ from the predicted sales by about 5.35.3 drinks.

The confidence interval

t interval for the slope

b±t∗⋅SEb,df=n−2b \pm t^* \cdot SE_b, \qquad df = n - 2

Find t∗t^* from a tt distribution with n−2n - 2 degrees of freedom. You lose two degrees of freedom because the sample is used to estimate two things: the slope and the intercept.

Worked example: Building the interval

Use the coffee cart output to build a 95%95\% confidence interval for the slope of the population regression line.

With n=10n = 10, df=10−2=8df = 10 - 2 = 8, and t∗=2.306t^* = 2.306.

2.3313±2.306(0.1944)=2.3313±0.44832.3313 \pm 2.306(0.1944) = 2.3313 \pm 0.4483

The interval is about (1.883, 2.780)(1.883,\ 2.780).

Worked example: Interpreting the interval

Interpret the interval from the previous example, and decide whether it gives convincing evidence of a linear relationship.

Interpretation: We are 95%95\% confident that the interval from 1.8831.883 to 2.7802.780 captures the slope of the population regression line relating daily high temperature to iced drinks sold. That is, for each 11°F increase in the daily high temperature, the mean number of iced drinks sold increases by somewhere between about 1.91.9 and 2.82.8 drinks.

Evidence: Every value in the interval is positive, and 00 is not a plausible value for β\beta. So the data give convincing evidence of a positive linear relationship between temperature and iced drink sales.

Common mistake

Interpret the slope as a change in the mean (or predicted) response, not in each individual's response. "Each extra degree increases sales by 2 drinks" is wrong: on any given day, sales bounce around the line. Also, the interval is about the population slope β\beta. You already know bb exactly; there is no need to be "confident" about it.

Checking conditions with a residual plot

Here is the residual plot for the coffee cart data.

x-axis: daily high temperature (°F). y-axis: residual (drinks).Open in grapher →

Worked example: Checking LINER

Check the conditions for the coffee cart interval. A dotplot of the residuals (not shown) is roughly symmetric with no outliers.

  • Linear: The residual plot shows no leftover curved pattern; points are scattered above and below 00.
  • Independent: 1010 days is less than 10%10\% of all days the cart operates.
  • Normal: The dotplot of residuals shows no strong skew or outliers.
  • Equal SD: The vertical spread of the residuals is about the same from left to right. There is no fan shape.
  • Random: The days were a random sample.

All conditions are met, so the tt interval is appropriate.

Worked example: Working from output only

A random sample of 2222 used cars of one model gives mileage xx (thousands of miles) and price yy (thousands of dollars). The output shows b=−0.0852b = -0.0852 and SEb=0.0118SE_b = 0.0118. Construct and interpret a 99%99\% confidence interval for β\beta. Assume the conditions are met.

df=22−2=20df = 22 - 2 = 20, so t∗=2.845t^* = 2.845.

−0.0852±2.845(0.0118)=−0.0852±0.0336-0.0852 \pm 2.845(0.0118) = -0.0852 \pm 0.0336

The interval is about (−0.1188, −0.0516)(-0.1188,\ -0.0516). We are 99%99\% confident that for each additional 1,0001{,}000 miles, the mean price of this model decreases by between about $52 and $119.

Tip

If you're given an interval and need to recover the pieces: the slope bb is the midpoint, and the margin of error is half the width. Dividing the margin of error by t∗t^* gives SEbSE_b.

Practice

Practice 1

A researcher computes a confidence interval for the slope using a random sample of 1515 individuals. How many degrees of freedom should she use for t∗t^*?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

For a random sample of 2525 individuals, which critical value should be used for a 90%90\% confidence interval for the slope?

Practice 3

A random sample of 1818 homes in a city gives the output below, where xx is floor area (hundreds of square feet) and yy is price (thousands of dollars).

PredictorCoefSE CoefTP
Constant41.3722.851.810.089
Area8.621.356.390.000

S = 28.41, R-Sq = 71.8%

Assuming the conditions are met, find the lower endpoint of a 95%95\% confidence interval for the population slope. Round to the nearest hundredth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A 95%95\% confidence interval for the slope relating a student's hours of weekly practice (xx) to typing speed in words per minute (yy) is (1.4, 3.1)(1.4,\ 3.1). Which is the best interpretation?

Practice 5

A residual plot for a regression shows residuals close to 00 for small values of xx and spreading out more and more as xx increases, in a fan shape. Which condition for inference is most clearly violated?

Practice 6

A 90%90\% confidence interval for a population slope is (1.2, 3.4)(1.2,\ 3.4). What is the sample slope bb?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A 95%95\% confidence interval for a population slope, based on a random sample of 2020 individuals, is (0.84, 1.56)(0.84,\ 1.56). Find the standard error of the slope, SEbSE_b. Round to the nearest thousandth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

A scientist plans a study to estimate the slope relating fertilizer amount (xx) to plant height (yy). Which change to her design would tend to produce a narrower confidence interval for the slope, keeping the confidence level the same?