Math Core

Lesson 9.2 · Inference for Slopes

Significance tests for slope

Every scatterplot of sample data shows some slope, even when the two variables have nothing to do with each other in the population. A significance test for slope asks the key question: is the sample slope far enough from 00 that chance alone is an unlikely explanation? If so, you have convincing evidence of a real linear relationship between the variables.

Hypotheses

If there is no linear relationship between xx and yy in the population, then the mean response μy\mu_y doesn't change as xx changes, so the population slope is β=0\beta = 0. That is almost always the null hypothesis.

Hypotheses for a test about slope

H0:β=0H_0: \beta = 0

Ha:β>0orHa:β<0orHa:β≠0H_a: \beta > 0 \quad \text{or} \quad H_a: \beta < 0 \quad \text{or} \quad H_a: \beta \ne 0

Here β\beta is the slope of the population regression line relating xx to yy. Define it in context, and choose the alternative before you look at the data, based on the research question.

A one-sided alternative such as β>0\beta > 0 matches a question like "Is there a positive linear relationship?" A two-sided alternative matches "Is there any linear relationship?"

The test statistic

The test follows the same pattern as every tt test you've seen: (statistic minus null value) divided by standard error.

t test for the slope

t=b−β0SEb,df=n−2t = \frac{b - \beta_0}{SE_b}, \qquad df = n - 2

where β0\beta_0 is the null value (usually 00). The P-value is the probability of getting a tt statistic at least this extreme, in the direction of HaH_a, if H0H_0 is true.

The conditions are the same LINER conditions from the previous lesson: Linear, Independent (10% condition), Normal residuals, Equal SD, and Random.

What the output already tells you

Computer output does most of the arithmetic. In the explanatory variable's row, the T column is b/SEbb / SE_b and the P column is the P-value for the test of H0:β=0H_0: \beta = 0 against the two-sided alternative Ha:β≠0H_a: \beta \ne 0.

  • For a two-sided test, use the P-value as printed.
  • For a one-sided test, if bb is on the side that HaH_a predicts, divide the printed P-value by 2.
  • If bb is on the opposite side from HaH_a (for example, bb is negative but Ha:β>0H_a: \beta > 0), the P-value is greater than 0.50.5 and the data give no support to HaH_a.

Common mistake

The P value in the output is almost always two-sided. Forgetting to halve it for a one-sided test is the most common mistake on these questions. Also, only use the printed T and P when the null value is 00. For any other null value, compute tt yourself.

A complete significance test

Worked example: Social media and sleep

A school counselor suspects that students who spend more time on social media get less sleep. She selects a random sample of 1414 students at her large high school and records each student's average daily social media use (xx, hours) and average nightly sleep (yy, hours).

x-axis: daily social media use (hours). y-axis: nightly sleep (hours).Open in grapher →
PredictorCoefSE CoefTP
Constant7.76080.258829.990.000
Social media-0.23790.0735-3.240.007

S = 0.4563, R-Sq = 46.6%

Do the data give convincing evidence at the α=0.05\alpha = 0.05 level of a negative linear relationship? A residual plot shows no pattern and a dotplot of residuals shows no strong skew.

State. H0:β=0H_0: \beta = 0 and Ha:β<0H_a: \beta < 0, where β\beta is the slope of the population regression line relating nightly sleep to daily social media use for all students at the school. Use α=0.05\alpha = 0.05.

Plan. tt test for the slope.

  • Linear: The scatterplot is roughly linear and the residual plot shows no curve.
  • Independent: 1414 is less than 10%10\% of the students at a large high school.
  • Normal: The dotplot of residuals shows no strong skew or outliers.
  • Equal SD: The residual plot shows roughly equal spread for all xx.
  • Random: Random sample of students.

Do. From the output, t=−0.2379−00.0735=−3.24t = \dfrac{-0.2379 - 0}{0.0735} = -3.24 with df=14−2=12df = 14 - 2 = 12. The printed P-value 0.0070.007 is two-sided. Since bb is negative, as HaH_a predicts, the one-sided P-value is 0.007/2≈0.00350.007 / 2 \approx 0.0035.

Conclude. Because the P-value 0.00350.0035 is less than α=0.05\alpha = 0.05, reject H0H_0. There is convincing evidence of a negative linear relationship between daily social media use and nightly sleep for students at this school.

Because this was a random sample, not an experiment, the conclusion is about association. It does not show that social media use causes less sleep.

When the evidence isn't convincing

Worked example: Practice and free throws

A coach takes a random sample of 1010 players in a large youth league and records hours of free-throw practice per week (xx) and free throws made out of 2020 attempts (yy). She wants to know whether there is any linear relationship. Assume the conditions are met.

PredictorCoefSE CoefTP
Constant11.64851.83916.330.000
Practice0.48480.25881.870.098

S = 2.3507, R-Sq = 30.5%

H0:β=0H_0: \beta = 0, Ha:β≠0H_a: \beta \ne 0, where β\beta is the slope of the population regression line relating free throws made to weekly practice hours. With α=0.05\alpha = 0.05:

t=0.48480.2588≈1.87t = \dfrac{0.4848}{0.2588} \approx 1.87, df=8df = 8, and the two-sided P-value is 0.0980.098.

Since 0.098>0.050.098 > 0.05, fail to reject H0H_0. There is not convincing evidence of a linear relationship between practice hours and free throws made for players in this league.

Notice what "fail to reject" does not say. It doesn't prove that β=0\beta = 0; with only 1010 players, the test may simply lack power to detect a real but modest relationship.

Tip

Switching to a one-sided test after seeing the data would halve the P-value here to about 0.0490.049, just under 0.050.05. That is exactly why the alternative must be chosen from the research question before looking at the results.

Testing a nonzero slope

Worked example: Checking a new thermometer

An engineer compares a new digital thermometer (yy) with a lab standard (xx) on a random sample of 1616 water baths. If the new thermometer is accurate, the slope should be 11. The output gives b=1.08b = 1.08 and SEb=0.035SE_b = 0.035. Test H0:β=1H_0: \beta = 1 against Ha:β≠1H_a: \beta \ne 1 at α=0.05\alpha = 0.05. Assume the conditions are met.

The printed T and P test β=0\beta = 0, so compute tt yourself:

t=1.08−10.035≈2.29,df=14t = \frac{1.08 - 1}{0.035} \approx 2.29, \qquad df = 14

The two-sided P-value is about 0.0380.038. Since 0.038<0.050.038 < 0.05, reject H0H_0. There is convincing evidence that the true slope differs from 11, so the new thermometer is not tracking the standard perfectly.

Tests and intervals agree

A two-sided test at level α\alpha and a C%C\% confidence interval with C=100(1−α)C = 100(1 - \alpha) give matching conclusions. If the interval does not contain 00, a two-sided test of H0:β=0H_0: \beta = 0 rejects H0H_0. If the interval does contain 00, the test fails to reject. The interval gives extra information: a set of plausible values for the size of the slope.

Worked example: Using an interval to decide

A 95%95\% confidence interval for the slope relating a car's engine size (liters) to its highway fuel economy (mpg) is (−6.8, −2.1)(-6.8,\ -2.1). What does it say about a test of H0:β=0H_0: \beta = 0 versus Ha:β≠0H_a: \beta \ne 0 at α=0.05\alpha = 0.05?

The interval does not contain 00; every plausible value of β\beta is negative. So the test would reject H0H_0: there is convincing evidence of a (negative) linear relationship between engine size and highway fuel economy.

Practice

Practice 1

A researcher wants to test whether there is a positive linear relationship between the number of hours a phone has been charging (xx) and its battery percentage (yy). Which hypotheses are correct?

Practice 2

Computer output for a regression shows a slope of 0.4150.415 with a standard error of 0.1620.162. Find the test statistic for H0:β=0H_0: \beta = 0. Round to the nearest hundredth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A significance test for slope is based on a random sample of 2727 individuals. How many degrees of freedom does the tt distribution have?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

Regression output relating xx to yy shows a slope of 1.361.36 in the explanatory variable's row, with P =0.084= 0.084. You are testing H0:β=0H_0: \beta = 0 against Ha:β>0H_a: \beta > 0. What is the P-value for your test?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A test of H0:β=0H_0: \beta = 0 versus Ha:β≠0H_a: \beta \ne 0, relating hours of daylight (xx) to daily electricity use (yy) for a random sample of homes, gives a P-value of 0.0230.023. Using α=0.01\alpha = 0.01, which conclusion is correct?

Practice 6

A biologist tests H0:β=0H_0: \beta = 0 versus Ha:β>0H_a: \beta > 0, where β\beta is the slope of the population regression line relating water temperature to the growth rate of a species of algae. Which describes a Type II error in this context?

Practice 7

A theory predicts that the population slope relating xx to yy is 22. From a random sample of 2525 individuals, b=2.46b = 2.46 and SEb=0.19SE_b = 0.19. Find the test statistic for H0:β=2H_0: \beta = 2 versus Ha:β≠2H_a: \beta \ne 2. Round to the nearest hundredth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

A 95%95\% confidence interval for the slope relating a runner's weekly mileage (xx) to her 5K time in minutes (yy) is (−0.31, 0.07)(-0.31,\ 0.07). What can you conclude about a test of H0:β=0H_0: \beta = 0 versus Ha:β≠0H_a: \beta \ne 0 at α=0.05\alpha = 0.05?