Math Core

Lesson 7.3 · Inference for Means

Significance tests for a mean

A pizza chain advertises a mean delivery time of 30 minutes. A group of customers suspects the real mean is longer. A confidence interval tells you what values of the mean are plausible; a significance test answers a sharper question: do the data give convincing evidence against a specific claimed value? This lesson carries the testing logic you used for proportions over to means.

Hypotheses about a mean

A significance test starts with two competing claims about the parameter μ\mu.

  • The null hypothesis H0:μ=μ0H_0: \mu = \mu_0 is the "nothing unusual" claim, where μ0\mu_0 is the claimed or historical value.
  • The alternative hypothesis HaH_a is what you suspect or hope to find evidence for. It is one of μ>μ0\mu > \mu_0, μ<μ0\mu < \mu_0 or μ≠μ0\mu \ne \mu_0.

Hypotheses are always about the population parameter μ\mu, never about xˉ\bar{x}. And the direction of HaH_a must come from the research question, decided before you look at the data. For the pizza chain, H0:μ=30H_0: \mu = 30 and Ha:μ>30H_a: \mu > 30, where μ\mu is the true mean delivery time in minutes.

The test statistic

The test asks: if H0H_0 were true, how surprising would a sample mean like ours be? You measure "how far" in standard errors.

One-sample t test for a mean

To test H0:μ=μ0H_0: \mu = \mu_0, compute

t=xˉ−μ0s/n.t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}}.

When H0H_0 is true and the conditions are met, tt follows a tt-distribution with df=n−1df = n - 1. The P-value is the probability, assuming H0H_0 is true, of getting a tt statistic at least as extreme as the one observed, in the direction(s) of HaH_a.

  • For Ha:μ>μ0H_a: \mu > \mu_0, the P-value is the area to the right of tt.
  • For Ha:μ<μ0H_a: \mu < \mu_0, it is the area to the left of tt.
  • For Ha:μ≠μ0H_a: \mu \ne \mu_0, it is twice the area in the tail beyond ∣t∣|t|.

The conditions are the same as for the one-sample tt interval: Random, 10% and Normal/Large Sample (n≥30n \ge 30, or a graph of the data shows no strong skewness or outliers).

Making a conclusion

Compare the P-value to the significance level α\alpha (usually 0.050.05 unless stated otherwise).

  • If P-value ≤α\le \alpha: reject H0H_0. There is convincing evidence for HaH_a.
  • If P-value >α> \alpha: fail to reject H0H_0. There is not convincing evidence for HaH_a.

Always write the conclusion in context, and always refer to the alternative hypothesis.

Common mistake

Failing to reject H0H_0 does not prove H0H_0 is true. It means the data are consistent with H0H_0, not that H0H_0 is correct. Never write "we accept H0H_0" or "the data prove the mean is 30 minutes."

A complete test: State, Plan, Do, Conclude

Worked example: Pizza delivery times

A random sample of 2020 deliveries from the pizza chain had a mean delivery time of xˉ=32.4\bar{x} = 32.4 minutes with s=4.8s = 4.8 minutes. A dotplot of the times shows no strong skewness or outliers. Is there convincing evidence at α=0.05\alpha = 0.05 that the true mean delivery time is greater than 30 minutes?

State. H0:μ=30H_0: \mu = 30 and Ha:μ>30H_a: \mu > 30, where μ\mu is the true mean delivery time (minutes) for this chain. Use α=0.05\alpha = 0.05.

Plan. One-sample tt test for μ\mu.

  • Random: the deliveries were randomly selected.
  • 10%: 2020 is less than 10% of all the chain's deliveries.
  • Normal/Large Sample: n=20<30n = 20 < 30, but the dotplot shows no strong skewness or outliers.

Do. The standard error is 4.820≈1.073\dfrac{4.8}{\sqrt{20}} \approx 1.073.

t=32.4−301.073≈2.236,df=19.t = \frac{32.4 - 30}{1.073} \approx 2.236, \qquad df = 19.

P-value =P(t>2.236)≈0.0188= P(t > 2.236) \approx 0.0188. (In the table, 2.2362.236 lies between 2.0932.093 and 2.5392.539 in the df=19df = 19 row, so the P-value is between 0.010.01 and 0.0250.025.)

Conclude. Because the P-value of 0.01880.0188 is less than α=0.05\alpha = 0.05, we reject H0H_0. There is convincing evidence that the true mean delivery time for this chain is greater than 30 minutes.

Here is a table excerpt for this lesson's examples and practice.

dfdftail 0.10tail 0.05tail 0.025tail 0.01tail 0.005
141.3451.7612.1452.6242.977
191.3281.7292.0932.5392.861
351.3061.6902.0302.4382.724
391.3041.6852.0232.4262.708

Interpreting a P-value

A P-value is a conditional probability. For the pizza example: "Assuming the true mean delivery time is 30 minutes, there is about a 0.01880.0188 probability of getting a sample mean of 32.432.4 minutes or more, just by chance, in a random sample of 2020 deliveries."

That's not the probability that H0H_0 is true. The calculation assumes H0H_0 is true.

Two-sided tests and confidence intervals

When the question asks whether the mean is different from a claimed value, use a two-sided alternative.

Worked example: Soda fill amounts

A bottling machine is supposed to fill cans with a mean of 1212 ounces. An inspector measures a random sample of 4040 cans and finds xˉ=11.92\bar{x} = 11.92 oz and s=0.21s = 0.21 oz. Test H0:μ=12H_0: \mu = 12 against Ha:μ≠12H_a: \mu \ne 12 at α=0.05\alpha = 0.05.

Solution. Conditions: random sample, 4040 cans is less than 10% of production, and n=40≥30n = 40 \ge 30.

SE=0.2140≈0.0332,t=11.92−120.0332≈−2.409,df=39.SE = \frac{0.21}{\sqrt{40}} \approx 0.0332, \qquad t = \frac{11.92 - 12}{0.0332} \approx -2.409, \qquad df = 39.

The P-value is 2⋅P(t<−2.409)≈2(0.0104)=0.02082 \cdot P(t < -2.409) \approx 2(0.0104) = 0.0208. Since 0.0208<0.050.0208 < 0.05, reject H0H_0. There is convincing evidence that the machine's true mean fill amount differs from 12 ounces.

A 95% confidence interval tells the same story: 11.92±2.023(0.0332)≈(11.853, 11.987)11.92 \pm 2.023(0.0332) \approx (11.853,\ 11.987). The value 1212 is not in the interval, which matches rejecting H0H_0 at α=0.05\alpha = 0.05. The interval adds something the test doesn't: the machine appears to be under-filling by roughly 0.010.01 to 0.150.15 ounces.

Tip

A two-sided test at significance level α\alpha and a confidence interval at level 1−α1 - \alpha always agree. If μ0\mu_0 is outside the interval, reject H0H_0; if it's inside, fail to reject. Use this to check your work.

Errors in context

As with proportions, a test can be wrong in two ways.

  • A Type I error is rejecting H0H_0 when it is actually true. Its probability is α\alpha.
  • A Type II error is failing to reject H0H_0 when HaH_a is actually true.

For the soda machine, a Type I error would mean concluding the machine's mean fill differs from 12 ounces when it really is 12, so the company might stop production for no reason. A Type II error would mean missing a real problem with the machine. Larger samples reduce the chance of a Type II error, which increases the power of the test.

Practice

Practice 1

A random sample of 3636 students has a mean reaction-time score of xˉ=103.2\bar{x} = 103.2 with s=9.6s = 9.6. Compute the test statistic for H0:μ=100H_0: \mu = 100.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

For the test in the previous problem, the alternative is Ha:μ>100H_a: \mu > 100 and t=2.0t = 2.0 with df=35df = 35. Find the P-value to 4 decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A city says its average household water use is 280 gallons per day. An environmental group believes the true average is lower and collects data from a random sample of households. Which hypotheses should the group test?

Practice 4

In a two-sided test of H0:μ=50H_0: \mu = 50 versus Ha:μ≠50H_a: \mu \ne 50, a random sample of 2020 gives t=−1.85t = -1.85. Find the P-value to 4 decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A one-sample tt test of H0:μ=8H_0: \mu = 8 versus Ha:μ>8H_a: \mu > 8 gives a P-value of 0.0310.031. Which conclusion is correct at α=0.05\alpha = 0.05?

Practice 6

A 95% confidence interval for the mean time (in minutes) customers spend in a store is (18.2, 23.6)(18.2,\ 23.6). Based only on this interval, what can you conclude about a two-sided test of H0:μ=25H_0: \mu = 25 at α=0.05\alpha = 0.05?

Practice 7

A sleep researcher claims teens average 88 hours of sleep on school nights. A random sample of 1515 teens averages xˉ=7.3\bar{x} = 7.3 hours with s=1.1s = 1.1 hours, and a dotplot shows no strong skew or outliers. For H0:μ=8H_0: \mu = 8 versus Ha:μ<8H_a: \mu < 8, find the tt statistic to 2 decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

Continue the sleep study from the previous problem (t≈−2.46t \approx -2.46, df=14df = 14). What is the conclusion at α=0.01\alpha = 0.01?