Lesson 7.3 · Inference for Means
Significance tests for a mean
A pizza chain advertises a mean delivery time of 30 minutes. A group of customers suspects the real mean is longer. A confidence interval tells you what values of the mean are plausible; a significance test answers a sharper question: do the data give convincing evidence against a specific claimed value? This lesson carries the testing logic you used for proportions over to means.
Hypotheses about a mean
A significance test starts with two competing claims about the parameter .
- The null hypothesis is the "nothing unusual" claim, where is the claimed or historical value.
- The alternative hypothesis is what you suspect or hope to find evidence for. It is one of , or .
Hypotheses are always about the population parameter , never about . And the direction of must come from the research question, decided before you look at the data. For the pizza chain, and , where is the true mean delivery time in minutes.
The test statistic
The test asks: if were true, how surprising would a sample mean like ours be? You measure "how far" in standard errors.
One-sample t test for a mean
To test , compute
When is true and the conditions are met, follows a -distribution with . The P-value is the probability, assuming is true, of getting a statistic at least as extreme as the one observed, in the direction(s) of .
- For , the P-value is the area to the right of .
- For , it is the area to the left of .
- For , it is twice the area in the tail beyond .
The conditions are the same as for the one-sample interval: Random, 10% and Normal/Large Sample (, or a graph of the data shows no strong skewness or outliers).
Making a conclusion
Compare the P-value to the significance level (usually unless stated otherwise).
- If P-value : reject . There is convincing evidence for .
- If P-value : fail to reject . There is not convincing evidence for .
Always write the conclusion in context, and always refer to the alternative hypothesis.
Common mistake
Failing to reject does not prove is true. It means the data are consistent with , not that is correct. Never write "we accept " or "the data prove the mean is 30 minutes."
A complete test: State, Plan, Do, Conclude
Worked example: Pizza delivery times
A random sample of deliveries from the pizza chain had a mean delivery time of minutes with minutes. A dotplot of the times shows no strong skewness or outliers. Is there convincing evidence at that the true mean delivery time is greater than 30 minutes?
State. and , where is the true mean delivery time (minutes) for this chain. Use .
Plan. One-sample test for .
- Random: the deliveries were randomly selected.
- 10%: is less than 10% of all the chain's deliveries.
- Normal/Large Sample: , but the dotplot shows no strong skewness or outliers.
Do. The standard error is .
P-value . (In the table, lies between and in the row, so the P-value is between and .)
Conclude. Because the P-value of is less than , we reject . There is convincing evidence that the true mean delivery time for this chain is greater than 30 minutes.
Here is a table excerpt for this lesson's examples and practice.
| tail 0.10 | tail 0.05 | tail 0.025 | tail 0.01 | tail 0.005 | |
|---|---|---|---|---|---|
| 14 | 1.345 | 1.761 | 2.145 | 2.624 | 2.977 |
| 19 | 1.328 | 1.729 | 2.093 | 2.539 | 2.861 |
| 35 | 1.306 | 1.690 | 2.030 | 2.438 | 2.724 |
| 39 | 1.304 | 1.685 | 2.023 | 2.426 | 2.708 |
Interpreting a P-value
A P-value is a conditional probability. For the pizza example: "Assuming the true mean delivery time is 30 minutes, there is about a probability of getting a sample mean of minutes or more, just by chance, in a random sample of deliveries."
That's not the probability that is true. The calculation assumes is true.
Two-sided tests and confidence intervals
When the question asks whether the mean is different from a claimed value, use a two-sided alternative.
Worked example: Soda fill amounts
A bottling machine is supposed to fill cans with a mean of ounces. An inspector measures a random sample of cans and finds oz and oz. Test against at .
Solution. Conditions: random sample, cans is less than 10% of production, and .
The P-value is . Since , reject . There is convincing evidence that the machine's true mean fill amount differs from 12 ounces.
A 95% confidence interval tells the same story: . The value is not in the interval, which matches rejecting at . The interval adds something the test doesn't: the machine appears to be under-filling by roughly to ounces.
Tip
A two-sided test at significance level and a confidence interval at level always agree. If is outside the interval, reject ; if it's inside, fail to reject. Use this to check your work.
Errors in context
As with proportions, a test can be wrong in two ways.
- A Type I error is rejecting when it is actually true. Its probability is .
- A Type II error is failing to reject when is actually true.
For the soda machine, a Type I error would mean concluding the machine's mean fill differs from 12 ounces when it really is 12, so the company might stop production for no reason. A Type II error would mean missing a real problem with the machine. Larger samples reduce the chance of a Type II error, which increases the power of the test.
Practice
A random sample of students has a mean reaction-time score of with . Compute the test statistic for .
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
For the test in the previous problem, the alternative is and with . Find the P-value to 4 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A city says its average household water use is 280 gallons per day. An environmental group believes the true average is lower and collects data from a random sample of households. Which hypotheses should the group test?
In a two-sided test of versus , a random sample of gives . Find the P-value to 4 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A one-sample test of versus gives a P-value of . Which conclusion is correct at ?
A 95% confidence interval for the mean time (in minutes) customers spend in a store is . Based only on this interval, what can you conclude about a two-sided test of at ?
A sleep researcher claims teens average hours of sleep on school nights. A random sample of teens averages hours with hours, and a dotplot shows no strong skew or outliers. For versus , find the statistic to 2 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Continue the sleep study from the previous problem (, ). What is the conclusion at ?