Lesson 7.5 · Inference for Means
Inference for two means
Do students who study with music score differently from students who study in silence? Does a new fertilizer grow taller plants than the standard one? Questions like these compare the means of two separate populations or two treatment groups. In this lesson you'll build confidence intervals and significance tests for the difference using two independent samples.
Two samples, not pairs
The procedures here are for two independent groups: two separate random samples, or two groups formed by randomly assigning subjects to treatments. Nothing links a particular observation in group 1 to a particular observation in group 2. If there is a built-in link (the same person measured twice, or matched pairs), use the paired procedures from the previous lesson instead.
The sampling distribution of the difference
The natural statistic is . From Unit 5, when the samples are independent, its standard deviation is
Variances add, even though you're subtracting the means. Replacing each with its sample standard deviation gives the standard error.
Two-sample t procedures
The standard error of is
- Interval for : .
- Test of : .
Degrees of freedom
The two-sample statistic doesn't follow a -distribution exactly, but it is very close to one with a certain number of degrees of freedom. There are two accepted ways to choose .
- Technology (recommended): your calculator's
2-SampTIntand2-SampTTestuse a formula that usually gives a non-integer somewhere between the smaller of and and . Choose "Pooled: No." - Conservative: use equal to the smaller of and . This gives a slightly larger (a wider interval, a larger P-value), so your conclusions err on the cautious side.
Both approaches earn full credit on the AP exam, as long as you say which one you used.
Common mistake
Don't pool the standard deviations, and don't use by habit. Pooling assumes the two populations have equal standard deviations, which you rarely know. AP Statistics uses the unpooled standard error shown above.
Conditions
Check the conditions for each group, plus independence between groups.
- Random: two independent random samples, or two groups from a randomized experiment.
- 10%: when sampling without replacement, each sample is at most 10% of its population. (Not needed for a randomized experiment that doesn't sample from a population.)
- Normal/Large Sample: for each group, , or the population is approximately normal, or a graph of that group's data shows no strong skewness or outliers.
Here are the critical values used below.
| 90% () | 95% () | 99% () | |
|---|---|---|---|
| 11 | 1.796 | 2.201 | 3.106 |
| 17 | 1.740 | 2.110 | 2.898 |
| 29 | 1.699 | 2.045 | 2.756 |
| 60 | 1.671 | 2.000 | 2.660 |
A two-sample confidence interval
Worked example: Music while studying
A teacher randomly assigns volunteers to study a passage either with music () or in silence (), then gives them the same quiz. The silence group averages points with ; the music group averages with . Construct and interpret a 95% confidence interval for .
State. Estimate , where is the true mean quiz score for students like these who study in silence and is the true mean for those who study with music, with 95% confidence.
Plan. Two-sample interval for .
- Random: students were randomly assigned to the two groups.
- Normal/Large Sample: and .
Do.
Technology gives and :
(With the conservative , and the interval is about .)
Conclude. We are 95% confident that the interval from about to points captures the true difference in mean quiz scores (silence − music). Because the interval is entirely above , studying in silence appears to produce a higher mean score. And since this was a randomized experiment, we can say studying in silence causes the higher mean for students like these.
Tip
If a confidence interval for contains , then "no difference" is a plausible value, and a two-sided test at the matching would fail to reject . If is outside the interval, the test would reject .
A two-sample significance test
Worked example: Fertilizer and plant growth
A gardener randomly assigns seedlings to a new fertilizer () or the standard fertilizer (). After six weeks, the new-fertilizer plants average cm with cm, and the standard-fertilizer plants average cm with cm. Dotplots of both groups show no strong skewness or outliers. Is there convincing evidence at that the new fertilizer produces taller plants on average?
State. and , where and are the true mean heights (cm) of seedlings like these grown with the new and standard fertilizers.
Plan. Two-sample test for .
- Random: seedlings were randomly assigned to fertilizers.
- Normal/Large Sample: both samples are under 30, but both dotplots show no strong skewness or outliers.
Do.
Technology gives and P-value . (Conservatively, gives P-value . In the table, lies between and in the row, so the P-value is between and .)
Conclude. Because the P-value of about is less than , we reject . There is convincing evidence that the new fertilizer produces a greater mean height than the standard fertilizer for seedlings like these.
Scope of conclusions
Two questions decide what you can say:
- Random sampling lets you generalize to the populations the samples came from.
- Random assignment lets you conclude that the treatment caused the difference.
Both examples above were experiments with random assignment, so cause-and-effect conclusions are justified, but only for subjects like the ones in the study. If two groups are compared in an observational study (say, a random sample of city residents and a random sample of rural residents), you can generalize to the two populations, but you cannot claim causation.
Practice
Which situation calls for a two-sample procedure rather than a paired procedure?
Group 1 has and . Group 2 has and . Find the standard error of to 3 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A random sample of students from School A averaged on a statewide exam with . An independent random sample of students from School B averaged with . Find the two-sample statistic for , to 2 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
For the school comparison in the previous problem, the alternative is , and technology reports . Find the P-value to 4 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A researcher compares battery life (hours) for two tablet models. Model 1: , , . Model 2: , , . The conditions are met, and technology reports , so for 90% confidence. Construct a 90% confidence interval for . Enter the endpoints to 2 decimal places, separated by a comma.
Separate answers with commas, e.g. 2, -5
For the tablet comparison in the previous problem (, ), how many degrees of freedom does the conservative approach use?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A 95% confidence interval for , the difference in mean daily screen time (hours) between randomly selected 10th graders and 12th graders, is . Which conclusion is correct?
A health researcher takes a random sample of adults who drink coffee daily and a separate random sample of adults who never drink coffee. A two-sample test finds that the coffee drinkers have a significantly higher mean resting heart rate. Which conclusion is justified?