Math Core

Lesson 7.5 · Inference for Means

Inference for two means

Do students who study with music score differently from students who study in silence? Does a new fertilizer grow taller plants than the standard one? Questions like these compare the means of two separate populations or two treatment groups. In this lesson you'll build confidence intervals and significance tests for the difference μ1−μ2\mu_1 - \mu_2 using two independent samples.

Two samples, not pairs

The procedures here are for two independent groups: two separate random samples, or two groups formed by randomly assigning subjects to treatments. Nothing links a particular observation in group 1 to a particular observation in group 2. If there is a built-in link (the same person measured twice, or matched pairs), use the paired procedures from the previous lesson instead.

The sampling distribution of the difference

The natural statistic is xˉ1−xˉ2\bar{x}_1 - \bar{x}_2. From Unit 5, when the samples are independent, its standard deviation is

σxˉ1−xˉ2=σ12n1+σ22n2.\sigma_{\bar{x}_1 - \bar{x}_2} = \sqrt{\frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}}.

Variances add, even though you're subtracting the means. Replacing each σ\sigma with its sample standard deviation gives the standard error.

Two-sample t procedures

The standard error of xˉ1−xˉ2\bar{x}_1 - \bar{x}_2 is

SE=s12n1+s22n2.SE = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}.
  • Interval for μ1−μ2\mu_1 - \mu_2: (xˉ1−xˉ2)±t∗⋅SE(\bar{x}_1 - \bar{x}_2) \pm t^* \cdot SE.
  • Test of H0:μ1−μ2=0H_0: \mu_1 - \mu_2 = 0: t=(xˉ1−xˉ2)−0SEt = \dfrac{(\bar{x}_1 - \bar{x}_2) - 0}{SE}.

Degrees of freedom

The two-sample tt statistic doesn't follow a tt-distribution exactly, but it is very close to one with a certain number of degrees of freedom. There are two accepted ways to choose dfdf.

  1. Technology (recommended): your calculator's 2-SampTInt and 2-SampTTest use a formula that usually gives a non-integer dfdf somewhere between the smaller of n1−1n_1 - 1 and n2−1n_2 - 1 and n1+n2−2n_1 + n_2 - 2. Choose "Pooled: No."
  2. Conservative: use dfdf equal to the smaller of n1−1n_1 - 1 and n2−1n_2 - 1. This gives a slightly larger t∗t^* (a wider interval, a larger P-value), so your conclusions err on the cautious side.

Both approaches earn full credit on the AP exam, as long as you say which one you used.

Common mistake

Don't pool the standard deviations, and don't use df=n1+n2−2df = n_1 + n_2 - 2 by habit. Pooling assumes the two populations have equal standard deviations, which you rarely know. AP Statistics uses the unpooled standard error shown above.

Conditions

Check the conditions for each group, plus independence between groups.

  • Random: two independent random samples, or two groups from a randomized experiment.
  • 10%: when sampling without replacement, each sample is at most 10% of its population. (Not needed for a randomized experiment that doesn't sample from a population.)
  • Normal/Large Sample: for each group, n≥30n \ge 30, or the population is approximately normal, or a graph of that group's data shows no strong skewness or outliers.

Here are the critical values used below.

dfdf90% (t∗t^*)95% (t∗t^*)99% (t∗t^*)
111.7962.2013.106
171.7402.1102.898
291.6992.0452.756
601.6712.0002.660

A two-sample confidence interval

Worked example: Music while studying

A teacher randomly assigns volunteers to study a passage either with music (n1=30n_1 = 30) or in silence (n2=35n_2 = 35), then gives them the same quiz. The silence group averages xˉ1=78.4\bar{x}_1 = 78.4 points with s1=9.2s_1 = 9.2; the music group averages xˉ2=73.1\bar{x}_2 = 73.1 with s2=11.5s_2 = 11.5. Construct and interpret a 95% confidence interval for μ1−μ2\mu_1 - \mu_2.

State. Estimate μ1−μ2\mu_1 - \mu_2, where μ1\mu_1 is the true mean quiz score for students like these who study in silence and μ2\mu_2 is the true mean for those who study with music, with 95% confidence.

Plan. Two-sample tt interval for μ1−μ2\mu_1 - \mu_2.

  • Random: students were randomly assigned to the two groups.
  • Normal/Large Sample: n1=30≥30n_1 = 30 \ge 30 and n2=35≥30n_2 = 35 \ge 30.

Do.

SE=9.2230+11.5235≈2.569,xˉ1−xˉ2=5.3.SE = \sqrt{\frac{9.2^2}{30} + \frac{11.5^2}{35}} \approx 2.569, \qquad \bar{x}_1 - \bar{x}_2 = 5.3.

Technology gives df≈62.7df \approx 62.7 and t∗≈1.999t^* \approx 1.999:

5.3±1.999(2.569)≈5.3±5.13=(0.17, 10.43).5.3 \pm 1.999(2.569) \approx 5.3 \pm 5.13 = (0.17,\ 10.43).

(With the conservative df=29df = 29, t∗=2.045t^* = 2.045 and the interval is about (0.05, 10.55)(0.05,\ 10.55).)

Conclude. We are 95% confident that the interval from about 0.170.17 to 10.4310.43 points captures the true difference in mean quiz scores (silence − music). Because the interval is entirely above 00, studying in silence appears to produce a higher mean score. And since this was a randomized experiment, we can say studying in silence causes the higher mean for students like these.

Tip

If a confidence interval for μ1−μ2\mu_1 - \mu_2 contains 00, then "no difference" is a plausible value, and a two-sided test at the matching α\alpha would fail to reject H0H_0. If 00 is outside the interval, the test would reject H0H_0.

A two-sample significance test

Worked example: Fertilizer and plant growth

A gardener randomly assigns 3838 seedlings to a new fertilizer (n1=18n_1 = 18) or the standard fertilizer (n2=20n_2 = 20). After six weeks, the new-fertilizer plants average xˉ1=24.6\bar{x}_1 = 24.6 cm with s1=4.1s_1 = 4.1 cm, and the standard-fertilizer plants average xˉ2=21.9\bar{x}_2 = 21.9 cm with s2=5.0s_2 = 5.0 cm. Dotplots of both groups show no strong skewness or outliers. Is there convincing evidence at α=0.05\alpha = 0.05 that the new fertilizer produces taller plants on average?

State. H0:μ1−μ2=0H_0: \mu_1 - \mu_2 = 0 and Ha:μ1−μ2>0H_a: \mu_1 - \mu_2 > 0, where μ1\mu_1 and μ2\mu_2 are the true mean heights (cm) of seedlings like these grown with the new and standard fertilizers.

Plan. Two-sample tt test for μ1−μ2\mu_1 - \mu_2.

  • Random: seedlings were randomly assigned to fertilizers.
  • Normal/Large Sample: both samples are under 30, but both dotplots show no strong skewness or outliers.

Do.

SE=4.1218+5.0220≈1.478,t=24.6−21.91.478≈1.83.SE = \sqrt{\frac{4.1^2}{18} + \frac{5.0^2}{20}} \approx 1.478, \qquad t = \frac{24.6 - 21.9}{1.478} \approx 1.83.

Technology gives df≈35.7df \approx 35.7 and P-value ≈0.038\approx 0.038. (Conservatively, df=17df = 17 gives P-value ≈0.043\approx 0.043. In the table, 1.831.83 lies between 1.7401.740 and 2.1102.110 in the df=17df = 17 row, so the P-value is between 0.0250.025 and 0.050.05.)

Conclude. Because the P-value of about 0.0380.038 is less than 0.050.05, we reject H0H_0. There is convincing evidence that the new fertilizer produces a greater mean height than the standard fertilizer for seedlings like these.

Scope of conclusions

Two questions decide what you can say:

  • Random sampling lets you generalize to the populations the samples came from.
  • Random assignment lets you conclude that the treatment caused the difference.

Both examples above were experiments with random assignment, so cause-and-effect conclusions are justified, but only for subjects like the ones in the study. If two groups are compared in an observational study (say, a random sample of city residents and a random sample of rural residents), you can generalize to the two populations, but you cannot claim causation.

Practice

Practice 1

Which situation calls for a two-sample tt procedure rather than a paired tt procedure?

Practice 2

Group 1 has n1=20n_1 = 20 and s1=6s_1 = 6. Group 2 has n2=25n_2 = 25 and s2=8s_2 = 8. Find the standard error of xˉ1−xˉ2\bar{x}_1 - \bar{x}_2 to 3 decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A random sample of 4040 students from School A averaged 88.588.5 on a statewide exam with s=12.4s = 12.4. An independent random sample of 4545 students from School B averaged 83.283.2 with s=10.8s = 10.8. Find the two-sample tt statistic for H0:μA−μB=0H_0: \mu_A - \mu_B = 0, to 2 decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

For the school comparison in the previous problem, the alternative is Ha:μA−μB≠0H_a: \mu_A - \mu_B \ne 0, and technology reports df≈77.9df \approx 77.9. Find the P-value to 4 decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A researcher compares battery life (hours) for two tablet models. Model 1: n1=15n_1 = 15, xˉ1=18.7\bar{x}_1 = 18.7, s1=3.2s_1 = 3.2. Model 2: n2=12n_2 = 12, xˉ2=16.1\bar{x}_2 = 16.1, s2=2.6s_2 = 2.6. The conditions are met, and technology reports df≈25.0df \approx 25.0, so t∗≈1.708t^* \approx 1.708 for 90% confidence. Construct a 90% confidence interval for μ1−μ2\mu_1 - \mu_2. Enter the endpoints to 2 decimal places, separated by a comma.

Separate answers with commas, e.g. 2, -5

Practice 6

For the tablet comparison in the previous problem (n1=15n_1 = 15, n2=12n_2 = 12), how many degrees of freedom does the conservative approach use?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A 95% confidence interval for μ1−μ2\mu_1 - \mu_2, the difference in mean daily screen time (hours) between randomly selected 10th graders and 12th graders, is (−0.8, 0.5)(-0.8,\ 0.5). Which conclusion is correct?

Practice 8

A health researcher takes a random sample of 5050 adults who drink coffee daily and a separate random sample of 5050 adults who never drink coffee. A two-sample tt test finds that the coffee drinkers have a significantly higher mean resting heart rate. Which conclusion is justified?