Math Core

Lesson 6.4 · Inference for Proportions

Inference for two proportions

Many of the most interesting questions in statistics are comparisons. Does a new reminder text raise vaccination rates compared with the old one? Are seniors more likely than juniors to have a job? To answer questions like these, you estimate and test the difference between two population proportions, p1−p2p_1 - p_2.

The sampling distribution of a difference

Suppose you take independent random samples from two populations (or randomly assign subjects to two treatments). The natural statistic is p^1−p^2\hat{p}_1 - \hat{p}_2. From Unit 5, its sampling distribution has:

  • center p1−p2p_1 - p_2 (so p^1−p^2\hat{p}_1 - \hat{p}_2 is unbiased),
  • standard deviation p1(1−p1)n1+p2(1−p2)n2\sqrt{\dfrac{p_1(1-p_1)}{n_1} + \dfrac{p_2(1-p_2)}{n_2}} (variances add, standard deviations do not),
  • approximately Normal shape when all four expected counts are at least 10.

Everything in this lesson builds on those three facts.

Conditions

  1. Random: two independent random samples, or two groups formed by random assignment in an experiment.
  2. 10%: when sampling without replacement, n1≤0.10N1n_1 \le 0.10N_1 and n2≤0.10N2n_2 \le 0.10N_2. (Not needed for a randomized experiment that does not sample from a population.)
  3. Large Counts: at least 10 successes and 10 failures in each group. For an interval, use the observed counts n1p^1n_1\hat{p}_1, n1(1−p^1)n_1(1-\hat{p}_1), n2p^2n_2\hat{p}_2, n2(1−p^2)n_2(1-\hat{p}_2). For a test, use the pooled proportion (defined below).

Confidence interval for p1−p2p_1 - p_2

Two-sample z-interval for a difference in proportions

(p^1−p^2)±z∗p^1(1−p^1)n1+p^2(1−p^2)n2(\hat{p}_1 - \hat{p}_2) \pm z^* \sqrt{\frac{\hat{p}_1(1-\hat{p}_1)}{n_1} + \frac{\hat{p}_2(1-\hat{p}_2)}{n_2}}

The critical values are the same as before: z∗=1.645z^* = 1.645 for 90%, 1.9601.960 for 95%, and 2.5762.576 for 99%.

The most important question to ask of a two-proportion interval is whether it contains 0. If every value in the interval is positive, you have convincing evidence that p1>p2p_1 > p_2. If every value is negative, p1<p2p_1 < p_2. If the interval contains 0, "no difference" is plausible.

Worked example: A 95% interval for a difference

A researcher randomly selects 150 students from a large urban district and 160 from a large rural district. In the urban sample, 45 students walk or bike to school; in the rural sample, 30 do. Construct and interpret a 95% confidence interval for the difference in the proportions of all students who walk or bike (urban minus rural).

State: Estimate pU−pRp_U - p_R, where pUp_U and pRp_R are the proportions of all students in the urban and rural districts who walk or bike to school, with 95% confidence.

Plan: Two-sample zz-interval for pU−pRp_U - p_R.

  • Random: independent random samples from each district.
  • 10%: 150 and 160 are each less than 10% of the students in a large district.
  • Large Counts: 45,105,30,13045, 105, 30, 130 are all at least 10.

Do: p^U=45150=0.30\hat{p}_U = \dfrac{45}{150} = 0.30 and p^R=30160=0.1875\hat{p}_R = \dfrac{30}{160} = 0.1875.

SE=0.30(0.70)150+0.1875(0.8125)160≈0.04850,(0.30−0.1875)±1.96(0.04850)=0.1125±0.0951,⇒(0.017, 0.208).\begin{aligned} SE &= \sqrt{\frac{0.30(0.70)}{150} + \frac{0.1875(0.8125)}{160}} \approx 0.04850, \\ (0.30 - 0.1875) &\pm 1.96(0.04850) = 0.1125 \pm 0.0951, \\ &\Rightarrow (0.017,\ 0.208). \end{aligned}

Conclude: We are 95% confident that the interval from 0.017 to 0.208 captures the true difference in proportions (urban minus rural) of students who walk or bike to school. Because the entire interval is above 0, there is convincing evidence that a higher proportion of urban students walk or bike.

Significance test for p1−p2p_1 - p_2

The null hypothesis is almost always "no difference": H0:p1−p2=0H_0: p_1 - p_2 = 0, which is the same as H0:p1=p2H_0: p_1 = p_2. The alternative can be p1−p2>0p_1 - p_2 > 0, p1−p2<0p_1 - p_2 < 0, or p1−p2≠0p_1 - p_2 \ne 0.

If H0H_0 is true, the two populations share one common proportion. Your best estimate of it combines both samples into a pooled (combined) proportion:

p^C=X1+X2n1+n2=total successestotal sample size.\hat{p}_C = \frac{X_1 + X_2}{n_1 + n_2} = \frac{\text{total successes}}{\text{total sample size}}.

Two-sample z-test for a difference in proportions

z=(p^1−p^2)−0p^C(1−p^C)(1n1+1n2)z = \frac{(\hat{p}_1 - \hat{p}_2) - 0}{\sqrt{\hat{p}_C(1-\hat{p}_C)\left(\dfrac{1}{n_1} + \dfrac{1}{n_2}\right)}}

Check Large Counts with n1p^Cn_1\hat{p}_C, n1(1−p^C)n_1(1-\hat{p}_C), n2p^Cn_2\hat{p}_C, n2(1−p^C)n_2(1-\hat{p}_C).

Common mistake

Use the pooled proportion only for a test, where H0H_0 says the proportions are equal. For a confidence interval, you are not assuming they are equal, so use the separate p^1\hat{p}_1 and p^2\hat{p}_2 in the standard error. Mixing these up is the most common error on this topic.

Worked example: A test in a randomized experiment

A clinic randomly assigns 400 patients who are due for a checkup to receive one of two reminder messages. Of the 210 who got a personalized text, 63 booked an appointment within a week. Of the 190 who got a standard text, 42 booked. Is there convincing evidence at α=0.05\alpha = 0.05 that the personalized text leads to a higher booking rate?

State: H0:pP−pS=0H_0: p_P - p_S = 0 and Ha:pP−pS>0H_a: p_P - p_S > 0, where pPp_P and pSp_S are the true proportions of patients like these who would book within a week after a personalized or a standard text. α=0.05\alpha = 0.05.

Plan: Two-sample zz-test for pP−pSp_P - p_S.

  • Random: patients were randomly assigned to the two messages.
  • 10%: not needed, since this is an experiment rather than a sample from a population.
  • Large Counts: p^C=63+42210+190=105400=0.2625\hat{p}_C = \dfrac{63 + 42}{210 + 190} = \dfrac{105}{400} = 0.2625. The expected counts 210(0.2625)≈55.1210(0.2625) \approx 55.1, 210(0.7375)≈154.9210(0.7375) \approx 154.9, 190(0.2625)≈49.9190(0.2625) \approx 49.9, and 190(0.7375)≈140.1190(0.7375) \approx 140.1 are all at least 10.

Do: p^P=0.30\hat{p}_P = 0.30 and p^S=42190≈0.2211\hat{p}_S = \dfrac{42}{190} \approx 0.2211.

z=0.30−0.22110.2625(0.7375)(1210+1190)=0.07890.04405≈1.79.z = \frac{0.30 - 0.2211}{\sqrt{0.2625(0.7375)\left(\dfrac{1}{210} + \dfrac{1}{190}\right)}} = \frac{0.0789}{0.04405} \approx 1.79.

P-value =P(Z≥1.79)≈0.0366= P(Z \ge 1.79) \approx 0.0366.

Conclude: Because 0.0366≤0.050.0366 \le 0.05, reject H0H_0. There is convincing evidence that the personalized text causes a higher booking rate than the standard text for patients like these. A causal conclusion is justified because the treatments were randomly assigned.

Scope of inference

What you can conclude depends on how the data were produced:

  • Random assignment lets you conclude that a difference was caused by the treatment.
  • Random sampling lets you generalize to the populations sampled.

In the reminder study, the patients were not a random sample of all patients everywhere, so the conclusion applies to patients like those in the study. In the walking example, the students were randomly sampled but not assigned to districts, so you can generalize to each district but cannot say that living in a city causes more walking.

Tip

Always define which group is "1" and which is "2" and keep the order consistent. If you subtract in the other order, the interval flips sign, for example (−0.208, −0.017)(-0.208,\ -0.017), and a one-sided alternative flips direction. Either order is fine as long as your conclusion matches.

Practice

Practice 1

A researcher wants to test whether the proportion of adults who own an electric vehicle differs between two states. Which formula should she use for the standard deviation in her test statistic?

Practice 2

In a random sample of 200 seniors at a large high school, 84 have a part-time job. In an independent random sample of 220 juniors at the same school, 66 have a part-time job. Assume the conditions are met. Find a 95% confidence interval for pS−pJp_S - p_J, the difference in the proportions of all seniors and all juniors with a part-time job. Give both endpoints, rounded to three decimal places.

Separate answers with commas, e.g. 2, -5

Practice 3

Based on the interval (0.029, 0.211)(0.029,\ 0.211) from the previous problem, which conclusion is best?

Practice 4

A company randomly assigns 800 website visitors to see one of two checkout pages. Of 400 who saw page A, 156 completed a purchase. Of 400 who saw page B, 124 completed a purchase. Find the pooled proportion p^C\hat{p}_C for testing whether the purchase rates differ.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

Continuing the checkout experiment, compute the test statistic zz for H0:pA−pB=0H_0: p_A - p_B = 0 versus Ha:pA−pB≠0H_a: p_A - p_B \ne 0. Round to two decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

Using z≈2.37z \approx 2.37 and the two-sided alternative Ha:pA−pB≠0H_a: p_A - p_B \ne 0, find the P-value. Round to four decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

For the checkout experiment (P-value ≈0.018\approx 0.018), which conclusion is correct at α=0.05\alpha = 0.05?

Practice 8

A 99% confidence interval for p1−p2p_1 - p_2, the difference in the proportions of left-handed people in two large countries, is (−0.018, 0.031)(-0.018,\ 0.031). Which statement is correct?