Math Core

Lesson 6.2 · Inference for Proportions

Significance tests for a proportion

A confidence interval asks "what values of pp are plausible?" A significance test asks a sharper question: "Is this particular claim about pp believable, given my data?" Tests are how researchers decide whether an effect is real or could just be sampling variability.

The logic of a significance test

A company claims that 30% of its customers use its mobile app. A manager suspects the true figure is higher. She takes a random sample of 200 customers and finds that 74 use the app, so p^=0.37\hat{p} = 0.37.

Is 0.37 convincing evidence that p>0.30p > 0.30? Even if the true proportion were exactly 0.30, random samples would sometimes give p^\hat{p} values as large as 0.37. The test measures how unusual 0.37 would be if the claim were true. If it would be very unusual, the claim looks doubtful.

This is proof by contradiction with probability: assume the claim is true, then ask whether the data are surprising under that assumption.

Hypotheses

Definition

Null and alternative hypotheses

The null hypothesis H0H_0 is the claim being tested, usually a statement of "no difference" or "no change." For one proportion it has the form H0:p=p0H_0: p = p_0.

The alternative hypothesis HaH_a is what you are trying to find evidence for. It takes one of three forms: Ha:p>p0H_a: p > p_0, Ha:p<p0H_a: p < p_0 (one-sided), or Ha:p≠p0H_a: p \ne p_0 (two-sided).

Hypotheses are always about the parameter pp, never about p^\hat{p}. You decide on HaH_a from the research question before looking at the data. For the app example, H0:p=0.30H_0: p = 0.30 and Ha:p>0.30H_a: p > 0.30, where pp is the proportion of all the company's customers who use the app.

The test statistic

To measure how far p^\hat{p} is from p0p_0, standardize it. Because you are assuming H0H_0 is true, use p0p_0 (not p^\hat{p}) in the standard deviation:

z=p^−p0p0(1−p0)n.z = \frac{\hat{p} - p_0}{\sqrt{\dfrac{p_0(1-p_0)}{n}}}.

The test statistic tells you how many standard deviations p^\hat{p} is from the null value.

The P-value

Definition

P-value

The P-value is the probability, computed assuming H0H_0 is true, of getting a statistic at least as extreme as the one observed, in the direction(s) given by HaH_a.

  • Ha:p>p0H_a: p > p_0: P-value =P(Z≥z)= P(Z \ge z).
  • Ha:p<p0H_a: p < p_0: P-value =P(Z≤z)= P(Z \le z).
  • Ha:p≠p0H_a: p \ne p_0: P-value =2P(Z≥∣z∣)= 2P(Z \ge |z|).

A small P-value means the observed result would rarely happen by chance alone if H0H_0 were true, which is evidence against H0H_0. Here is an excerpt of a standard Normal table, giving the area to the left of zz:

zz1.801.902.002.092.162.302.32
Area left0.96410.97130.97720.98170.98460.98930.9898

By symmetry, the area to the left of −z-z equals the area to the right of zz, which is 11 minus the table value.

Making a decision

Before collecting data, choose a significance level α\alpha, commonly 0.05. Then compare:

Decision rule

  • If P-value ≤α\le \alpha: reject H0H_0. There is convincing evidence for HaH_a.
  • If P-value >α> \alpha: fail to reject H0H_0. There is not convincing evidence for HaH_a.

Common mistake

Failing to reject H0H_0 does not prove H0H_0 is true. It only means the data were not strong enough to rule it out. Never write "we accept H0H_0" or "this proves p=0.30p = 0.30." Write "there is not convincing evidence that…"

Conditions and the four steps

The conditions match those for intervals, with one change: check Large Counts using p0p_0, because the test assumes H0H_0 is true.

  1. Random: random sample or randomized experiment.
  2. 10%: n≤0.10Nn \le 0.10N when sampling without replacement.
  3. Large Counts: np0≥10np_0 \ge 10 and n(1−p0)≥10n(1 - p_0) \ge 10.

The four steps become:

  • State: hypotheses, parameter in context, and α\alpha.
  • Plan: name the one-sample zz-test for pp and check conditions.
  • Do: compute p^\hat{p}, zz, and the P-value.
  • Conclude: compare the P-value to α\alpha, make a decision, and state it in context.

Worked example: A one-sided test, start to finish

Use the app data: 74 of 200 randomly selected customers use the app. Is there convincing evidence at α=0.05\alpha = 0.05 that more than 30% of customers use it?

State: H0:p=0.30H_0: p = 0.30, Ha:p>0.30H_a: p > 0.30, where pp is the proportion of all the company's customers who use the app. Use α=0.05\alpha = 0.05.

Plan: One-sample zz-test for pp.

  • Random: random sample of customers.
  • 10%: 200 is less than 10% of all customers (assume the company has more than 2,000).
  • Large Counts: 200(0.30)=60≥10200(0.30) = 60 \ge 10 and 200(0.70)=140≥10200(0.70) = 140 \ge 10.

Do: p^=0.37\hat{p} = 0.37 and

z=0.37−0.300.30(0.70)200=0.070.0324≈2.16.z = \frac{0.37 - 0.30}{\sqrt{\dfrac{0.30(0.70)}{200}}} = \frac{0.07}{0.0324} \approx 2.16.

P-value =P(Z≥2.16)=1−0.9846=0.0154= P(Z \ge 2.16) = 1 - 0.9846 = 0.0154.

Conclude: Because 0.0154≤0.050.0154 \le 0.05, we reject H0H_0. There is convincing evidence that more than 30% of the company's customers use the mobile app.

Worked example: A two-sided test

A national report says 60% of high school seniors have a driver's license. A researcher wonders whether the proportion in her state is different. In a random sample of 420 seniors in the state, 231 have a license. Test at α=0.05\alpha = 0.05.

State: H0:p=0.60H_0: p = 0.60, Ha:p≠0.60H_a: p \ne 0.60, where pp is the proportion of all seniors in the state with a license.

Plan: One-sample zz-test. Random sample; 420 is less than 10% of seniors in a state; 420(0.6)=252420(0.6) = 252 and 420(0.4)=168420(0.4) = 168 are both at least 10.

Do: p^=231420=0.55\hat{p} = \dfrac{231}{420} = 0.55 and

z=0.55−0.600.6(0.4)420=−0.050.0239≈−2.09.z = \frac{0.55 - 0.60}{\sqrt{\dfrac{0.6(0.4)}{420}}} = \frac{-0.05}{0.0239} \approx -2.09.

P-value =2P(Z≤−2.09)=2(1−0.9817)≈0.0365= 2P(Z \le -2.09) = 2(1 - 0.9817) \approx 0.0365.

Conclude: Since 0.0365≤0.050.0365 \le 0.05, reject H0H_0. There is convincing evidence that the proportion of seniors in this state with a driver's license differs from 0.60 (in fact, it appears to be lower).

Interpreting a P-value

On the AP exam you may be asked to interpret a P-value in context. Use this template: "Assuming [H0H_0 in context] is true, there is a [P-value] probability of getting a sample proportion of [p^\hat{p}] or [more extreme direction] by chance alone." For the app example: assuming 30% of customers use the app, there is about a 0.0154 probability of getting a sample proportion of 0.37 or higher in a random sample of 200.

Tests and intervals agree

A two-sided test at significance level α\alpha and a C=1−αC = 1 - \alpha confidence interval tell a consistent story: if p0p_0 lies outside the interval, a two-sided test would reject H0:p=p0H_0: p = p_0; if p0p_0 lies inside, the test would fail to reject. The interval gives extra information, a whole range of plausible values, so it is often worth reporting both.

Tip

Keep your p^\hat{p}, standard deviation and zz unrounded on your calculator until the end. Rounding zz to two decimals is fine for reading a table, but rounding the standard deviation too early can shift the P-value in the third decimal place.

Practice

Practice 1

A city once found that 45% of residents recycle regularly. After a new education campaign, officials want to know if the proportion has increased. Which hypotheses are appropriate?

Practice 2

A website says 30% of visitors click on its daily deal. After a redesign, a random sample of 150 visitors shows that 58 clicked. For testing H0:p=0.30H_0: p = 0.30 versus Ha:p>0.30H_a: p > 0.30, compute the test statistic zz. Round to two decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

Continuing the previous problem (z≈2.32z \approx 2.32, Ha:p>0.30H_a: p > 0.30), find the P-value. Round to four decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

For the website test (H0:p=0.30H_0: p = 0.30, Ha:p>0.30H_a: p > 0.30, P-value ≈0.010\approx 0.010), which conclusion is correct at α=0.05\alpha = 0.05?

Practice 5

A coin-flipping app claims to be fair. In 800 simulated flips it produced 412 heads. Test H0:p=0.5H_0: p = 0.5 versus Ha:p≠0.5H_a: p \ne 0.5, where pp is the long-run proportion of heads. Find the P-value, rounded to three decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

A study tested H0:p=0.6H_0: p = 0.6 against Ha:p<0.6H_a: p < 0.6, where pp is the proportion of adults who exercise weekly, and got p^=0.55\hat{p} = 0.55 with a P-value of 0.04. Which is the correct interpretation of the P-value?

Practice 7

A national survey found that 20% of adults had donated blood in the past five years. A random sample of 250 adults in one county found 38 donors (p^=0.152\hat{p} = 0.152). For H0:p=0.20H_0: p = 0.20 versus Ha:p<0.20H_a: p < 0.20, the test gives z≈−1.90z \approx -1.90 and P-value ≈0.029\approx 0.029. What is the correct decision at α=0.01\alpha = 0.01?

Practice 8

A 95% confidence interval for the proportion of voters who favor a ballot measure is (0.46, 0.53)(0.46,\ 0.53). Based on this interval, what can you conclude about the two-sided test of H0:p=0.5H_0: p = 0.5 at α=0.05\alpha = 0.05?