Math Core

Lesson 6.3 · Inference for Proportions

Errors and power

Every significance test ends in a decision, and every decision can be wrong. Because you only see a sample, you can reject a null hypothesis that is actually true, or fail to reject one that is actually false. This lesson names those two mistakes, shows how likely each one is, and explains how to design a study that catches real effects.

Two kinds of error

There are four possible outcomes of a test, depending on the truth (which you never see) and your decision:

H0H_0 is trueH0H_0 is false
Reject H0H_0Type I errorCorrect decision
Fail to reject H0H_0Correct decisionType II error

Definition

Type I and Type II errors

A Type I error occurs when you reject H0H_0 even though H0H_0 is true (a "false alarm").

A Type II error occurs when you fail to reject H0H_0 even though HaH_a is true (a "missed effect").

On the AP exam you almost always need to describe errors in context and often to give a consequence. A generic definition earns little credit.

Worked example: Describing errors in context

A snack company tests whether more than 8% of the chips in its bags are broken, in which case it will adjust its packaging line. The hypotheses are H0:p=0.08H_0: p = 0.08 and Ha:p>0.08H_a: p > 0.08, where pp is the proportion of all chips produced that are broken. Describe a Type I and a Type II error and a consequence of each.

  • Type I error: The company finds convincing evidence that more than 8% of chips are broken when, in reality, the proportion is 8%. Consequence: it spends money adjusting a packaging line that did not need it.
  • Type II error: The company does not find convincing evidence that more than 8% of chips are broken when, in reality, more than 8% are. Consequence: customers keep getting bags with too many broken chips, and the company may lose sales.

The probability of a Type I error

If H0H_0 is true, you reject it exactly when the P-value is at most α\alpha. When conditions are met, that happens with probability α\alpha.

Error probabilities

  • P(Type I error)=αP(\text{Type I error}) = \alpha, the significance level.
  • P(Type II error)=βP(\text{Type II error}) = \beta, which depends on the true value of pp.
  • Power =1−β= 1 - \beta is the probability of correctly rejecting H0H_0 when a specific alternative value of pp is true.

So choosing α\alpha is choosing how often you are willing to raise a false alarm. If a Type I error would be very costly (say, approving an ineffective drug), choose a smaller α\alpha such as 0.01. But there is a trade-off: making α\alpha smaller makes it harder to reject H0H_0, which increases β\beta for any fixed sample size. If a Type II error is the more serious mistake (say, missing a dangerous defect), a larger α\alpha such as 0.10 may be appropriate.

Power

Power answers the question "if the effect I care about is real, how likely is my study to detect it?" A study with low power often fails to reject H0H_0 even when HaH_a is true, which wastes time and money. Researchers aim for power of 0.80 or higher.

Computing power for a proportion

The AP exam rarely asks for a full power calculation, but doing one makes the idea concrete. There are two steps:

  1. Assuming H0H_0 is true, find the values of p^\hat{p} that would lead you to reject H0H_0.
  2. Assuming a specific alternative value of pp is true, find the probability that p^\hat{p} lands in that rejection region.

Worked example: Computing power

A basketball player has historically made 50% of her three-point attempts. After summer training, she will take 100 shots and test H0:p=0.5H_0: p = 0.5 versus Ha:p>0.5H_a: p > 0.5 at α=0.05\alpha = 0.05. If her true success rate is now p=0.6p = 0.6, what is the power of the test?

Step 1: rejection region. Under H0H_0, p^\hat{p} is approximately Normal with mean 0.5 and standard deviation 0.5(0.5)/100=0.05\sqrt{0.5(0.5)/100} = 0.05. For a one-sided test at α=0.05\alpha = 0.05, reject when z≥1.645z \ge 1.645, that is, when

p^≥0.5+1.645(0.05)=0.58225.\hat{p} \ge 0.5 + 1.645(0.05) = 0.58225.

Step 2: probability under the alternative. If p=0.6p = 0.6, then p^\hat{p} has mean 0.6 and standard deviation 0.6(0.4)/100≈0.04899\sqrt{0.6(0.4)/100} \approx 0.04899. So

Power=P(p^≥0.58225)=P(Z≥0.58225−0.60.04899)=P(Z≥−0.36)≈0.64.\text{Power} = P(\hat{p} \ge 0.58225) = P\left(Z \ge \frac{0.58225 - 0.6}{0.04899}\right) = P(Z \ge -0.36) \approx 0.64.

There is about a 64% chance the test detects her improvement, so β≈0.36\beta \approx 0.36. That power is fairly low, so taking more shots would be a better plan.

What increases power

Power goes up when it becomes easier to tell the null distribution of p^\hat{p} apart from the true distribution. Four things do this:

  1. Larger sample size nn. Both distributions get narrower, so they overlap less. This is the most common way to boost power.
  2. Larger significance level α\alpha. The rejection region grows, so you reject more often (at the cost of more Type I errors).
  3. True value farther from p0p_0. A bigger effect is easier to detect.
  4. Less variability in general, such as through better measurement or a better design.

Worked example: Reasoning about power

In the basketball example, suppose the player takes 200 shots instead of 100. Describe the effect on the power of the test and on the probability of a Type I error.

With n=200n = 200, both sampling distributions are narrower. The rejection cutoff moves closer to 0.5: 0.5+1.6450.25/200≈0.5580.5 + 1.645\sqrt{0.25/200} \approx 0.558. If p=0.6p = 0.6 is the truth, p^\hat{p} is now much more likely to exceed that cutoff, and power rises to about 0.89. The probability of a Type I error stays at α=0.05\alpha = 0.05, because α\alpha is chosen by the researcher, not determined by nn.

Common mistake

Increasing the sample size does not change the probability of a Type I error; that is always α\alpha. It increases power, which means it decreases the probability of a Type II error. Students often write that a larger sample "reduces both errors." At a fixed α\alpha, it only reduces β\beta.

Tip

When an FRQ asks "which error could you have made?", look at the decision first. If you rejected H0H_0, the only possible error is Type I. If you failed to reject, the only possible error is Type II.

Practice

Practice 1

A school tests H0:p=0.25H_0: p = 0.25 versus Ha:p<0.25H_a: p < 0.25, where pp is the proportion of students who are chronically absent, to see if a new attendance program worked. Which describes a Type I error?

Practice 2

A water utility tests whether more than 2% of its pipes have dangerous lead levels, using H0:p=0.02H_0: p = 0.02 and Ha:p>0.02H_a: p > 0.02. Which is a consequence of a Type II error?

Practice 3

A researcher tests a claim about a proportion using α=0.01\alpha = 0.01. If the null hypothesis is actually true, what is the probability that she makes a Type I error?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

For a particular alternative value of pp, the probability of a Type II error for a test is 0.23. What is the power of the test against that alternative?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A health department believes 20% of adults in a region smoke. After an anti-smoking campaign, it will survey a random sample of 400 adults and test H0:p=0.20H_0: p = 0.20 versus Ha:p<0.20H_a: p < 0.20 at α=0.05\alpha = 0.05. Assuming H0H_0 is true, below what value of p^\hat{p} will the department reject H0H_0? Round to three decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

Continuing the previous problem, suppose the true proportion of smokers after the campaign is p=0.15p = 0.15. Find the power of the test. Round to two decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A researcher plans a test of H0:p=0.4H_0: p = 0.4 versus Ha:p>0.4H_a: p > 0.4. Which change would increase the power of the test against the alternative p=0.5p = 0.5?

Practice 8

A pharmaceutical company tests whether a new drug cures a condition in more than 70% of patients, H0:p=0.70H_0: p = 0.70 versus Ha:p>0.70H_a: p > 0.70. Regulators worry most about approving a drug that is no better than the current one. Which significance level is the most appropriate, and why?