Lesson 6.3 · Inference for Proportions
Errors and power
Every significance test ends in a decision, and every decision can be wrong. Because you only see a sample, you can reject a null hypothesis that is actually true, or fail to reject one that is actually false. This lesson names those two mistakes, shows how likely each one is, and explains how to design a study that catches real effects.
Two kinds of error
There are four possible outcomes of a test, depending on the truth (which you never see) and your decision:
| is true | is false | |
|---|---|---|
| Reject | Type I error | Correct decision |
| Fail to reject | Correct decision | Type II error |
Definition
Type I and Type II errors
A Type I error occurs when you reject even though is true (a "false alarm").
A Type II error occurs when you fail to reject even though is true (a "missed effect").
On the AP exam you almost always need to describe errors in context and often to give a consequence. A generic definition earns little credit.
Worked example: Describing errors in context
A snack company tests whether more than 8% of the chips in its bags are broken, in which case it will adjust its packaging line. The hypotheses are and , where is the proportion of all chips produced that are broken. Describe a Type I and a Type II error and a consequence of each.
- Type I error: The company finds convincing evidence that more than 8% of chips are broken when, in reality, the proportion is 8%. Consequence: it spends money adjusting a packaging line that did not need it.
- Type II error: The company does not find convincing evidence that more than 8% of chips are broken when, in reality, more than 8% are. Consequence: customers keep getting bags with too many broken chips, and the company may lose sales.
The probability of a Type I error
If is true, you reject it exactly when the P-value is at most . When conditions are met, that happens with probability .
Error probabilities
- , the significance level.
- , which depends on the true value of .
- Power is the probability of correctly rejecting when a specific alternative value of is true.
So choosing is choosing how often you are willing to raise a false alarm. If a Type I error would be very costly (say, approving an ineffective drug), choose a smaller such as 0.01. But there is a trade-off: making smaller makes it harder to reject , which increases for any fixed sample size. If a Type II error is the more serious mistake (say, missing a dangerous defect), a larger such as 0.10 may be appropriate.
Power
Power answers the question "if the effect I care about is real, how likely is my study to detect it?" A study with low power often fails to reject even when is true, which wastes time and money. Researchers aim for power of 0.80 or higher.
Computing power for a proportion
The AP exam rarely asks for a full power calculation, but doing one makes the idea concrete. There are two steps:
- Assuming is true, find the values of that would lead you to reject .
- Assuming a specific alternative value of is true, find the probability that lands in that rejection region.
Worked example: Computing power
A basketball player has historically made 50% of her three-point attempts. After summer training, she will take 100 shots and test versus at . If her true success rate is now , what is the power of the test?
Step 1: rejection region. Under , is approximately Normal with mean 0.5 and standard deviation . For a one-sided test at , reject when , that is, when
Step 2: probability under the alternative. If , then has mean 0.6 and standard deviation . So
There is about a 64% chance the test detects her improvement, so . That power is fairly low, so taking more shots would be a better plan.
What increases power
Power goes up when it becomes easier to tell the null distribution of apart from the true distribution. Four things do this:
- Larger sample size . Both distributions get narrower, so they overlap less. This is the most common way to boost power.
- Larger significance level . The rejection region grows, so you reject more often (at the cost of more Type I errors).
- True value farther from . A bigger effect is easier to detect.
- Less variability in general, such as through better measurement or a better design.
Worked example: Reasoning about power
In the basketball example, suppose the player takes 200 shots instead of 100. Describe the effect on the power of the test and on the probability of a Type I error.
With , both sampling distributions are narrower. The rejection cutoff moves closer to 0.5: . If is the truth, is now much more likely to exceed that cutoff, and power rises to about 0.89. The probability of a Type I error stays at , because is chosen by the researcher, not determined by .
Common mistake
Increasing the sample size does not change the probability of a Type I error; that is always . It increases power, which means it decreases the probability of a Type II error. Students often write that a larger sample "reduces both errors." At a fixed , it only reduces .
Tip
When an FRQ asks "which error could you have made?", look at the decision first. If you rejected , the only possible error is Type I. If you failed to reject, the only possible error is Type II.
Practice
A school tests versus , where is the proportion of students who are chronically absent, to see if a new attendance program worked. Which describes a Type I error?
A water utility tests whether more than 2% of its pipes have dangerous lead levels, using and . Which is a consequence of a Type II error?
A researcher tests a claim about a proportion using . If the null hypothesis is actually true, what is the probability that she makes a Type I error?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
For a particular alternative value of , the probability of a Type II error for a test is 0.23. What is the power of the test against that alternative?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A health department believes 20% of adults in a region smoke. After an anti-smoking campaign, it will survey a random sample of 400 adults and test versus at . Assuming is true, below what value of will the department reject ? Round to three decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Continuing the previous problem, suppose the true proportion of smokers after the campaign is . Find the power of the test. Round to two decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A researcher plans a test of versus . Which change would increase the power of the test against the alternative ?
A pharmaceutical company tests whether a new drug cures a condition in more than 70% of patients, versus . Regulators worry most about approving a drug that is no better than the current one. Which significance level is the most appropriate, and why?