Lesson 6.2 · Inference for Proportions
Significance tests for a proportion
A confidence interval asks "what values of are plausible?" A significance test asks a sharper question: "Is this particular claim about believable, given my data?" Tests are how researchers decide whether an effect is real or could just be sampling variability.
The logic of a significance test
A company claims that 30% of its customers use its mobile app. A manager suspects the true figure is higher. She takes a random sample of 200 customers and finds that 74 use the app, so .
Is 0.37 convincing evidence that ? Even if the true proportion were exactly 0.30, random samples would sometimes give values as large as 0.37. The test measures how unusual 0.37 would be if the claim were true. If it would be very unusual, the claim looks doubtful.
This is proof by contradiction with probability: assume the claim is true, then ask whether the data are surprising under that assumption.
Hypotheses
Definition
Null and alternative hypotheses
The null hypothesis is the claim being tested, usually a statement of "no difference" or "no change." For one proportion it has the form .
The alternative hypothesis is what you are trying to find evidence for. It takes one of three forms: , (one-sided), or (two-sided).
Hypotheses are always about the parameter , never about . You decide on from the research question before looking at the data. For the app example, and , where is the proportion of all the company's customers who use the app.
The test statistic
To measure how far is from , standardize it. Because you are assuming is true, use (not ) in the standard deviation:
The test statistic tells you how many standard deviations is from the null value.
The P-value
Definition
P-value
The P-value is the probability, computed assuming is true, of getting a statistic at least as extreme as the one observed, in the direction(s) given by .
- : P-value .
- : P-value .
- : P-value .
A small P-value means the observed result would rarely happen by chance alone if were true, which is evidence against . Here is an excerpt of a standard Normal table, giving the area to the left of :
| 1.80 | 1.90 | 2.00 | 2.09 | 2.16 | 2.30 | 2.32 | |
|---|---|---|---|---|---|---|---|
| Area left | 0.9641 | 0.9713 | 0.9772 | 0.9817 | 0.9846 | 0.9893 | 0.9898 |
By symmetry, the area to the left of equals the area to the right of , which is minus the table value.
Making a decision
Before collecting data, choose a significance level , commonly 0.05. Then compare:
Decision rule
- If P-value : reject . There is convincing evidence for .
- If P-value : fail to reject . There is not convincing evidence for .
Common mistake
Failing to reject does not prove is true. It only means the data were not strong enough to rule it out. Never write "we accept " or "this proves ." Write "there is not convincing evidence that…"
Conditions and the four steps
The conditions match those for intervals, with one change: check Large Counts using , because the test assumes is true.
- Random: random sample or randomized experiment.
- 10%: when sampling without replacement.
- Large Counts: and .
The four steps become:
- State: hypotheses, parameter in context, and .
- Plan: name the one-sample -test for and check conditions.
- Do: compute , , and the P-value.
- Conclude: compare the P-value to , make a decision, and state it in context.
Worked example: A one-sided test, start to finish
Use the app data: 74 of 200 randomly selected customers use the app. Is there convincing evidence at that more than 30% of customers use it?
State: , , where is the proportion of all the company's customers who use the app. Use .
Plan: One-sample -test for .
- Random: random sample of customers.
- 10%: 200 is less than 10% of all customers (assume the company has more than 2,000).
- Large Counts: and .
Do: and
P-value .
Conclude: Because , we reject . There is convincing evidence that more than 30% of the company's customers use the mobile app.
Worked example: A two-sided test
A national report says 60% of high school seniors have a driver's license. A researcher wonders whether the proportion in her state is different. In a random sample of 420 seniors in the state, 231 have a license. Test at .
State: , , where is the proportion of all seniors in the state with a license.
Plan: One-sample -test. Random sample; 420 is less than 10% of seniors in a state; and are both at least 10.
Do: and
P-value .
Conclude: Since , reject . There is convincing evidence that the proportion of seniors in this state with a driver's license differs from 0.60 (in fact, it appears to be lower).
Interpreting a P-value
On the AP exam you may be asked to interpret a P-value in context. Use this template: "Assuming [ in context] is true, there is a [P-value] probability of getting a sample proportion of [] or [more extreme direction] by chance alone." For the app example: assuming 30% of customers use the app, there is about a 0.0154 probability of getting a sample proportion of 0.37 or higher in a random sample of 200.
Tests and intervals agree
A two-sided test at significance level and a confidence interval tell a consistent story: if lies outside the interval, a two-sided test would reject ; if lies inside, the test would fail to reject. The interval gives extra information, a whole range of plausible values, so it is often worth reporting both.
Tip
Keep your , standard deviation and unrounded on your calculator until the end. Rounding to two decimals is fine for reading a table, but rounding the standard deviation too early can shift the P-value in the third decimal place.
Practice
A city once found that 45% of residents recycle regularly. After a new education campaign, officials want to know if the proportion has increased. Which hypotheses are appropriate?
A website says 30% of visitors click on its daily deal. After a redesign, a random sample of 150 visitors shows that 58 clicked. For testing versus , compute the test statistic . Round to two decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Continuing the previous problem (, ), find the P-value. Round to four decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
For the website test (, , P-value ), which conclusion is correct at ?
A coin-flipping app claims to be fair. In 800 simulated flips it produced 412 heads. Test versus , where is the long-run proportion of heads. Find the P-value, rounded to three decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A study tested against , where is the proportion of adults who exercise weekly, and got with a P-value of 0.04. Which is the correct interpretation of the P-value?
A national survey found that 20% of adults had donated blood in the past five years. A random sample of 250 adults in one county found 38 donors (). For versus , the test gives and P-value . What is the correct decision at ?
A 95% confidence interval for the proportion of voters who favor a ballot measure is . Based on this interval, what can you conclude about the two-sided test of at ?