Math Core

Lesson 6.1 · Inference for Proportions

Confidence intervals for a proportion

A single sample proportion is almost never exactly right. In the last unit you learned how much p^\hat{p} wobbles from sample to sample; now you turn that knowledge around and use one sample to build a range of plausible values for the unknown population proportion pp.

From a point estimate to an interval

Suppose a random sample of 500 adults finds that 285 of them get most of their news from a phone. The point estimate is p^=285500=0.57\hat{p} = \dfrac{285}{500} = 0.57. It is our single best guess for pp, but a different sample would give a different number. Reporting "57%" alone hides that uncertainty.

A confidence interval fixes this by attaching a margin of error to the estimate:

statistic±(critical value)(standard error of statistic).\text{statistic} \pm (\text{critical value})(\text{standard error of statistic}).

The idea comes straight from the sampling distribution. When the conditions below are met, p^\hat{p} is approximately Normal with mean pp and standard deviation p(1−p)n\sqrt{\dfrac{p(1-p)}{n}}. About 95% of all samples produce a p^\hat{p} within 1.96 standard deviations of pp. So if you go 1.96 standard deviations out from your p^\hat{p} in both directions, you will capture pp for about 95% of samples.

Because pp is unknown, you cannot compute that standard deviation directly. You replace pp with p^\hat{p} and call the result the standard error:

SEp^=p^(1−p^)n.SE_{\hat{p}} = \sqrt{\frac{\hat{p}(1-\hat{p})}{n}}.

One-sample z-interval for a proportion

p^±z∗p^(1−p^)n\hat{p} \pm z^* \sqrt{\frac{\hat{p}(1-\hat{p})}{n}}

Here z∗z^* is the critical value from the standard Normal distribution that leaves the stated confidence level CC in the middle, with 1−C2\dfrac{1-C}{2} in each tail. The quantity z∗p^(1−p^)/nz^* \sqrt{\hat{p}(1-\hat{p})/n} is the margin of error.

Critical values

You find z∗z^* from the standard Normal distribution. For a 95% interval, the middle 95% leaves 2.5% in each tail, so z∗z^* is the value with area 0.975 to its left. These are the values you will use most often:

Confidence level CCTail area on each sidez∗z^*
90%0.051.645
95%0.0251.960
98%0.012.326
99%0.0052.576

On a calculator, z∗=invNorm(0.975)≈1.960z^* = \text{invNorm}(0.975) \approx 1.960 for 95%. A higher confidence level needs a larger z∗z^*, which makes the interval wider.

Conditions

Every AP inference procedure comes with conditions. For a one-sample zz-interval for pp, check all three and say why each is met.

  1. Random: the data come from a random sample from the population of interest (or a randomized experiment).
  2. 10% condition: when sampling without replacement, n≤0.10Nn \le 0.10N, so that observations are approximately independent.
  3. Large Counts: np^≥10n\hat{p} \ge 10 and n(1−p^)≥10n(1-\hat{p}) \ge 10. In words, at least 10 successes and 10 failures in the sample. This makes the sampling distribution of p^\hat{p} approximately Normal.

For a confidence interval you check Large Counts with the observed counts of successes and failures, because you do not know pp.

The four-step process

AP free-response questions reward a complete, organized answer. Use State, Plan, Do, Conclude.

  • State: name the parameter in context and the confidence level.
  • Plan: name the procedure and check the conditions.
  • Do: compute the interval.
  • Conclude: interpret the interval in context.

Worked example: A complete 95% interval

A random sample of 500 adults from a large city finds that 285 get most of their news from a phone. Construct and interpret a 95% confidence interval for the proportion of all adults in the city who get most of their news from a phone.

State: We want to estimate pp, the proportion of all adults in the city who get most of their news from a phone, with 95% confidence.

Plan: One-sample zz-interval for pp.

  • Random: the adults were a random sample.
  • 10%: 500 is less than 10% of all adults in a large city.
  • Large Counts: 285≥10285 \ge 10 successes and 215≥10215 \ge 10 failures.

Do: p^=0.57\hat{p} = 0.57 and

SE=0.57(0.43)500≈0.02214,ME=1.96(0.02214)≈0.0434,0.57±0.0434  ⇒  (0.5266, 0.6134).\begin{aligned} SE &= \sqrt{\frac{0.57(0.43)}{500}} \approx 0.02214, \\ \text{ME} &= 1.96(0.02214) \approx 0.0434, \\ 0.57 &\pm 0.0434 \;\Rightarrow\; (0.5266,\ 0.6134). \end{aligned}

Conclude: We are 95% confident that the interval from 0.527 to 0.613 captures the true proportion of all adults in this city who get most of their news from a phone.

What "95% confident" means

The confidence level describes the method, not one particular interval. If you took many random samples of the same size and built a 95% interval from each, about 95% of those intervals would capture pp. Your one interval either contains pp or it does not; you just don't know which.

Common mistake

Never say "there is a 95% probability that pp is between 0.527 and 0.613." The parameter pp is a fixed number, not a random one. Say instead that you are 95% confident the interval captures pp, or that 95% of intervals built this way capture pp. Also avoid saying 95% of sample proportions fall in the interval.

What controls the margin of error

The margin of error z∗p^(1−p^)/nz^*\sqrt{\hat{p}(1-\hat{p})/n} gets:

  • larger when the confidence level goes up (bigger z∗z^*),
  • smaller when the sample size goes up (bigger nn in the denominator).

Because nn sits under a square root, you must quadruple the sample size to cut the margin of error in half. Note that the margin of error only accounts for sampling variability. It does nothing to protect you from bias such as undercoverage or nonresponse.

Worked example: Changing the confidence level

In a random sample of 1,200 teens, 372 say they have used a budgeting app. Compare the 90% and 99% confidence intervals for the true proportion.

p^=3721200=0.31\hat{p} = \dfrac{372}{1200} = 0.31 and SE=0.31(0.69)1200≈0.01335SE = \sqrt{\dfrac{0.31(0.69)}{1200}} \approx 0.01335.

  • 90%: 0.31±1.645(0.01335)=0.31±0.02200.31 \pm 1.645(0.01335) = 0.31 \pm 0.0220, giving (0.288, 0.332)(0.288,\ 0.332).
  • 99%: 0.31±2.576(0.01335)=0.31±0.03440.31 \pm 2.576(0.01335) = 0.31 \pm 0.0344, giving (0.276, 0.344)(0.276,\ 0.344).

The 99% interval is wider. More confidence costs you precision.

Choosing a sample size

Before collecting data, you can choose nn to guarantee a desired margin of error ME\text{ME}. Solve

z∗p∗(1−p∗)n≤ME⟹n≥(z∗ME)2p∗(1−p∗),z^* \sqrt{\frac{p^*(1-p^*)}{n}} \le \text{ME} \quad\Longrightarrow\quad n \ge \left(\frac{z^*}{\text{ME}}\right)^2 p^*(1-p^*),

where p∗p^* is a guess for pp from a past study. If you have no guess, use p∗=0.5p^* = 0.5, which gives the largest possible value of p∗(1−p∗)p^*(1-p^*) and so the safest (largest) nn. Always round up.

Worked example: How many to survey

A school newspaper wants a 95% confidence interval for the proportion of students who support a later start time, with a margin of error of at most 3 percentage points. How many students should they survey?

With no prior estimate, use p∗=0.5p^* = 0.5:

n≥(1.960.03)2(0.5)(0.5)≈1067.1.n \ge \left(\frac{1.96}{0.03}\right)^2 (0.5)(0.5) \approx 1067.1.

Round up: they need at least 1,068 students. Rounding down to 1,067 would give a margin of error slightly larger than 0.03.

Tip

If a problem gives you an interval such as (0.42, 0.56)(0.42,\ 0.56), the point estimate is the midpoint and the margin of error is half the width: p^=0.49\hat{p} = 0.49 and ME=0.07\text{ME} = 0.07.

Practice

Practice 1

Find the critical value z∗z^* for a 98% confidence interval. Round to three decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

A random sample of 400 registered voters in a large county finds that 132 plan to vote by mail. Assume the conditions are met. Find the endpoints of a 95% confidence interval for the proportion of all registered voters in the county who plan to vote by mail. Give both endpoints, rounded to three decimal places.

Separate answers with commas, e.g. 2, -5

Practice 3

In a random sample of 250 customers, 160 said they would recommend a store to a friend. Find the margin of error for a 99% confidence interval for the true proportion of customers who would recommend the store. Round to three decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A random sample of students produced the 90% confidence interval (0.41, 0.49)(0.41,\ 0.49) for the proportion of students at a university who work a part-time job. Which is a correct interpretation of the confidence level?

Practice 5

A quality inspector takes a random sample of 150 light bulbs from a shipment of 20,000 and finds 7 defective bulbs. She wants to build a 95% zz-interval for the proportion of defective bulbs in the shipment. Which condition is not met?

Practice 6

A news report gives a 95% confidence interval of (0.42, 0.56)(0.42,\ 0.56) for the proportion of residents who support a new transit tax. What was the sample proportion p^\hat{p}?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A previous study estimated that about 20% of adults in a region have tried a meal-kit service. How large a random sample is needed to estimate this proportion within 0.04 with 90% confidence?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

A researcher computes a 95% confidence interval for a proportion from a random sample of 300 people. Which change would produce a narrower interval, assuming p^\hat{p} stays about the same?