Math Core

Lesson 5.2 · Sampling Distributions

Sampling distributions of proportions

Polls, quality-control checks and medical studies all report a sample proportion. To judge whether a sample proportion is surprising, you need to know how much p^\hat{p} naturally varies from sample to sample. This lesson gives you the center, spread and shape of the sampling distribution of p^\hat{p}, and shows how to use it to find probabilities.

Where the formulas come from

Suppose a population has a proportion pp of "successes," and you take a random sample of size nn. Let XX be the number of successes in the sample. If the sample is taken with replacement (or the population is so large that it hardly matters), XX is a binomial random variable, which you studied in Unit 4:

μX=np,σX=np(1−p).\mu_X = np, \qquad \sigma_X = \sqrt{np(1-p)}.

The sample proportion is p^=Xn\hat{p} = \dfrac{X}{n}, which is just XX multiplied by the constant 1n\dfrac{1}{n}. Multiplying a random variable by a constant multiplies its mean and its standard deviation by that constant, so

μp^=npn=p,σp^=np(1−p)n=p(1−p)n.\mu_{\hat{p}} = \frac{np}{n} = p, \qquad \sigma_{\hat{p}} = \frac{\sqrt{np(1-p)}}{n} = \sqrt{\frac{p(1-p)}{n}}.

Sampling distribution of a sample proportion

For a random sample of size nn from a population with proportion of successes pp:

  • Center: μp^=p\mu_{\hat{p}} = p. The sample proportion is an unbiased estimator of pp.
  • Spread: σp^=p(1−p)n\sigma_{\hat{p}} = \sqrt{\dfrac{p(1-p)}{n}}, as long as the 10% condition holds: n≤0.10Nn \le 0.10N, where NN is the population size.
  • Shape: approximately normal when the Large Counts condition holds: np≥10np \ge 10 and n(1−p)≥10n(1-p) \ge 10.

The two conditions

Each formula comes with a condition, and AP graders expect you to check them by name, with numbers.

The 10% condition justifies the standard deviation formula. When you sample without replacement, the draws are not quite independent: removing one success slightly changes the chance that the next person is a success. If the sample is no more than 10% of the population, this effect is tiny and the formula p(1−p)/n\sqrt{p(1-p)/n} is accurate. For example, a sample of 150 students from a school of 2,400 is fine because 150≤0.10(2400)=240150 \le 0.10(2400) = 240.

The Large Counts condition justifies using a normal model. A binomial distribution is skewed when the expected number of successes or failures is small. When both npnp and n(1−p)n(1-p) are at least 10, the distribution of p^\hat{p} is close enough to normal to use normal calculations.

Common mistake

The two conditions do different jobs. The 10% condition is about the standard deviation (independence). The Large Counts condition is about the shape (normality). Checking one does not check the other, and "n≥30n \ge 30" is not the condition for proportions.

Finding probabilities about p^\hat{p}

Once the conditions are met, you find a probability the same way you did for any normal distribution: standardize, then use a table or technology.

z=p^−pp(1−p)nz = \frac{\hat{p} - p}{\sqrt{\dfrac{p(1-p)}{n}}}

Here is an excerpt of the standard normal table (area to the left of zz) with the values used in this lesson.

zz−2.00-2.00−1.50-1.50−1.25-1.25−1.00-1.001.001.001.251.251.501.502.002.00
Area to the left0.02280.02280.06680.06680.10560.10560.15870.15870.84130.84130.89440.89440.93320.93320.97720.9772

Worked example: Describing the sampling distribution

About 60% of the 3,000 students at a large high school say they get less than 8 hours of sleep on school nights. A random sample of 150 students is selected. Describe the sampling distribution of p^\hat{p}, the proportion in the sample who get less than 8 hours of sleep.

Center: μp^=p=0.60\mu_{\hat{p}} = p = 0.60.

Spread: The 10% condition holds because 150≤0.10(3000)=300150 \le 0.10(3000) = 300. So

σp^=0.60(0.40)150=0.24150=0.0016=0.04.\sigma_{\hat{p}} = \sqrt{\frac{0.60(0.40)}{150}} = \sqrt{\frac{0.24}{150}} = \sqrt{0.0016} = 0.04.

Shape: np=150(0.60)=90≥10np = 150(0.60) = 90 \ge 10 and n(1−p)=150(0.40)=60≥10n(1-p) = 150(0.40) = 60 \ge 10, so the Large Counts condition holds and the sampling distribution is approximately normal.

In context: across many random samples of 150 students, the sample proportion would be approximately normal, centered at 0.60, and would typically vary from 0.60 by about 0.04.

Worked example: A normal probability for p-hat

Using the school from the previous example, find the probability that fewer than 55% of the 150 sampled students get less than 8 hours of sleep.

The conditions were checked above, so p^\hat{p} is approximately normal with mean 0.60 and standard deviation 0.04.

z=0.55−0.600.04=−1.25z = \frac{0.55 - 0.60}{0.04} = -1.25

From the table, P(Z<−1.25)=0.1056P(Z < -1.25) = 0.1056. There is about a 10.56% chance that a random sample of 150 students gives a sample proportion below 0.55.

Sampling distribution of p̂: mean 0.60, standard deviation 0.04. The shaded area to the left of 0.55 is about 0.1056.Open in grapher →

Worked example: Is a sample result surprising?

A company claims that 80% of its customers are satisfied. In a random sample of 100 of its many thousands of customers, only 72% say they are satisfied. If the claim is true, how surprising is a result this low?

Conditions. The sample of 100 is far less than 10% of the customers. Also, np=80≥10np = 80 \ge 10 and n(1−p)=20≥10n(1-p) = 20 \ge 10. So the distribution of p^\hat{p} is approximately normal with

μp^=0.80,σp^=0.80(0.20)100=0.0016=0.04.\mu_{\hat{p}} = 0.80, \qquad \sigma_{\hat{p}} = \sqrt{\frac{0.80(0.20)}{100}} = \sqrt{0.0016} = 0.04.

Probability.

z=0.72−0.800.04=−2.00,P(p^≤0.72)=P(Z≤−2.00)=0.0228.z = \frac{0.72 - 0.80}{0.04} = -2.00, \qquad P(\hat{p} \le 0.72) = P(Z \le -2.00) = 0.0228.

If 80% of customers really were satisfied, only about 2.3% of random samples of 100 would give a sample proportion of 0.72 or lower. That is unusual, so the sample gives some reason to doubt the company's claim. You will formalize this reasoning with significance tests in Unit 6.

Sample size and spread

Because nn sits under a square root in the denominator, the spread of p^\hat{p} shrinks slowly as the sample grows. Quadrupling nn cuts σp^\sigma_{\hat{p}} in half. For a fixed nn, the spread is largest when p=0.5p = 0.5, because p(1−p)p(1-p) is maximized there. That is why pollsters who don't know pp plan their sample sizes using p=0.5p = 0.5: it gives the most conservative (largest) standard deviation.

Tip

When you check conditions, write the actual numbers: "np=90≥10np = 90 \ge 10 and n(1−p)=60≥10n(1-p) = 60 \ge 10," not just "Large Counts is met." And always use the parameter pp in these formulas when it is known, not a sample value.

Practice

Use the table excerpt above where needed. Give probabilities to four decimal places.

Practice 1

In a large city, 25% of households have a dog. A random sample of 300 households is selected. Find the standard deviation of the sampling distribution of p^\hat{p}, the proportion of sampled households with a dog.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

A factory knows that 4% of the light bulbs it produces are defective. An inspector selects a random sample of 200 bulbs from a day's production of 50,000. Can the sampling distribution of p^\hat{p} be modeled as approximately normal?

Practice 3

A small college has 1,200 students, and 45% of them live on campus. A researcher plans to select a random sample of 300 students without replacement. Which statement about the formula σp^=p(1−p)/n\sigma_{\hat{p}} = \sqrt{p(1-p)/n} is correct?

Practice 4

In a large state, 20% of adults have a library card. A random sample of 400 adults is selected. Find the probability that more than 24% of the sample have a library card.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

Return to the high school where 60% of the 3,000 students get less than 8 hours of sleep, with random samples of 150 students. Find the probability that the sample proportion is between 0.55 and 0.65.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

A pollster wants the standard deviation of p^\hat{p} to be at most 0.02. She does not know pp, so she uses the conservative value p=0.5p = 0.5. What is the smallest sample size that achieves her goal?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A company claims that 80% of its customers are satisfied. For random samples of 100 customers, p^\hat{p} is approximately normal with mean 0.80 and standard deviation 0.04. A consultant surveys 100 randomly selected customers and finds that 76% are satisfied. Which conclusion is most reasonable?