Math Core

Lesson 5.4 · Sampling Distributions

The central limit theorem

Most real populations are not normal. Incomes, wait times and repair costs are skewed; test scores bunch up against a maximum. Yet statisticians use normal calculations for sample means all the time. The reason is one of the most important results in statistics: the central limit theorem.

The theorem

In the last lesson, the population was normal, so xˉ\bar{x} was exactly normal. What if the population is skewed, or has several peaks, or is shaped like nothing you've seen before? The center and spread of xˉ\bar{x} do not change: μxˉ=μ\mu_{\bar{x}} = \mu and σxˉ=σ/n\sigma_{\bar{x}} = \sigma/\sqrt{n} still hold. The remaining question is the shape.

The central limit theorem (CLT)

When you take a random sample of size nn from any population with mean μ\mu and finite standard deviation σ\sigma, the sampling distribution of xˉ\bar{x} becomes approximately normal as nn gets large, regardless of the shape of the population.

In AP Statistics, n≥30n \ge 30 is considered large enough to use a normal model for xˉ\bar{x} unless the population is extremely skewed or has outliers.

Why does averaging produce a normal shape? When you average many independent values, unusually large values tend to be offset by unusually small ones. Getting an extreme average would require many extreme values at once, which is rare. The averages pile up in the middle and thin out symmetrically in both directions.

How large is large enough?

The more a population departs from normal, the larger nn must be before xˉ\bar{x} looks normal.

  • Population normal: xˉ\bar{x} is normal for any nn (previous lesson).
  • Population roughly symmetric, no outliers: xˉ\bar{x} is close to normal even for fairly small nn, such as 10 or 15.
  • Population skewed: you need a larger nn. The AP guideline is n≥30n \ge 30.
  • Population shape unknown and n<30n < 30: you cannot assume xˉ\bar{x} is normal. (In Unit 7 you will learn to check a graph of the sample data in this case.)

Together with the 10% condition, this gives the checklist for using a normal model for xˉ\bar{x}:

ConditionWhat it justifiesHow to check
RandomResults generalize; formulas applyRandom sample or random assignment
10%σxˉ=σ/n\sigma_{\bar{x}} = \sigma/\sqrt{n}n≤0.10Nn \le 0.10N
Normal / Large SampleNormal shape of xˉ\bar{x}Population normal, or n≥30n \ge 30 (CLT)

Common mistake

The CLT is about the sampling distribution of xˉ\bar{x}, not about the data. Taking a larger sample does not make the population, or the distribution of values in your sample, more normal. A sample of 500 incomes will still be right-skewed, just like the population. Only the distribution of the sample means becomes normal.

Here is a table excerpt with the values used in this lesson.

zz−2.00-2.00−1.00-1.001.251.251.501.501.651.652.002.00
Area to the left0.02280.02280.15870.15870.89440.89440.93320.93320.95050.95050.97720.9772

Worked example: Skewed wait times

The time customers wait on hold at a help line is strongly right-skewed, with mean μ=8\mu = 8 minutes and standard deviation σ=8\sigma = 8 minutes. A supervisor selects a random sample of 64 calls from the thousands received this month. Find the probability that the mean wait time for the sample is more than 9.5 minutes.

The population of individual wait times is strongly right-skewed.Open in grapher →

Conditions. The sample is random. The 10% condition holds because 64 is far less than 10% of thousands of calls. The population is skewed, but n=64≥30n = 64 \ge 30, so by the CLT the sampling distribution of xˉ\bar{x} is approximately normal.

Sampling distribution.

μxˉ=8,σxˉ=864=1.\mu_{\bar{x}} = 8, \qquad \sigma_{\bar{x}} = \frac{8}{\sqrt{64}} = 1.

Probability.

z=9.5−81=1.50,P(xˉ>9.5)=1−0.9332=0.0668.z = \frac{9.5 - 8}{1} = 1.50, \qquad P(\bar{x} > 9.5) = 1 - 0.9332 = 0.0668.
The sampling distribution of x̄ for n = 64 is approximately normal: mean 8, SD 1. The shaded area above 9.5 is about 0.0668.Open in grapher →

The population is skewed, but the sample means are approximately normal.

Totals are means in disguise

Many AP questions ask about a total rather than a mean: the total weight in an elevator, the total amount spent by a group of customers. Since the total equals nxˉn\bar{x}, a statement about the total can be rewritten as a statement about the mean. Divide the cutoff by nn and use the sampling distribution of xˉ\bar{x}.

Worked example: A total becomes a mean

A snack company fills bags whose weights have mean 20 ounces and standard deviation 3 ounces; the distribution of individual weights is not known to be normal. A store orders a random sample of 36 bags. Find the probability that the total weight of the 36 bags is less than 702 ounces.

The total is less than 702 exactly when the mean is less than 70236=19.5\dfrac{702}{36} = 19.5 ounces.

Conditions: the bags are a random sample from a very large production, so the 10% condition holds, and n=36≥30n = 36 \ge 30, so the CLT says xˉ\bar{x} is approximately normal with

μxˉ=20,σxˉ=336=0.5.\mu_{\bar{x}} = 20, \qquad \sigma_{\bar{x}} = \frac{3}{\sqrt{36}} = 0.5.z=19.5−200.5=−1.00,P(total<702)=P(xˉ<19.5)=0.1587.z = \frac{19.5 - 20}{0.5} = -1.00, \qquad P(\text{total} < 702) = P(\bar{x} < 19.5) = 0.1587.

Worked example: When the CLT does not apply

Annual medical costs for employees at a company are strongly right-skewed with a few very large values. An analyst takes a random sample of 12 employees and wants to find the probability that the sample mean cost exceeds a certain amount using a normal model. Is this appropriate?

No. The population is strongly skewed, and n=12n = 12 is too small for the CLT to guarantee an approximately normal sampling distribution. The sampling distribution of xˉ\bar{x} will still be noticeably right-skewed, so a normal calculation could be quite inaccurate. The formulas for the mean and standard deviation of xˉ\bar{x} still hold; only the normal shape is in doubt.

Tip

When you invoke the CLT on the AP exam, say all three pieces: the population is not known to be normal (or is skewed), nn is at least 30, and therefore the sampling distribution of xˉ\bar{x} is approximately normal. Naming the sampling distribution, not "the data," is what earns credit.

Practice

Use the table excerpt above where needed. Give probabilities to four decimal places.

Practice 1

The number of songs on a student's phone is right-skewed across all students at a university, with mean 420 and standard deviation 300. A random sample of 40 students is selected. Which statement about the sampling distribution of xˉ\bar{x} is correct?

Practice 2

A population of home sale prices is strongly right-skewed. For which sample size is it least appropriate to model the sampling distribution of xˉ\bar{x} as approximately normal?

Practice 3

The number of pets per household in a large town has a right-skewed distribution with mean 3.2 and standard deviation 1.6. A random sample of 64 households is selected. Find the probability that the sample mean is greater than 3.53.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

The lifetimes of a type of battery have an unknown distribution with standard deviation σ=2.5\sigma = 2.5 hours. An engineer tests a random sample of 100 batteries from a large shipment. What is the probability that the sample mean lifetime is within 0.5 hour of the true mean μ\mu?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A student says, "Because of the central limit theorem, if I take a big enough sample of incomes, the incomes in my sample will be normally distributed." Which response is correct?

Practice 6

A small plane's cargo hold is rated for a total of 4,125 pounds of luggage. The weights of passengers' checked luggage have mean 40 pounds and standard deviation 10 pounds, and the distribution is skewed to the right. For a random sample of 100 bags, find the probability that the total weight exceeds 4,125 pounds.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A population is uniform on the integers 1 through 6 (like rolling a fair die). Simulations compute the sampling distribution of xˉ\bar{x} for n=2n = 2 and for n=25n = 25. Compared with n=2n = 2, the sampling distribution for n=25n = 25 has