Lesson 5.1 · Sampling Distributions
Sampling variability
Every time you take a random sample, you get a slightly different answer. A poll of 500 voters might say 52% support a measure, and a second poll of 500 different voters might say 49%. This lesson is about that wobble: where it comes from, how to describe it, and why it is the foundation for everything in inference.
Parameters and statistics
In Unit 3 you learned how to collect data with random sampling. Now we connect the sample back to the population it came from. Two words do most of the work.
Definition
Parameter and statistic
A parameter is a number that describes a population. It is usually unknown, because measuring the whole population is too expensive or impossible.
A statistic is a number that describes a sample. You calculate it from your data, and you use it to estimate a parameter.
A handy memory trick: parameter goes with population, statistic goes with sample. The notation follows a pattern too. Parameters get Greek letters or plain letters, and statistics get "hats" or bars.
| Quantity | Parameter (population) | Statistic (sample) |
|---|---|---|
| Proportion | ||
| Mean | ||
| Standard deviation |
For example, suppose 38% of all students at a large high school walk to school. That 38% is the parameter . If you survey a random sample of 80 students and 27 of them walk, your statistic is .
Sampling variability
If a classmate takes a different random sample of 80 students, they will almost certainly not get 27 walkers. Maybe they get 33, so . Neither of you did anything wrong. Different random samples contain different individuals, so they produce different values of the statistic.
Sampling variability
The value of a statistic changes from one random sample to the next. This is called sampling variability. The parameter stays fixed; the statistic is the thing that varies.
Sampling variability is not a mistake to be eliminated. It is a predictable feature of random sampling, and the whole point of this unit is that you can describe it with precision.
The sampling distribution of a statistic
Imagine taking every possible random sample of a given size from a population, computing the statistic for each one, and graphing all those values. The resulting distribution is the sampling distribution of the statistic.
Definition
Sampling distribution
The sampling distribution of a statistic is the distribution of the values the statistic takes in all possible samples of the same size from the same population.
Keep three distributions separate in your head:
- The population distribution: the values of the variable for every individual in the population.
- The distribution of sample data: the values of the variable for the individuals in one sample.
- The sampling distribution: the values of a statistic (like or ) across many samples.
In practice you rarely list every possible sample. Instead, you simulate: draw many random samples with software, compute the statistic each time, and plot the results. The line plot below shows 20 simulated values of , each from a random sample of 10 students from the school where .
The values pile up near 0.38, but some samples gave 0.1 and some gave 0.6. That spread is sampling variability made visible.
Worked example: A sampling distribution you can list completely
A tiny population has four values: . Its mean is . List the sampling distribution of for random samples of size 2 taken without replacement, and find the mean of that distribution.
There are possible samples:
| Sample | ||||||
|---|---|---|---|---|---|---|
Each sample is equally likely, so the sampling distribution gives probability to and to . The mean of the six sample means is
The sampling distribution of is centered exactly at the population mean .
Bias and variability
When you use a statistic to estimate a parameter, you care about two different properties of its sampling distribution: where it is centered and how spread out it is.
Definition
Unbiased estimator
A statistic is an unbiased estimator of a parameter if the mean of its sampling distribution equals the value of the parameter. If the sampling distribution is centered above or below the parameter, the statistic is biased.
The example above showed that is an unbiased estimator of . It turns out that is an unbiased estimator of as well. Not every statistic is unbiased, though. The sample range, for instance, can never be larger than the population range, and it is usually smaller, so its sampling distribution is centered below the population range. The sample range is a biased estimator.
Think of a dartboard with the parameter at the bullseye. Each throw is one sample's statistic.
- Low bias, low variability: throws cluster tightly around the bullseye. This is what you want.
- Low bias, high variability: throws are centered on the bullseye but scattered widely.
- High bias, low variability: throws cluster tightly, but in the wrong spot.
- High bias, high variability: scattered and off-center.
Bias comes from the method, such as a poorly designed sample or a statistic whose average misses the target. Variability comes largely from sample size.
Larger samples, less variability
Sample size and spread
For a statistic from a random sample, larger samples produce less variability. The sampling distribution gets narrower as grows. Increasing the sample size does not reduce bias.
The spread shrinks in proportion to , a fact you will see in formulas in the next two lessons. That square root has a practical consequence: to cut the standard deviation of a statistic in half, you need four times as many observations. To cut it to one third, you need nine times as many.
Worked example: Comparing two estimators
Two statistics, A and B, are proposed to estimate a population parameter . Simulation shows that A's sampling distribution has mean 50 and standard deviation 6, while B's has mean 53 and standard deviation 2. Describe each estimator.
Statistic A is unbiased (centered at 50) but has high variability. Statistic B is biased (centered at 53, not 50) but has low variability. Neither is ideal. Which one to prefer depends on the context: B will usually land within a few units of 53, but it will never be centered on the truth. A larger sample would shrink A's spread while keeping it unbiased, which is often the better fix.
Common mistake
A larger sample does not fix a biased sampling method. If a survey only reaches people who choose to respond online, surveying 10,000 of them instead of 1,000 just gives you a more precise estimate of the wrong number. Sample size controls variability; good design controls bias.
Tip
On AP free-response questions, when asked whether an estimator is biased, refer to the center of its sampling distribution compared with the parameter. When asked about precision, refer to its spread. Keep the two ideas in separate sentences.
Practice
A state agency reports that the mean commute time for all workers in the state is 27.4 minutes. A reporter surveys a random sample of 150 workers and finds a mean commute of 29.1 minutes. Which statement is correct?
Two researchers each take a random sample of 200 adults from the same city and compute the proportion who own a bicycle. One gets and the other gets . What is the best explanation for the difference?
A population consists of the four values . A random sample of size 2 is selected without replacement. What is the probability that the sample mean is at least 6? Give your answer as a fraction or a decimal.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A simulation draws 1,000 random samples of size 25 from a population and records the sample range each time. The population range is 40. The mean of the 1,000 sample ranges is 31. Which statement is best supported?
A polling firm is planning a survey. Sample sizes of 400 and 1,600 are both being considered, and both samples would be selected at random from the same population. Compared with the sampling distribution of for , the sampling distribution for will
A researcher's current plan uses a random sample of 100. She wants the standard deviation of her sample mean to be one third as large as it would be with the current plan. What sample size does she need?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A website asks visitors to click a button if they support a new city park, and 2,300 of 2,500 respondents click "support." A second group takes a simple random sample of 120 city residents and finds that 58% support the park. Which statement is most accurate?