Math Core

Lesson 3.1 · Collecting Data

Sampling methods

In Units 1 and 2 you described data that was already sitting in front of you. Now you step back and ask where data comes from. How a sample is chosen decides whether its results can tell you anything trustworthy about the larger group you actually care about.

Populations, samples and censuses

Every statistical question is about some population: the entire group of individuals you want information about. The population might be all 2,300 students at a high school, every registered voter in a state, or every battery produced at a factory this week.

Measuring every individual in the population is called a census. A census sounds ideal, but it is usually too slow, too expensive or simply impossible. (If testing a battery's lifetime means running it until it dies, a census would destroy the whole week's production.) So instead you collect data from a sample, a subset of the population, and use it to draw conclusions about the whole population.

A number that describes a population, like the true proportion of all voters who support a ballot measure, is a parameter. A number computed from a sample, like the proportion of 500 surveyed voters who support it, is a statistic. You use statistics to estimate parameters.

The central question of this lesson is: how should you pick the sample? The answer, over and over, is use chance.

Simple random samples

Definition

Simple random sample (SRS)

A simple random sample of size nn is a sample chosen so that every group of nn individuals in the population has an equal chance of being the sample that is selected.

Notice the definition is about groups, not just individuals. It is not enough for every individual to have the same chance of being chosen; every possible set of nn individuals must be equally likely.

To take an SRS you need a sampling frame, a list of the individuals in the population. Then you can select the sample in a few standard ways:

  • Slips of paper (the "hat" method). Write each name on an identical slip, mix the slips thoroughly in a container, and draw nn slips without looking.
  • Technology. Number the individuals 11 to NN. Use a random number generator to produce integers from 11 to NN, ignoring repeats, until you have nn different numbers. The individuals with those numbers form the sample.
  • A table of random digits. Give each individual a label with the same number of digits (for N=850N = 850, use 001001 to 850850). Read groups of three digits from the table, skipping 000000, numbers above 850850, and repeats.

On the AP exam, a complete description of an SRS names the labeling, the random mechanism, what to do with repeats, and how many individuals to select. Sampling in this course is almost always without replacement: once someone is chosen, they can't be chosen again.

Worked example: Describing an SRS in context

A principal wants to survey 40 of the school's 850 students about a new bell schedule. Describe how to select a simple random sample.

Solution. Obtain an alphabetical list of all 850 students and number them from 1 to 850. Use a random number generator to produce integers from 1 to 850. Ignore any repeated numbers and keep going until 40 different numbers are produced. Survey the 40 students whose numbers were selected.

Each part matters: the list (sampling frame), the labels, the random mechanism, the rule for repeats, and the sample size.

Other random sampling methods

An SRS is the baseline, but it isn't always the most practical or the most precise design. Three other random methods show up constantly in AP Statistics.

Stratified random sampling

First divide the population into groups called strata (singular stratum) of individuals that are similar in some way that might affect the response, such as grade level or income bracket. Then take a separate SRS from each stratum and combine them into one sample.

Stratifying works best when individuals within each stratum are similar to one another and the strata differ from each other. It guarantees every stratum is represented, and it usually produces a more precise estimate (less sample-to-sample variability) than an SRS of the same size.

A common approach is proportional allocation: each stratum's share of the sample matches its share of the population.

sample size from a stratum=stratum sizepopulation size×n\text{sample size from a stratum} = \frac{\text{stratum size}}{\text{population size}} \times n

Cluster sampling

Divide the population into groups called clusters, often based on location, such as classrooms, city blocks or apartment buildings. Randomly select some of the clusters, then collect data from every individual in the chosen clusters.

Cluster sampling is used because it is cheaper and more convenient. Surveying every student in 8 randomly chosen homerooms is far easier than tracking down 200 students scattered across the building. Ideally each cluster looks like a small version of the population, so clusters should be mixed (heterogeneous) inside.

Systematic random sampling

Choose a starting point at random from the first kk individuals in an ordered list or line, then select every kkth individual after that. If you want a sample of nn from a population of NN, use k=Nnk = \dfrac{N}{n} (rounded down if needed).

Systematic sampling is easy to carry out in the field, for example selecting every 20th customer leaving a store after picking a random start from 1 to 20. It can go wrong if the list has a repeating pattern.

Strata versus clusters

In a stratified sample, you sample some individuals from every group. Strata should be alike inside and different from each other.

In a cluster sample, you sample every individual from some groups. Clusters should each look like the whole population.

Worked example: Identifying the sampling method

A state park wants feedback from campers. Name the sampling method in each plan.

  1. The park has 120 campsites. A ranger randomly selects 12 campsites and surveys every camper at those sites.
  2. The park separates campers into tent campers and RV campers and randomly selects 30 from each group.
  3. As cars enter the park gate, a ranger rolls a die to pick one of the first 6 cars, then surveys every 6th car after that.
  4. A ranger puts the registration numbers of all 480 current campers into a random number generator and surveys the 50 selected.

Solution.

  1. Cluster sample. The campsites are clusters; some are chosen at random, and everyone in them is surveyed.
  2. Stratified random sample. The strata are tent campers and RV campers, and an SRS is taken from each.
  3. Systematic random sample with k=6k = 6 and a random start.
  4. Simple random sample. Every group of 50 campers is equally likely to be selected.

Worked example: Proportional allocation

A high school has 1,200 students: 360 ninth graders, 324 tenth graders, 288 eleventh graders and 228 twelfth graders. A stratified random sample of 100 students, stratified by grade with proportional allocation, will be selected. How many students should come from each grade?

Solution. Multiply each grade's fraction of the school by 100.

9th: 3601200×100=3010th: 3241200×100=2711th: 2881200×100=2412th: 2281200×100=19\begin{aligned} \text{9th: } & \tfrac{360}{1200} \times 100 = 30 \\ \text{10th: } & \tfrac{324}{1200} \times 100 = 27 \\ \text{11th: } & \tfrac{288}{1200} \times 100 = 24 \\ \text{12th: } & \tfrac{228}{1200} \times 100 = 19 \end{aligned}

Check: 30+27+24+19=10030 + 27 + 24 + 19 = 100. Select an SRS of 30 ninth graders, 27 tenth graders, 24 eleventh graders and 19 twelfth graders.

Worked example: Choosing a design

A researcher wants to estimate the mean number of hours per week that students at a large university spend on paid work. She suspects that commuter students work many more hours than students who live on campus. About 40% of students commute. Should she use a stratified sample or a cluster sample of dormitory floors? Explain.

Solution. A stratified sample, with commuters and residents as strata, is better. Hours worked probably differ a lot between the two groups and are more similar within each group, so stratifying will give a more precise estimate. A cluster sample of dormitory floors would miss commuters entirely, because commuters don't live in the dorms, so it could not represent the whole population.

Common mistake

A stratified sample is not a simple random sample, even though it uses SRSs inside each stratum. In a stratified sample, a group of nn students that all come from one stratum can never be selected, so not every group of nn is equally likely. The same is true of cluster and systematic samples.

Practice

Practice 1

A grocery chain wants to know how customers rate its checkout speed. At one store, a manager randomly picks a number from 1 to 15, surveys that customer in line, and then surveys every 15th customer after that for the rest of the day. What type of sample is this?

Practice 2

A school has 45 homerooms. To learn how students feel about the cafeteria, the student council randomly selects 8 homerooms and surveys every student in those homerooms. What type of sample is this?

Practice 3

A company has 2,400 employees: 1,500 hourly workers, 700 salaried staff and 200 managers. It will select a stratified random sample of 120 employees, using proportional allocation by job type. How many managers should be in the sample?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A county clerk has an ordered list of 1,800 property owners and wants a systematic random sample of 60 of them. After choosing a random starting point, every kkth owner on the list will be selected. What value of kk should the clerk use?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A club has 30 juniors and 20 seniors. The advisor randomly selects 6 juniors and, separately, 4 seniors to attend a meeting. Every club member has a 15\dfrac{1}{5} chance of being chosen. Why is this not a simple random sample of 10 members?

Practice 6

A polling firm estimating support for a new city tax stratifies voters by neighborhood income level (low, middle, high) instead of taking an SRS of the same size. Under which condition does stratifying give the biggest improvement in precision?

Practice 7

A city health department wants to estimate the proportion of households with a working smoke detector. It numbers all 600 residential blocks, randomly selects 25 blocks, and inspects every household on those blocks. What is the main advantage of this design compared with an SRS of the same number of households from across the city?