Lesson 3.2 · Collecting Data
Bias in sampling
A sample of 10,000 people sounds impressive, but size alone can't rescue a badly designed survey. In this lesson you'll learn to spot the ways a sample can systematically miss the truth, and to explain in context which way the results are likely to be off.
What bias means
Definition
Bias
A sampling method is biased if it tends to produce estimates that are systematically too high or systematically too low compared with the true population parameter. Bias is a property of the method, not of one unlucky sample.
Every random sample differs a little from the population just by chance. That is sampling variability, and it is not bias. Sampling variability goes both ways and shrinks as the sample gets larger. Bias pushes estimates in the same direction every time you repeat the method, and taking a larger sample does not fix it.
On the AP exam, a good answer about bias does three things: names the problem, explains why the people in the sample are likely to differ from the population, and states the direction of the bias (overestimate or underestimate) in context.
Bias from how the sample is chosen
Convenience sampling
A convenience sample consists of individuals who are easiest to reach. Asking the first 50 people who walk into the student center, or surveying your own friends, are convenience samples. The people who are easy to reach are rarely typical of the population.
Voluntary response sampling
A voluntary response sample consists of people who choose themselves by responding to a general invitation: an online poll on a website, a call-in radio question, a comment card at a restaurant. People with strong opinions, especially strong negative opinions, are the most likely to respond, so the results tend to exaggerate strong views.
Undercoverage
Undercoverage occurs when some members of the population have a smaller chance (or no chance) of being included in the sample, often because the sampling frame leaves them out. A telephone survey that dials only landline numbers undercovers younger adults, who are much less likely to have a landline. Undercoverage can happen even when the sample is chosen at random from the frame, because the frame itself is incomplete.
Bias after the sample is chosen
Randomly selecting the sample protects you from the problems above, but two more kinds of bias can creep in when you try to get answers from the people you selected.
Nonresponse bias
Nonresponse occurs when an individual chosen for the sample can't be contacted or refuses to participate. If the people who don't respond differ in an important way from those who do, the results are biased. A survey about how busy people are, conducted by phone at 6 p.m., will hear mostly from people who are home and not busy.
The response rate is the fraction of selected individuals who actually respond:
Response bias
Response bias is any systematic pattern of inaccurate answers. Common causes include:
- Question wording. Leading or loaded questions push people toward one answer. "Do you support the reckless plan to cut the library's hours?" will get fewer yes answers than a neutral version.
- Sensitive topics and social desirability. People underreport behaviors they are embarrassed about and overreport behaviors that sound good, such as voting or exercising.
- Interviewer effects. The interviewer's appearance, tone or reactions can change what people are willing to say.
- Faulty memory. People asked how much they spent on snacks last month often guess badly, and not randomly.
Undercoverage vs. nonresponse
Undercoverage: some people were never given a chance to be selected.
Nonresponse: people were selected but didn't give an answer.
Ask yourself, "Were these people in the sample and just didn't answer, or were they never in the sample at all?"
Worked example: Voluntary response
A local news website asks, "Should the city build a new downtown parking garage?" and invites readers to click Yes or No. Of 3,200 votes, 78% say No. Explain why this result is likely biased and give the direction of the bias.
Solution. This is a voluntary response sample. Readers who feel strongly, especially residents who are angry about the cost or construction, are much more likely to take the time to vote. The 78% likely overestimates the proportion of all city residents who oppose the garage. The large number of votes does not help, because the method, not the sample size, is the problem.
Worked example: Undercoverage and its direction
To estimate the proportion of adults in a county who have a college degree, researchers take a random sample of names from a list of people who own homes in the county. Identify a source of bias and its likely direction.
Solution. The sampling frame leaves out renters, so this is undercoverage. Renters are, on average, younger and have lower incomes than homeowners, and as a group they are less likely to have a college degree. Because the sample includes only homeowners, the estimate will probably overestimate the proportion of all county adults with a college degree.
Worked example: Nonresponse versus response bias
A school district randomly selects 400 parents and mails them a survey that includes the question "How many hours per week do you spend helping your child with homework?" Only 132 parents return it.
(a) Find the response rate.
(b) Describe one possible source of nonresponse bias and one possible source of response bias, with directions.
Solution.
(a) , a response rate of 33%.
(b) Nonresponse bias: parents who are very involved in their children's schoolwork are probably more likely to fill out and return a survey about homework. The 132 respondents likely overestimate the mean homework-help time for all parents in the district.
Response bias: parents may overstate their hours because helping with homework is seen as what "good parents" do. This would also push the estimate too high.
Common mistake
"Get a bigger sample" is never a fix for bias. A voluntary response sample of 50,000 is just as biased as one of 500; it merely gives a more precise estimate of the wrong value. Larger samples reduce variability, not bias. The fix for bias is a better method: random selection, a complete sampling frame, follow-ups with nonrespondents, and neutral question wording.
Tip
Random selection fixes convenience and voluntary response bias. It does not fix undercoverage in the frame, nonresponse or response bias. Those can affect even a perfectly random sample.
Practice
A sports-talk radio host asks listeners to call in and say whether the city's football coach should be fired. Of the 210 callers, 81% say yes. What is the most serious source of bias?
A polling organization selects a random sample of landline telephone numbers in a state to estimate the proportion of adults who stream music every day. Which statement best describes the likely bias?
A researcher randomly selects 800 customers and emails them a satisfaction survey. 232 customers respond. What is the response rate, as a percent?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A random sample of residents is asked: "Given the dangerous condition of our crumbling bridges, do you support a small tax increase to repair them?" 84% answer yes. What is the most likely problem with this survey?
A voluntary online poll about a proposed school uniform policy received 600 responses. The organizers plan to leave the poll open longer so they can collect 6,000 responses. What effect will the larger sample have?
To estimate the average number of hours per week that adults in a town exercise, a researcher surveys 150 people as they leave the town's largest gym. How will this method most likely affect the estimate?
A university randomly selects 500 students from its complete enrollment list and asks them to complete a 45-minute in-person interview about their study habits. Only 190 students show up. Which type of bias is the biggest concern?