Lesson 9.3 · Statistics and Probability
Experimental design
People who drink more coffee live longer, according to some studies. Does that mean coffee makes you live longer? Maybe, or maybe coffee drinkers differ in other ways. To show that one thing causes another, you need a well-designed experiment. This lesson shows what makes an experiment trustworthy and how to judge whether its result is real or just luck.
Three ways to collect data
Definition
Survey, observational study and experiment
- A survey asks people questions about themselves or their opinions.
- In an observational study, researchers observe and measure individuals without influencing them.
- In an experiment, researchers impose a treatment on the subjects and then measure a response.
The variable the researcher changes is the explanatory variable (or factor), and each value it takes is a treatment. The outcome that is measured is the response variable.
Only experiments can establish cause and effect. Observational studies can show that two variables are related, but not why. The trouble is confounding.
Definition
Confounding variable
A confounding variable is related to both the explanatory variable and the response, so its effect can't be separated from the effect of the explanatory variable.
In the coffee example, people who drink coffee may also have higher incomes, and higher income is linked to better health care. Is it the coffee or the income? An observational study can't tell.
Worked example: Classifying studies
Classify each study and say whether it could show cause and effect.
- Researchers review medical records of adults and compare heart disease rates between those who exercise regularly and those who don't.
- volunteers are randomly split into two groups. One group takes a new allergy medicine for a month; the other takes an identical-looking sugar pill. Symptoms are recorded for both.
- A school asks students how many hours of sleep they usually get.
Solution.
- Observational study: the researchers don't choose who exercises. It can show an association but not causation. (People who exercise may also eat better, for instance.)
- Experiment: the researchers assign the treatment. Because assignment is random, it can show cause and effect.
- Survey. It describes the students but tests no treatment.
Principles of experimental design
Control, randomize, replicate
- Control: keep other variables the same for all groups, and compare the treatment to a control group that gets no treatment, a standard treatment, or a placebo.
- Randomize: use chance to assign subjects to treatments. This balances confounding variables, known and unknown, across the groups.
- Replicate: use enough subjects in each group that natural variation between individuals averages out.
Random assignment is different from random sampling. Random sampling (last lesson) lets you generalize from a sample to a population. Random assignment lets you conclude that the treatment caused the difference.
When subjects are people, their expectations can affect results. A placebo is a fake treatment, like a sugar pill, that looks identical to the real one. People often improve simply because they believe they are being treated; this is the placebo effect. In a single-blind experiment, the subjects don't know which treatment they receive. In a double-blind experiment, neither the subjects nor the people measuring the response know. Blinding keeps expectations from biasing the results.
Common designs
- Completely randomized design: every subject is randomly assigned to one of the treatments.
- Randomized block design: subjects are first split into blocks of similar individuals (for example, by age group), and then randomly assigned to treatments within each block. This reduces variation caused by the blocking variable.
- Matched pairs design: subjects are paired up by similarity, and one member of each pair gets each treatment at random. Often each subject is their own "pair": they receive both treatments in random order.
Worked example: Designing an experiment
A company wants to know whether a new sunscreen protects better than its old formula. It has volunteers. Describe a matched pairs design.
Solution. Use each volunteer as their own pair. For each person, flip a coin to decide whether the new sunscreen goes on the left arm or the right arm; the old formula goes on the other arm. After a set time in the sun, a technician who doesn't know which arm got which sunscreen rates the redness of each arm (this makes the measurement blind). Compare the redness scores within each person. Pairing controls for skin type, and the coin flip prevents a left-arm or right-arm effect from being confused with the sunscreen.
Is the difference real?
Suppose the treatment group does better on average. Even if the treatment did nothing, the two groups would rarely have exactly equal means, just because of how the subjects happened to be split. The question is whether the observed difference is bigger than chance alone would usually produce.
A randomization test answers this with a simulation:
- Pool all the response values together.
- Randomly re-split them into two groups of the original sizes and compute the difference in means.
- Repeat many times (hundreds or thousands).
- Find the proportion of re-randomizations with a difference at least as large as the one actually observed.
If that proportion is very small (a common cutoff is less than ), the result is called statistically significant: chance alone rarely produces such a large difference, so the treatment probably had an effect.
Worked example: A plant growth experiment
Ten seedlings are randomly assigned to a new fertilizer or to plain water. After three weeks, their heights in centimeters are:
| Group | Heights (cm) |
|---|---|
| Fertilizer | |
| Water |
Find the difference in means. Then interpret a randomization test in which re-randomizations produced a difference at least this large times.
Solution. Fertilizer mean: . Water mean: . The difference is cm.
Chance alone produced a difference of cm or more in only of re-randomizations. Since , the difference is statistically significant: the evidence suggests the fertilizer increases growth.
Common mistake
"Not statistically significant" does not mean "the treatment has no effect." It means the experiment didn't give strong enough evidence. A larger experiment might.
Practice
Researchers record the diets of adults for ten years and find that those who eat more fish have fewer heart attacks. Which statement is correct?
What is the main purpose of randomly assigning subjects to treatment groups?
A study finds that students who own more books tend to have higher reading scores. Which is the most likely confounding variable?
To test two keyboard layouts, each of volunteers types a passage on both layouts, with the order chosen by a coin flip. Typing speeds on the two layouts are compared for each person. What design is this?
In an experiment, a treatment group scored and a control group scored . Find the difference in means (treatment minus control).
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In a randomization test with re-randomizations, produced a difference in means at least as large as the observed difference. What percent of re-randomizations is this?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In a trial of a new pain reliever, patients don't know whether they got the drug or a placebo, and the nurses who rate patients' pain also don't know. What is this called, and why does it matter?