Math Core

Lesson 9.3 · Statistics and Probability

Experimental design

People who drink more coffee live longer, according to some studies. Does that mean coffee makes you live longer? Maybe, or maybe coffee drinkers differ in other ways. To show that one thing causes another, you need a well-designed experiment. This lesson shows what makes an experiment trustworthy and how to judge whether its result is real or just luck.

Three ways to collect data

Definition

Survey, observational study and experiment

  • A survey asks people questions about themselves or their opinions.
  • In an observational study, researchers observe and measure individuals without influencing them.
  • In an experiment, researchers impose a treatment on the subjects and then measure a response.

The variable the researcher changes is the explanatory variable (or factor), and each value it takes is a treatment. The outcome that is measured is the response variable.

Only experiments can establish cause and effect. Observational studies can show that two variables are related, but not why. The trouble is confounding.

Definition

Confounding variable

A confounding variable is related to both the explanatory variable and the response, so its effect can't be separated from the effect of the explanatory variable.

In the coffee example, people who drink coffee may also have higher incomes, and higher income is linked to better health care. Is it the coffee or the income? An observational study can't tell.

Worked example: Classifying studies

Classify each study and say whether it could show cause and effect.

  1. Researchers review medical records of 5,0005{,}000 adults and compare heart disease rates between those who exercise regularly and those who don't.
  2. 6060 volunteers are randomly split into two groups. One group takes a new allergy medicine for a month; the other takes an identical-looking sugar pill. Symptoms are recorded for both.
  3. A school asks 300300 students how many hours of sleep they usually get.

Solution.

  1. Observational study: the researchers don't choose who exercises. It can show an association but not causation. (People who exercise may also eat better, for instance.)
  2. Experiment: the researchers assign the treatment. Because assignment is random, it can show cause and effect.
  3. Survey. It describes the students but tests no treatment.

Principles of experimental design

Control, randomize, replicate

  • Control: keep other variables the same for all groups, and compare the treatment to a control group that gets no treatment, a standard treatment, or a placebo.
  • Randomize: use chance to assign subjects to treatments. This balances confounding variables, known and unknown, across the groups.
  • Replicate: use enough subjects in each group that natural variation between individuals averages out.

Random assignment is different from random sampling. Random sampling (last lesson) lets you generalize from a sample to a population. Random assignment lets you conclude that the treatment caused the difference.

When subjects are people, their expectations can affect results. A placebo is a fake treatment, like a sugar pill, that looks identical to the real one. People often improve simply because they believe they are being treated; this is the placebo effect. In a single-blind experiment, the subjects don't know which treatment they receive. In a double-blind experiment, neither the subjects nor the people measuring the response know. Blinding keeps expectations from biasing the results.

Common designs

  • Completely randomized design: every subject is randomly assigned to one of the treatments.
  • Randomized block design: subjects are first split into blocks of similar individuals (for example, by age group), and then randomly assigned to treatments within each block. This reduces variation caused by the blocking variable.
  • Matched pairs design: subjects are paired up by similarity, and one member of each pair gets each treatment at random. Often each subject is their own "pair": they receive both treatments in random order.

Worked example: Designing an experiment

A company wants to know whether a new sunscreen protects better than its old formula. It has 4040 volunteers. Describe a matched pairs design.

Solution. Use each volunteer as their own pair. For each person, flip a coin to decide whether the new sunscreen goes on the left arm or the right arm; the old formula goes on the other arm. After a set time in the sun, a technician who doesn't know which arm got which sunscreen rates the redness of each arm (this makes the measurement blind). Compare the redness scores within each person. Pairing controls for skin type, and the coin flip prevents a left-arm or right-arm effect from being confused with the sunscreen.

Is the difference real?

Suppose the treatment group does better on average. Even if the treatment did nothing, the two groups would rarely have exactly equal means, just because of how the subjects happened to be split. The question is whether the observed difference is bigger than chance alone would usually produce.

A randomization test answers this with a simulation:

  1. Pool all the response values together.
  2. Randomly re-split them into two groups of the original sizes and compute the difference in means.
  3. Repeat many times (hundreds or thousands).
  4. Find the proportion of re-randomizations with a difference at least as large as the one actually observed.

If that proportion is very small (a common cutoff is less than 5%5\%), the result is called statistically significant: chance alone rarely produces such a large difference, so the treatment probably had an effect.

Worked example: A plant growth experiment

Ten seedlings are randomly assigned to a new fertilizer or to plain water. After three weeks, their heights in centimeters are:

GroupHeights (cm)
Fertilizer14,17,15,18,1614, 17, 15, 18, 16
Water13,15,12,14,1113, 15, 12, 14, 11

Find the difference in means. Then interpret a randomization test in which 1,0001{,}000 re-randomizations produced a difference at least this large 2121 times.

Solution. Fertilizer mean: 14+17+15+18+165=805=16\dfrac{14 + 17 + 15 + 18 + 16}{5} = \dfrac{80}{5} = 16. Water mean: 13+15+12+14+115=655=13\dfrac{13 + 15 + 12 + 14 + 11}{5} = \dfrac{65}{5} = 13. The difference is 16−13=316 - 13 = 3 cm.

Chance alone produced a difference of 33 cm or more in only 211000=2.1%\dfrac{21}{1000} = 2.1\% of re-randomizations. Since 2.1%<5%2.1\% < 5\%, the difference is statistically significant: the evidence suggests the fertilizer increases growth.

Common mistake

"Not statistically significant" does not mean "the treatment has no effect." It means the experiment didn't give strong enough evidence. A larger experiment might.

Practice

Practice 1

Researchers record the diets of 2,0002{,}000 adults for ten years and find that those who eat more fish have fewer heart attacks. Which statement is correct?

Practice 2

What is the main purpose of randomly assigning subjects to treatment groups?

Practice 3

A study finds that students who own more books tend to have higher reading scores. Which is the most likely confounding variable?

Practice 4

To test two keyboard layouts, each of 2020 volunteers types a passage on both layouts, with the order chosen by a coin flip. Typing speeds on the two layouts are compared for each person. What design is this?

Practice 5

In an experiment, a treatment group scored 12,15,11,14,1312, 15, 11, 14, 13 and a control group scored 10,12,9,11,1310, 12, 9, 11, 13. Find the difference in means (treatment minus control).

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

In a randomization test with 500500 re-randomizations, 99 produced a difference in means at least as large as the observed difference. What percent of re-randomizations is this?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

In a trial of a new pain reliever, patients don't know whether they got the drug or a placebo, and the nurses who rate patients' pain also don't know. What is this called, and why does it matter?