Lesson 3.4 · Collecting Data
Experimental design
Knowing that experiments can show cause and effect is only half the story. A careless experiment can be just as misleading as an observational study. This lesson covers the principles that make an experiment trustworthy and the three designs you need to recognize and describe on the AP exam.
Four principles of experimental design
Principles of a well-designed experiment
- Comparison. Use a design that compares two or more treatments.
- Random assignment. Use chance to assign experimental units to treatments.
- Control. Keep other variables that might affect the response the same for all groups.
- Replication. Use enough experimental units in each group that differences caused by the treatments can be distinguished from chance differences between the groups.
Comparison matters because a single group tells you little. If 70% of patients feel better after a new medicine, you don't know how many would have felt better anyway. The comparison group is often a control group that receives an inactive treatment or the current standard treatment.
Random assignment is what makes causal conclusions possible. By letting chance decide who gets which treatment, you create groups that are roughly equivalent at the start, with variables like age, motivation and health balanced out between them. Then, if the groups differ at the end, the treatment is the only systematic difference that could explain it.
Control means holding other variables constant where you can: same dosage schedule, same testing room, same instructions. This reduces variability in the responses and keeps extra variables from becoming confounded with the treatment.
Replication means each treatment is given to many units. With only two subjects per group, one unusual person can swing the results.
Placebos and blinding
People often respond to any treatment, even a fake one, simply because they expect to improve. This is the placebo effect. A placebo is a treatment with no active ingredient, such as a sugar pill or a sham procedure, given to the control group so that both groups have the same expectations.
- In a single-blind experiment, either the subjects or the people who interact with them and measure the response don't know which treatment each subject received.
- In a double-blind experiment, neither the subjects nor those who interact with them and measure the response know which treatment a subject received.
Blinding the subjects prevents the placebo effect from differing between groups. Blinding the evaluators prevents their expectations from influencing how they treat subjects or judge responses.
Completely randomized design
In a completely randomized design, the experimental units are assigned to the treatments completely at random.
A full description of the random assignment names the method and the group sizes. For example, with 60 subjects and 3 treatments: write each subject's name on an identical slip of paper, mix the slips in a container, and draw 20 slips. Those subjects get Treatment A. Draw 20 more for Treatment B, and the remaining 20 get Treatment C. With technology: number the subjects 1 to 60, use a random number generator to pick 20 different numbers for Treatment A, then 20 more different numbers for Treatment B, and assign the rest to Treatment C.
Randomized block design
Sometimes you know in advance that some variable will strongly affect the response. You can form blocks: groups of experimental units that are similar in a way that is expected to affect the response. Then you randomly assign treatments separately within each block.
Definition
Randomized block design
In a randomized block design, the random assignment of experimental units to treatments is carried out separately within each block. Comparing treatments within each block removes the variability caused by the blocking variable, making the treatment effect easier to detect.
Blocking is to experiments what stratifying is to samples: form groups of similar individuals first, then randomize inside each group. Notice that blocking does not replace random assignment; it is done in addition to it.
Matched pairs
A matched pairs design is a special block design with blocks of size 2. Either:
- two very similar units are paired, and chance decides which member of each pair gets each treatment, or
- each subject receives both treatments, in a random order, so each subject acts as their own control.
Worked example: Designing a completely randomized experiment
A researcher wants to know whether a new caffeine-free energy drink improves reaction time. She has 50 volunteers. Describe a completely randomized design with a placebo that is double-blind.
Solution. Number the volunteers 1 to 50. Use a random number generator to select 25 different numbers; those volunteers get the energy drink, and the other 25 get a placebo drink that looks and tastes the same but has no active ingredients. Serve both drinks in identical unlabeled cups prepared by an assistant who does not interact with the subjects. Thirty minutes later, measure each volunteer's reaction time with the same computer test. Neither the subjects nor the person running the reaction test knows who received which drink, so the experiment is double-blind. Compare the mean reaction times of the two groups.
Worked example: Counting units in a block design
A study of a new tutoring program has 60 students: 24 are in Algebra 1 and 36 are in Geometry. The course is expected to affect test scores, so the researcher blocks by course. There are 3 treatments: the new program, the old program and no tutoring. Within each block, students are split equally among the treatments.
(a) How many Algebra 1 students receive each treatment?
(b) How many students in total receive the new program?
Solution.
(a) Algebra 1 students per treatment.
(b) Each treatment gets Geometry students, so the new program has students in total.
The random assignment is done twice: once among the 24 Algebra 1 students and once among the 36 Geometry students.
Worked example: Why block?
A company is testing two formulas of running-shoe soles to see which wears down more slowly. Ten runners will each wear shoes for a month. Explain why a matched pairs design, with each runner wearing formula A on one foot and formula B on the other, is better than giving 5 runners formula A and 5 runners formula B.
Solution. Runners differ a lot in how much and how hard they run, which strongly affects wear. In the completely randomized design, that runner-to-runner variability could hide a real difference between the formulas. With matched pairs, both formulas experience the same runner, the same miles and the same terrain, so comparing the two feet of each runner removes that variability. The foot that gets formula A should be decided at random (for example, by a coin flip for each runner), because many runners wear one foot more than the other.
Statistical significance
Even with no real treatment effect, random assignment will produce groups with slightly different results just by chance. An observed difference is called statistically significant if it is so large that it would rarely occur by chance alone. Only then do we have convincing evidence that the treatment caused the difference. You will learn how to measure "rarely" precisely in the inference units.
Common mistake
Don't confuse random selection with random assignment. Random selection chooses who is in the study (a sampling idea). Random assignment decides which treatment each unit gets (an experimental design idea). Many experiments use volunteers, with no random selection at all, and still use random assignment.
Tip
When you describe a block design, say what the blocks are, why that variable is expected to affect the response, and that random assignment is done within each block.
Practice
To test whether a new hand lotion reduces dryness, each of 40 volunteers applies the new lotion to one hand and a standard lotion to the other. A coin flip determines which hand gets the new lotion. What design is this?
What is the main purpose of randomly assigning subjects to treatments in an experiment?
An agronomist has 90 plots spread across 3 fields, with 30 plots in each field. Soil quality differs from field to field, so she blocks by field and tests 5 fertilizer treatments, assigning the treatments at random within each field so that every treatment appears equally often in each field. How many plots in each field receive each fertilizer?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In the agronomist's experiment from the previous problem, how many plots in total receive fertilizer A?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In a study of a new migraine medication, patients receive either the medication or an identical-looking placebo. Neither the patients nor the doctors who assess their pain levels know which pill each patient took. This experiment is
A gym tests two strength-training programs on 80 members: 40 have lifted weights for years and 40 are beginners. The researchers randomly assign 20 experienced members and 20 beginners to each program. Why did they block by experience?
Why do medical experiments often give the control group a placebo instead of no treatment at all?
In a randomized experiment, students who used a new reading app had a mean reading-score gain that was 4 points higher than students who did not. The researchers say the difference is statistically significant. What does that mean?