Lesson 7.4 · Inference for Means
Paired data
Does a typing course actually make people faster? One approach is to measure each person's speed before and after the course. Because each "before" value is linked to one specific "after" value, the data come in pairs, and the smart move is to analyze the differences. This lesson shows why that works so well and how it turns a two-measurement problem into the one-sample procedures you already know.
What makes data paired
Definition
Paired data
Data are paired when each observation in one group is naturally matched with exactly one observation in the other group. Common sources of pairing:
- the same individual measured twice (before and after, left hand and right hand, two treatments in random order);
- matched pairs of different individuals who are similar in important ways (twins, siblings, subjects matched by age and fitness), with one of each pair getting each treatment.
The test for pairing is simple: could you draw a line connecting each value in one list to one specific value in the other list, for a reason built into the study design? If yes, the data are paired. If the two groups were chosen or assigned separately, the data are two independent samples (next lesson).
Why analyze differences?
People differ a lot from one another. Some typists are fast and some are slow, and that person-to-person variation is much bigger than the improvement a course might produce. If you compared the whole "before" group to the whole "after" group, that large variation would drown out the effect.
When you subtract within each pair, each person serves as their own control. A fast typist's high speed shows up in both measurements and cancels out. What's left is the change, which is exactly what you care about.
Paired t procedures
For paired data, compute the difference for each pair (always in the same order), and then use one-sample procedures on the differences.
- Parameter: , the true mean difference.
- Interval: .
- Test statistic: , usually testing .
Here is the number of pairs, and .
Conditions for paired data
The conditions are the one-sample conditions, applied to the differences.
- Random: the pairs are a random sample, or the treatments were randomly assigned within each pair (for example, a random order for the two treatments).
- 10%: when sampling without replacement, the number of pairs is at most 10% of the population.
- Normal/Large Sample: the population of differences is approximately normal, or the number of pairs is at least 30. With fewer than 30 pairs, graph the differences and check for strong skewness and outliers.
Common mistake
Check normality on the differences, not on the "before" and "after" lists separately. Two skewed lists can have perfectly well-behaved differences, and the differences are the only data the procedure uses.
A complete paired test
Worked example: Does the typing course help?
Eight randomly selected employees took a typing course. Their speeds in words per minute:
| Employee | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Before | 42 | 51 | 38 | 47 | 55 | 40 | 49 | 44 |
| After | 46 | 53 | 43 | 47 | 60 | 45 | 50 | 49 |
| Difference (after − before) | 4 | 2 | 5 | 0 | 5 | 5 | 1 | 5 |
Do the data give convincing evidence at that the course increases mean typing speed?
State. and , where is the true mean increase in typing speed (after − before, in wpm) for employees who take the course.
Plan. Paired test.
- Random: the employees were randomly selected.
- 10%: is less than 10% of all employees at the company.
- Normal/Large Sample: there are only pairs, so graph the differences.
The dotplot is somewhat left-skewed but has no outliers, so the paired test is reasonable.
Do. For the differences, and .
P-value .
Conclude. Because , we reject . There is convincing evidence that the course increases mean typing speed for employees like these.
Here is the -table excerpt you'll need.
| tail 0.05 | tail 0.025 | tail 0.01 | tail 0.005 | |
|---|---|---|---|---|
| confidence | 90% | 95% | 98% | 99% |
| 5 | 2.015 | 2.571 | 3.365 | 4.032 |
| 7 | 1.895 | 2.365 | 2.998 | 3.499 |
| 14 | 1.761 | 2.145 | 2.624 | 2.977 |
Worked example: Estimating the size of the effect
Construct a 95% confidence interval for the mean increase in typing speed from the previous example.
Solution. With , .
We are 95% confident that the interval from about to words per minute captures the true mean increase in typing speed for employees who take the course. The interval lies entirely above , which is consistent with the test.
Notice what would happen if you ignored the pairing. The "before" speeds have a standard deviation of about wpm, far larger than the differences' . A two-sample analysis would use that big person-to-person variation and could easily miss the improvement.
Tip
Say the order of subtraction out loud, and define with it: "after − before." Then check that the sign of matches. If the course helps, after − before should be positive, so .
Practice
Which study produces paired data?
Six randomly chosen students solved a logic puzzle before and after a week of practice. Their times in minutes:
| Student | A | B | C | D | E | F |
|---|---|---|---|---|---|---|
| Before | 12.1 | 10.4 | 13.8 | 9.9 | 11.5 | 12.6 |
| After | 11.3 | 10.6 | 12.5 | 9.1 | 11.0 | 11.8 |
Find the mean difference (before − after), to 3 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In a paired study of patients, the mean drop in systolic blood pressure after a new diet was mmHg with mmHg. Find the statistic for , to 2 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
For the blood pressure study in the previous problem, and . Find the P-value to 4 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Using the same study ( pairs, , ), construct a 95% confidence interval for the true mean drop in blood pressure. Enter the endpoints to 2 decimal places, separated by a comma.
Separate answers with commas, e.g. 2, -5
The P-value for the blood pressure study is about . Which is the best conclusion at ?
A student analyzes the typing-course data from this lesson by treating "before" and "after" as two independent samples. Why is that a poor choice?
A researcher has paired data from subjects. Which graph should she examine to check the Normal/Large Sample condition for a paired test?