Lesson 2.1 · Exploring Two-Variable Data
Two-way tables
Unit 1 was about one variable at a time. Now the question changes: is one variable related to another? When both variables are categorical, such as a student's grade level and their favorite lunch, a two-way table organizes the counts, and comparing the right proportions tells you whether the variables are associated.
Two-way tables and relative frequencies
A researcher surveyed households about where they live and what kind of pet they have. Each household falls into exactly one cell.
| Dog | Cat | No pet | Total | |
|---|---|---|---|---|
| House | ||||
| Apartment | ||||
| Total |
Raw counts are hard to compare when groups have different sizes, so statisticians convert counts to relative frequencies (proportions or percents). There are three kinds, and the denominator is what tells them apart.
Definition
Joint, marginal and conditional relative frequencies
- A joint relative frequency is a cell count divided by the grand total. The proportion of all households that live in an apartment and own a cat is .
- A marginal relative frequency is a row or column total divided by the grand total. The proportion of all households with a dog is .
- A conditional relative frequency is a cell count divided by the total of its own row or column. The proportion of house dwellers who own a dog is .
The collection of marginal relative frequencies for one variable is its marginal distribution. For pet type: dog, cat, no pet. The collection of conditional relative frequencies within one row (or column) is a conditional distribution.
Conditional distributions and association
To decide whether two categorical variables are related, compare the conditional distribution of one variable across the categories of the other.
| Dog | Cat | No pet | Total | |
|---|---|---|---|---|
| House | ||||
| Apartment |
Each row now sums to . Households in houses are about twice as likely to own a dog ( versus ), while apartment households are much more likely to have no pet ( versus ).
Association between categorical variables
Two categorical variables are associated if knowing the value of one changes the distribution of the other. In practice: if the conditional distributions of one variable are different across the categories of the other, there is an association. If they are the same (or nearly the same), there is no association.
A segmented bar graph shows this at a glance. Each category of the explanatory variable (housing type) gets one bar of height , and the bar is divided into segments whose heights are the conditional relative frequencies. If the segments line up across bars, there's no association; if they are noticeably different, there is. A mosaic plot goes one step further and makes each bar's width proportional to the size of its group.
Common mistake
Watch the denominator. "The proportion of dog owners who live in a house" is (condition on dog owners, so divide by the dog column). "The proportion of house dwellers who own a dog" is (condition on houses). The phrase right after "of" names the group you divide by.
Worked example: Reading all three kinds of frequency
Use the pet table. Find (a) the proportion of all households that live in a house and have no pet, (b) the proportion of households that live in an apartment, and (c) the proportion of cat owners who live in an apartment.
(a) This is joint: .
(b) This is marginal: .
(c) This is conditional on cat owners, so divide by the cat column total: .
Worked example: Is there an association?
A school asked students whether they take the bus to school and whether they were late at least once last month.
| Late at least once | Never late | Total | |
|---|---|---|---|
| Bus | |||
| No bus | |||
| Total |
Is there an association between riding the bus and being late?
Compare the conditional proportions of lateness:
The proportions are identical, so in this sample there is no association between riding the bus and being late. Knowing how a student gets to school tells you nothing extra about whether they were late.
Simpson's paradox
Combining groups can hide, or even reverse, an association. This surprising effect is called Simpson's paradox, and it is caused by a lurking variable that is related to both variables in the table.
Worked example: Which hospital is better?
Two hospitals each performed a certain surgery on patients. Patients arrived in either good or poor condition.
| Survived (good) | Total (good) | Survived (poor) | Total (poor) | |
|---|---|---|---|---|
| Hospital A | ||||
| Hospital B |
Overall, Hospital A saved of patients () and Hospital B saved of (). Hospital B looks better.
Now compare within each condition:
- Good condition: A saved and B saved .
- Poor condition: A saved and B saved .
Hospital A does better for both kinds of patient. The overall comparison reverses because Hospital A treats far more patients in poor condition ( of its patients versus at B), and those patients survive less often no matter where they go. Patient condition is the lurking variable.
Tip
Before comparing groups, ask whether some other variable differs between them. When it does, break the table apart by that variable and compare within each piece.
Practice
The next five problems use this table from a survey of adults.
| Drinks coffee daily | Does not | Total | |
|---|---|---|---|
| Age 18–34 | |||
| Age 35–54 | |||
| Age 55+ | |||
| Total |
What proportion of all the adults surveyed drink coffee daily? Give the exact decimal.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
What proportion of all the adults surveyed are age or older and do not drink coffee daily?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
What proportion of daily coffee drinkers are age –? Round to three decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Which statement is best supported by the table?
Suppose age group and coffee drinking had no association at all, and the overall proportion of daily coffee drinkers stayed at . How many of the adults age – would you then expect to drink coffee daily?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A gym surveyed members about their favorite type of workout. Fill in the missing entries mentally and find the number of evening members who prefer weights.
| Cardio | Weights | Total | |
|---|---|---|---|
| Morning | |||
| Evening | |||
| Total |
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Two tutoring programs report pass rates. Program P has a higher pass rate than Program Q for students who started behind and for students who started on track, yet Program Q has a higher pass rate overall. Which explanation is most plausible?