Math Core

Lesson 2.1 · Exploring Two-Variable Data

Two-way tables

Unit 1 was about one variable at a time. Now the question changes: is one variable related to another? When both variables are categorical, such as a student's grade level and their favorite lunch, a two-way table organizes the counts, and comparing the right proportions tells you whether the variables are associated.

Two-way tables and relative frequencies

A researcher surveyed 250250 households about where they live and what kind of pet they have. Each household falls into exactly one cell.

DogCatNo petTotal
House727230301818120120
Apartment383842425050130130
Total11011072726868250250

Raw counts are hard to compare when groups have different sizes, so statisticians convert counts to relative frequencies (proportions or percents). There are three kinds, and the denominator is what tells them apart.

Definition

Joint, marginal and conditional relative frequencies

  • A joint relative frequency is a cell count divided by the grand total. The proportion of all households that live in an apartment and own a cat is 42250=0.168\dfrac{42}{250} = 0.168.
  • A marginal relative frequency is a row or column total divided by the grand total. The proportion of all households with a dog is 110250=0.44\dfrac{110}{250} = 0.44.
  • A conditional relative frequency is a cell count divided by the total of its own row or column. The proportion of house dwellers who own a dog is 72120=0.60\dfrac{72}{120} = 0.60.

The collection of marginal relative frequencies for one variable is its marginal distribution. For pet type: 0.440.44 dog, 0.2880.288 cat, 0.2720.272 no pet. The collection of conditional relative frequencies within one row (or column) is a conditional distribution.

Conditional distributions and association

To decide whether two categorical variables are related, compare the conditional distribution of one variable across the categories of the other.

DogCatNo petTotal
House0.6000.6000.2500.2500.1500.1501.0001.000
Apartment0.2920.2920.3230.3230.3850.3851.0001.000

Each row now sums to 11. Households in houses are about twice as likely to own a dog (60%60\% versus 29.2%29.2\%), while apartment households are much more likely to have no pet (38.5%38.5\% versus 15%15\%).

Association between categorical variables

Two categorical variables are associated if knowing the value of one changes the distribution of the other. In practice: if the conditional distributions of one variable are different across the categories of the other, there is an association. If they are the same (or nearly the same), there is no association.

A segmented bar graph shows this at a glance. Each category of the explanatory variable (housing type) gets one bar of height 100%100\%, and the bar is divided into segments whose heights are the conditional relative frequencies. If the segments line up across bars, there's no association; if they are noticeably different, there is. A mosaic plot goes one step further and makes each bar's width proportional to the size of its group.

Common mistake

Watch the denominator. "The proportion of dog owners who live in a house" is 72110≈0.655\dfrac{72}{110} \approx 0.655 (condition on dog owners, so divide by the dog column). "The proportion of house dwellers who own a dog" is 72120=0.60\dfrac{72}{120} = 0.60 (condition on houses). The phrase right after "of" names the group you divide by.

Worked example: Reading all three kinds of frequency

Use the pet table. Find (a) the proportion of all households that live in a house and have no pet, (b) the proportion of households that live in an apartment, and (c) the proportion of cat owners who live in an apartment.

(a) This is joint: 18250=0.072\dfrac{18}{250} = 0.072.

(b) This is marginal: 130250=0.52\dfrac{130}{250} = 0.52.

(c) This is conditional on cat owners, so divide by the cat column total: 4272≈0.583\dfrac{42}{72} \approx 0.583.

Worked example: Is there an association?

A school asked 300300 students whether they take the bus to school and whether they were late at least once last month.

Late at least onceNever lateTotal
Bus36368484120120
No bus5454126126180180
Total9090210210300300

Is there an association between riding the bus and being late?

Compare the conditional proportions of lateness:

Bus: 36120=0.30No bus: 54180=0.30\text{Bus: } \frac{36}{120} = 0.30 \qquad \text{No bus: } \frac{54}{180} = 0.30

The proportions are identical, so in this sample there is no association between riding the bus and being late. Knowing how a student gets to school tells you nothing extra about whether they were late.

Simpson's paradox

Combining groups can hide, or even reverse, an association. This surprising effect is called Simpson's paradox, and it is caused by a lurking variable that is related to both variables in the table.

Worked example: Which hospital is better?

Two hospitals each performed a certain surgery on 1,0001{,}000 patients. Patients arrived in either good or poor condition.

Survived (good)Total (good)Survived (poor)Total (poor)
Hospital A590590600600210210400400
Hospital B8708709009004040100100

Overall, Hospital A saved 590+210=800590 + 210 = 800 of 1,0001{,}000 patients (80%80\%) and Hospital B saved 870+40=910870 + 40 = 910 of 1,0001{,}000 (91%91\%). Hospital B looks better.

Now compare within each condition:

  • Good condition: A saved 590600≈98.3%\dfrac{590}{600} \approx 98.3\% and B saved 870900≈96.7%\dfrac{870}{900} \approx 96.7\%.
  • Poor condition: A saved 210400=52.5%\dfrac{210}{400} = 52.5\% and B saved 40100=40%\dfrac{40}{100} = 40\%.

Hospital A does better for both kinds of patient. The overall comparison reverses because Hospital A treats far more patients in poor condition (40%40\% of its patients versus 10%10\% at B), and those patients survive less often no matter where they go. Patient condition is the lurking variable.

Tip

Before comparing groups, ask whether some other variable differs between them. When it does, break the table apart by that variable and compare within each piece.

Practice

The next five problems use this table from a survey of 400400 adults.

Drinks coffee dailyDoes notTotal
Age 18–3484845656140140
Age 35–541171173333150150
Age 55+66664444110110
Total267267133133400400
Practice 1

What proportion of all the adults surveyed drink coffee daily? Give the exact decimal.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

What proportion of all the adults surveyed are age 5555 or older and do not drink coffee daily?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

What proportion of daily coffee drinkers are age 1818–3434? Round to three decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

Which statement is best supported by the table?

Practice 5

Suppose age group and coffee drinking had no association at all, and the overall proportion of daily coffee drinkers stayed at 0.66750.6675. How many of the 140140 adults age 1818–3434 would you then expect to drink coffee daily?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

A gym surveyed 200200 members about their favorite type of workout. Fill in the missing entries mentally and find the number of evening members who prefer weights.

CardioWeightsTotal
Morning54549090
Evening110110
Total128128200200

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

Two tutoring programs report pass rates. Program P has a higher pass rate than Program Q for students who started behind and for students who started on track, yet Program Q has a higher pass rate overall. Which explanation is most plausible?