Math Core

Lesson 8.2 · Inference for Categorical Data: Chi-Square

Chi-square tests for homogeneity and independence

In Unit 2 you described association in a two-way table by comparing conditional distributions. But a sample can show differences just by chance. The chi-square tests for homogeneity and independence use the same χ2\chi^2 statistic you met in the goodness-of-fit test to decide whether the pattern in a two-way table is strong enough to be real.

Two questions, one table

Two-way tables arise from two different study designs, and each design asks its own question.

Definition

Homogeneity versus independence

  • A chi-square test for homogeneity compares the distribution of one categorical variable across two or more populations or treatments. The data come from separate random samples (or from groups in a randomized experiment).
  • A chi-square test for independence checks whether two categorical variables are associated in one population. The data come from a single random sample, with each individual classified on both variables.

The quickest way to tell them apart is to ask: how many samples were taken? Several samples (or treatment groups), one variable measured: homogeneity. One sample, two variables measured: independence.

The hypotheses follow the design.

TestH0H_0HaH_a
HomogeneityThe distribution of the variable is the same for every population.The distribution is not the same for all populations.
IndependenceThere is no association between the two variables (they are independent).There is an association between the two variables.

The mechanics that follow are identical for both tests.

Expected counts in a two-way table

A school district randomly samples students from three high schools and asks how they usually get to school.

BusCarWalk/BikeTotal
School A424238382020100100
School B303058583232120120
School C484844442828120120
Total1201201401408080340340

If H0H_0 is true and the commute distribution is the same at all three schools, then each school should match the overall distribution. Overall, 120340\dfrac{120}{340} of students ride the bus, so School A, with 100100 students, should have about 100×120340≈35.29100 \times \dfrac{120}{340} \approx 35.29 bus riders. That reasoning gives a general formula.

Expected counts and degrees of freedom

For each cell of a two-way table,

expected count=(row total)(column total)grand total.\text{expected count} = \frac{(\text{row total})(\text{column total})}{\text{grand total}}.

The statistic is χ2=∑(O−E)2E\chi^2 = \sum \dfrac{(O - E)^2}{E} over every cell, and

df=(number of rows−1)(number of columns−1).\text{df} = (\text{number of rows} - 1)(\text{number of columns} - 1).

Expected counts don't have to be whole numbers; keep a few decimal places. For the commute data, df =(3−1)(3−1)=4= (3-1)(3-1) = 4.

Conditions

The conditions match the goodness-of-fit test, adjusted for the design:

  • Random: separate random samples from each population, or groups formed by random assignment (homogeneity); a single random sample (independence).
  • 10%: each sample is at most 10%10\% of its population when sampling without replacement.
  • Large counts: all expected counts are at least 55.

Worked example: Test for homogeneity: commuting at three schools

Use the commute table. Do the data give convincing evidence at α=0.05\alpha = 0.05 that the distribution of commute method differs among the three schools?

State. H0H_0: the distribution of commute method is the same at all three schools. HaH_a: the distribution is not the same at all three schools.

Plan. Chi-square test for homogeneity. Random: separate random samples from each school. 10%10\%: assume each sample is less than 10%10\% of its school's enrollment (a high school with at least 1,2001{,}200 students). Large counts: the expected counts are

BusCarWalk/Bike
School A35.2935.2941.1841.1823.5323.53
School B42.3542.3549.4149.4128.2428.24
School C42.3542.3549.4149.4128.2428.24

All are at least 55.

Do. The nine contributions are

BusCarWalk/Bike
School A1.2741.2740.2450.2450.5290.529
School B3.6033.6031.4931.4930.5020.502
School C0.7530.7530.5930.5930.0020.002

Summing, χ2≈8.994\chi^2 \approx 8.994 with df =4= 4. In the df =4= 4 row of the table, 8.9948.994 is between 7.7797.779 (0.100.10) and 9.4889.488 (0.050.05). Technology gives P≈0.061P \approx 0.061.

Conclude. Because 0.061>0.050.061 > 0.05, fail to reject H0H_0. There is not convincing evidence that the distribution of commute method differs among the three schools.

Common mistake

Use counts, never row or column percents, in the table you analyze. Also, don't mix up the two tests in your conclusion. A homogeneity conclusion talks about whether distributions differ among populations; an independence conclusion talks about whether two variables are associated. And neither test shows causation unless the data came from a randomized experiment.

Worked example: Test for independence: age and texting

A random sample of 180180 adults in a city was asked their age group and whether they prefer to reach friends by texting or calling.

TextCallTotal
18–29363624246060
30–49292941417070
50+151535355050
Total8080100100180180

Is there convincing evidence at α=0.01\alpha = 0.01 of an association between age group and communication preference in this city? Which cells contribute most?

State. H0H_0: there is no association between age group and preference. HaH_a: there is an association.

Plan. Chi-square test for independence. One random sample of adults; 180180 is less than 10%10\% of the city's adults. Expected counts: 60⋅80180≈26.67\dfrac{60 \cdot 80}{180} \approx 26.67, 33.3333.33; 31.1131.11, 38.8938.89; 22.2222.22, 27.7827.78. All are at least 55.

Do.

TextCall
18–293.2673.2672.6132.613
30–490.1430.1430.1150.115
50+2.3472.3471.8781.878

χ2≈10.363\chi^2 \approx 10.363 with df =(3−1)(2−1)=2= (3-1)(2-1) = 2. In the table, 10.36310.363 is between 9.2109.210 (0.010.01) and 10.59710.597 (0.0050.005). Technology: P≈0.0056P \approx 0.0056.

Conclude. Because 0.0056<0.010.0056 < 0.01, reject H0H_0. There is convincing evidence of an association between age group and communication preference among adults in this city.

The largest contributions come from the youngest group, who texted more than expected (3636 versus 26.6726.67), and the oldest group, who texted less than expected (1515 versus 22.2222.22).

A 2-by-2 table and the two-proportion z test

When the table has just two rows and two columns, a chi-square test for homogeneity asks the same question as a two-sided two-proportion zz test from Unit 6, and the two tests always agree: χ2=z2\chi^2 = z^2, and the P-values are equal.

Worked example: Two ways to test the same 2-by-2 table

In a randomized experiment, 100100 students received a reminder text about a scholarship deadline and 100100 did not. Of those reminded, 4545 applied early; of those not reminded, 3030 applied early. Compare the chi-square statistic with the two-proportion zz statistic.

The table is

Applied earlyDid notTotal
Reminder45455555100100
No reminder30307070100100
Total7575125125200200

Expected counts: 100⋅75200=37.5\dfrac{100 \cdot 75}{200} = 37.5 and 62.562.5 in each row. So

χ2=7.5237.5+7.5262.5+7.5237.5+7.5262.5=1.5+0.9+1.5+0.9=4.8,\chi^2 = \frac{7.5^2}{37.5} + \frac{7.5^2}{62.5} + \frac{7.5^2}{37.5} + \frac{7.5^2}{62.5} = 1.5 + 0.9 + 1.5 + 0.9 = 4.8,

with df =1= 1 and P≈0.028P \approx 0.028.

For the zz test, the pooled proportion is p^C=75200=0.375\hat{p}_C = \dfrac{75}{200} = 0.375, and

z=0.45−0.300.375(0.625)(1100+1100)≈0.150.06847≈2.191.z = \frac{0.45 - 0.30}{\sqrt{0.375(0.625)\left(\dfrac{1}{100} + \dfrac{1}{100}\right)}} \approx \frac{0.15}{0.06847} \approx 2.191.

Indeed 2.1912≈4.802.191^2 \approx 4.80, and the two-sided P-value is also 0.0280.028. Because the reminders were randomly assigned, this result supports a cause-and-effect conclusion.

Tip

When a test is significant, compare the observed and expected counts in the cells with the largest contributions. That turns "there is an association" into a specific, useful statement about how the groups differ.

Practice

Problems 1, 4, 5 and 8 use this table. A random sample of 6060 ninth graders and a separate random sample of 100100 twelfth graders at a large school were asked what setting they prefer while studying.

SilenceMusicBackground noiseTotal
9th grade3030202010106060
12th grade252535354040100100
Total555555555050160160
Practice 1

What is the expected count of ninth graders who prefer silence, assuming the distribution of study setting is the same for both grades?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

A researcher builds a two-way table with 44 rows and 33 columns. How many degrees of freedom does the chi-square test have?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A polling group selects one random sample of 1,0001{,}000 adults and records each person's region of the country and their preferred type of news source. Which test is appropriate for determining whether region and news source are related?

Practice 4

For the study-setting table, calculate the chi-square statistic. Round to two decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

Continue the study-setting problem (χ2≈13.38\chi^2 \approx 13.38). Find the P-value using technology. Round to four decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

Which pair of hypotheses is correct for a chi-square test for independence between pet ownership (dog, cat, none) and housing type (house, apartment) among adults in a city?

Practice 7

A two-way table with 33 rows and 44 columns gives χ2=11.2\chi^2 = 11.2 in a chi-square test for independence. Which conclusion is correct at α=0.05\alpha = 0.05?

Practice 8

The study-setting test is significant. Which statement best describes the cell with the largest contribution to χ2\chi^2?