Math Core

Lesson 8.1 · Inference for Categorical Data: Chi-Square

Chi-square goodness of fit

In Unit 6 you tested a claim about one proportion. But many claims involve a whole distribution across several categories at once: "Our bags are 30% cherry, 25% lime, 25% orange and 20% grape," or "Library visits are spread evenly across the school week." A chi-square goodness-of-fit test checks whether sample counts across several categories are consistent with a claimed distribution.

Observed versus expected counts

Suppose a school librarian believes that checkouts are uniformly distributed across the five weekdays. She takes a random sample of 170170 checkouts from the past year and records the day of each.

DayMonTueWedThuFriTotal
Observed40402828303025254747170170

If the claim is true, each day should get 20%20\% of the checkouts. The expected count for each day is

expected=n×pi=170×0.20=34.\text{expected} = n \times p_i = 170 \times 0.20 = 34.

The observed counts will never match the expected counts exactly, because of sampling variability. The real question is the same one you have asked all year: are the differences bigger than chance alone would reasonably produce?

To answer that, you need one number that measures the total distance between the observed and expected counts.

Definition

Chi-square statistic

For a table of counts,

χ2=∑(Observed−Expected)2Expected,\chi^2 = \sum \frac{(\text{Observed} - \text{Expected})^2}{\text{Expected}},

where the sum is over all categories. Each term (O−E)2E\dfrac{(O - E)^2}{E} is called that category's contribution (or component) to χ2\chi^2.

Squaring makes every difference count positively, and dividing by the expected count scales each difference: being off by 66 matters much more when you expected 1010 than when you expected 1,0001{,}000. A value of χ2\chi^2 near 00 means the data fit the claim well. A large value means the data fit poorly, which is evidence against the claim.

Worked example: Computing the statistic for the library data

Find χ2\chi^2 for the librarian's sample.

Every expected count is 3434. Compute each contribution:

DayOOEE(O−E)2/E(O - E)^2 / E
Mon4040343436/34≈1.05936/34 \approx 1.059
Tue2828343436/34≈1.05936/34 \approx 1.059
Wed3030343416/34≈0.47116/34 \approx 0.471
Thu2525343481/34≈2.38281/34 \approx 2.382
Fri47473434169/34≈4.971169/34 \approx 4.971

Add the contributions: χ2≈1.059+1.059+0.471+2.382+4.971≈9.94\chi^2 \approx 1.059 + 1.059 + 0.471 + 2.382 + 4.971 \approx 9.94. (Adding the unrounded fractions, 33834\dfrac{338}{34}, gives 9.9419.941.)

Friday contributes the most by far: many more checkouts happened on Fridays than the uniform claim predicts.

The chi-square distributions

When the null hypothesis is true, the sampling distribution of the χ2\chi^2 statistic is approximately a chi-square distribution. There is a different chi-square distribution for each number of degrees of freedom. For a goodness-of-fit test,

df=(number of categories)−1.\text{df} = (\text{number of categories}) - 1.

Why minus one? Once you know nn and all but one of the counts, the last count is forced.

Chi-square distributions have three features worth remembering:

  • They take only values χ2≥0\chi^2 \ge 0 and are skewed right.
  • The mean of a chi-square distribution equals its df.
  • As df increases, the peak moves right and the curve becomes more symmetric.
Chi-square density curves with df = 3 (highest peak), df = 5, and df = 8 (lowest, farthest right). Each is skewed right, and the center moves right as df grows.Open in grapher →

Because only large values of χ2\chi^2 count as evidence against H0H_0, the P-value is always the area in the right tail: P-value=P(χdf2≥observed χ2)P\text{-value} = P(\chi^2_{\text{df}} \ge \text{observed } \chi^2).

Finding P-values: technology and the table

On a calculator, χ2cdf(lower=χ2,upper=∞,df)\chi^2\text{cdf}(\text{lower}=\chi^2, \text{upper}=\infty, \text{df}) gives the P-value directly (many calculators also run the whole test with χ2GOF-Test\chi^2\text{GOF-Test}). On the AP exam you may also use a table of critical values. Here is an excerpt. Each entry is the value with the given right-tail area.

df0.100.100.050.050.0250.0250.010.010.0050.0050.0010.001
112.7062.7063.8413.8415.0245.0246.6356.6357.8797.87910.82810.828
224.6054.6055.9915.9917.3787.3789.2109.21010.59710.59713.81613.816
336.2516.2517.8157.8159.3489.34811.34511.34512.83812.83816.26616.266
447.7797.7799.4889.48811.14311.14313.27713.27714.86014.86018.46718.467
559.2369.23611.07011.07012.83312.83315.08615.08616.75016.75020.51520.515
6610.64510.64512.59212.59214.44914.44916.81216.81218.54818.54822.45822.458

For the library data, df =5−1=4= 5 - 1 = 4 and χ2=9.941\chi^2 = 9.941. In the df =4= 4 row, 9.9419.941 falls between 9.4889.488 (tail area 0.050.05) and 11.14311.143 (tail area 0.0250.025). So 0.025<P-value<0.050.025 < P\text{-value} < 0.05. Technology gives P≈0.041P \approx 0.041.

The chi-square distribution with df = 3. The shaded right tail beyond 7.815 has area 0.05.Open in grapher →

Conditions and the full test

Like every significance test, a goodness-of-fit test needs conditions before the chi-square distribution is a trustworthy model.

Chi-square goodness-of-fit test

Hypotheses. H0H_0: the claimed distribution is correct (list the proportions). HaH_a: at least one of the claimed proportions is incorrect.

Conditions.

  • Random: the data come from a random sample (or a randomized experiment).
  • 10%: when sampling without replacement, n≤10%n \le 10\% of the population.
  • Large counts: all expected counts are at least 55.

Statistic. χ2=∑(O−E)2E\chi^2 = \sum \dfrac{(O - E)^2}{E} with df == number of categories −1- 1.

P-value. The right-tail area under the chi-square curve.

Common mistake

The Large Counts condition is about expected counts, not observed counts. An observed count of 33 is fine as long as the matching expected count is at least 55. Also, always use counts, never percents or proportions, in the χ2\chi^2 formula. Plugging in percents changes the value of the statistic and gives a wrong P-value.

Worked example: A full four-step test: fruit chew flavors

A candy company claims its fruit chews are 30%30\% cherry, 25%25\% lime, 25%25\% orange and 20%20\% grape. A student buys a random sample of 200200 chews from many stores and counts 5050 cherry, 5858 lime, 4242 orange and 5050 grape. Is there convincing evidence at α=0.05\alpha = 0.05 that the company's claim is wrong?

State. H0H_0: pcherry=0.30p_{\text{cherry}} = 0.30, plime=0.25p_{\text{lime}} = 0.25, porange=0.25p_{\text{orange}} = 0.25, pgrape=0.20p_{\text{grape}} = 0.20. HaH_a: at least one of these proportions is wrong.

Plan. Chi-square goodness-of-fit test. Random: random sample of chews. 10%10\%: 200200 is less than 10%10\% of all the company's chews. Large counts: the expected counts are 200(0.30)=60200(0.30) = 60, 200(0.25)=50200(0.25) = 50, 5050, and 200(0.20)=40200(0.20) = 40, all at least 55.

Do.

χ2=(50−60)260+(58−50)250+(42−50)250+(50−40)240=1.667+1.28+1.28+2.5≈6.727\chi^2 = \frac{(50-60)^2}{60} + \frac{(58-50)^2}{50} + \frac{(42-50)^2}{50} + \frac{(50-40)^2}{40} = 1.667 + 1.28 + 1.28 + 2.5 \approx 6.727

with df =4−1=3= 4 - 1 = 3. In the table, 6.7276.727 is between 6.2516.251 and 7.8157.815, so 0.05<P<0.100.05 < P < 0.10. Technology: P≈0.081P \approx 0.081.

Conclude. Because 0.081>0.050.081 > 0.05, fail to reject H0H_0. There is not convincing evidence that the company's claimed flavor distribution is wrong.

Notice what the conclusion does not say. Failing to reject H0H_0 does not prove the claim is true; the data are simply consistent with it.

Worked example: Following up a significant result

For the library data, χ2≈9.941\chi^2 \approx 9.941, df =4= 4, and P≈0.041P \approx 0.041. (Conditions: the sample is random, 170170 is less than 10%10\% of a year's checkouts, and all expected counts are 3434.) What should the librarian conclude at α=0.05\alpha = 0.05, and where is the claim failing?

Since 0.041<0.050.041 < 0.05, reject H0H_0. There is convincing evidence that checkouts are not uniformly distributed across the weekdays.

To see why, look at the contributions. Friday's contribution (4.9714.971) is half of the total, and Friday had more checkouts than expected (4747 versus 3434). Thursday is next (2.3822.382), with fewer checkouts than expected (2525 versus 3434). The biggest departure from the claim is a surplus of Friday checkouts.

Tip

A quick sanity check: the expected counts must add up to the same total nn as the observed counts. If they don't, you made an arithmetic slip or used proportions that don't sum to 11.

Practice

Practice 1

A textbook claims that blood types in a region are 42%42\% type O, 40%40\% type A, 11%11\% type B and 7%7\% type AB. In a random sample of 200200 blood donors, what is the expected count of type B donors if the claim is true?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

To test whether a six-sided die is fair, you roll it 120120 times and record how many times each face appears. How many degrees of freedom does the chi-square goodness-of-fit test have?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

In one category of a goodness-of-fit test, the observed count is 5858 and the expected count is 5050. What is this category's contribution to the chi-square statistic?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A seed company says its wildflower mix is 40%40\% poppies, 40%40\% daisies and 20%20\% lupines. A random sample of 100100 seedlings contains 3131 poppies, 4545 daisies and 2424 lupines. Calculate the chi-square statistic.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

Continue the wildflower problem: χ2=3.45\chi^2 = 3.45. Use technology to find the P-value of the goodness-of-fit test. Round to three decimal places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

A researcher tests whether births at a large hospital are equally likely in all four seasons. In a random sample of 200200 births, χ2=3.88\chi^2 = 3.88 with df =3= 3, giving P≈0.275P \approx 0.275. Which conclusion is correct at α=0.05\alpha = 0.05?

Practice 7

A club wants to test whether its members' favorite of five activities matches a national distribution of 35%35\%, 30%30\%, 20%20\%, 12%12\% and 3%3\%. It takes a random sample of 120120 of its 2,0002{,}000 members. Which condition for a chi-square goodness-of-fit test is not met?

Practice 8

A goodness-of-fit test with 66 categories produces χ2=13.1\chi^2 = 13.1. Using the table of critical values, which statement about the P-value is correct?