Lesson 8.1 · Inference for Categorical Data: Chi-Square
Chi-square goodness of fit
In Unit 6 you tested a claim about one proportion. But many claims involve a whole distribution across several categories at once: "Our bags are 30% cherry, 25% lime, 25% orange and 20% grape," or "Library visits are spread evenly across the school week." A chi-square goodness-of-fit test checks whether sample counts across several categories are consistent with a claimed distribution.
Observed versus expected counts
Suppose a school librarian believes that checkouts are uniformly distributed across the five weekdays. She takes a random sample of checkouts from the past year and records the day of each.
| Day | Mon | Tue | Wed | Thu | Fri | Total |
|---|---|---|---|---|---|---|
| Observed |
If the claim is true, each day should get of the checkouts. The expected count for each day is
The observed counts will never match the expected counts exactly, because of sampling variability. The real question is the same one you have asked all year: are the differences bigger than chance alone would reasonably produce?
To answer that, you need one number that measures the total distance between the observed and expected counts.
Definition
Chi-square statistic
For a table of counts,
where the sum is over all categories. Each term is called that category's contribution (or component) to .
Squaring makes every difference count positively, and dividing by the expected count scales each difference: being off by matters much more when you expected than when you expected . A value of near means the data fit the claim well. A large value means the data fit poorly, which is evidence against the claim.
Worked example: Computing the statistic for the library data
Find for the librarian's sample.
Every expected count is . Compute each contribution:
| Day | |||
|---|---|---|---|
| Mon | |||
| Tue | |||
| Wed | |||
| Thu | |||
| Fri |
Add the contributions: . (Adding the unrounded fractions, , gives .)
Friday contributes the most by far: many more checkouts happened on Fridays than the uniform claim predicts.
The chi-square distributions
When the null hypothesis is true, the sampling distribution of the statistic is approximately a chi-square distribution. There is a different chi-square distribution for each number of degrees of freedom. For a goodness-of-fit test,
Why minus one? Once you know and all but one of the counts, the last count is forced.
Chi-square distributions have three features worth remembering:
- They take only values and are skewed right.
- The mean of a chi-square distribution equals its df.
- As df increases, the peak moves right and the curve becomes more symmetric.
Because only large values of count as evidence against , the P-value is always the area in the right tail: .
Finding P-values: technology and the table
On a calculator, gives the P-value directly (many calculators also run the whole test with ). On the AP exam you may also use a table of critical values. Here is an excerpt. Each entry is the value with the given right-tail area.
| df | ||||||
|---|---|---|---|---|---|---|
For the library data, df and . In the df row, falls between (tail area ) and (tail area ). So . Technology gives .
Conditions and the full test
Like every significance test, a goodness-of-fit test needs conditions before the chi-square distribution is a trustworthy model.
Chi-square goodness-of-fit test
Hypotheses. : the claimed distribution is correct (list the proportions). : at least one of the claimed proportions is incorrect.
Conditions.
- Random: the data come from a random sample (or a randomized experiment).
- 10%: when sampling without replacement, of the population.
- Large counts: all expected counts are at least .
Statistic. with df number of categories .
P-value. The right-tail area under the chi-square curve.
Common mistake
The Large Counts condition is about expected counts, not observed counts. An observed count of is fine as long as the matching expected count is at least . Also, always use counts, never percents or proportions, in the formula. Plugging in percents changes the value of the statistic and gives a wrong P-value.
Worked example: A full four-step test: fruit chew flavors
A candy company claims its fruit chews are cherry, lime, orange and grape. A student buys a random sample of chews from many stores and counts cherry, lime, orange and grape. Is there convincing evidence at that the company's claim is wrong?
State. : , , , . : at least one of these proportions is wrong.
Plan. Chi-square goodness-of-fit test. Random: random sample of chews. : is less than of all the company's chews. Large counts: the expected counts are , , , and , all at least .
Do.
with df . In the table, is between and , so . Technology: .
Conclude. Because , fail to reject . There is not convincing evidence that the company's claimed flavor distribution is wrong.
Notice what the conclusion does not say. Failing to reject does not prove the claim is true; the data are simply consistent with it.
Worked example: Following up a significant result
For the library data, , df , and . (Conditions: the sample is random, is less than of a year's checkouts, and all expected counts are .) What should the librarian conclude at , and where is the claim failing?
Since , reject . There is convincing evidence that checkouts are not uniformly distributed across the weekdays.
To see why, look at the contributions. Friday's contribution () is half of the total, and Friday had more checkouts than expected ( versus ). Thursday is next (), with fewer checkouts than expected ( versus ). The biggest departure from the claim is a surplus of Friday checkouts.
Tip
A quick sanity check: the expected counts must add up to the same total as the observed counts. If they don't, you made an arithmetic slip or used proportions that don't sum to .
Practice
A textbook claims that blood types in a region are type O, type A, type B and type AB. In a random sample of blood donors, what is the expected count of type B donors if the claim is true?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
To test whether a six-sided die is fair, you roll it times and record how many times each face appears. How many degrees of freedom does the chi-square goodness-of-fit test have?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In one category of a goodness-of-fit test, the observed count is and the expected count is . What is this category's contribution to the chi-square statistic?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A seed company says its wildflower mix is poppies, daisies and lupines. A random sample of seedlings contains poppies, daisies and lupines. Calculate the chi-square statistic.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Continue the wildflower problem: . Use technology to find the P-value of the goodness-of-fit test. Round to three decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A researcher tests whether births at a large hospital are equally likely in all four seasons. In a random sample of births, with df , giving . Which conclusion is correct at ?
A club wants to test whether its members' favorite of five activities matches a national distribution of , , , and . It takes a random sample of of its members. Which condition for a chi-square goodness-of-fit test is not met?
A goodness-of-fit test with categories produces . Using the table of critical values, which statement about the P-value is correct?