Lesson 7.1 · Inference for Means
The t-distribution
In Unit 6 you built confidence intervals and tests for proportions using critical values and the standard normal curve. Means are different in one important way: you almost never know the population standard deviation , so you have to estimate it from the sample. That extra estimate adds extra uncertainty, and the -distributions are the tool that accounts for it.
The problem with standardizing a sample mean
From Unit 5, the sampling distribution of has mean and standard deviation . When that sampling distribution is approximately normal, the standardized statistic
follows the standard normal distribution. The catch is . If you don't know the population mean , you almost certainly don't know the population standard deviation either.
The natural fix is to replace with the sample standard deviation . The result is called the standard error of the sample mean.
Definition
Standard error of the mean
When is unknown, the standard error of is
where is the sample standard deviation and is the sample size. It estimates how far typically falls from .
Now the standardized statistic is
This is not a -score anymore. The numerator varies from sample to sample, and so does the denominator, because changes from sample to sample too. Sometimes underestimates , which makes larger in size than would have been. The result is a distribution that is more spread out than the standard normal, especially for small samples.
Meet the t-distributions
William Gosset worked out the distribution of this statistic in 1908 while studying small samples at a brewery. There isn't just one -distribution. There is a whole family, and each member is identified by its degrees of freedom.
Properties of the t-distributions
When you draw an SRS of size from a normal population, the statistic has a -distribution with degrees of freedom. Every -distribution:
- is symmetric, single-peaked and centered at , like the standard normal;
- has heavier tails and a lower peak than the standard normal, so more of its area lies far from ;
- gets closer and closer to the standard normal as increases.
Why ? The sample standard deviation is built from the deviations , and those deviations always add to . Once you know of them, the last one is forced. Only deviations are free to vary, so that's how much independent information contains.
The heavier tails are the whole point. Because is only an estimate, you need to go farther from the center to capture the same middle area. For example, the middle 95% of the standard normal lies between and . The middle 95% of the -distribution with degrees of freedom lies between and .
Reading a t-table
A -table lists critical values . Each row is a number of degrees of freedom, and each column is an upper-tail area (or, equivalently, a confidence level for the middle area). Here is an excerpt.
| tail 0.10 | tail 0.05 | tail 0.025 | tail 0.01 | tail 0.005 | |
|---|---|---|---|---|---|
| confidence | 80% | 90% | 95% | 98% | 99% |
| 4 | 1.533 | 2.132 | 2.776 | 3.747 | 4.604 |
| 7 | 1.415 | 1.895 | 2.365 | 2.998 | 3.499 |
| 9 | 1.383 | 1.833 | 2.262 | 2.821 | 3.250 |
| 11 | 1.363 | 1.796 | 2.201 | 2.718 | 3.106 |
| 14 | 1.345 | 1.761 | 2.145 | 2.624 | 2.977 |
| 15 | 1.341 | 1.753 | 2.131 | 2.602 | 2.947 |
| 19 | 1.328 | 1.729 | 2.093 | 2.539 | 2.861 |
| 20 | 1.325 | 1.725 | 2.086 | 2.528 | 2.845 |
| 24 | 1.318 | 1.711 | 2.064 | 2.492 | 2.797 |
| 29 | 1.311 | 1.699 | 2.045 | 2.462 | 2.756 |
| 30 | 1.310 | 1.697 | 2.042 | 2.457 | 2.750 |
| 40 | 1.303 | 1.684 | 2.021 | 2.423 | 2.704 |
| 60 | 1.296 | 1.671 | 2.000 | 2.390 | 2.660 |
| () | 1.282 | 1.645 | 1.960 | 2.326 | 2.576 |
Read down any column and the critical values shrink toward the value in the last row. That's the "approaches the normal" property in numbers.
For a confidence interval with confidence level , the area sits in the middle and the leftover area is split between the two tails. So a 95% interval uses the column with upper-tail area .
Worked example: Finding a critical value
You plan a 95% confidence interval for a mean from a random sample of observations. What critical value should you use?
Solution. The degrees of freedom are . A 95% confidence level leaves in the two tails together, so in the upper tail. In the row for and the column, .
On a TI-83/84, invT(0.975, 19) gives the same value: . (You enter the area to the left of , which is .)
Common mistake
Don't use . The degrees of freedom for one-sample procedures are . Also watch the column: for a 95% interval you need upper-tail area , not . Using the column gives a 90% interval.
Tail areas and p-values
In a significance test you go the other direction: you have a statistic and need the area beyond it. A table can only bracket that area. Technology gives the exact value.
Worked example: Finding a tail area
Find for a -distribution with degrees of freedom.
Solution with the table. In the row, falls between (tail ) and (tail ). So the upper-tail area is between and .
Solution with technology. tcdf(2.1, 1E99, 9) gives about . That fits inside the bracket from the table.
Worked example: A lower tail
Find when .
Solution. By symmetry, . In the row, lies between (tail ) and (tail ), so the area is between and . Technology, tcdf(-1E99, -1.5, 15), gives about .
Compare with the standard normal: . The area is larger because of the heavier tails.
Tip
A quick sanity check: for the same confidence level, is always a bit bigger than . If your comes out smaller than for a 95% interval, you've used the wrong column or the wrong row.
Why this matters for the rest of the unit
Every procedure in this unit (one-sample intervals, one-sample tests, paired data and two-sample comparisons) uses the same recipe:
What changes from lesson to lesson is the statistic, the standard error and the degrees of freedom. Once you're comfortable with and tail areas, the rest is careful bookkeeping.
Practice
A researcher takes a random sample of bags of trail mix and plans to use a one-sample procedure. How many degrees of freedom does the procedure use?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A random sample of has sample standard deviation . What is the standard error of the sample mean?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Which statement correctly compares a -distribution with degrees of freedom to the standard normal distribution?
Find the critical value for a 95% confidence interval based on a random sample of observations.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Find for a 90% confidence interval when .
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Find for a -distribution with . Give your answer to 4 decimal places.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
As the degrees of freedom increase, what happens to the critical value for a 95% confidence interval?
A study uses a random sample of students. Find the critical value for a 99% confidence interval for the mean.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.