Lesson 1.4 · Exploring One-Variable Data
Describing distributions
Making a graph is only half the job. On the AP exam, and in any real analysis, you have to put into words what the graph shows. Statisticians describe the distribution of a quantitative variable with four features (shape, outliers, center and spread), always tied to the context of the data. This lesson teaches the vocabulary and the habits of a complete description.
SOCS: the four features
Describe a distribution with SOCS, in context
- Shape: the overall pattern. How many peaks? Symmetric or skewed?
- Outliers: individual values that fall outside the overall pattern.
- Center: a typical value, such as the median or mean.
- Spread: how much the values vary, such as the range.
Every sentence should mention the context: the variable, its units, and who or what was measured.
For now you'll judge outliers by eye and estimate the center and spread from the graph. The next two lessons make all four features precise with formulas and rules.
Shape
Describe shape with two kinds of words.
Number of peaks. A distribution with one clear peak is unimodal. One with two clear peaks is bimodal, and one with more is multimodal. A distribution with no noticeable peaks, where every interval holds about the same number of values, is uniform.
Symmetry. Imagine folding the graph at its center.
Definition
Symmetric and skewed
- A distribution is roughly symmetric if the left and right halves are approximately mirror images.
- A distribution is skewed right if its right side (larger values) stretches out much farther than its left side. The long, thin part is called the tail.
- A distribution is skewed left if its left side (smaller values) has the long tail.
Real data are never perfectly symmetric, so say "roughly symmetric" or "approximately symmetric." Common examples: incomes, home prices and waiting times are usually skewed right (they can't go below zero, but a few values are very large). Scores on an easy test are usually skewed left (most students score near the top, and a few score much lower).
Also watch for gaps (stretches with no data) and clusters (groups of values separated by gaps). These often signal that the data mix different kinds of individuals.
Outliers, center and spread by eye
An outlier is a value that falls well outside the overall pattern, usually separated from the rest of the data by a gap. Always point out apparent outliers, and give their value in context. Don't throw them away: an outlier might be a recording error, or it might be the most interesting individual in the data.
For center, a good estimate from a graph is the median, the value with about half the data on each side. For spread, give the smallest and largest values (the range), or describe where most of the data lie.
Worked example: A complete description
The dot plot shows the number of hours of homework that AP students reported doing last week. Describe the distribution.
Shape. Apart from one student, the distribution of homework hours is unimodal and roughly symmetric, with a single peak at hours.
Outliers. One student reported hours, separated from the rest by a gap of hours. That value is a possible outlier.
Center. The median is the 13th value in order. Counting from the left: student at , at , at (that's ), then at (that's ). The median is 5 hours.
Spread. The number of hours ranges from to , a range of hours. Without the possible outlier, all students did between and hours.
Put together: The distribution of last week's homework time for these 25 AP students is roughly symmetric and unimodal, apart from one possible high outlier at 15 hours. The median is 5 hours, and the times range from 2 to 15 hours, with every other student between 2 and 9 hours.
Shape tells a story
When a distribution has an unusual shape, ask why. Two peaks usually mean two different groups were measured together.
Worked example: A bimodal distribution
A bakery recorded the amount spent by each of customers during one morning.
Describe the shape, and suggest a reason for it.
The distribution is bimodal, with one peak in the $4 to $8 range and a second peak in the $16 to $20 range. A likely explanation is that the bakery has two kinds of customers: people who stop in for a coffee and a pastry, and people who buy a full breakfast or a box of baked goods for a group. Summarizing these data with a single "typical" value would hide that story. The median falls in the $12 to $16 bin, between the two peaks, where relatively few customers actually are.
Worked example: Weak versus strong descriptions
A student describes a histogram of the prices of used laptops listed online: "It's skewed. The middle is around 300 and it goes from 80 to 1,400."
What is missing?
- Direction of skew. "Skewed" alone is not a shape; it must be skewed right or left. Prices have a floor but a few can be very high, so this is likely skewed right.
- Outliers. The student didn't say whether any prices stand apart, such as the $1,400 laptop.
- Context and units. "300" what? A stronger version: The distribution of prices for these 150 used laptops is skewed right, with a median of about $300. Prices range from $80 to $1,400; the $1,400 laptop appears to be a high outlier.
Common mistake
The direction of skew is the direction of the tail, not the side where the data pile up. If most values are on the left and a few stretch far to the right, the distribution is skewed right. Students often get this backward, so look for the long, thin side.
Tip
On free-response questions, graders look for all four features and context. Before you finish, check that your description names the variable and units, gives a direction for any skew, and mentions outliers (even if only to say there are none).
Practice
The dot plot shows the number of times each of people visited a gym last month. Which best describes its shape?
Use the dot plot of gym visits from the previous problem. What is the median number of visits?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A histogram of the heights of all players on a school's basketball team and its gymnastics team combined shows two distinct peaks. What is the most reasonable explanation?
The last digits of randomly chosen phone numbers are recorded. Which shape would you expect for the distribution of these digits?
Scores on a very easy -point quiz had a median of . A few students who missed the review session scored in the s. Which description is the most complete?
The stemplot shows the ages of people at a community meeting. One age falls clearly outside the overall pattern. What is that age?
| Stem | Leaves |
|---|---|
| 4 | 1 |
| 5 | |
| 6 | 2 5 8 |
| 7 | 0 1 3 4 6 8 9 |
| 8 | 0 2 3 5 7 |
| 9 | 1 4 6 |
Key: means years old.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Which variable is most likely to have a distribution that is skewed right?