Math Core

Lesson 1.4 · Exploring One-Variable Data

Describing distributions

Making a graph is only half the job. On the AP exam, and in any real analysis, you have to put into words what the graph shows. Statisticians describe the distribution of a quantitative variable with four features (shape, outliers, center and spread), always tied to the context of the data. This lesson teaches the vocabulary and the habits of a complete description.

SOCS: the four features

Describe a distribution with SOCS, in context

  • Shape: the overall pattern. How many peaks? Symmetric or skewed?
  • Outliers: individual values that fall outside the overall pattern.
  • Center: a typical value, such as the median or mean.
  • Spread: how much the values vary, such as the range.

Every sentence should mention the context: the variable, its units, and who or what was measured.

For now you'll judge outliers by eye and estimate the center and spread from the graph. The next two lessons make all four features precise with formulas and rules.

Shape

Describe shape with two kinds of words.

Number of peaks. A distribution with one clear peak is unimodal. One with two clear peaks is bimodal, and one with more is multimodal. A distribution with no noticeable peaks, where every interval holds about the same number of values, is uniform.

Symmetry. Imagine folding the graph at its center.

Definition

Symmetric and skewed

  • A distribution is roughly symmetric if the left and right halves are approximately mirror images.
  • A distribution is skewed right if its right side (larger values) stretches out much farther than its left side. The long, thin part is called the tail.
  • A distribution is skewed left if its left side (smaller values) has the long tail.

Real data are never perfectly symmetric, so say "roughly symmetric" or "approximately symmetric." Common examples: incomes, home prices and waiting times are usually skewed right (they can't go below zero, but a few values are very large). Scores on an easy test are usually skewed left (most students score near the top, and a few score much lower).

Also watch for gaps (stretches with no data) and clusters (groups of values separated by gaps). These often signal that the data mix different kinds of individuals.

Outliers, center and spread by eye

An outlier is a value that falls well outside the overall pattern, usually separated from the rest of the data by a gap. Always point out apparent outliers, and give their value in context. Don't throw them away: an outlier might be a recording error, or it might be the most interesting individual in the data.

For center, a good estimate from a graph is the median, the value with about half the data on each side. For spread, give the smallest and largest values (the range), or describe where most of the data lie.

Worked example: A complete description

The dot plot shows the number of hours of homework that 2525 AP students reported doing last week. Describe the distribution.

0246810121416✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕
Hours of homework last week for 25 AP students

Shape. Apart from one student, the distribution of homework hours is unimodal and roughly symmetric, with a single peak at 55 hours.

Outliers. One student reported 1515 hours, separated from the rest by a gap of 66 hours. That value is a possible outlier.

Center. The median is the 13th value in order. Counting from the left: 11 student at 22, 22 at 33, 44 at 44 (that's 77), then 66 at 55 (that's 1313). The median is 5 hours.

Spread. The number of hours ranges from 22 to 1515, a range of 1313 hours. Without the possible outlier, all students did between 22 and 99 hours.

Put together: The distribution of last week's homework time for these 25 AP students is roughly symmetric and unimodal, apart from one possible high outlier at 15 hours. The median is 5 hours, and the times range from 2 to 15 hours, with every other student between 2 and 9 hours.

Shape tells a story

When a distribution has an unusual shape, ask why. Two peaks usually mean two different groups were measured together.

Worked example: A bimodal distribution

A bakery recorded the amount spent by each of 6060 customers during one morning.

Amount spent by 60 bakery customers (dollars)

Describe the shape, and suggest a reason for it.

The distribution is bimodal, with one peak in the $4 to $8 range and a second peak in the $16 to $20 range. A likely explanation is that the bakery has two kinds of customers: people who stop in for a coffee and a pastry, and people who buy a full breakfast or a box of baked goods for a group. Summarizing these data with a single "typical" value would hide that story. The median falls in the $12 to $16 bin, between the two peaks, where relatively few customers actually are.

Worked example: Weak versus strong descriptions

A student describes a histogram of the prices of 150150 used laptops listed online: "It's skewed. The middle is around 300 and it goes from 80 to 1,400."

What is missing?

  • Direction of skew. "Skewed" alone is not a shape; it must be skewed right or left. Prices have a floor but a few can be very high, so this is likely skewed right.
  • Outliers. The student didn't say whether any prices stand apart, such as the $1,400 laptop.
  • Context and units. "300" what? A stronger version: The distribution of prices for these 150 used laptops is skewed right, with a median of about $300. Prices range from $80 to $1,400; the $1,400 laptop appears to be a high outlier.

Common mistake

The direction of skew is the direction of the tail, not the side where the data pile up. If most values are on the left and a few stretch far to the right, the distribution is skewed right. Students often get this backward, so look for the long, thin side.

Tip

On free-response questions, graders look for all four features and context. Before you finish, check that your description names the variable and units, gives a direction for any skew, and mentions outliers (even if only to say there are none).

Practice

Practice 1

The dot plot shows the number of times each of 2020 people visited a gym last month. Which best describes its shape?

024681012✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕✕
Gym visits last month
Practice 2

Use the dot plot of gym visits from the previous problem. What is the median number of visits?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A histogram of the heights of all players on a school's basketball team and its gymnastics team combined shows two distinct peaks. What is the most reasonable explanation?

Practice 4

The last digits of 500500 randomly chosen phone numbers are recorded. Which shape would you expect for the distribution of these digits?

Practice 5

Scores on a very easy 100100-point quiz had a median of 9292. A few students who missed the review session scored in the 5050s. Which description is the most complete?

Practice 6

The stemplot shows the ages of 1919 people at a community meeting. One age falls clearly outside the overall pattern. What is that age?

StemLeaves
41
5
62 5 8
70 1 3 4 6 8 9
80 2 3 5 7
91 4 6

Key: 6∣26 \mid 2 means 6262 years old.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

Which variable is most likely to have a distribution that is skewed right?