Math Core

Lesson 10.1 · Data and Statistics

Measures of center and spread

Two classes can have the same average test score and still be completely different: in one, everyone scored close to the average; in the other, scores were all over the place. To describe a data set honestly you need two numbers: one for its center (a typical value) and one for its spread (how much the values vary). This lesson reviews the tools you met in middle school and adds the one statisticians use most, the standard deviation.

Measures of center

The two main measures of center are the mean and the median.

  • The mean is the sum of the values divided by how many there are. For values x1,x2,…,xnx_1, x_2, \dots, x_n it is written xˉ\bar{x} ("x-bar"):
xˉ=x1+x2+⋯+xnn.\bar{x} = \frac{x_1 + x_2 + \cdots + x_n}{n}.
  • The median is the middle value when the data are in order. With an even number of values, it is the mean of the two middle values.

The mean uses the actual size of every value, so one extreme value can drag it a long way. The median only cares about position, so it barely moves. Statisticians say the median is resistant to extreme values and the mean is not.

Worked example: One long commute

Eight students recorded how many minutes it takes them to get to school:

12, 15, 15, 18, 20, 22, 25, 6112, \ 15, \ 15, \ 18, \ 20, \ 22, \ 25, \ 61

Find the mean and the median. Which better describes a typical commute?

Mean. The sum is 12+15+15+18+20+22+25+61=18812 + 15 + 15 + 18 + 20 + 22 + 25 + 61 = 188, so xˉ=1888=23.5\bar{x} = \dfrac{188}{8} = 23.5 minutes.

Median. The two middle values are the 4th and 5th, 1818 and 2020, so the median is 18+202=19\dfrac{18 + 20}{2} = 19 minutes.

Seven of the eight students take 2525 minutes or less, yet the mean is 23.523.5, higher than six of them. The single 6161-minute commute pulled it up. The median, 19 minutes, is the better description of a typical commute.

Measures of spread

The range is maximum minus minimum. It is quick, but it depends only on the two most extreme values.

The interquartile range is IQR=Q3−Q1\text{IQR} = Q_3 - Q_1, the width of the middle half of the data. To find the quartiles, put the data in order and split it at the median. Q1Q_1 is the median of the lower half and Q3Q_3 is the median of the upper half. When the number of values is odd, leave the median itself out of both halves. Like the median, the IQR is resistant to extreme values.

Standard deviation

The IQR ignores the actual sizes of most values. The standard deviation uses every value. It measures how far the values typically are from the mean.

Definition

Standard deviation

The standard deviation of a data set is

σ=(x1−xˉ)2+(x2−xˉ)2+⋯+(xn−xˉ)2n.\sigma = \sqrt{\frac{(x_1 - \bar{x})^2 + (x_2 - \bar{x})^2 + \cdots + (x_n - \bar{x})^2}{n}}.

Each x−xˉx - \bar{x} is a deviation: how far a value is above (++) or below (−-) the mean. The symbol σ\sigma is the Greek letter sigma. The quantity under the square root, the mean of the squared deviations, is called the variance.

Why square the deviations? Because the deviations always add up to 00. The positive ones cancel the negative ones exactly, so their plain average tells you nothing. Squaring makes every deviation positive, and the square root at the end brings the answer back to the original units.

Computing a standard deviation

  1. Find the mean xˉ\bar{x}.
  2. Subtract the mean from each value to get the deviations. (Check: they add to 00.)
  3. Square each deviation.
  4. Find the mean of the squares: add them and divide by nn.
  5. Take the square root.

Worked example: Standard deviation by hand

Find the standard deviation of 4, 7, 8, 10, 114, \ 7, \ 8, \ 10, \ 11.

The mean is xˉ=405=8\bar{x} = \dfrac{40}{5} = 8. A table keeps the work organized:

xxx−xˉx - \bar{x}(x−xˉ)2(x - \bar{x})^2
44−4-41616
77−1-111
880000
10102244
11113399
total003030

The deviations add to 00, as they should. The variance is 305=6\dfrac{30}{5} = 6, so

σ=6≈2.45.\sigma = \sqrt{6} \approx 2.45.

The values are typically about 2.452.45 units away from the mean of 88.

A standard deviation of 00 means every value is the same. The more the values scatter away from the mean, the larger the standard deviation gets. For example, 49,50,5149, 50, 51 and 10,50,9010, 50, 90 both have mean 5050, but the first has σ≈0.82\sigma \approx 0.82 and the second has σ≈32.7\sigma \approx 32.7.

Tip

Graphing calculators and spreadsheets report two standard deviations. σx\sigma_x divides by nn, as above; it describes the data you have. sxs_x divides by n−1n - 1 instead; statisticians use it when the data are a sample from a larger population. For the data in the example, sx=30/4≈2.74s_x = \sqrt{30/4} \approx 2.74. In this course, "standard deviation" means σ\sigma unless a problem says otherwise.

Outliers and the 1.5 × IQR rule

An outlier is a value that is unusually far from the rest of the data. "Unusually far" needs a precise meaning, so statisticians use a standard test.

The 1.5 × IQR rule

Compute the fences

lower fence=Q1−1.5⋅IQR,upper fence=Q3+1.5⋅IQR.\text{lower fence} = Q_1 - 1.5 \cdot \text{IQR}, \qquad \text{upper fence} = Q_3 + 1.5 \cdot \text{IQR}.

A value is an outlier if it is below the lower fence or above the upper fence.

Worked example: Finding outliers

Here are the numbers of text messages ten friends sent in one hour:

3, 5, 6, 6, 7, 8, 9, 10, 11, 253, \ 5, \ 6, \ 6, \ 7, \ 8, \ 9, \ 10, \ 11, \ 25

Are there any outliers?

Quartiles. There are 1010 values, so the median is 7+82=7.5\dfrac{7 + 8}{2} = 7.5. The lower half is 3,5,6,6,73, 5, 6, 6, 7, so Q1=6Q_1 = 6. The upper half is 8,9,10,11,258, 9, 10, 11, 25, so Q3=10Q_3 = 10. The IQR is 10−6=410 - 6 = 4.

Fences. 1.5⋅4=61.5 \cdot 4 = 6. The lower fence is 6−6=06 - 6 = 0 and the upper fence is 10+6=1610 + 6 = 16.

Check. No value is below 00, but 25>1625 > 16. So 25 is an outlier, and it is the only one.

Outliers deserve a second look. Sometimes they are recording mistakes (someone typed 2525 instead of 2.52.5), and sometimes they are the most interesting values in the data. Don't delete one just because it's inconvenient.

Choosing the right summary

Because the mean and the standard deviation use every value, one outlier affects them both. The median and the IQR resist outliers. That gives a simple rule:

the data are…describe center withdescribe spread with
roughly symmetric, no outliersmeanstandard deviation
skewed, or have outliersmedianIQR

Changing every value

What happens to these measures if you change every value in the same way?

  • Adding a constant kk to every value shifts the whole data set. The mean and median go up by kk, but the values are just as spread out, so the range, IQR and standard deviation do not change.
  • Multiplying by a positive constant kk stretches the data set. The mean, median, range, IQR and standard deviation are all multiplied by kk.

Worked example: Curving a test

A test has a mean of 6868 points and a standard deviation of 99 points.

  1. The teacher adds 55 points to every score. What are the new mean and standard deviation?
  2. Instead, the teacher multiplies every original score by 1.251.25. What are the new mean and standard deviation?

Solutions.

  1. Adding 55 shifts everything: the new mean is 68+5=7368 + 5 = 73. The spread is unchanged, so the standard deviation is still 99.
  2. Multiplying by 1.251.25 scales everything: the new mean is 1.25⋅68=851.25 \cdot 68 = 85 and the new standard deviation is 1.25⋅9=11.251.25 \cdot 9 = 11.25.

Common mistake

Adding the same number to every value does not change the standard deviation. Students often add it to σ\sigma as well. The distances between the values, and their distances from the mean, stay exactly the same.

Practice

Practice 1

Find the mean of 14, 9, 12, 20, 1514, \ 9, \ 12, \ 20, \ 15.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

Find the median of 31, 25, 40, 28, 35, 22, 38, 3031, \ 25, \ 40, \ 28, \ 35, \ 22, \ 38, \ 30.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

A town reports the yearly incomes of its households. A few households earn many millions of dollars. Which pair of measures best describes a typical household income and the spread of incomes?

Practice 4

Find the standard deviation σ\sigma of 3, 5, 6, 9, 123, \ 5, \ 6, \ 9, \ 12. Round to the nearest hundredth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

All four data sets have a mean of 5050. Which one has the largest standard deviation?

Practice 6

A data set has Q1=18Q_1 = 18 and Q3=30Q_3 = 30. What is the upper fence for outliers?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A data set has a mean of 2020 and a standard deviation of 44. Every value is doubled, and then 55 is added to each result. What is the new standard deviation?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

How many outliers does this data set have, according to the 1.5×IQR1.5 \times \text{IQR} rule?

12, 30, 31, 33, 34, 35, 36, 38, 39, 41, 5812, \ 30, \ 31, \ 33, \ 34, \ 35, \ 36, \ 38, \ 39, \ 41, \ 58

Enter a number. Fractions like 3/4 and sqrt(2) are OK.