Lesson 10.1 · Data and Statistics
Measures of center and spread
Two classes can have the same average test score and still be completely different: in one, everyone scored close to the average; in the other, scores were all over the place. To describe a data set honestly you need two numbers: one for its center (a typical value) and one for its spread (how much the values vary). This lesson reviews the tools you met in middle school and adds the one statisticians use most, the standard deviation.
Measures of center
The two main measures of center are the mean and the median.
- The mean is the sum of the values divided by how many there are. For values it is written ("x-bar"):
- The median is the middle value when the data are in order. With an even number of values, it is the mean of the two middle values.
The mean uses the actual size of every value, so one extreme value can drag it a long way. The median only cares about position, so it barely moves. Statisticians say the median is resistant to extreme values and the mean is not.
Worked example: One long commute
Eight students recorded how many minutes it takes them to get to school:
Find the mean and the median. Which better describes a typical commute?
Mean. The sum is , so minutes.
Median. The two middle values are the 4th and 5th, and , so the median is minutes.
Seven of the eight students take minutes or less, yet the mean is , higher than six of them. The single -minute commute pulled it up. The median, 19 minutes, is the better description of a typical commute.
Measures of spread
The range is maximum minus minimum. It is quick, but it depends only on the two most extreme values.
The interquartile range is , the width of the middle half of the data. To find the quartiles, put the data in order and split it at the median. is the median of the lower half and is the median of the upper half. When the number of values is odd, leave the median itself out of both halves. Like the median, the IQR is resistant to extreme values.
Standard deviation
The IQR ignores the actual sizes of most values. The standard deviation uses every value. It measures how far the values typically are from the mean.
Definition
Standard deviation
The standard deviation of a data set is
Each is a deviation: how far a value is above () or below () the mean. The symbol is the Greek letter sigma. The quantity under the square root, the mean of the squared deviations, is called the variance.
Why square the deviations? Because the deviations always add up to . The positive ones cancel the negative ones exactly, so their plain average tells you nothing. Squaring makes every deviation positive, and the square root at the end brings the answer back to the original units.
Computing a standard deviation
- Find the mean .
- Subtract the mean from each value to get the deviations. (Check: they add to .)
- Square each deviation.
- Find the mean of the squares: add them and divide by .
- Take the square root.
Worked example: Standard deviation by hand
Find the standard deviation of .
The mean is . A table keeps the work organized:
| total |
The deviations add to , as they should. The variance is , so
The values are typically about units away from the mean of .
A standard deviation of means every value is the same. The more the values scatter away from the mean, the larger the standard deviation gets. For example, and both have mean , but the first has and the second has .
Tip
Graphing calculators and spreadsheets report two standard deviations. divides by , as above; it describes the data you have. divides by instead; statisticians use it when the data are a sample from a larger population. For the data in the example, . In this course, "standard deviation" means unless a problem says otherwise.
Outliers and the 1.5 × IQR rule
An outlier is a value that is unusually far from the rest of the data. "Unusually far" needs a precise meaning, so statisticians use a standard test.
The 1.5 × IQR rule
Compute the fences
A value is an outlier if it is below the lower fence or above the upper fence.
Worked example: Finding outliers
Here are the numbers of text messages ten friends sent in one hour:
Are there any outliers?
Quartiles. There are values, so the median is . The lower half is , so . The upper half is , so . The IQR is .
Fences. . The lower fence is and the upper fence is .
Check. No value is below , but . So 25 is an outlier, and it is the only one.
Outliers deserve a second look. Sometimes they are recording mistakes (someone typed instead of ), and sometimes they are the most interesting values in the data. Don't delete one just because it's inconvenient.
Choosing the right summary
Because the mean and the standard deviation use every value, one outlier affects them both. The median and the IQR resist outliers. That gives a simple rule:
| the data are… | describe center with | describe spread with |
|---|---|---|
| roughly symmetric, no outliers | mean | standard deviation |
| skewed, or have outliers | median | IQR |
Changing every value
What happens to these measures if you change every value in the same way?
- Adding a constant to every value shifts the whole data set. The mean and median go up by , but the values are just as spread out, so the range, IQR and standard deviation do not change.
- Multiplying by a positive constant stretches the data set. The mean, median, range, IQR and standard deviation are all multiplied by .
Worked example: Curving a test
A test has a mean of points and a standard deviation of points.
- The teacher adds points to every score. What are the new mean and standard deviation?
- Instead, the teacher multiplies every original score by . What are the new mean and standard deviation?
Solutions.
- Adding shifts everything: the new mean is . The spread is unchanged, so the standard deviation is still .
- Multiplying by scales everything: the new mean is and the new standard deviation is .
Common mistake
Adding the same number to every value does not change the standard deviation. Students often add it to as well. The distances between the values, and their distances from the mean, stay exactly the same.
Practice
Find the mean of .
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
Find the median of .
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A town reports the yearly incomes of its households. A few households earn many millions of dollars. Which pair of measures best describes a typical household income and the spread of incomes?
Find the standard deviation of . Round to the nearest hundredth.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
All four data sets have a mean of . Which one has the largest standard deviation?
A data set has and . What is the upper fence for outliers?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A data set has a mean of and a standard deviation of . Every value is doubled, and then is added to each result. What is the new standard deviation?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
How many outliers does this data set have, according to the rule?
Enter a number. Fractions like 3/4 and sqrt(2) are OK.