Math Core

Lesson 9.1 · Statistics and Probability

The normal distribution

Heights of adults, weights of cereal boxes, lifetimes of light bulbs, errors in careful measurements: when you graph data like these, the same bell shape shows up again and again. That shape is the normal distribution, and once you know its mean and standard deviation, you can estimate what fraction of the data falls in any range.

The shape of a normal curve

A histogram shows how data are spread out. If you collect more and more data and make the bars narrower and narrower, the tops of the bars start to trace a smooth curve. For many real quantities that curve is symmetric and bell-shaped.

Definition

Normal distribution

A normal distribution is a symmetric, bell-shaped distribution described by two numbers:

  • the mean μ\mu (the center, where the peak is), and
  • the standard deviation σ\sigma (how spread out the data are).

The total area under a normal curve is 11, or 100%100\%. The area over an interval equals the proportion of the data in that interval.

Some facts follow straight from the shape:

  • The curve is symmetric about the mean, so the mean, median and mode are all equal to μ\mu.
  • Exactly half the area lies on each side of μ\mu.
  • The curve never touches the horizontal axis, but it gets very close to it once you are more than about 33 standard deviations from the mean.
  • A larger σ\sigma gives a wider, flatter curve. A smaller σ\sigma gives a taller, narrower one.

Below is the standard normal curve, the normal distribution with mean 00 and standard deviation 11. The shaded region runs from one standard deviation below the mean to one standard deviation above it.

The standard normal curve. The shaded area between −1 and 1 is about 68% of the total.Open in grapher →

The 68–95–99.7 rule

Every normal distribution, no matter its mean or standard deviation, splits up the same way.

The 68–95–99.7 rule (empirical rule)

In a normal distribution:

  • about 68%68\% of the data lie within 11 standard deviation of the mean, between μ−σ\mu - \sigma and μ+σ\mu + \sigma;
  • about 95%95\% lie within 22 standard deviations, between μ−2σ\mu - 2\sigma and μ+2σ\mu + 2\sigma;
  • about 99.7%99.7\% lie within 33 standard deviations, between μ−3σ\mu - 3\sigma and μ+3σ\mu + 3\sigma.

Because the curve is symmetric, you can cut these regions in half. That gives the percentages in each band:

Bandμ−3σ\mu - 3\sigma to μ−2σ\mu - 2\sigmaμ−2σ\mu - 2\sigma to μ−σ\mu - \sigmaμ−σ\mu - \sigma to μ\muμ\mu to μ+σ\mu + \sigmaμ+σ\mu + \sigma to μ+2σ\mu + 2\sigmaμ+2σ\mu + 2\sigma to μ+3σ\mu + 3\sigma
Percent2.35%2.35\%13.5%13.5\%34%34\%34%34\%13.5%13.5\%2.35%2.35\%

The remaining 0.3%0.3\% is split between the two tails beyond 3σ3\sigma, 0.15%0.15\% on each side. The value 13.5%13.5\% comes from 95−682\dfrac{95 - 68}{2}, and 2.35%2.35\% comes from 99.7−952\dfrac{99.7 - 95}{2}.

Worked example: Using the 68–95–99.7 rule

The heights of adult women in a large city are approximately normal with mean 6464 inches and standard deviation 2.52.5 inches. Estimate the percent of women who are

  1. between 61.561.5 and 66.566.5 inches tall;
  2. taller than 6969 inches;
  3. between 5959 and 66.566.5 inches tall.

Solution. First mark the key values: μ−2σ=59\mu - 2\sigma = 59, μ−σ=61.5\mu - \sigma = 61.5, μ=64\mu = 64, μ+σ=66.5\mu + \sigma = 66.5, μ+2σ=69\mu + 2\sigma = 69.

  1. 61.561.5 to 66.566.5 is within 11 standard deviation of the mean: about 68%68\%.
  2. 6969 is 22 standard deviations above the mean. About 95%95\% lie within 2σ2\sigma, so 5%5\% lie outside, split evenly into two tails: 5%2=2.5%\dfrac{5\%}{2} = 2.5\%.
  3. 5959 to 6464 is half of the middle 95%95\%, which is 47.5%47.5\%. 6464 to 66.566.5 is half of the middle 68%68\%, which is 34%34\%. Total: 47.5%+34%=81.5%47.5\% + 34\% = 81.5\%.
Heights with mean 64 and standard deviation 2.5. The shaded region from 59 to 66.5 holds about 81.5% of the women.Open in grapher →

z-scores

The 68–95–99.7 rule only works at whole numbers of standard deviations. To handle any value, measure how far it is from the mean in units of standard deviations.

Definition

z-score

The z-score of a data value xx is

z=x−μσ.z = \frac{x - \mu}{\sigma}.

A positive zz means xx is above the mean; a negative zz means it is below. For example, z=−1.5z = -1.5 means one and a half standard deviations below the mean.

z-scores also let you compare values from different distributions: the value whose z-score is farther from 00 is more unusual relative to its own group.

Worked example: Comparing with z-scores

Maya scored 8282 on a history test where the mean was 7474 and the standard deviation was 66. Jordan scored 8888 on a chemistry test where the mean was 8080 and the standard deviation was 88. Who did better compared with their class?

zMaya=82−746=86≈1.33,zJordan=88−808=1.z_{\text{Maya}} = \frac{82 - 74}{6} = \frac{8}{6} \approx 1.33, \qquad z_{\text{Jordan}} = \frac{88 - 80}{8} = 1.

Maya's score is 1.331.33 standard deviations above her class mean, while Jordan's is only 11 standard deviation above. Relative to her class, Maya did better, even though Jordan's raw score is higher.

Areas from a z-table

A z-table lists the area under the standard normal curve to the left of a z-score, which is the proportion of data below that value. Here is a short table.

zz000.250.250.50.50.750.75111.251.251.51.51.751.75222.52.533
Area to the left0.50000.50000.59870.59870.69150.69150.77340.77340.84130.84130.89440.89440.93320.93320.95990.95990.97720.97720.99380.99380.99870.9987

For a negative z-score, use symmetry: the area to the left of −z-z equals the area to the right of zz, which is 11 minus the table value. For example, the area to the left of −1.5-1.5 is 1−0.9332=0.06681 - 0.9332 = 0.0668.

Three kinds of area questions

  • Below a value: P(X<x)=P(X < x) = the table value for its z-score.
  • Above a value: P(X>x)=1−P(X > x) = 1 - the table value.
  • Between two values: subtract the smaller table value from the larger one.

Worked example: Battery lifetimes

The lifetime of a certain battery is normally distributed with mean 4040 hours and standard deviation 44 hours. Find the probability that a randomly chosen battery lasts

  1. less than 4545 hours;
  2. more than 3434 hours;
  3. between 3838 and 4646 hours.

Solution.

  1. z=45−404=1.25z = \dfrac{45 - 40}{4} = 1.25. The table gives P(X<45)=0.8944P(X < 45) = 0.8944.
  2. z=34−404=−1.5z = \dfrac{34 - 40}{4} = -1.5. The area to the left of −1.5-1.5 is 1−0.9332=0.06681 - 0.9332 = 0.0668, so P(X>34)=1−0.0668=0.9332P(X > 34) = 1 - 0.0668 = 0.9332. (By symmetry, the area to the right of −1.5-1.5 equals the area to the left of 1.51.5.)
  3. z=38−404=−0.5z = \dfrac{38 - 40}{4} = -0.5 and z=46−404=1.5z = \dfrac{46 - 40}{4} = 1.5. The area left of −0.5-0.5 is 1−0.6915=0.30851 - 0.6915 = 0.3085. So
P(38<X<46)=0.9332−0.3085=0.6247.P(38 < X < 46) = 0.9332 - 0.3085 = 0.6247.

Common mistake

The table gives the area to the left of zz. For "more than" questions, don't read the table value directly; subtract it from 11. A quick sketch of the curve with the region shaded will tell you whether your answer should be more or less than 0.50.5.

Tip

You can also work backward. Since 0.84130.8413 of the area is below z=1z = 1, a value one standard deviation above the mean is at about the 8484th percentile. In the battery example, 40+4=4440 + 4 = 44 hours is the 8484th percentile.

Practice

Use the 68–95–99.7 rule or the z-table in this lesson.

Practice 1

Which statement is not true of every normal distribution?

Practice 2

SAT section scores are approximately normal with mean 500500 and standard deviation 100100. About what percent of scores fall between 400400 and 600600?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 3

Using the same scores (mean 500500, standard deviation 100100), about what percent of scores are above 700700?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A population has mean 5050 and standard deviation 44. What is the z-score of the value 5757?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

The time a pizza shop takes to deliver an order is normally distributed with mean 3030 minutes and standard deviation 55 minutes. Use the z-table to find the probability that a delivery takes more than 4040 minutes. Give a decimal to four places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

Cereal boxes are filled with a mean of 1616 ounces and a standard deviation of 0.20.2 ounces, normally distributed. Use the z-table to find the probability that a box contains between 15.815.8 and 16.516.5 ounces. Give a decimal to four places.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

Scores on a statewide exam are normal with mean 7070 and standard deviation 88. Students in the top 2.5%2.5\% earn an award. Use the 68–95–99.7 rule to find the minimum score needed for the award.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 8

A runner's 5K time is 2121 minutes in a race where times have mean 2626 and standard deviation 2.52.5 minutes. A swimmer's 100-meter time is 5858 seconds in a meet where times have mean 6464 and standard deviation 44 seconds. In both sports a lower time is better. Who performed better relative to their competition?