Math Core

Lesson 1.1 · Exploring One-Variable Data

Variables and frequency tables

Statistics is the science of learning from data, and every data set starts the same way: a list of individuals and the variables measured on them. Before you can draw a graph or compute a single number, you need to know what kind of variable you are looking at, because that choice decides which displays and summaries make sense. This lesson sets up the vocabulary you will use for the rest of AP Statistics.

Individuals and variables

The individuals are the objects described by a data set. They can be people, but they can also be animals, cars, schools, days or anything else. A variable is any characteristic that can take different values for different individuals.

A data set is usually organized as a table with one row per individual and one column per variable. Here is a small data set about five used cars on a dealer's lot.

CarMakeModel yearColorMileagePrice (dollars)Doors
1Honda2019Silver48,20017,9004
2Ford2016Blue91,55011,4002
3Toyota2021White22,08023,6504
4Honda2018Black63,90015,2004
5Kia2020White35,31019,8004

The individuals are the five cars. The variables are make, model year, color, mileage, price and number of doors.

Categorical and quantitative variables

Definition

Categorical and quantitative variables

  • A categorical variable places each individual into one of several groups, or categories.
  • A quantitative variable takes numerical values for which arithmetic (adding, averaging, finding differences) makes sense.

In the car data, make and color are categorical. Mileage and price are quantitative: it makes sense to say one car costs $3,500 more than another, or to find the average mileage.

The test is not "is it a number?" but "does arithmetic on it mean anything?" Zip codes, jersey numbers, phone area codes and student ID numbers are all written with digits, yet they are categorical. The average of two zip codes is not a place.

Quantitative variables come in two kinds:

  • Discrete variables take a countable set of values, often whole-number counts: number of doors, number of siblings, goals in a game.
  • Continuous variables can take any value in an interval, limited only by how precisely you measure: mileage, price, height, time.

Worked example: Classifying variables

A researcher records the following for each of 200200 dogs at a shelter: breed, weight in pounds, age in months, whether the dog has been microchipped (yes or no), and the number of days the dog has been at the shelter. Identify the individuals and classify each variable.

Individuals: the 200200 dogs.

  • Breed: categorical.
  • Weight: quantitative, continuous.
  • Age in months: quantitative. (Recorded in whole months, it behaves like a discrete count, but age itself is continuous. Either description is reasonable; what matters is that it is quantitative.)
  • Microchipped: categorical, with just two categories.
  • Days at the shelter: quantitative, discrete.

Frequency tables

The distribution of a variable tells you what values the variable takes and how often it takes each value. For a categorical variable, the simplest way to show a distribution is a table.

  • A frequency table lists each category with its frequency, the count of individuals in that category.
  • A relative frequency table lists each category with its relative frequency, the proportion (or percent) of individuals in that category.

Relative frequency

relative frequency=frequency of the categorytotal number of individuals\text{relative frequency} = \frac{\text{frequency of the category}}{\text{total number of individuals}}

The relative frequencies of all the categories add to 11 (or 100%100\%), apart from small rounding errors.

Relative frequencies are what let you compare groups of different sizes. Knowing that 3030 students at one school and 4545 at another walk to school tells you little until you know how many students each school has.

Worked example: Building a relative frequency table

Forty students were asked how they usually get to school. The frequency table is below. Complete the relative frequency column.

ModeFrequencyRelative frequency
Bus1414
Car1111
Walk66
Bike55
Other44
Total4040

Divide each frequency by 4040:

ModeFrequencyRelative frequency
Bus14140.3500.350
Car11110.2750.275
Walk660.1500.150
Bike550.1250.125
Other440.1000.100
Total40401.0001.000

The relative frequencies add to 11, as they must. You could also report them as percents: 35%35\% of these students ride the bus.

Worked example: Working backward from percents

In a survey of 250250 adults, 32%32\% said their main streaming service is Service A, 28%28\% said Service B, 18%18\% said Service C, and the rest said they don't use a streaming service. How many adults said they don't use one?

The percents must add to 100%100\%, so the missing category is 100%−32%−28%−18%=22%100\% - 32\% - 28\% - 18\% = 22\%. That's 0.22⋅250=550.22 \cdot 250 = 55 adults.

Frequency tables for quantitative data

A quantitative variable with only a few distinct values, such as number of siblings, can go straight into a frequency table. When a quantitative variable has many different values, group the values into classes of equal width first.

Worked example: Grouping into classes

A coffee shop recorded the number of minutes each of 2525 customers waited for their order. The grouped frequency table is below. What proportion of customers waited at least 66 minutes?

Wait (minutes)Frequency
00 to less than 2244
22 to less than 4499
44 to less than 6677
66 to less than 8833
88 to less than 101022

"At least 66" means the last two classes: 3+2=53 + 2 = 5 customers. The proportion is 525=0.20\dfrac{5}{25} = 0.20.

Notice what the table can't tell you: the exact wait of any single customer. Grouping trades detail for a clearer overall picture.

Common mistake

Don't decide the type of a variable by whether it is recorded with digits. Ask whether averaging or subtracting the values would make sense in context. Area codes, zip codes and ID numbers are categorical; a survey answer coded 1=1 = "agree," 2=2 = "neutral," 3=3 = "disagree" is also categorical, even though the codes are numbers.

Tip

Write every relative frequency as a decimal to the same number of places, then check that they add to 11. If they add to 0.990.99 or 1.011.01, that's rounding. If they add to 0.850.85, you've missed a category.

Practice

Practice 1

A school records these variables for each student. Which one is categorical?

Practice 2

Which variable is quantitative and continuous?

Practice 3

In a sample of 7575 cars in a parking lot, 1818 were white. What is the relative frequency of white cars? Give a decimal.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A relative frequency table shows the blood types of 400400 donors at a blood drive. The relative frequencies are type O: 0.440.44, type A: 0.420.42, type B: 0.020.02. The only other type is AB. How many donors had type AB?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

A study of commercial flights out of a regional airport recorded, for each flight in March, the airline, the number of passengers, and the number of minutes the departure was delayed. What are the individuals in this study?

Practice 6

The grouped frequency table shows the lengths, in inches, of 5050 fish caught in a lake. What proportion of the fish were at least 1010 inches but less than 1616 inches long?

Length (inches)Frequency
66 to less than 8844
88 to less than 101099
1010 to less than 12121313
1212 to less than 14141111
1414 to less than 161688
1616 to less than 181855

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

In a survey, 3636 people chose "Option C," and they made up 15%15\% of everyone surveyed. The relative frequency of "Option A" was 0.400.40. How many people chose Option A?

Enter a number. Fractions like 3/4 and sqrt(2) are OK.