Lesson 2.6 · Exploring Two-Variable Data
Transforming data for linearity
Many real relationships are curved: populations grow exponentially, and the period of a pendulum grows with the square root of its length. The least-squares line can't fit a curve directly. But if you transform one or both variables, for example by taking logarithms, a curved pattern can become a straight one. Then all the linear tools from this unit work again, and you undo the transformation at the end to make predictions.
When a line isn't enough
A lab technician counts the bacteria in a culture (in thousands) every hour.
| Hour, | ||||||
|---|---|---|---|---|---|---|
| Count, |
A linear fit gives with . That looks respectable, but the residuals are : positive, then negative, then positive. That U-shaped residual pattern says the linear model is wrong. The counts don't go up by the same amount each hour; they go up by roughly the same factor (about times). That is exponential growth.
Linearizing exponential growth: log of y
If is exponential, then taking the base- logarithm of both sides gives
That's a linear equation in with intercept and slope . So if the data are roughly exponential, a scatterplot of against should be roughly linear.
| Hour, | ||||||
|---|---|---|---|---|---|---|
The LSRL for the transformed data is
The residual plot for the transformed data shows no pattern:
Choosing a transformation
- Exponential pattern ( multiplies by a constant factor as goes up by a constant amount): plot against .
- Power pattern (): plot against .
- Other curves: try , , or similar re-expressions.
Whichever you try, judge success the same way as always: the transformed scatterplot should look linear and the residual plot should show no leftover pattern.
Making predictions: undo the transformation
The model predicts , not , so the last step is always to reverse the transformation.
Worked example: Predicting from an exponential model
Use to predict the bacteria count at hours.
First predict the logarithm:
Then undo the log by raising to that power:
The predicted count is about bacteria ( thousand). That sits sensibly between the observed at hour and at hour .
The slope also has a multiplicative meaning. Each additional hour adds to the predicted , which multiplies the predicted count by . So the model says the culture grows by about per hour.
Common mistake
Don't stop at . A prediction of " bacteria" is the most common error on transformation questions. Always ask: is my answer in the original units? If the model used instead of , undo it with , not .
Linearizing a power model: log of both
A physics class times the period (seconds) of a pendulum for several lengths (centimeters).
| (cm) | ||||||
|---|---|---|---|---|---|---|
| (s) |
The plot rises but bends downward, the shape of a power function with an exponent less than . Taking logs of both variables:
The log-log plot is linear, and technology gives with .
Worked example: Predicting from a power model
Predict the period of a cm pendulum.
A slope of about in the log-log model means , which matches the physics: the period is proportional to the square root of the length.
Worked example: Choosing between models
A student fits three models to the same data. Which should she use?
| Model | Residual plot | |
|---|---|---|
| on | clear curve | |
| on | random scatter | |
| on | clear curve |
She should use the on (exponential) model. It is the only one whose residual plot has no leftover pattern, and it also has the largest . If a model with a slightly higher had a curved residual plot, the residual plot would still win: a curved residual plot means the model has the wrong form.
Tip
Before transforming, ask what kind of growth makes sense. Constant amounts per step suggest linear. Constant factors per step suggest exponential (). Doubling multiplies by a constant factor suggests a power model ( and ).
Practice
Use the bacteria model to predict the count (in thousands) at hours. Round to the nearest whole number.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
The number of subscribers to a video channel roughly triples every year. Which transformation is most likely to produce a linear scatterplot?
A power model is . Predict when . Round to the nearest tenth.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A model uses the natural logarithm: . Predict when . Round to the nearest hundredth.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
For stopping-distance data, a researcher fits the model , where is speed in meters per second and is stopping distance in meters. Predict the stopping distance at meters per second.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
In the bacteria model , by what factor does the predicted count multiply for each additional hour? Round to the nearest hundredth.
Enter a number. Fractions like 3/4 and sqrt(2) are OK.
A linear regression of on has and a residual plot with a clear U shape. A regression of on has and a residual plot with random scatter. Which conclusion is best?
A scatterplot of against is strongly linear with slope . Which model describes the original data?