Math Core

Lesson 2.6 · Exploring Two-Variable Data

Transforming data for linearity

Many real relationships are curved: populations grow exponentially, and the period of a pendulum grows with the square root of its length. The least-squares line can't fit a curve directly. But if you transform one or both variables, for example by taking logarithms, a curved pattern can become a straight one. Then all the linear tools from this unit work again, and you undo the transformation at the end to make predictions.

When a line isn't enough

A lab technician counts the bacteria in a culture (in thousands) every hour.

Hour, xx001122334455
Count, yy2020313144446868101101150150
x-axis: hour. y-axis: bacteria count (thousands). The LSRL ŷ = 5.857 + 25.257x misses the curve.Open in grapher →

A linear fit gives y^=5.857+25.257x\hat{y} = 5.857 + 25.257x with r2≈0.926r^2 \approx 0.926. That r2r^2 looks respectable, but the residuals are 14.1,−0.1,−12.4,−13.6,−5.9,17.914.1, -0.1, -12.4, -13.6, -5.9, 17.9: positive, then negative, then positive. That U-shaped residual pattern says the linear model is wrong. The counts don't go up by the same amount each hour; they go up by roughly the same factor (about 1.51.5 times). That is exponential growth.

Linearizing exponential growth: log of y

If y=a⋅bxy = a \cdot b^x is exponential, then taking the base-1010 logarithm of both sides gives

log⁡y=log⁡a+xlog⁡b.\log y = \log a + x \log b.

That's a linear equation in xx with intercept log⁡a\log a and slope log⁡b\log b. So if the data are roughly exponential, a scatterplot of log⁡y\log y against xx should be roughly linear.

Hour, xx001122334455
log⁡y\log y1.3011.3011.4911.4911.6431.6431.8331.8332.0042.0042.1762.176
x-axis: hour. y-axis: log(count). The transformed points lie almost exactly on a line.Open in grapher →

The LSRL for the transformed data is

log⁡y^=1.306+0.174x,r2≈0.9995.\widehat{\log y} = 1.306 + 0.174x, \qquad r^2 \approx 0.9995.

The residual plot for the transformed data shows no pattern:

Residuals from the log model. x-axis: hour. y-axis: residual in log(count). Random scatter.Open in grapher →

Choosing a transformation

  • Exponential pattern (yy multiplies by a constant factor as xx goes up by a constant amount): plot log⁡y\log y against xx.
  • Power pattern (y=axpy = a x^p): plot log⁡y\log y against log⁡x\log x.
  • Other curves: try y\sqrt{y}, y2y^2, 1x\dfrac{1}{x} or similar re-expressions.

Whichever you try, judge success the same way as always: the transformed scatterplot should look linear and the residual plot should show no leftover pattern.

Making predictions: undo the transformation

The model predicts log⁡y\log y, not yy, so the last step is always to reverse the transformation.

Worked example: Predicting from an exponential model

Use log⁡y^=1.306+0.174x\widehat{\log y} = 1.306 + 0.174x to predict the bacteria count at x=3.5x = 3.5 hours.

First predict the logarithm:

log⁡y^=1.306+0.174(3.5)=1.306+0.609=1.915.\widehat{\log y} = 1.306 + 0.174(3.5) = 1.306 + 0.609 = 1.915.

Then undo the log by raising 1010 to that power:

y^=101.915≈82.2.\hat{y} = 10^{1.915} \approx 82.2.

The predicted count is about 82,20082{,}200 bacteria (82.282.2 thousand). That sits sensibly between the observed 6868 at hour 33 and 101101 at hour 44.

The slope also has a multiplicative meaning. Each additional hour adds 0.1740.174 to the predicted log⁡y\log y, which multiplies the predicted count by 100.174≈1.4910^{0.174} \approx 1.49. So the model says the culture grows by about 49%49\% per hour.

Common mistake

Don't stop at log⁡y^\widehat{\log y}. A prediction of "1.9151.915 bacteria" is the most common error on transformation questions. Always ask: is my answer in the original units? If the model used ln⁡y\ln y instead of log⁡y\log y, undo it with e(⋅)e^{(\cdot)}, not 10(⋅)10^{(\cdot)}.

Linearizing a power model: log of both

A physics class times the period TT (seconds) of a pendulum for several lengths LL (centimeters).

LL (cm)10102020404060608080100100
TT (s)0.630.630.900.901.271.271.551.551.791.792.012.01
x-axis: length (cm). y-axis: period (s). The pattern bends and levels off.Open in grapher →

The plot rises but bends downward, the shape of a power function with an exponent less than 11. Taking logs of both variables:

log⁡L\log L1.0001.0001.3011.3011.6021.6021.7781.7781.9031.9032.0002.000
log⁡T\log T−0.201-0.201−0.046-0.0460.1040.1040.1900.1900.2530.2530.3030.303
x-axis: log(length). y-axis: log(period). The log-log plot is linear.Open in grapher →

The log-log plot is linear, and technology gives log⁡T^=−0.701+0.502log⁡L\widehat{\log T} = -0.701 + 0.502 \log L with r2≈0.9999r^2 \approx 0.9999.

Worked example: Predicting from a power model

Predict the period of a 5050 cm pendulum.

log⁡T^=−0.701+0.502log⁡50=−0.701+0.502(1.699)=−0.701+0.853=0.152\widehat{\log T} = -0.701 + 0.502 \log 50 = -0.701 + 0.502(1.699) = -0.701 + 0.853 = 0.152T^=100.152≈1.42 seconds.\hat{T} = 10^{0.152} \approx 1.42 \text{ seconds.}

A slope of about 0.50.5 in the log-log model means T≈a⋅L0.5=aLT \approx a \cdot L^{0.5} = a\sqrt{L}, which matches the physics: the period is proportional to the square root of the length.

Worked example: Choosing between models

A student fits three models to the same data. Which should she use?

Modelr2r^2Residual plot
yy on xx0.930.93clear curve
log⁡y\log y on xx0.970.97random scatter
log⁡y\log y on log⁡x\log x0.950.95clear curve

She should use the log⁡y\log y on xx (exponential) model. It is the only one whose residual plot has no leftover pattern, and it also has the largest r2r^2. If a model with a slightly higher r2r^2 had a curved residual plot, the residual plot would still win: a curved residual plot means the model has the wrong form.

Tip

Before transforming, ask what kind of growth makes sense. Constant amounts per step suggest linear. Constant factors per step suggest exponential (log⁡y\log y). Doubling xx multiplies yy by a constant factor suggests a power model (log⁡y\log y and log⁡x\log x).

Practice

Practice 1

Use the bacteria model log⁡y^=1.306+0.174x\widehat{\log y} = 1.306 + 0.174x to predict the count (in thousands) at x=6x = 6 hours. Round to the nearest whole number.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 2

The number of subscribers to a video channel roughly triples every year. Which transformation is most likely to produce a linear scatterplot?

Practice 3

A power model is log⁡y^=0.5+2log⁡x\widehat{\log y} = 0.5 + 2\log x. Predict yy when x=10x = 10. Round to the nearest tenth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 4

A model uses the natural logarithm: ln⁡y^=2.1+0.3x\widehat{\ln y} = 2.1 + 0.3x. Predict yy when x=4x = 4. Round to the nearest hundredth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 5

For stopping-distance data, a researcher fits the model y^=1.2+0.5x\widehat{\sqrt{y}} = 1.2 + 0.5x, where xx is speed in meters per second and yy is stopping distance in meters. Predict the stopping distance at x=6x = 6 meters per second.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 6

In the bacteria model log⁡y^=1.306+0.174x\widehat{\log y} = 1.306 + 0.174x, by what factor does the predicted count multiply for each additional hour? Round to the nearest hundredth.

Enter a number. Fractions like 3/4 and sqrt(2) are OK.

Practice 7

A linear regression of yy on xx has r2=0.96r^2 = 0.96 and a residual plot with a clear U shape. A regression of log⁡y\log y on xx has r2=0.94r^2 = 0.94 and a residual plot with random scatter. Which conclusion is best?

Practice 8

A scatterplot of log⁡y\log y against log⁡x\log x is strongly linear with slope 33. Which model describes the original data?