Linear regression: a trend with a procedure
How well does porosity predict permeability? Every petrophysicist has drawn a line through core data. Here's how a computer decides which line is best, how it finds it, and how a line can fit too well.
Fit the line yourself
These are core plugs from Sand A in Wells A and B: 281 small cylinders of rock, each with its porosity and permeability measured in the laboratory. Permeability goes on a log scale because it spans orders of magnitude, so a straight line here means log₁₀ k = a + b·φ. Drag the line to fit the plugs as well as you can.In the teaching field, permeability comes from Pore Lab's Kozeny-Carman model, so it depends on grain size and sorting as well as porosity. That's where the scatter comes from.
What makes one line better than another? Each plug's miss is the vertical gap between it and the line. Square every miss and add them up: the line with the smallest total is the least-squares line. Squaring makes every miss count as positive and makes big misses count for much more than small ones. Switch to squares in the figure and watch the big ones dominate as you move the line.Misses are measured up and down only, in decades of permeability, as if porosity were exact. Regress porosity on permeability instead and you get a different line.
For a straight line there's a formula for the best answer, so the computer doesn't need to search. Most models don't have one.
Rolling downhill
Every possible line has an error, so all the lines together make a landscape: two directions (the line's level and its slope) and a height (the error). The least-squares line sits at the bottom of the valley. A model with no formula has to find the bottom by walking: work out which way is downhill from where it is, take a step, and repeat. That's gradient descent, and it's how neural networks learn in chapter 8.
The one thing you have to choose is the size of each step, the learning rate. Too small and it crawls. Too big and every step overshoots the valley floor by more than the last, until the error blows up. In between, it zig-zags across the valley and settles at the same answer the formula gives.
How bendy should the line be?
A straight line is a choice. A curve can follow the plugs more closely. Below, the model learns from just 16 plugs from Well A, then it's tested on 139 plugs from Well C, a well it has never seen. Raise the degree of the curve and watch both errors.Why test on a whole new well rather than on plugs left out of the same well? Plugs a foot apart are near-copies of each other. Chapter 4 shows what goes wrong if you mix them.
The error on the training plugs only ever goes down: a bendier curve can always get closer to the points it learned from. The error on Well C improves a little, then explodes. By degree 7 the curve is threading between 16 plugs and swinging wildly where it has no plugs to hold it. That's overfitting: the model has learned the noise in its training data, not the rock.
From core to log
The point of a poro-perm line is to predict permeability where there's no core. Apply the least-squares line from Figure 2.1 to Well C's density porosity log and you get a permeability log, which you can check against Well C's own core and, because this is the teaching field, the true permeability.
The prediction follows the big swings and gets the detail wrong, because porosity is only part of the story. In the teaching field, as in real rock, grain size and sorting change permeability at the same porosity. A model can only use what it's given: if it needs to know grain size, it needs a log that responds to it.
Try it on real wells
Chapter 2's lines ran through core plugs from three close wells. A real line has to carry from the wells it was fitted on to the next one.
Words you'll meet
Regression comes with its own vocabulary, and scikit-learn's LinearRegression uses it. Here they are in plain words, with scikit-learn's name where it has one.
- Slope and intercept
- The two numbers of a straight line: how much the answer changes per unit of the input, and where it starts.
coef_,intercept_ - Residual
- The miss for one sample: the true value minus the prediction.
- Least squares
- Choosing the line that makes the sum of the squared residuals as small as possible. Also called ordinary least squares, or OLS.
- Loss
- The number training tries to make small; here the mean squared error. Also called the cost or objective.
- Gradient descent
- Finding the bottom of the loss by repeatedly stepping downhill: work out which way the loss falls fastest, and take a small step that way.
- Learning rate
- How big each downhill step is. Too small and it crawls; too big and it overshoots and can blow up.
- Polynomial features
- Powers of the input (porosity squared, cubed and so on) added as extra inputs, so a straight-line method can fit a curve.
PolynomialFeatures,degree - Overfitting and underfitting
- Too flexible a model chases the noise in its training samples; too stiff a one misses the real shape. Both show up as poor scores on new samples.
- Training and test error
- The miss on the samples the model learned from, and on samples it didn't. Only the second tells you how it will do.
- Log transform
- Fitting log10 of permeability instead of permeability, because permeability spans orders of magnitude and its scatter grows with its size.
- Extrapolation
- Predicting outside the range of the training data, where the fitted line has no evidence either way.
What to remember
- Least squares picks the line with the smallest sum of squared misses. For a straight line there's a formula.
- Gradient descent finds the same answer by walking downhill. The learning rate decides whether it crawls, settles or blows up.
- More flexibility always fits the training data better. Only data the model hasn't seen tells you whether it helped.