Decision trees: cutoffs, learned
Could a computer choose our cutoffs? A decision tree is a stack of yes-or-no questions about the logs, each one a cutoff. It's the most petrophysical of the algorithms, and the easiest to read.
A crossplot zoned by cutoffs
Petrophysicists zone crossplots with cutoffs all the time: above so many API it's shale, below so much density it's porous sand. Stack a few of those questions and you have a decision tree. Try it: three cutoffs in two levels, on gamma ray and density.
Each box is named after the facies most of its training samples are, which is exactly how a tree labels its leaves. Whether or not you beat the algorithm, it did something you didn't: it tried every possible cutoff.The algorithm here is CART (classification and regression trees), the one behind scikit-learn's DecisionTreeClassifier, random forests and most gradient boosting.
How the algorithm chooses
At each box, the algorithm tries every cutoff on every log, halfway between each pair of neighbouring values. For each one it measures how mixed the two halves would be, and it keeps the cutoff that unmixes them most. Then it does the same again inside each half.
That's a greedy strategy: each cutoff is the best it can be on its own, with no thought for the cutoffs that follow. It's fast, and it's why you can sometimes beat it by eye. And because each cutoff compares values within one log at a time, the units never matter: unlike k-means and k-nearest neighbours, a tree doesn't need scaled logs.The finished tree can be read as rules, as at the end of Figure 6.2. That's the main reason trees are popular: you can check every rule against what you know about the rock.
How deep should it grow?
Left alone, the algorithm keeps splitting until every box is pure. Depth is the tree's version of k: the one setting you have to choose.
Too shallow and there aren't enough boxes for six facies. Too deep and the tree carves out a sliver for every odd sample, scores nearly 100% on the wells it learned from, and does no better on Well C. Choose the depth the chapter 4 way, on a validation well. Or grow lots of deep trees and let them vote: that's the next chapter.
Try it on real wells
Cutoffs that sort the teaching field's crest wells cleanly. What do learned cutoffs make of real wells?
Words you'll meet
Trees have their own vocabulary, and scikit-learn's DecisionTreeClassifier uses it. Here they are in plain words, with scikit-learn's name where it has one.
- Node, split and leaf
- A node is a box of samples; a split is a cutoff on one log that divides it in two; a leaf is a box that isn't split again. The first node is the root.
- Threshold
- The cutoff value of a split. scikit-learn tries the midpoints between neighbouring values.
- Depth
- How many questions deep the tree goes. Deeper means more boxes, and sooner or later boxes around single samples.
max_depth - Gini impurity
- How mixed a box is: the chance two samples drawn from it at random are different facies. 0 is pure.
criterion='gini' - Entropy
- Another measure of how mixed a box is, from information theory. It usually picks very similar splits.
criterion='entropy' - Greedy
- Choosing each split as the best one on its own, without looking ahead to the splits that follow.
- Leaf size
- A minimum number of samples in a leaf, or needed to split, which stops the tree carving out single samples.
min_samples_leaf,min_samples_split - Pruning
- Growing a big tree, then cutting back branches that add little.
ccp_alpha - Rules
- A tree read as a list of if-then statements, one per leaf.
export_text - Feature importance
- How much each log's splits reduced the impurity, totalled over the tree. Chapter 7 shows how it can mislead.
feature_importances_ - Regression tree
- The same idea for numbers: split to make each box's values alike, and predict the box's average.
DecisionTreeRegressor - Variance
- How much a model changes when the training data changes a little. Deep trees have a lot of it, which chapter 7's forests fix.
What to remember
- A decision tree is a stack of cutoffs. Each box is named after the commonest facies of its training samples.
- The algorithm picks each cutoff by trying them all and keeping the one that leaves the two halves least mixed. It never looks ahead.
- Deep trees memorise: choose the depth on a validation well. Trees don't need scaled logs.