Random forests: many trees, one vote
Do many rough trees beat one careful one? Grow a hundred trees, each on a shuffled copy of the data, and let them vote. On a new well the forest does better than a single tree, and it doesn't depend on luck.
One tree is fickle
Chapter 6 grew trees on Wells A and B and tested them on Well C, a close neighbour. Here's a weakness that test hid. Train on the three crest wells, test on the step-out wells, and tiny, arbitrary choices change a deep tree's answers.
A deep tree fits its training wells so closely that a coin toss deep in the tree, which log to use when two tie, changes how it treats a new well. The forest fixes that by averaging: it scores better on average, and it gives nearly the same answer whatever its random seed.Ties are common deep in a tree, where boxes hold a handful of samples and several cutoffs separate them equally well.
Inside the forest
Averaging only helps if the trees make different mistakes, so a forest makes its trees different on purpose, in two ways. Each tree learns from a bootstrap sample: as many samples as the data, drawn at random with replacement, so some appear twice or more and about a third are left out. And each split looks at only a random few of the logs. Step through the trees below, on gamma ray and density.
No single tree is much good. The vote is. It's the wisdom of crowds, and it works for the same reason: the trees' errors are different, so they cancel. This recipe, each model trained on its own bootstrap sample and then a vote, is called bagging, short for bootstrap aggregating. A random forest is bagged trees, plus the random choice of logs at each split, which makes the trees more different still.
Out of bag: a free test, with a catch
Every tree leaves about a third of the training samples out of its bag. Those samples are an unseen test for that tree, and between them the trees cover every sample: each one is left out by about 37 of the 100 trees. Ask just those trees to vote on it and you have a prediction for a training sample from trees that never saw it. Do that for every sample, count how often the vote is right, and you have the out-of-bag score, without setting aside a single sample.
That makes the out-of-bag score useful for quick comparisons on the training wells: how many trees, how many logs per split. It isn't the score to report. The samples it judges sit between their neighbours in the same wells, which is exactly chapter 4's random split. The honest test is still a well the forest has never seen.
Which logs matter?
A forest keeps count of how much each log helped: the total drop in impurity from every split on it. It's quick and it's popular, and it's easy to over-read.
Filling in a missing sonic
Forests predict numbers too. Each tree predicts the average sonic of the samples in its leaf, and the forest averages its trees. Well F had no sonic run in its top 40 m. Train a forest on Wells A to E to predict sonic from gamma ray, density, neutron and Pe, and fill the gap.Predicting a missing log from the others is one of the commonest jobs for machine learning in petrophysics, and chapter 8 does it again with a neural network.
A forest can only predict an average of values it has seen, so it can't produce a sonic faster than the fastest rock in its training wells. A straight line can, but here it does no better, because the bed doesn't look fast on the other logs either: it looks like nothing the models know. No model can be right about a rock it has never seen. What you can do is notice when you're asking it to try, by checking how far each new sample is from the training data.
Try it on real wells
Does the vote still help on real wells, and does the out-of-bag score still flatter?
Words you'll meet
Forests come with a vocabulary, and scikit-learn's settings use it. Here it is in plain words, with scikit-learn's name where it has one.
- Ensemble
- A model made of many models whose answers are combined. A forest is an ensemble of trees.
- Bootstrap sample
- As many samples as the training data, drawn at random with replacement: some twice or more, about a third not at all.
bootstrap=True - Bagging
- Bootstrap aggregating: train each model on its own bootstrap sample, then let them vote (or average, for numbers).
- In bag, out of bag
- For one tree, the samples it learned from, and the ones it never saw.
- Out-of-bag score
- Each training sample predicted only by the trees that left it out, then scored. Free, and optimistic when samples have near-copies in the training data, as neighbouring depths do.
oob_score=True - Number of trees
- More trees never make a forest overfit; the vote just settles, and the forest gets slower. 100 is the default and usually plenty.
n_estimators - Logs tried per split
- How many logs, picked at random, each split may choose from. Fewer makes the trees more different from each other. For facies the default is the square root of the number of logs; for a number like sonic, all of them.
max_features - Depth and leaf size
- Forests usually grow every tree to full depth and let the vote do the smoothing. Setting a minimum number of samples in a leaf, or a maximum depth, smooths each tree too.
max_depth,min_samples_leaf - Hard and soft voting
- A hard vote counts each tree's facies; a soft vote averages the share of each facies in each tree's leaf. scikit-learn's forests vote soft.
- Class probabilities
- The forest's vote shares, one per facies. Useful for ranking and for ROC curves (chapter 5), but not true probabilities without calibration.
predict_proba - Impurity importance
- How much each log's splits cleaned up the boxes, totalled over the forest (Figure 7.4). Shared among similar logs.
feature_importances_ - Permutation importance
- Shuffle one log in a test well and see how far the score falls. Slower, but it answers the more useful question.
permutation_importance - Bias and variance
- Bias is being wrong in the same way every time (too simple a model); variance is giving different answers for small changes to the data (too flexible a model). Deep trees have low bias and high variance, and averaging cuts the variance.
- Random seed
- The number that fixes the random draws, so a forest can be grown again exactly. Change it and a good forest's score barely moves.
random_state - Extra trees
- A cousin of the forest that also picks each cutoff at random, rather than the best one.
ExtraTreesClassifier - Gradient boosting
- A different family: trees grown one after another, each fixing the last one's mistakes, rather than side by side. XGBoost and LightGBM are boosting, not forests.
What to remember
- A random forest grows many deep trees, each on a bootstrap sample with a random choice of logs at each split, and lets them vote.
- On a new well it beats a single tree, and it's steadier. Its out-of-bag score is a free check, but a random split in disguise: report a blind well.
- It only predicts what it has seen, and importance is shared among similar logs. Check where a new well is unlike the training wells before trusting what it predicts there.