Chapter 4 · Testing honestly

Train, validate and test

Why can 98% accuracy be useless? A score is only worth the test behind it. Here's how to split your wells so the number you report is the number you'll get.

12 minutesAll six wells3,672 samples

Six wells, one score

Chapter 3's best model, seven neighbours on five logs, got 98.5% of Well C right. Now all six wells are in. Wells A, B and C were drilled close together on the crest of the field; D, E and F are step-outs, further away. The quickest way to test a model is to shuffle every sample and hold back a random 30%.In the teaching field the step-out wells differ the way real wells do: Well D has a washed-out shale and more compacted shales, Well E was logged decades earlier with older tools, and Well F's lower sand is feldspathic.

Nothing about the model changed between those two numbers. Only the test did.

Why the random split leaks

Look at where Well D goes wrong.

Samples from the same well share its hole, its mud, its tools and its rock. Shuffle them and the test set is full of near-copies of the training set: the model is marked on questions it has already seen the answers to. Hold out whole wells instead and the test asks the question you actually care about: how will the model do on a well it has never seen?

Train, validate, test

A model has settings, and choosing them is learning too. Pick k by looking at a well's score, and the model now knows something about that well. So the wells get three jobs: train wells the model learns from, validation wells for choosing settings, and test wells, sealed until every choice is made. Try different roles below.With only six wells, any one validation well is a small sample. Cross-validation by well lets each non-test well take a turn and shows how much the score varies from well to well.

The rule that makes the test worth anything: open it once. Look at the test score, change k, and look again, and you're using the test well to choose k. It has quietly become a validation well, and its score will flatter you just as the random split did.

Testing in blocks

Holding out whole wells needs wells to spare. With one or two, you can't afford to. The fallback is to hold out blocks of depth: cut each well into slabs a few metres thick, hide some slabs, train on the rest, and rotate. The question is how thick the slabs should be.Permeability here is predicted the same way facies were in chapter 3, by asking the seven most similar samples and taking the average of their permeabilities.

Too thin and the blocks leak like the random split: every test sample has bed-mates in training. Too thick and they go wrong the other way: a block holding a whole formation takes every example of its rock with it, and the model is tested on something it was never shown. Aim for blocks thicker than the beds and the tools' vertical resolution, thinner than the formations, with a small gap at each edge. And remember what blocks can't do: they test the model on the well you have. Only another well can tell you about a washout or an old tool your well doesn't contain.

Accuracy hides the rare facies

One number can't describe a model that predicts six facies. Chapter 3 showed that a big k outvotes the rare ones. Here's how that looks in the scores.

Report recall for each facies, especially the rare ones that matter: the coal that changes your net-to-gross, the anhydrite that seals the trap. And report each test well on its own: an average across wells can hide one well that went badly wrong.

The well logged with old tools

Wells logged years apart, with different tools and different calibrations, rarely agree exactly. A model trained on one set of tools can be thrown by another.

Normalisation is the standard fix for a calibration difference, and it works on Well E. On Well D it backfires, because the trouble there isn't calibration, it's bad hole. Know which problem you have before you fix it: check the caliper before you rescale the logs.

Try it on real wells

The leak is easy to show on synthetic wells, where it was built in. Is it as big on real ones?

Words you'll meet

Validation has its own vocabulary, and scikit-learn's model_selection tools use it. Here they are in plain words, with scikit-learn's name where it has one.

Training set
The samples a model learns from.
Validation set
The samples used to choose settings, such as k, and to compare models. Used often, so it gets used up.
Test set, or blind well
Samples kept back until the very end and used once, to report how the model will do on new data.
Random split
Dealing samples into sets at random. Fine for independent samples; for logs it puts near-copies on both sides. train_test_split
Leakage
Information from the test samples reaching the model during training, so the score flatters it. Neighbouring depths are the commonest leak in well data.
Hyperparameter
A setting chosen by you, not learned from the data: k, a tree's depth, the number of trees.
Cross-validation
Splitting the training data into folds, holding each out in turn and averaging the scores, so every sample is used for both training and checking. KFold, cross_val_score
Grouped cross-validation
Folds made of whole groups, here whole wells, so no well is on both sides. GroupKFold, LeaveOneGroupOut
Block cross-validation
Folds made of depth blocks within a well, when there are too few wells to hold any out. Blocks must be thick enough not to leak, and thin enough not to starve the training data.
Tuning
Trying settings and keeping the best on validation data. GridSearchCV, with a grouped splitter for wells
Confusion matrix
True facies against predicted, counting every combination. Rows give recall, columns precision (chapter 5). confusion_matrix
Dataset shift
New data that differs from the training data, such as a well logged with older tools. Also called domain shift.
Normalisation
Rescaling one well's log to match a reference, for example matching its 5th and 95th percentiles, to remove tool and calibration differences before modelling.

What to remember

  1. Test on whole wells the model has never seen. A random split of samples tests the model on near-copies of its training data.
  2. Use three sets: train to fit, validate to choose settings, and test once, at the end. Short of wells, test in blocks thicker than the beds and thinner than the formations.
  3. Report recall for each facies and a score for each well. One number can hide a rare facies that's never found and a well that went wrong.

Try any combination

A blind well against a random fifth, on the teaching field or the real wells, in the workbench.

Read more

Data Quality Considerations for Machine Learning Models

Next: metrics, plainly

Which number tells you a model is any good? Accuracy, precision, recall, F1, RMSE and R², without the jargon.