Chapter 8 · Supervised learning

Neural networks: logs from logs

Can we predict a missing sonic log? A neural network mixes the logs you give it into new logs of its own, then mixes those into the answer. It learns the mixing by trial and error, and the hard part is knowing when to stop.

12 minutesGamma ray, density, neutron and resistivityLearns from Wells A, B and D, checks on Well C, tested on Well F

Draw the sonic

Sonic is the log most often missing, and the one most often predicted: it's needed for synthetics and rock mechanics, and it's frequently not run in the top hole. Before any model, try it yourself. Below is 30 m of Well F with the sonic hidden. You have the other logs and the facies; draw the sonic you'd expect.Well F's sonic exists below 2060 m, so any prediction there can be checked. Above it, nobody knows.

You worked from rules of thumb: shale is slow, limestone fast, coal very slow. The network had no rules, only 918 examples from three wells, and it learned the relationship well enough to be hard to beat by hand. Neither of you had seen a rock like the dolomitised bed, though, and it shows.

Inside a network

A neural network is layers of simple units. Each unit in the hidden layer takes a weighted sum of the logs, adds a constant, and cuts off anything below zero. That makes each one a new log, invented by the network: a blend of gamma ray, density, neutron and resistivity that switches on in some rocks and off in others. The output adds those invented logs up, with weights of its own, into the sonic. Training starts from random weights and nudges them, batch by batch, towards a smaller error. Step through it below; this one has six hidden units, so you can see every weight.

Nobody told the units what to look for. Some end up as shale or hard-rock detectors you could name; others are blends that don't map onto any rock, and some switch off for good. And a unit only makes sense alongside the others: one that slows the sonic in anhydrite may be correcting another that speeds it up too far. That's the trade: a network finds useful logs on its own, and they're hard to read afterwards."Neural" is a loose borrowing from biology, where a neuron fires when its inputs add up past a threshold. The resemblance stops about there.

When to stop

Train long enough and a network will fit its training samples almost exactly, the way chapter 2's degree-9 polynomial bent itself around 16 core plugs. Whether that's a problem depends on how much it has to learn from. Watch a bigger network, two layers of 32 units, train on 47 samples and then on ten times as many, scored after every epoch on the training samples and on Well C.

The training curve always falls; it's the validation curve that tells you when to stop. This is chapter 4's validation well put to work over time rather than over settings, and the same rule applies: it has to be a well the network doesn't learn from, not a random handful of samples from the training wells.

The blind well

Now the real test. Train five networks from five random starts, two layers of 16 units each, stopping each one when Well C stops improving, and predict the whole of Well F, including the top 40 m where no sonic was run.

Chapter 7's forest missed Well F by 3.9 µs/ft. Here a forest does much better, because it's given resistivity instead of Pe and is spared Well E's older logs, and the network does a little better again. The choice of logs and wells moved the score more than the choice of model. That's the usual story with well logs: the petrophysics before the machine learning.

Try it on real wells

Filling in a missing sonic is the commonest real job for a network. Here it is on real wells, against the plainest model there is.

Words you'll meet

Networks bring their own vocabulary, and scikit-learn's MLPRegressor and the deep learning libraries use it. Here it is in plain words, with scikit-learn's name where it has one.

Unit, or neuron
One weighted sum of its inputs, plus a constant, passed through an activation. A hidden unit's output is a log the network invents.
Layer
A set of units that all take the same inputs. The input layer is the logs, the output layer the prediction; the layers between are hidden. hidden_layer_sizes
Weights and biases
The numbers the network learns: a weight for every connection, and a constant (a bias, or intercept) for every unit. coefs_, intercepts_
Activation
What a unit does to its sum. ReLU keeps positive values and cuts negatives to zero; tanh squashes everything between −1 and 1; the logistic function between 0 and 1. activation
Multilayer perceptron
The classic network of fully connected layers, the kind in this chapter. MLPRegressor, MLPClassifier
Loss
The error training tries to shrink. For predicting a log, the mean squared error (scikit-learn halves it). loss_curve_
Gradient and learning rate
The gradient says which way to move each weight to lower the loss; the learning rate says how far. Too big and training jumps about or blows up; too small and it crawls, as in chapter 2. learning_rate_init
Backpropagation
The bookkeeping that works out the gradient for every weight at once, by applying the chain rule from the output back through the layers.
Batch and epoch
A batch is the handful of samples used for one step; an epoch is one pass through all of them. batch_size, max_iter
Adam and SGD
Ways of taking the steps. Plain stochastic gradient descent steps by the learning rate times the gradient; Adam adapts the step for each weight from its recent gradients, and usually needs less tuning. solver
Overfitting
Fitting the training samples' quirks instead of the rock: training error keeps falling while validation error rises.
Early stopping
Score a validation set after every epoch, keep the best weights, and stop after a set number of epochs without improvement (the patience). early_stopping, n_iter_no_change
Regularisation
A penalty on large weights, which keeps the network smoother. Also called L2 or weight decay. alpha
Scaling
Networks train badly on raw logs with wildly different units, so the inputs, and here the sonic too, are put in standard units first, as in chapter 1. StandardScaler
Random start
The random weights a network begins from. Different starts end in different places; averaging a few networks is cheap insurance. random_state
Deep learning
Networks with many layers, often of special kinds (convolutional, recurrent, transformers), built in libraries such as Keras or PyTorch. The ideas here carry over.

What to remember

  1. A neural network builds new logs from the ones you give it (each a weighted blend, cut off at zero) and adds them up into the answer. Training nudges every weight, batch by batch, towards a smaller error.
  2. The training error always falls. Watch a validation well and stop when it stops improving; more data is the better cure.
  3. Train a few networks from different random starts. Where they disagree, they're guessing, and the choice of input logs and training wells matters more than the model.

Read more

Well Log Measurement Prediction Using Neural Networks with Keras

How to Create a Simple Neural Network Model in Python

Next: isolation forest

Which samples are bad hole? Finding the strange, and deciding what it means.