Chapter 3 · Supervised learning

k-nearest neighbours: lithology by vote

Which lithology is this sample, going by the wells we've already described? Find the samples most like it and let them vote. That's k-nearest neighbours, the simplest supervised method there is.

8 minutesLearns from Wells A and B, tested on Well C1,312 labelled samples

Labels change everything

In chapter 1, k-means found groups without knowing what any of them were. Now suppose the facies in Wells A and B have been described from core and cuttings. Every sample from those wells carries a label, and the question for Well C becomes: which of the labelled samples does each new sample look like?Supervised learning means learning from examples that come with the answers. Chapter 1 was unsupervised: no answers, just groups.

Below are 1,312 labelled samples from Wells A and B on the density-neutron crossplot, coloured by facies. A blue sample from Well C is waiting. Decide what it is from where it falls, then see what its seven nearest neighbours say.

If you looked at the colours around the blue sample and went with the majority, you did exactly what the algorithm does. There's no fitting and no equation: the model is the labelled samples themselves, and every prediction is a fresh search for the nearest ones.

How many neighbours?

The one choice to make is k, how many neighbours get a vote. The background below shows what the vote would decide at every point on the plot: the decision map. Drag the blue probe to watch a vote, then change k.

With one neighbour the map is a patchwork. Every odd sample claims a patch of its own, and the model scores 100% on the wells it learned from, because every sample is its own nearest neighbour. That's memorising, not learning. With a handful of neighbours the patchwork smooths over and Well C does slightly better. With too many, the vote drowns out anything rare: by 75, coal and anhydrite have vanished from the map, because there are only 17 and 22 samples of them to learn from.Rare classes losing every vote is a symptom of imbalanced data. Chapter 4 shows how a single accuracy figure hides it.

Units, again

Nearest means a distance, and as in chapter 1, distance depends on units. Here's gamma ray against density, unscaled and scaled.

More logs, or a better k?

Two logs fit on a crossplot, but the algorithm doesn't need to see a plot. Give it five logs (gamma ray, density, neutron, Pe and sonic) and it measures distance in five dimensions in exactly the same way.

Tuning k is worth about a point. Adding three logs is worth about three: shaly sand and limestone overlap on density and neutron but separate on Pe and sonic. The misses that remain sit almost all at bed boundaries, where the tools' vertical resolution blends one bed into the next. No choice of k can fix a sample whose logs are a mixture of two rocks.

One caution before trusting numbers this good. Well C was logged with the same tools as Wells A and B, and it cuts the same beds. Chapter 4 is about what happens when the new well isn't so kind.

Try it on real wells

Seven neighbours scored 98.9% on the teaching field's Well C. Here they face real wells, one of them from another part of the North Sea.

Words you'll meet

Nearest neighbours has a short vocabulary, and scikit-learn's KNeighborsClassifier uses it. Here they are in plain words, with scikit-learn's name where it has one.

Neighbours
The training samples closest to the sample being predicted, measured across all the logs at once.
k
How many neighbours vote. Small k follows every wrinkle; big k smooths, and can outvote rare classes. n_neighbors
Distance
Straight-line (Euclidean) distance across the logs, unless you choose another. metric
Majority vote
The prediction is the facies most of the neighbours have; ties go to the class listed first.
Distance weighting
Closer neighbours get a bigger say. weights='distance'
Decision map
The facies the model would predict at every point of a crossplot. Its edges are the decision boundaries.
Scaling
Putting the logs in standard units first, so a log with big numbers doesn't dominate the distances. StandardScaler
Lazy learner
kNN doesn't learn a formula; it keeps every training sample and does the work at prediction time, which makes it slow on big datasets.
Vote shares
The share of neighbours of each facies: a rough probability, used for ROC curves in chapter 5. predict_proba
Curse of dimensionality
With many logs, all samples end up far from each other, and "nearest" stops meaning much. More logs help only if they carry information.
kNN regression
The same idea for numbers: predict the average of the neighbours' values. KNeighborsRegressor

What to remember

  1. k-nearest neighbours labels a sample by a vote of the k most similar labelled samples. There's no training step: the model is the data.
  2. k = 1 memorises and a large k outvotes rare facies. A handful of neighbours is usually best, but the choice matters less than the logs you give it.
  3. Scale the logs. Distance-based methods treat the biggest numbers as the most important.

Try any combination

Neighbours on any logs, on the teaching field or the real wells, in the workbench.

Read more

k-Nearest Neighbors for Lithology Classification from Well Logs Using Python

Next: train, validate and test

Why can 98% accuracy be useless? Testing a model on wells it has never seen.