Chapter 1 · Unsupervised learning

k-means: finding electrofacies

Can the logs find the facies on their own? Give k-means a well and a number, and it will split the samples into that many groups without being told what any of them are.

8 minutesWell A, teaching field656 samples

A crossplot you already know

Every petrophysicist has zoned a density-neutron crossplot by eye: the sands over here, the shales over there, a few odd points that turn out to be coal. k-means does the same job with a rule instead of an eye, and the rule is simple enough to follow by hand.

Below are 656 samples from Well A, one every 30 cm through 200 m of the teaching field. Each point is one depth. The four numbered circles are your centroids: think of them as the centres of four groups. Every sample belongs to whichever centroid is nearest. Move them and see.The teaching field is six synthetic wells built from the same rock physics as the other Labs, so the true facies of every sample is known.

If you followed the tip, you just ran k-means by hand. Move each centroid to the middle of its own samples, let every sample pick its nearest centroid again, and repeat until nothing changes. That's the whole algorithm."Spread" is what scikit-learn calls inertia: the total squared distance from each sample to its centroid.

Assign, update, repeat

The algorithm starts from random positions, so it needs no skill, just patience. Each round has two moves. Assign: every sample joins its nearest centroid. Update: every centroid moves to the mean of its samples. Neither move can make the spread bigger, so the loop always settles, usually within a dozen rounds.

Watch the log as you step through. The cluster track fills in at the first assign step and changes less each round. Drag down the log to pick out a bed and see where its samples sit on the crossplot, or drag a box on the crossplot to find those samples in the well.

Try a new random start a few times. Most starts end in the same place, but not all of them: k-means only finds the best grouping near where it began. That's why scikit-learn starts it with k-means++, which spreads the first centroids out and avoids most bad starts, and why you can ask it for several starts and keep the tightest.

What "close" means

There's a catch hidden in the word nearest. Distance adds up the differences on each log, and logs come in very different units. A shale and a sand might differ by 80 API on the gamma ray and 0.15 g/cm³ on density. Square those and add them, and the gamma ray wins by a factor of about 280,000.

The density-neutron plots above were scaled behind the scenes: each log was converted into standard deviations from its mean before any distance was measured. Here's what happens with gamma ray and density if you skip that step.Standardising is one choice among several. Min-max and robust scaling (median and interquartile range) give different answers, especially with outliers like coal.

Unscaled, the clusters are bands of gamma ray, and density, the log that sets coal and anhydrite apart from everything else, is ignored. Scale your logs before any method that measures distance: k-means, k-nearest neighbours and support vector machines all care. Decision trees, which come in chapter 6, don't.

Clusters are not facies

k-means needs you to choose k, and it will return as many clusters as you ask for. The elbow plot shows the spread for each k. It always falls as k rises, so look for the point where the gains flatten off. It's a guide, not an answer.

Then comes the step no algorithm can do: deciding what each cluster is. This time k-means used five logs (gamma ray, density, neutron, Pe and sonic), more than any one crossplot can show. Name the clusters from their logs, then compare your names with the true facies. In a real well you would never see that last track. Here, because the teaching field was built from a rock model, you can.

At k = 4 even the best names fall short: there's no cluster left for the sandstone/shale beds or the thin anhydrites, so they're absorbed into their neighbours. Try k = 6. The elbow plot never asked for it, but the rock did, and agreement jumps to 97%. Most of what's left sits at bed boundaries and in thin beds, where the tools' vertical resolution blends one rock into the next. k-means groups samples by what the logs say. How many groups the geology needs is still your call.

Try it on real wells

In the teaching field, six clusters matched the facies almost exactly. Real wells are less obliging.

Words you'll meet

Clustering has its own vocabulary, and scikit-learn's KMeans uses it. Here they are in plain words, with scikit-learn's name where it has one.

Cluster
A group of samples the algorithm puts together because they read alike. It has no name until a person gives it one.
Centroid
The middle of a cluster: the average of its samples on every log. cluster_centers_
k
The number of clusters you ask for. k-means always returns exactly k. n_clusters
Assign and update
The two moves k-means repeats: every sample joins its nearest centroid, then every centroid moves to the mean of its samples. Also called Lloyd's algorithm.
Spread, or inertia
The total squared distance from each sample to its centroid. Lower is tighter. inertia_
Converged
Stopped because the centroids have stopped moving (by less than a tolerance), or after a maximum number of rounds. tol, max_iter
Random start
Where the centroids begin. Different starts can settle in different places. random_state
k-means++
A smarter start: each new centroid is picked from the samples, favouring ones far from the centroids already chosen. scikit-learn's default. init='k-means++'
Several starts
Run k-means more than once and keep the tightest result. With k-means++ scikit-learn runs one start by default; with random starts, ten. n_init
Elbow plot
Spread against k. It always falls; look for where the gains flatten.
Silhouette score
Another guide to k: how much closer each sample is to its own cluster than to the next one, from −1 to 1. silhouette_score
Standardising
Putting every log in standard units (subtract the mean, divide by the standard deviation) so no log dominates the distances because of its units. StandardScaler
Electrofacies
Clusters of log responses, named afterwards. They're groups of readings, which may or may not match the geologist's facies.

What to remember

  1. k-means repeats two steps: assign each sample to its nearest centroid, then move each centroid to the mean of its samples.
  2. Scale your logs first. Unscaled, the log with the biggest numbers decides the groups.
  3. Clusters are electrofacies, not facies, until someone who knows the rock names them.

Beat k-means

Place four centroids yourself and try to finish tighter than k-means does, in challenge 1.

Read more

K-Means Clustering: An Introduction

Clustering well log data with Python

Next: linear regression

How well does porosity predict permeability? A trend line with a procedure behind it, and the first look at overfitting.