Chapter 9 · Unsupervised learning

Isolation forest: finding the strange

Which samples are bad hole? An isolation forest doesn't know what bad hole is. It finds the samples that are easiest to separate from the rest, which catches a washout, and every rock that's simply rare.

10 minutesGamma ray, density, neutron and PeAll six wells, no labels

Spot the bad hole

Where the borehole washes out, the pad tools lose contact with the formation and read the mud instead. Density reads low, Pe reads like mud, and any model trained on those samples learns nonsense. Chapter 4 showed what Well D's washout did to a classifier. Finding bad hole is usually the first job on a new well, so try it yourself first.Real wells have more help than this one: a density correction curve (DRHO) and caliper-based flags. Here the caliper is hidden until you've made your call.

You used petrophysics: a washout is a shale whose density and Pe have gone wrong while its gamma ray and sonic haven't. An isolation forest has no petrophysics. It never sees a label either, so it can't learn what bad hole looks like; it can only look for what's unusual. That makes it unsupervised, like k-means in chapter 1.

Few and different

The idea is simple. Pick a log at random and cut it at a random value; keep cutting the piece a sample is in until it's alone. A sample out on its own gets cut off in a few cuts. One in the middle of a crowd takes many, because most cuts fall where its neighbours are. Count the cuts, average over many random trees, and you have a score of how strange each sample is, without ever saying what strange means.

The forest never measures distance or density of points, which makes it fast and indifferent to units: like a decision tree, it only ever compares values within one log. Each tree sees only 256 samples, which is plenty to tell strange from ordinary, and it stops at a depth of 8, because the ordinary samples are the ones it doesn't need to finish isolating.

What the forest flags

Now the whole field. One forest, every second sample of all six wells, four logs. It ranks all 3,936 samples by how strange they are; you choose how many to flag.

So an isolation forest doesn't find bad hole; it finds candidates. Every flag is a question for a petrophysicist: bad hole, a rare rock, or a log that needs normalising? Used that way, it's a quick first pass over a big dataset, and it's honest about what it can't know. Given a measurement of the hole itself, it gets much closer, which is the lesson of the whole lab in small: the logs you give a model matter more than the model.

Try it on real wells

Real bad hole doesn't come labelled. Here it's marked where the density correction is large, the usual sign of poor pad contact.

Words you'll meet

Anomaly detection has its own vocabulary, and scikit-learn's IsolationForest uses it. Here it is in plain words, with scikit-learn's name where it has one.

Anomaly, or outlier
A sample unlike most of the others. Unlike doesn't mean wrong: a coal is an outlier among shales.
Novelty
A new sample unlike the training data, as opposed to an odd one inside it. Chapter 7's unfamiliarity check was novelty detection.
Unsupervised
Learning without labels: the model is never told which samples are bad hole, so it can only find structure, never meaning.
Isolation tree
A tree of random cuts, each on a random log at a random value within the box, grown until samples are alone or a height limit is reached.
Path length
How many cuts it took to reach a sample's leaf. Short means strange.
Anomaly score
Two to the power of minus the average path length over its expected value: near 1 is very strange, about 0.5 or below is ordinary. scikit-learn reports its negative. score_samples
Contamination
The share of samples you expect to be outliers, which sets the cutoff for flagging. Left at 'auto', samples scoring above 0.5 are flagged. contamination
Samples per tree
Each tree sees a random subset, 256 by default, which keeps it fast and stops dense clusters of outliers hiding each other. max_samples
Number of trees
As in a random forest, more trees make the score steadier; 100 is the default. n_estimators
Flags
scikit-learn's predict returns −1 for an outlier and 1 for an inlier; decision_function is the score shifted so that 0 is the cutoff.
Masking and swamping
Masking: a big cluster of outliers looks normal to itself, so it's missed. Swamping: ordinary samples near outliers get flagged too. Small samples per tree reduce both.
Local outlier factor
Another detector: compares each sample's crowding with its neighbours', so it finds samples strange for their neighbourhood. LocalOutlierFactor
One-class SVM
Another detector: draws a boundary around the bulk of the data and flags what falls outside. OneClassSVM
Bad hole
Where the borehole is enlarged or rough enough that pad tools (density, Pe, often neutron) read wrong. Usually flagged from the caliper and the density correction.

What to remember

  1. An isolation forest scores how easy each sample is to cut off from the rest with random cuts: strange samples are isolated in few cuts.
  2. It finds the unusual, not the wrong. Bad hole is unusual, but so is every rare rock and every badly calibrated log.
  3. Treat its flags as questions for a petrophysicist, and give it the measurements that describe the problem (hole size here) if you want it to find that problem.

Read more

Isolation Forest: Auto Anomaly Detection with Python

Well Log Data Outlier Detection With Machine Learning and Python

That's the chapters

Back to the contents. The workbench, where you can mix the algorithms and wells yourself, is coming soon.