Glossary

Words you'll meet

Every term from every chapter, in one place: 134 entries in plain words, with scikit-learn's name where it has one and a link to the chapter that explains it.

A

Accuracy
How many samples did it get right? Watch out: rewards always saying the common answer. Chapter 5
Activation
What a unit does to its sum. ReLU keeps positive values and cuts negatives to zero; tanh squashes everything between −1 and 1; the logistic function between 0 and 1. activation Chapter 8
Adam and SGD
Ways of taking the steps. Plain stochastic gradient descent steps by the learning rate times the gradient; Adam adapts the step for each weight from its recent gradients, and usually needs less tuning. solver Chapter 8
Anomaly, or outlier
A sample unlike most of the others. Unlike doesn't mean wrong: a coal is an outlier among shales. Chapter 9
Anomaly score
Two to the power of minus the average path length over its expected value: near 1 is very strange, about 0.5 or below is ordinary. scikit-learn reports its negative. score_samples Chapter 9
Assign and update
The two moves k-means repeats: every sample joins its nearest centroid, then every centroid moves to the mean of its samples. Also called Lloyd's algorithm. Chapter 1
Average precision
How precise is it as it finds more of the class? Watch out: its floor is the class's share, not 50%, so compare it with that. Chapter 5

B

Backpropagation
The bookkeeping that works out the gradient for every weight at once, by applying the chain rule from the output back through the layers. Chapter 8
Bad hole
Where the borehole is enlarged or rough enough that pad tools (density, Pe, often neutron) read wrong. Usually flagged from the caliper and the density correction. Chapter 9
Bagging
Bootstrap aggregating: train each model on its own bootstrap sample, then let them vote (or average, for numbers). Chapter 7
Balanced accuracy, macro F1
How well does it do on every class, rare ones included? Watch out: a tiny class with a few errors can swing them. Chapter 5
Batch and epoch
A batch is the handful of samples used for one step; an epoch is one pass through all of them. batch_size, max_iter Chapter 8
Bias
Does it read high or low on average? Watch out: zero bias can hide big errors that cancel out. Chapter 5
Bias and variance
Bias is being wrong in the same way every time (too simple a model); variance is giving different answers for small changes to the data (too flexible a model). Deep trees have low bias and high variance, and averaging cuts the variance. Chapter 7
Block cross-validation
Folds made of depth blocks within a well, when there are too few wells to hold any out. Blocks must be thick enough not to leak, and thin enough not to starve the training data. Chapter 4
Bootstrap sample
As many samples as the training data, drawn at random with replacement: some twice or more, about a third not at all. bootstrap=True Chapter 7

C

Centroid
The middle of a cluster: the average of its samples on every log. cluster_centers_ Chapter 1
Classification
A supervised question whose answer is a category: which facies, pay or not. Chapter 0
Class probabilities
The forest's vote shares, one per facies. Useful for ranking and for ROC curves (chapter 5), but not true probabilities without calibration. predict_proba Chapter 7
Cluster
A group of samples the algorithm puts together because they read alike. It has no name until a person gives it one. Chapter 1
Confusion matrix
True facies against predicted, counting every combination. Rows give recall, columns precision (chapter 5). confusion_matrix Chapter 4
Confusion matrix
Which classes does it mistake for which? Watch out: rows give recall, columns precision. Look at the counts as well as the shares. Chapter 5
Contamination
The share of samples you expect to be outliers, which sets the cutoff for flagging. Left at 'auto', samples scoring above 0.5 are flagged. contamination Chapter 9
Converged
Stopped because the centroids have stopped moving (by less than a tolerance), or after a maximum number of rounds. tol, max_iter Chapter 1
Cross-validation
Splitting the training data into folds, holding each out in turn and averaging the scores, so every sample is used for both training and checking. KFold, cross_val_score Chapter 4
Curse of dimensionality
With many logs, all samples end up far from each other, and "nearest" stops meaning much. More logs help only if they carry information. Chapter 3

D

DataFrame
pandas's table, the usual way to hold logs in Python once they're read from a LAS file: one row per depth, one column per curve. Chapter 0
Dataset shift
New data that differs from the training data, such as a well logged with older tools. Also called domain shift. Chapter 4
Decision map
The facies the model would predict at every point of a crossplot. Its edges are the decision boundaries. Chapter 3
Deep learning
Networks with many layers, often of special kinds (convolutional, recurrent, transformers), built in libraries such as Keras or PyTorch. The ideas here carry over. Chapter 8
Depth
How many questions deep the tree goes. Deeper means more boxes, and sooner or later boxes around single samples. max_depth Chapter 6
Depth and leaf size
Forests usually grow every tree to full depth and let the vote do the smoothing. Setting a minimum number of samples in a leaf, or a maximum depth, smooths each tree too. max_depth, min_samples_leaf Chapter 7
Distance
Straight-line (Euclidean) distance across the logs, unless you choose another. metric Chapter 3
Distance weighting
Closer neighbours get a bigger say. weights='distance' Chapter 3

E

Early stopping
Score a validation set after every epoch, keep the best weights, and stop after a set number of epochs without improvement (the patience). early_stopping, n_iter_no_change Chapter 8
Elbow plot
Spread against k. It always falls; look for where the gains flatten. Chapter 1
Electrofacies
Clusters of log responses, named afterwards. They're groups of readings, which may or may not match the geologist's facies. Chapter 1
Ensemble
A model made of many models whose answers are combined. A forest is an ensemble of trees. Chapter 7
Entropy
Another measure of how mixed a box is, from information theory. It usually picks very similar splits. criterion='entropy' Chapter 6
Extrapolation
Predicting outside the range of the training data, where the fitted line has no evidence either way. Chapter 2
Extra trees
A cousin of the forest that also picks each cutoff at random, rather than the best one. ExtraTreesClassifier Chapter 7

F

F1
Precision and recall together? Watch out: treats a miss and a false alarm as equally bad, which your job may not. Chapter 5
Feature
A column that goes into the model: here a log. Also called an input, a predictor, a variable or an attribute. scikit-learn calls the table of features X. Chapter 0
Feature importance
How much each log's splits reduced the impurity, totalled over the tree. Chapter 7 shows how it can mislead. feature_importances_ Chapter 6
Feature space
The space with one axis per feature, where each sample is a point. With two logs it's a crossplot; with five it can't be drawn, but the algorithms don't mind. Chapter 0
Flags
scikit-learn's predict returns −1 for an outlier and 1 for an inlier; decision_function is the score shifted so that 0 is the cutoff. Chapter 9

G

Gini impurity
How mixed a box is: the chance two samples drawn from it at random are different facies. 0 is pure. criterion='gini' Chapter 6
Gradient and learning rate
The gradient says which way to move each weight to lower the loss; the learning rate says how far. Too big and training jumps about or blows up; too small and it crawls, as in chapter 2. learning_rate_init Chapter 8
Gradient boosting
A different family: trees grown one after another, each fixing the last one's mistakes, rather than side by side. XGBoost and LightGBM are boosting, not forests. Chapter 7
Gradient descent
Finding the bottom of the loss by repeatedly stepping downhill: work out which way the loss falls fastest, and take a small step that way. Chapter 2
Greedy
Choosing each split as the best one on its own, without looking ahead to the splits that follow. Chapter 6
Grouped cross-validation
Folds made of whole groups, here whole wells, so no well is on both sides. GroupKFold, LeaveOneGroupOut Chapter 4

H

Hard and soft voting
A hard vote counts each tree's facies; a soft vote averages the share of each facies in each tree's leaf. scikit-learn's forests vote soft. Chapter 7
Hyperparameter
A setting chosen by you, not learned from the data: k, a tree's depth, the number of trees. Chapter 4

I

Impurity importance
How much each log's splits cleaned up the boxes, totalled over the forest (Figure 7.4). Shared among similar logs. feature_importances_ Chapter 7
In bag, out of bag
For one tree, the samples it learned from, and the ones it never saw. Chapter 7
Isolation tree
A tree of random cuts, each on a random log at a random value within the box, grown until samples are alone or a height limit is reached. Chapter 9

K

k
The number of clusters you ask for. k-means always returns exactly k. n_clusters Chapter 1
k
How many neighbours vote. Small k follows every wrinkle; big k smooths, and can outvote rare classes. n_neighbors Chapter 3
k-means++
A smarter start: each new centroid is picked from the samples, favouring ones far from the centroids already chosen. scikit-learn's default. init='k-means++' Chapter 1
kNN regression
The same idea for numbers: predict the average of the neighbours' values. KNeighborsRegressor Chapter 3

L

Label
The column the model is asked to predict. Also called the target, the response or the output. scikit-learn calls it y. Chapter 0
Layer
A set of units that all take the same inputs. The input layer is the logs, the output layer the prediction; the layers between are hidden. hidden_layer_sizes Chapter 8
Lazy learner
kNN doesn't learn a formula; it keeps every training sample and does the work at prediction time, which makes it slow on big datasets. Chapter 3
Leaf size
A minimum number of samples in a leaf, or needed to split, which stops the tree carving out single samples. min_samples_leaf, min_samples_split Chapter 6
Leakage
Information from the test samples reaching the model during training, so the score flatters it. Neighbouring depths are the commonest leak in well data. Chapter 4
Learning rate
How big each downhill step is. Too small and it crawls; too big and it overshoots and can blow up. Chapter 2
Least squares
Choosing the line that makes the sum of the squared residuals as small as possible. Also called ordinary least squares, or OLS. Chapter 2
Local outlier factor
Another detector: compares each sample's crowding with its neighbours', so it finds samples strange for their neighbourhood. LocalOutlierFactor Chapter 9
Logs tried per split
How many logs, picked at random, each split may choose from. Fewer makes the trees more different from each other. For facies the default is the square root of the number of logs; for a number like sonic, all of them. max_features Chapter 7
Log transform
Fitting log10 of permeability instead of permeability, because permeability spans orders of magnitude and its scatter grows with its size. Chapter 2
Loss
The number training tries to make small; here the mean squared error. Also called the cost or objective. Chapter 2
Loss
The error training tries to shrink. For predicting a log, the mean squared error (scikit-learn halves it). loss_curve_ Chapter 8

M

MAE
How far off is a typical prediction? Watch out: says nothing about the worst misses. Chapter 5
Majority vote
The prediction is the facies most of the neighbours have; ties go to the class listed first. Chapter 3
Masking and swamping
Masking: a big cluster of outliers looks normal to itself, so it's missed. Swamping: ordinary samples near outliers get flagged too. Small samples per tree reduce both. Chapter 9
Micro average
Pool every class's boxes, then score? Watch out: with one label per sample it's just accuracy. Chapter 5
Missing value
A gap in the table, such as a sonic that wasn't run. Most algorithms can't use a row with a gap, so the row is dropped or the gap is filled. In pandas it's NaN. Chapter 0
Model
The recipe that turns features into a prediction, with numbers learned from the data. fit learns them; predict uses them. Chapter 0
Multilayer perceptron
The classic network of fully connected layers, the kind in this chapter. MLPRegressor, MLPClassifier Chapter 8

N

Neighbours
The training samples closest to the sample being predicted, measured across all the logs at once. Chapter 3
Node, split and leaf
A node is a box of samples; a split is a cutoff on one log that divides it in two; a leaf is a box that isn't split again. The first node is the root. Chapter 6
Normalisation
Rescaling one well's log to match a reference, for example matching its 5th and 95th percentiles, to remove tool and calibration differences before modelling. Chapter 4
Novelty
A new sample unlike the training data, as opposed to an odd one inside it. Chapter 7's unfamiliarity check was novelty detection. Chapter 9
Number of trees
More trees never make a forest overfit; the vote just settles, and the forest gets slower. 100 is the default and usually plenty. n_estimators Chapter 7
Number of trees
As in a random forest, more trees make the score steadier; 100 is the default. n_estimators Chapter 9

O

One-class SVM
Another detector: draws a boundary around the bulk of the data and flags what falls outside. OneClassSVM Chapter 9
One-vs-rest AUC, macro
Does it rank every class well against the others? Watch out: can be near perfect while a class is never called. Chapter 5
Out-of-bag score
Each training sample predicted only by the trees that left it out, then scored. Free, and optimistic when samples have near-copies in the training data, as neighbouring depths do. oob_score=True Chapter 7
Overfitting
Fitting the training samples' quirks instead of the rock: training error keeps falling while validation error rises. Chapter 8
Overfitting and underfitting
Too flexible a model chases the noise in its training samples; too stiff a one misses the real shape. Both show up as poor scores on new samples. Chapter 2

P

Path length
How many cuts it took to reach a sample's leaf. Short means strange. Chapter 9
Permutation importance
Shuffle one log in a test well and see how far the score falls. Slower, but it answers the more useful question. permutation_importance Chapter 7
Polynomial features
Powers of the input (porosity squared, cubed and so on) added as extra inputs, so a straight-line method can fit a curve. PolynomialFeatures, degree Chapter 2
Precision
When it says "sand", how often is it sand? Watch out: can be perfect while missing most of the sand. Chapter 5
Pruning
Growing a big tree, then cutting back branches that add little. ccp_alpha Chapter 6

R

R²
How much better than predicting the average? Watch out: no units; depends on how varied the truth is; can hide a bias. Chapter 5
Random seed
The number that fixes the random draws, so a forest can be grown again exactly. Change it and a good forest's score barely moves. random_state Chapter 7
Random split
Dealing samples into sets at random. Fine for independent samples; for logs it puts near-copies on both sides. train_test_split Chapter 4
Random start
Where the centroids begin. Different starts can settle in different places. random_state Chapter 1
Random start
The random weights a network begins from. Different starts end in different places; averaging a few networks is cheap insurance. random_state Chapter 8
Recall
Of all the sand, how much did it find? Watch out: can be perfect by calling everything sand. Chapter 5
Regression
A supervised question whose answer is a number: what porosity, what sonic. Chapter 0
Regression tree
The same idea for numbers: split to make each box's values alike, and predict the box's average. DecisionTreeRegressor Chapter 6
Regularisation
A penalty on large weights, which keeps the network smoother. Also called L2 or weight decay. alpha Chapter 8
Residual
The miss for one sample: the true value minus the prediction. Chapter 2
RMSE
How far off, with big misses counting extra? Watch out: a handful of bad samples can dominate it. Chapter 5
ROC AUC
Does the score rank the class above the rest, whatever the cutoff? Watch out: says nothing about the cutoff you use; flatters when the class is rare. Chapter 5
Rules
A tree read as a list of if-then statements, one per leaf. export_text Chapter 6

S

Sample
One row of the table: one depth in one well. Also called an observation, an instance or a record. Chapter 0
Samples per tree
Each tree sees a random subset, 256 by default, which keeps it fast and stops dense clusters of outliers hiding each other. max_samples Chapter 9
Scaling
Putting the logs in standard units first, so a log with big numbers doesn't dominate the distances. StandardScaler Chapter 3
Scaling
Networks train badly on raw logs with wildly different units, so the inputs, and here the sonic too, are put in standard units first, as in chapter 1. StandardScaler Chapter 8
Several starts
Run k-means more than once and keep the tightest result. With k-means++ scikit-learn runs one start by default; with random starts, ten. n_init Chapter 1
Silhouette score
Another guide to k: how much closer each sample is to its own cluster than to the next one, from −1 to 1. silhouette_score Chapter 1
Slope and intercept
The two numbers of a straight line: how much the answer changes per unit of the input, and where it starts. coef_, intercept_ Chapter 2
Spread, or inertia
The total squared distance from each sample to its centroid. Lower is tighter. inertia_ Chapter 1
Standardising
Putting every log in standard units (subtract the mean, divide by the standard deviation) so no log dominates the distances because of its units. StandardScaler Chapter 1
Supervised
Learning from rows where the label is known, to predict it where it isn't. Chapter 0
Support
How many true samples of each class are there? Watch out: a score from a handful of samples is mostly luck. Chapter 5

T

Test set, or blind well
Samples kept back until the very end and used once, to report how the model will do on new data. Chapter 4
Threshold
The cutoff value of a split. scikit-learn tries the midpoints between neighbouring values. Chapter 6
Training and test error
The miss on the samples the model learned from, and on samples it didn't. Only the second tells you how it will do. Chapter 2
Training set
The samples a model learns from. Chapter 4
Training, validation and test
The rows a model learns from, the rows used to choose its settings, and the rows kept back to score it honestly. Chapter 4 is about getting these right. Chapter 0
Tuning
Trying settings and keeping the best on validation data. GridSearchCV, with a grouped splitter for wells Chapter 4

U

Unit, or neuron
One weighted sum of its inputs, plus a constant, passed through an activation. A hidden unit's output is a log the network invents. Chapter 8
Unsupervised
Learning without labels: finding groups or oddities in the features alone. Chapter 0
Unsupervised
Learning without labels: the model is never told which samples are bad hole, so it can only find structure, never meaning. Chapter 9

V

Validation set
The samples used to choose settings, such as k, and to compare models. Used often, so it gets used up. Chapter 4
Variance
How much a model changes when the training data changes a little. Deep trees have a lot of it, which chapter 7's forests fix. Chapter 6
Vote shares
The share of neighbours of each facies: a rough probability, used for ROC curves in chapter 5. predict_proba Chapter 3

W

Weighted F1
How well does it do, sample by sample? Watch out: hides rare classes, like accuracy. Chapter 5
Weights and biases
The numbers the network learns: a weight for every connection, and a constant (a bias, or intercept) for every unit. coefs_, intercepts_ Chapter 8