Glossary
Words you'll meet
Every term from every chapter, in one place: 134 entries in plain words, with scikit-learn's name where it has one and a link to the chapter that explains it.
A
- Accuracy
- How many samples did it get right? Watch out: rewards always saying the common answer. Chapter 5
- Activation
- What a unit does to its sum. ReLU keeps positive values and cuts negatives to zero; tanh squashes everything between −1 and 1; the logistic function between 0 and 1.
activationChapter 8 - Adam and SGD
- Ways of taking the steps. Plain stochastic gradient descent steps by the learning rate times the gradient; Adam adapts the step for each weight from its recent gradients, and usually needs less tuning.
solverChapter 8 - Anomaly, or outlier
- A sample unlike most of the others. Unlike doesn't mean wrong: a coal is an outlier among shales. Chapter 9
- Anomaly score
- Two to the power of minus the average path length over its expected value: near 1 is very strange, about 0.5 or below is ordinary. scikit-learn reports its negative.
score_samplesChapter 9 - Assign and update
- The two moves k-means repeats: every sample joins its nearest centroid, then every centroid moves to the mean of its samples. Also called Lloyd's algorithm. Chapter 1
- Average precision
- How precise is it as it finds more of the class? Watch out: its floor is the class's share, not 50%, so compare it with that. Chapter 5
B
- Backpropagation
- The bookkeeping that works out the gradient for every weight at once, by applying the chain rule from the output back through the layers. Chapter 8
- Bad hole
- Where the borehole is enlarged or rough enough that pad tools (density, Pe, often neutron) read wrong. Usually flagged from the caliper and the density correction. Chapter 9
- Bagging
- Bootstrap aggregating: train each model on its own bootstrap sample, then let them vote (or average, for numbers). Chapter 7
- Balanced accuracy, macro F1
- How well does it do on every class, rare ones included? Watch out: a tiny class with a few errors can swing them. Chapter 5
- Batch and epoch
- A batch is the handful of samples used for one step; an epoch is one pass through all of them.
batch_size,max_iterChapter 8 - Bias
- Does it read high or low on average? Watch out: zero bias can hide big errors that cancel out. Chapter 5
- Bias and variance
- Bias is being wrong in the same way every time (too simple a model); variance is giving different answers for small changes to the data (too flexible a model). Deep trees have low bias and high variance, and averaging cuts the variance. Chapter 7
- Block cross-validation
- Folds made of depth blocks within a well, when there are too few wells to hold any out. Blocks must be thick enough not to leak, and thin enough not to starve the training data. Chapter 4
- Bootstrap sample
- As many samples as the training data, drawn at random with replacement: some twice or more, about a third not at all.
bootstrap=TrueChapter 7
C
- Centroid
- The middle of a cluster: the average of its samples on every log.
cluster_centers_Chapter 1 - Classification
- A supervised question whose answer is a category: which facies, pay or not. Chapter 0
- Class probabilities
- The forest's vote shares, one per facies. Useful for ranking and for ROC curves (chapter 5), but not true probabilities without calibration.
predict_probaChapter 7 - Cluster
- A group of samples the algorithm puts together because they read alike. It has no name until a person gives it one. Chapter 1
- Confusion matrix
- True facies against predicted, counting every combination. Rows give recall, columns precision (chapter 5).
confusion_matrixChapter 4 - Confusion matrix
- Which classes does it mistake for which? Watch out: rows give recall, columns precision. Look at the counts as well as the shares. Chapter 5
- Contamination
- The share of samples you expect to be outliers, which sets the cutoff for flagging. Left at
'auto', samples scoring above 0.5 are flagged.contaminationChapter 9 - Converged
- Stopped because the centroids have stopped moving (by less than a tolerance), or after a maximum number of rounds.
tol,max_iterChapter 1 - Cross-validation
- Splitting the training data into folds, holding each out in turn and averaging the scores, so every sample is used for both training and checking.
KFold,cross_val_scoreChapter 4 - Curse of dimensionality
- With many logs, all samples end up far from each other, and "nearest" stops meaning much. More logs help only if they carry information. Chapter 3
D
- DataFrame
- pandas's table, the usual way to hold logs in Python once they're read from a LAS file: one row per depth, one column per curve. Chapter 0
- Dataset shift
- New data that differs from the training data, such as a well logged with older tools. Also called domain shift. Chapter 4
- Decision map
- The facies the model would predict at every point of a crossplot. Its edges are the decision boundaries. Chapter 3
- Deep learning
- Networks with many layers, often of special kinds (convolutional, recurrent, transformers), built in libraries such as Keras or PyTorch. The ideas here carry over. Chapter 8
- Depth
- How many questions deep the tree goes. Deeper means more boxes, and sooner or later boxes around single samples.
max_depthChapter 6 - Depth and leaf size
- Forests usually grow every tree to full depth and let the vote do the smoothing. Setting a minimum number of samples in a leaf, or a maximum depth, smooths each tree too.
max_depth,min_samples_leafChapter 7 - Distance
- Straight-line (Euclidean) distance across the logs, unless you choose another.
metricChapter 3 - Distance weighting
- Closer neighbours get a bigger say.
weights='distance'Chapter 3
E
- Early stopping
- Score a validation set after every epoch, keep the best weights, and stop after a set number of epochs without improvement (the patience).
early_stopping,n_iter_no_changeChapter 8 - Elbow plot
- Spread against k. It always falls; look for where the gains flatten. Chapter 1
- Electrofacies
- Clusters of log responses, named afterwards. They're groups of readings, which may or may not match the geologist's facies. Chapter 1
- Ensemble
- A model made of many models whose answers are combined. A forest is an ensemble of trees. Chapter 7
- Entropy
- Another measure of how mixed a box is, from information theory. It usually picks very similar splits.
criterion='entropy'Chapter 6 - Extrapolation
- Predicting outside the range of the training data, where the fitted line has no evidence either way. Chapter 2
- Extra trees
- A cousin of the forest that also picks each cutoff at random, rather than the best one.
ExtraTreesClassifierChapter 7
F
- F1
- Precision and recall together? Watch out: treats a miss and a false alarm as equally bad, which your job may not. Chapter 5
- Feature
- A column that goes into the model: here a log. Also called an input, a predictor, a variable or an attribute. scikit-learn calls the table of features
X. Chapter 0 - Feature importance
- How much each log's splits reduced the impurity, totalled over the tree. Chapter 7 shows how it can mislead.
feature_importances_Chapter 6 - Feature space
- The space with one axis per feature, where each sample is a point. With two logs it's a crossplot; with five it can't be drawn, but the algorithms don't mind. Chapter 0
- Flags
- scikit-learn's
predictreturns −1 for an outlier and 1 for an inlier;decision_functionis the score shifted so that 0 is the cutoff. Chapter 9
G
- Gini impurity
- How mixed a box is: the chance two samples drawn from it at random are different facies. 0 is pure.
criterion='gini'Chapter 6 - Gradient and learning rate
- The gradient says which way to move each weight to lower the loss; the learning rate says how far. Too big and training jumps about or blows up; too small and it crawls, as in chapter 2.
learning_rate_initChapter 8 - Gradient boosting
- A different family: trees grown one after another, each fixing the last one's mistakes, rather than side by side. XGBoost and LightGBM are boosting, not forests. Chapter 7
- Gradient descent
- Finding the bottom of the loss by repeatedly stepping downhill: work out which way the loss falls fastest, and take a small step that way. Chapter 2
- Greedy
- Choosing each split as the best one on its own, without looking ahead to the splits that follow. Chapter 6
- Grouped cross-validation
- Folds made of whole groups, here whole wells, so no well is on both sides.
GroupKFold,LeaveOneGroupOutChapter 4
H
- Hard and soft voting
- A hard vote counts each tree's facies; a soft vote averages the share of each facies in each tree's leaf. scikit-learn's forests vote soft. Chapter 7
- Hyperparameter
- A setting chosen by you, not learned from the data: k, a tree's depth, the number of trees. Chapter 4
I
- Impurity importance
- How much each log's splits cleaned up the boxes, totalled over the forest (Figure 7.4). Shared among similar logs.
feature_importances_Chapter 7 - In bag, out of bag
- For one tree, the samples it learned from, and the ones it never saw. Chapter 7
- Isolation tree
- A tree of random cuts, each on a random log at a random value within the box, grown until samples are alone or a height limit is reached. Chapter 9
K
- k
- The number of clusters you ask for. k-means always returns exactly k.
n_clustersChapter 1 - k
- How many neighbours vote. Small k follows every wrinkle; big k smooths, and can outvote rare classes.
n_neighborsChapter 3 - k-means++
- A smarter start: each new centroid is picked from the samples, favouring ones far from the centroids already chosen. scikit-learn's default.
init='k-means++'Chapter 1 - kNN regression
- The same idea for numbers: predict the average of the neighbours' values.
KNeighborsRegressorChapter 3
L
- Label
- The column the model is asked to predict. Also called the target, the response or the output. scikit-learn calls it
y. Chapter 0 - Layer
- A set of units that all take the same inputs. The input layer is the logs, the output layer the prediction; the layers between are hidden.
hidden_layer_sizesChapter 8 - Lazy learner
- kNN doesn't learn a formula; it keeps every training sample and does the work at prediction time, which makes it slow on big datasets. Chapter 3
- Leaf size
- A minimum number of samples in a leaf, or needed to split, which stops the tree carving out single samples.
min_samples_leaf,min_samples_splitChapter 6 - Leakage
- Information from the test samples reaching the model during training, so the score flatters it. Neighbouring depths are the commonest leak in well data. Chapter 4
- Learning rate
- How big each downhill step is. Too small and it crawls; too big and it overshoots and can blow up. Chapter 2
- Least squares
- Choosing the line that makes the sum of the squared residuals as small as possible. Also called ordinary least squares, or OLS. Chapter 2
- Local outlier factor
- Another detector: compares each sample's crowding with its neighbours', so it finds samples strange for their neighbourhood.
LocalOutlierFactorChapter 9 - Logs tried per split
- How many logs, picked at random, each split may choose from. Fewer makes the trees more different from each other. For facies the default is the square root of the number of logs; for a number like sonic, all of them.
max_featuresChapter 7 - Log transform
- Fitting log10 of permeability instead of permeability, because permeability spans orders of magnitude and its scatter grows with its size. Chapter 2
- Loss
- The number training tries to make small; here the mean squared error. Also called the cost or objective. Chapter 2
- Loss
- The error training tries to shrink. For predicting a log, the mean squared error (scikit-learn halves it).
loss_curve_Chapter 8
M
- MAE
- How far off is a typical prediction? Watch out: says nothing about the worst misses. Chapter 5
- Majority vote
- The prediction is the facies most of the neighbours have; ties go to the class listed first. Chapter 3
- Masking and swamping
- Masking: a big cluster of outliers looks normal to itself, so it's missed. Swamping: ordinary samples near outliers get flagged too. Small samples per tree reduce both. Chapter 9
- Micro average
- Pool every class's boxes, then score? Watch out: with one label per sample it's just accuracy. Chapter 5
- Missing value
- A gap in the table, such as a sonic that wasn't run. Most algorithms can't use a row with a gap, so the row is dropped or the gap is filled. In pandas it's
NaN. Chapter 0 - Model
- The recipe that turns features into a prediction, with numbers learned from the data.
fitlearns them;predictuses them. Chapter 0 - Multilayer perceptron
- The classic network of fully connected layers, the kind in this chapter.
MLPRegressor,MLPClassifierChapter 8
N
- Neighbours
- The training samples closest to the sample being predicted, measured across all the logs at once. Chapter 3
- Node, split and leaf
- A node is a box of samples; a split is a cutoff on one log that divides it in two; a leaf is a box that isn't split again. The first node is the root. Chapter 6
- Normalisation
- Rescaling one well's log to match a reference, for example matching its 5th and 95th percentiles, to remove tool and calibration differences before modelling. Chapter 4
- Novelty
- A new sample unlike the training data, as opposed to an odd one inside it. Chapter 7's unfamiliarity check was novelty detection. Chapter 9
- Number of trees
- More trees never make a forest overfit; the vote just settles, and the forest gets slower. 100 is the default and usually plenty.
n_estimatorsChapter 7 - Number of trees
- As in a random forest, more trees make the score steadier; 100 is the default.
n_estimatorsChapter 9
O
- One-class SVM
- Another detector: draws a boundary around the bulk of the data and flags what falls outside.
OneClassSVMChapter 9 - One-vs-rest AUC, macro
- Does it rank every class well against the others? Watch out: can be near perfect while a class is never called. Chapter 5
- Out-of-bag score
- Each training sample predicted only by the trees that left it out, then scored. Free, and optimistic when samples have near-copies in the training data, as neighbouring depths do.
oob_score=TrueChapter 7 - Overfitting
- Fitting the training samples' quirks instead of the rock: training error keeps falling while validation error rises. Chapter 8
- Overfitting and underfitting
- Too flexible a model chases the noise in its training samples; too stiff a one misses the real shape. Both show up as poor scores on new samples. Chapter 2
P
- Path length
- How many cuts it took to reach a sample's leaf. Short means strange. Chapter 9
- Permutation importance
- Shuffle one log in a test well and see how far the score falls. Slower, but it answers the more useful question.
permutation_importanceChapter 7 - Polynomial features
- Powers of the input (porosity squared, cubed and so on) added as extra inputs, so a straight-line method can fit a curve.
PolynomialFeatures,degreeChapter 2 - Precision
- When it says "sand", how often is it sand? Watch out: can be perfect while missing most of the sand. Chapter 5
- Pruning
- Growing a big tree, then cutting back branches that add little.
ccp_alphaChapter 6
R
- R²
- How much better than predicting the average? Watch out: no units; depends on how varied the truth is; can hide a bias. Chapter 5
- Random seed
- The number that fixes the random draws, so a forest can be grown again exactly. Change it and a good forest's score barely moves.
random_stateChapter 7 - Random split
- Dealing samples into sets at random. Fine for independent samples; for logs it puts near-copies on both sides.
train_test_splitChapter 4 - Random start
- Where the centroids begin. Different starts can settle in different places.
random_stateChapter 1 - Random start
- The random weights a network begins from. Different starts end in different places; averaging a few networks is cheap insurance.
random_stateChapter 8 - Recall
- Of all the sand, how much did it find? Watch out: can be perfect by calling everything sand. Chapter 5
- Regression
- A supervised question whose answer is a number: what porosity, what sonic. Chapter 0
- Regression tree
- The same idea for numbers: split to make each box's values alike, and predict the box's average.
DecisionTreeRegressorChapter 6 - Regularisation
- A penalty on large weights, which keeps the network smoother. Also called L2 or weight decay.
alphaChapter 8 - Residual
- The miss for one sample: the true value minus the prediction. Chapter 2
- RMSE
- How far off, with big misses counting extra? Watch out: a handful of bad samples can dominate it. Chapter 5
- ROC AUC
- Does the score rank the class above the rest, whatever the cutoff? Watch out: says nothing about the cutoff you use; flatters when the class is rare. Chapter 5
- Rules
- A tree read as a list of if-then statements, one per leaf.
export_textChapter 6
S
- Sample
- One row of the table: one depth in one well. Also called an observation, an instance or a record. Chapter 0
- Samples per tree
- Each tree sees a random subset, 256 by default, which keeps it fast and stops dense clusters of outliers hiding each other.
max_samplesChapter 9 - Scaling
- Putting the logs in standard units first, so a log with big numbers doesn't dominate the distances.
StandardScalerChapter 3 - Scaling
- Networks train badly on raw logs with wildly different units, so the inputs, and here the sonic too, are put in standard units first, as in chapter 1.
StandardScalerChapter 8 - Several starts
- Run k-means more than once and keep the tightest result. With k-means++ scikit-learn runs one start by default; with random starts, ten.
n_initChapter 1 - Silhouette score
- Another guide to k: how much closer each sample is to its own cluster than to the next one, from −1 to 1.
silhouette_scoreChapter 1 - Slope and intercept
- The two numbers of a straight line: how much the answer changes per unit of the input, and where it starts.
coef_,intercept_Chapter 2 - Spread, or inertia
- The total squared distance from each sample to its centroid. Lower is tighter.
inertia_Chapter 1 - Standardising
- Putting every log in standard units (subtract the mean, divide by the standard deviation) so no log dominates the distances because of its units.
StandardScalerChapter 1 - Supervised
- Learning from rows where the label is known, to predict it where it isn't. Chapter 0
- Support
- How many true samples of each class are there? Watch out: a score from a handful of samples is mostly luck. Chapter 5
T
- Test set, or blind well
- Samples kept back until the very end and used once, to report how the model will do on new data. Chapter 4
- Threshold
- The cutoff value of a split. scikit-learn tries the midpoints between neighbouring values. Chapter 6
- Training and test error
- The miss on the samples the model learned from, and on samples it didn't. Only the second tells you how it will do. Chapter 2
- Training set
- The samples a model learns from. Chapter 4
- Training, validation and test
- The rows a model learns from, the rows used to choose its settings, and the rows kept back to score it honestly. Chapter 4 is about getting these right. Chapter 0
- Tuning
- Trying settings and keeping the best on validation data.
GridSearchCV, with a grouped splitter for wells Chapter 4
U
- Unit, or neuron
- One weighted sum of its inputs, plus a constant, passed through an activation. A hidden unit's output is a log the network invents. Chapter 8
- Unsupervised
- Learning without labels: finding groups or oddities in the features alone. Chapter 0
- Unsupervised
- Learning without labels: the model is never told which samples are bad hole, so it can only find structure, never meaning. Chapter 9
V
- Validation set
- The samples used to choose settings, such as k, and to compare models. Used often, so it gets used up. Chapter 4
- Variance
- How much a model changes when the training data changes a little. Deep trees have a lot of it, which chapter 7's forests fix. Chapter 6
- Vote shares
- The share of neighbours of each facies: a rough probability, used for ROC curves in chapter 5.
predict_probaChapter 3