Chapter 5 · Scoring a model

Metrics, plainly

Which number tells you a model is any good? Accuracy, precision, recall, F1, ROC curves, averages over many classes, RMSE and R², without the jargon: what each one answers, and how each one can fool you.

16 minutesA gamma ray cutoff, two facies models and a porosity log

Two kinds of question

A metric is a question with a number for an answer, so start with the question. Models answer two kinds. Classification asks which one?: which facies, sand or not, pay or not. Regression asks how much?: how much porosity, what permeability, what sonic. They're scored differently, because being wrong means different things. Call a shale a sand and you're simply wrong; predict 21% porosity when it's 22% and you're nearly right.

Is it sand?

The simplest classification has two answers, and every petrophysicist has built one: a gamma ray cutoff. Below it, call the rock sand. Every sample then lands in one of four boxes: sand you called sand (a hit), sand you missed, something else you called sand (a false alarm), and something else you left alone. Set the cutoff where you think it does best.The samples are every 30 cm through Wells A, B and C: 1,968 of them, a quarter of them clean sandstone. "Sand" here means the sandstone facies; sandstone/shale counts as not sand.

Each score answers a different question, so each has a different best cutoff. Precision asks how much to trust a "sand" call: it's highest with a low cutoff that only calls the cleanest sand. Recall asks how much of the sand you found: it reaches 100% once the cutoff is high enough to catch everything, false alarms and all. F1 balances the two. Which matters depends on the job. Summing net sand for volumes, a missed sand costs you; picking depths for sidewall cores, a false alarm does.

And watch accuracy. Press "Call nothing sand": a rule that never finds a single sand scores 74%, because three quarters of the samples aren't sand. When one answer is much more common than the other, accuracy rewards saying the common one.

Every cutoff at once

Most models don't hand you a yes or a no. They give a score: a probability of sand, a share of the neighbours' votes, or, for the simplest model of all, the gamma ray itself. A cutoff turns the score into calls, and as Figure 5.1 showed, every cutoff scores differently. A ROC curve shows every cutoff at once: for each one, how much of the sand it finds (the hit rate, which is recall) against how much of everything else it wrongly calls sand (the false-alarm rate).ROC stands for receiver operating characteristic, a name left over from 1940s radar, where the question was whether a blip was an aircraft.

The area under the ROC curve, AUC, scores the ranking before anyone chooses a cutoff. That makes it good for comparing models, and blind to the cutoff you'll actually use: a model with a fine AUC can still make poor calls with a badly placed cutoff. It's also blind to how rare the class is, because both of its rates are shares of the truth. That's a strength when you compare wells with different amounts of sand, and a weakness when the class is rare: thousands of correct rejections make a handful of false alarms look tiny. The precision–recall curve and its summary, average precision, don't have that blind spot. When the class you're after is rare (coal, a thin pay zone, bad hole), look at precision–recall.

Six facies, one against the rest

With more than two classes, everything so far still works, one facies at a time. Pick a facies, and the six-by-six confusion matrix from chapter 4 splits into the same four boxes as the sand cutoff: its hits on the diagonal, its misses along its row, its false alarms down its column, and everything else. That gives each facies its own precision, recall and F1, and, using the model's vote share for it as a score, its own ROC curve and AUC. This is one against the rest (one-vs-rest), and it's how scikit-learn's classification report is built.

The report has one row per facies. Support is how many samples of that facies are really there; check it before trusting a row, because a precision worked out from four samples swings by 25 points with one more. The coal row with 75 neighbours makes the chapter's sharpest point: recall 0%, ROC AUC near 100%. The calls and the scores behind them are different things, and each metric judges only one of them.When a facies is never called, its precision has nothing to divide by. scikit-learn reports 0 and prints a warning; the figure says so instead.

Six facies, one number?

The report ends with averages, because reports usually want one number. There are several ways to average, and they disagree in exactly the case chapter 4 warned about.

Averages that weight each facies by its number of samples are dominated by shale and sand. Macro averages, and balanced accuracy, give every facies an equal say, so a facies that's never found drags them down hard. If the rare facies matter to you, report a macro average, or better, the recall for each facies. And don't let a macro AUC stand in for them: it judges the scores, so it can't see a facies that's never called.

Fill it in yourself

The quickest way to get a feel for these scores is to push them around. Change the matrix one tap at a time and watch which numbers move, and which don't.

A few taps make the pattern plain. A hit or a mistake in a big class moves accuracy and the weighted averages; in a small class it barely touches them but swings the macro averages, because each facies gets an equal vote there whatever its size. And a mistake always counts twice: one facies' miss is another's false alarm.

How far off?

For a regression, every prediction has an error, and the metrics summarise them in different ways. Here the "model" is density porosity, and the truth is the porosity measured on 420 core plugs. Break it in different ways and see which numbers notice.p.u. is a porosity unit: one percent of the rock's volume. 0.01 v/v is 1 p.u.

Bias tells you whether the model reads high or low on average, and it's the one error that's easy to fix. MAE is the typical miss, in the same units as the log. RMSE is like MAE but squares each miss first, so a few big misses count for a lot; it's always at least as big as MAE, and the gap between them tells you whether a few bad samples are doing the damage. R² compares the model with the laziest one possible, predicting the average every time. It doesn't have units, which makes it easy to quote and easy to misread: a high R² can still hide a bias, and R² depends on how varied the truth is, so the same model gets a lower R² in a well with a narrow porosity range.

Try it on real wells

A classification report from a real blind well, where the rare facies really are rare.

The cheat sheet

MetricThe question it answersWatch out
AccuracyHow many samples did it get right?Rewards always saying the common answer.
PrecisionWhen it says "sand", how often is it sand?Can be perfect while missing most of the sand.
RecallOf all the sand, how much did it find?Can be perfect by calling everything sand.
F1Precision and recall togetherTreats a miss and a false alarm as equally bad, which your job may not.
Balanced accuracy, macro F1How well does it do on every class, rare ones included?A tiny class with a few errors can swing them.
Weighted F1How well does it do, sample by sample?Hides rare classes, like accuracy.
Micro averagePool every class's boxes, then scoreWith one label per sample it's just accuracy.
Confusion matrixWhich classes does it mistake for which?Rows give recall, columns precision. Look at the counts as well as the shares.
SupportHow many true samples of each class are there?A score from a handful of samples is mostly luck.
ROC AUCDoes the score rank the class above the rest, whatever the cutoff?Says nothing about the cutoff you use; flatters when the class is rare.
Average precisionHow precise is it as it finds more of the class?Its floor is the class's share, not 50%, so compare it with that.
One-vs-rest AUC, macroDoes it rank every class well against the others?Can be near perfect while a class is never called.
BiasDoes it read high or low on average?Zero bias can hide big errors that cancel out.
MAEHow far off is a typical prediction?Says nothing about the worst misses.
RMSEHow far off, with big misses counting extra?A handful of bad samples can dominate it.
R²How much better than predicting the average?No units; depends on how varied the truth is; can hide a bias.

Other names you'll meet. Specificity is recall for the "not" class: the share of non-sand left alone, one minus the false-alarm rate. Cohen's kappa is accuracy corrected for the agreement you'd get by chance, so 0 means no better than guessing in the right proportions. The Matthews correlation coefficient condenses all four boxes into one number from −1 to 1 that stays fair when one class is rare. Log loss (cross-entropy) scores the probabilities themselves rather than the calls, and punishes confident mistakes hardest.

What to remember

  1. Choose the question first, then the metric that answers it. The best cutoff, or model, depends on which mistakes cost you more.
  2. Accuracy flatters when one answer is common. Read the report one class at a time with its support, and remember that ROC AUC judges the scores, not the calls: use precision–recall when the class is rare.
  3. For regression, quote bias and a typical miss (MAE or RMSE) in the log's own units, not just R².

Read more

Data Quality Considerations for Machine Learning Models

Next: decision trees

Could a computer choose our cutoffs? Facies by a tree of yes-or-no questions.