A log is a table
What does a model actually see? Not a log, not a well, not the geology: a table of numbers, one row per depth. Everything in the chapters that follow starts from that table.
What a model sees
A petrophysicist looks at a log and sees beds, boundaries and trends. A machine learning model is given something much plainer: a table, one row for each depth, one column for each log. Watch 6 m of Well A make the trip.The six wells in this lab are synthetic: fictional wells logged with the rock physics from Probe Lab, so the true facies is known at every depth. The logs are sampled every 15 cm, as real logs usually are.
Nothing about the rock survives the trip except the numbers. The depth is still there, but only as the row's address; the model isn't told that one row sits above another, that two rows come from the same bed, or that a log was run with an older tool. Each of those turns out to matter, and later chapters come back to every one.
Points on a crossplot
If each row is a point, then rocks that read alike sit together, and every crossplot a petrophysicist has ever drawn is already a picture of what a model works with. The log and the crossplot are two views of the same table, so a selection in one lights up the same rows in the other.
That's most of machine learning on logs in one figure. The facies form clouds; algorithms either find the clouds on their own (chapter 1), draw boundaries between them (chapters 3, 6 and 7), or fit a surface through them (chapters 2 and 8). The hard cases are the points between clouds, and the points that land in the wrong cloud because the hole or the tool misbehaved (chapters 4 and 9).
Features and labels
The table doesn't say what to do with it. The question does: which columns go in, which one is the answer, and which are left out.
Two kinds of question run through the lab. Supervised learning has a label: some rows come with the answer, from core, from an interpreter, or from a log run in another well, and the model learns to give it where it's missing. Unsupervised learning has none: the model looks for structure on its own, and a person decides what that structure means.
Try it on real wells
The rest of the lab uses the six synthetic wells, because their truth is known. Every chapter ends here instead, on real wells from the FORCE 2020 competition, where the tidy lessons get messier. This is what those real tables look like.
Words you'll meet
Machine learning renames familiar things, and different books use different names for the same one. Here they are in plain words, with the name in scikit-learn and pandas where it helps.
- Sample
- One row of the table: one depth in one well. Also called an observation, an instance or a record.
- Feature
- A column that goes into the model: here a log. Also called an input, a predictor, a variable or an attribute. scikit-learn calls the table of features
X. - Label
- The column the model is asked to predict. Also called the target, the response or the output. scikit-learn calls it
y. - Supervised
- Learning from rows where the label is known, to predict it where it isn't.
- Unsupervised
- Learning without labels: finding groups or oddities in the features alone.
- Classification
- A supervised question whose answer is a category: which facies, pay or not.
- Regression
- A supervised question whose answer is a number: what porosity, what sonic.
- Model
- The recipe that turns features into a prediction, with numbers learned from the data.
fitlearns them;predictuses them. - Training, validation and test
- The rows a model learns from, the rows used to choose its settings, and the rows kept back to score it honestly. Chapter 4 is about getting these right.
- Feature space
- The space with one axis per feature, where each sample is a point. With two logs it's a crossplot; with five it can't be drawn, but the algorithms don't mind.
- Missing value
- A gap in the table, such as a sonic that wasn't run. Most algorithms can't use a row with a gap, so the row is dropped or the gap is filled. In pandas it's
NaN. - DataFrame
- pandas's table, the usual way to hold logs in Python once they're read from a LAS file: one row per depth, one column per curve.
What to remember
- A model sees a table: one row per depth, one column per log. Beds, boundaries and wells aren't in it unless you put them there.
- Each row is a point; rocks that read alike sit together. A crossplot is already a picture of what a model works with.
- The question decides the roles: features go in, the label is the answer, and with no label the learning is unsupervised.