Assumptions
What the lab simplifies
Every simplification in the teaching field, every physics value still to check, and every place the lab's algorithms differ from scikit-learn's.
The teaching field
- Six fictional wells through the same layer-cake stratigraphy: shale, Sand A, a shale with coals, a limestone, a shale with anhydrite, Sand B and shale. Each well is 200 m, sampled every 15.24 cm, and vertical, so measured depth is true vertical depth.
- Six facies, all classes in the FORCE 2020 labels: sandstone, sandstone/shale, shale, limestone, coal and anhydrite. Coal and anhydrite are kept rare on purpose.
- Logs are Probe Lab's point responses for each rock, smoothed to each tool's vertical resolution, with noise matched to a FORCE 2020 interval. The deep and medium resistivity are induction curves.
- One mud for every well (fresh water based, 1.2 g/cm³, no barite), an 8.5 in bit and 25 cm of invasion in permeable beds. Sand A has an oil–water contact at 2045 m.
- Permeability comes from Pore Lab's Kozeny-Carman model, with grain size and sorting set bed by bed and about a quarter of a decade of variation from sample to sample. The 420 core plugs, every 30 cm through Sand A in Wells A, B and C, add a little laboratory measurement noise.
- Built-in problems, each placed for a chapter: a washout in Well D that corrupts density, neutron and Pe; no sonic in the top 40 m of Wells B and F; Well E logged with older tools (gamma ray 15 API high, neutron 0.04 low, density 0.035 g/cm³ high, Pe 0.4 high); a tight dolomitised limestone in Well F, harder than any rock in the other wells; beds thinner than the tools can resolve; and step-out wells: Well D with more compacted shales and density reading 0.02 g/cm³ high, and Well F with slightly compacted shales, a feldspathic Sand B and density reading 0.015 g/cm³ low.
Physics values still to check
The rock physics is copied from Probe Lab, whose research report marks some values as textbook figures still to be checked against a source. These are all the values so marked in the material table Fit Lab copies; not every one of these minerals appears in the teaching field. Two tool constants are also unverified: the density tool's mass attenuation coefficient at 662 keV, and the neutron tool's spacings and slowing-down lengths.
Where the lab differs from scikit-learn
- k-means
- From the same starting centroids, the same labels, centroids and inertia as
KMeans. The figures start from your centroids or from seeded random starts;KMeansstarts with k-means++ by default. - Linear regression
- The same coefficients as
LinearRegressionandPolynomialFeatures, to within 10−9. Gradient descent is the lab's own, for showing the steps. - k-nearest neighbours
- Brute-force search, the same predictions and vote shares as
KNeighborsClassifier, including ties, which go to the class listed first. - Decision trees
- CART with Gini impurity and 32-bit cutoffs, as
DecisionTreeClassifier. When two logs tie exactly for the best cut, the lab tries the logs in order (or a seeded shuffle); scikit-learn's order is its own random one, so deep trees can part ways after the first exact tie. - Random forests
- Bootstrap samples, the square root of the logs at each split and a soft vote, as
RandomForestClassifier, but with the lab's own random numbers: blind-well accuracy is within two points of scikit-learn's average over ten seeds, and the sonic regression's error within its range. - Neural networks
- The design of
MLPRegressor(Glorot starts, ReLU, Adam, half the squared error plus an L2 penalty) with the lab's own random numbers; the training loss and the blind-well error fall within scikit-learn's spread over ten seeds. The figures use a learning rate of 0.01 and batches of 64 (scikit-learn's defaults are 0.001 and up to 200) so they train in about a second, and early stopping uses a validation well where scikit-learn holds back a random tenth. - Isolation forest
- The design and defaults of
IsolationForest(100 trees of 256 samples, a height limit of 8), with the lab's own random numbers; the washout's ROC AUC is within two points of scikit-learn's average over ten seeds. - Metrics
- The same values as
sklearn.metrics, to twelve decimal places. Where a facies is never called, its precision has nothing to divide by: scikit-learn reports 0 with a warning, and so does the lab, with a note or a dash in the figures. - Validation
- Random splits are seeded. Block cross-validation assigns depth blocks to folds by a fixed rule, and per-well normalisation matches each well's 5th and 95th percentiles to a reference; both are the lab's own helpers, as scikit-learn has no direct equivalent.
Credits
- Rock physics and tool responses: Probe Lab. Permeability: Pore Lab. Vertical resolution smoothing: Bore Lab. All three are in the Labs series.
- Real wells: the FORCE 2020 Machine Learning Competition dataset. See the Data page for the citation and licences.