Assumptions

What the lab simplifies

Every simplification in the teaching field, every physics value still to check, and every place the lab's algorithms differ from scikit-learn's.

The teaching field

Physics values still to check

The rock physics is copied from Probe Lab, whose research report marks some values as textbook figures still to be checked against a source. These are all the values so marked in the material table Fit Lab copies; not every one of these minerals appears in the teaching field. Two tool constants are also unverified: the density tool's mass attenuation coefficient at 662 keV, and the neutron tool's spacings and slowing-down lengths.

Where the lab differs from scikit-learn

k-means
From the same starting centroids, the same labels, centroids and inertia as KMeans. The figures start from your centroids or from seeded random starts; KMeans starts with k-means++ by default.
Linear regression
The same coefficients as LinearRegression and PolynomialFeatures, to within 10−9. Gradient descent is the lab's own, for showing the steps.
k-nearest neighbours
Brute-force search, the same predictions and vote shares as KNeighborsClassifier, including ties, which go to the class listed first.
Decision trees
CART with Gini impurity and 32-bit cutoffs, as DecisionTreeClassifier. When two logs tie exactly for the best cut, the lab tries the logs in order (or a seeded shuffle); scikit-learn's order is its own random one, so deep trees can part ways after the first exact tie.
Random forests
Bootstrap samples, the square root of the logs at each split and a soft vote, as RandomForestClassifier, but with the lab's own random numbers: blind-well accuracy is within two points of scikit-learn's average over ten seeds, and the sonic regression's error within its range.
Neural networks
The design of MLPRegressor (Glorot starts, ReLU, Adam, half the squared error plus an L2 penalty) with the lab's own random numbers; the training loss and the blind-well error fall within scikit-learn's spread over ten seeds. The figures use a learning rate of 0.01 and batches of 64 (scikit-learn's defaults are 0.001 and up to 200) so they train in about a second, and early stopping uses a validation well where scikit-learn holds back a random tenth.
Isolation forest
The design and defaults of IsolationForest (100 trees of 256 samples, a height limit of 8), with the lab's own random numbers; the washout's ROC AUC is within two points of scikit-learn's average over ten seeds.
Metrics
The same values as sklearn.metrics, to twelve decimal places. Where a facies is never called, its precision has nothing to divide by: scikit-learn reports 0 with a warning, and so does the lab, with a note or a dash in the figures.
Validation
Random splits are seeded. Block cross-validation assigns depth blocks to folds by a fixed rule, and per-well normalisation matches each well's 5th and 95th percentiles to a reference; both are the lab's own helpers, as scikit-learn has no direct equivalent.

Credits