Machine-learning QI: neural networks for facies classification
Learning objectives
- Contrast rule-based (Section 7.5) vs data-driven (ML) facies classification
- Recognize overfitting: sparse training data + flexible model = wrong decision regions
- Understand the role of training-validation-test split for honest accuracy estimates
- Identify when ML adds value vs when simpler methods suffice
- Apply practical safeguards: data augmentation, regularization, blind-well validation
Section 7.5 classified facies by rule: a Gaussian density for each facies in the plane of and , a prior for each, and Bayes’ rule for the posterior. That works when each facies is described well by its density and the training samples are a fair picture of the ground. A machine-learning classifier drops the explicit densities and learns the boundaries between facies directly from labelled samples. It is more flexible, and for the same reason more at the mercy of what those samples are and of how it is tested.
Machine learning is now routine for facies classification from logs and seismic attributes (Hall 2016; Wrona et al. 2018), and it is spreading into inversion, horizon tracking and noise attenuation (Dramsch 2020). This section stays with the simplest case, a classifier on two elastic attributes, because what it teaches about training data, testing and probabilities carries over to every application.
What a classifier learns
Given labelled samples, attributes such as , and density with a facies for each, a classifier builds a function from attributes to facies. Its parameters are fitted by minimising a training loss, for classification usually the cross-entropy, the mean of for the true facies.
- Linear (logistic regression): straight boundaries. Robust and easy to read, but it cannot follow curved clusters.
- nearest neighbours: a vote among the training samples nearest in the attributes. Nothing is fitted; the boundaries follow the samples, and at every sample is memorised.
- Decision trees and random forests: splits along one attribute at a time. Robust and interpretable on tabular data.
- Gradient-boosted trees: often the strongest on tabular data such as well logs, and among the leading entries of the 2016 SEG facies contest (Hall and Hall 2017).
- Neural networks (multilayer perceptrons): smooth boundaries free to bend; a softmax output gives a score for every facies.
- Convolutional networks: for images and volumes rather than tables, such as seismic facies or salt bodies picked from the image itself.
Reading Figure 8.5
The figure trains real classifiers in your browser on the anticline of Figure 7.5: the same four facies, the same rocks built by the rules of Figure 5.5 with the hydrocarbon sands by Gassmann, and the same inversion error, except that here the error is band-limited, smooth over tens of metres down, so a sand in one well carries one error throughout. Wells are columns of the section, 64 samples each. The training wells come in a fixed drilling order, two flank wells first; two validation wells and two blind wells are held back; and every voxel no well drilled gives the truth, which only a synthetic earth can show. The network is small, two layers of 12 units, and trains by gradient descent for up to 600 epochs; a nearest-neighbour vote and the Gaussian rule of Figure 7.5 can take its place.
In its default state, four wells and 300 epochs, the network finds each facies on the blind wells 90.9 % of the time on average; the two validation wells said 74.1 % and the undrilled section, the truth, gives 82.8 %. Two wells are a small sample: here the validation pair is harder than the section and the blind pair easier. Real projects never see the truth, which is why they need more than one well held back, and why a blind-well score should be read with its spread in mind.
Training, validation and test, split by well
Never score a model on the data it was trained on. Three sets do three jobs:
- Training set: fits the weights.
- Validation set: watches the training and chooses the settings, such as the number of epochs, the size of the network or . Because the settings are chosen to suit it, it ends up a little optimistic.
- Test set: used once, after everything is fixed, to report the score. In quantitative interpretation these are the blind wells.
Samples from one well are not independent. Neighbouring samples share the rock and, above all, the error of the inversion that delivered them, which a wavelet and a low-frequency model spread over tens of metres. Split at random, the validation samples sit beside training samples that are near copies of themselves, so the validation set measures what the model remembers, not what it predicts. Split by well, the validation wells are as new to the model as the next well will be. Bergen et al. (2019) stress that a test set must be independent of the training data, which for geoscience data means separating it in space or time (Roberts et al. 2017 compare the ways of doing so), and the 2016 SEG contest scored every entry on two wells that no team could see (Hall 2016; Hall and Hall 2017).
Figure 8.5 measures the leak. Split at random, the network’s validation loss follows its training loss down (0.13 nats at 300 epochs, against 0.10) and is least at the last epoch, while the loss on the undrilled section rises from 0.19 at epoch 100 to 0.26 at 600; the random validation reports each facies found 94.3 % of the time, the blind wells 91.4 % and the section 84.2 %. The leak shows most in the loss: across the figure’s 56 states of error and wells, the random split’s validation loss lies below the section’s in 45. In balanced accuracy this small network on two attributes leaks little, because it cannot remember single samples: its random estimate exceeds the section’s in only 35 states, by 0.9 points on average. The nearest-neighbour vote with , a pure memoriser, shows the leak plainly: by well it estimates 78.8 % against a truth of 79.0 %; split at random it estimates 85.9 % while the blind wells and the section both give 78.1 %, and across the 56 states its random estimate overstates the blind wells’ score in 45, by 10.6 points on average. All these states share one earth and one set of wells, and the two wells V are a hard pair: their loss lies above the section’s in 54 of the 56.
Overfitting and the validation curve
Gradient descent drives the training loss down epoch after epoch, and past some point it does so by fitting what is particular to the training wells. The sign is a validation curve that stops falling and turns up while the training curve keeps going down. In Figure 8.5’s default state the training loss falls from 0.33 nats at epoch 10 to 0.03 at 300, while the validation wells’ loss climbs from 0.46 to 1.98 and the section’s, the truth, is least at epoch 40. Across the figure’s states, wherever the truth’s loss rises between epochs 100 and 600 (46 of 56), the held-back wells’ loss rises too in all but one; with the random split the section’s own loss rises in 45 states, and its validation loss in only 13 of them.
Here the balanced accuracy barely changes over the same epochs: what grows is confidence. Each answer becomes surer, the wrong ones included, so the loss on new wells rises while the share right stays put. Early stopping, ending the training where the validation loss is least, weight decay and smaller networks are the usual remedies (Goodfellow et al. 2016), and every one of them needs a validation set that is honest, which means whole wells.
Rare facies and a balanced score
Shale is three quarters of the section and of every well, and the gas sand the well was drilled to find is the rarest facies. A plain accuracy counts samples, so it rewards shale: calling every sample shale scores 75 %. The balanced accuracy averages each facies’ share found, its recall, so every facies counts alike and calling everything shale scores 25 %. Report both, with the recall of each facies, as Figures 7.5 and 8.5 do.
The difference is starkest when a facies was never drilled. Trained on the two flank wells, which cut only shale and brine sand, the network calls no oil or gas anywhere; on the blind wells it calls 77.3 % of the samples right yet finds each facies only 48.7 % of the time, every oil and gas sample called brine sand. No architecture can name a facies its wells never cut. The network’s loss weighs each facies by the square root of its inverse frequency, so the rare ones count more than their numbers; that helps a rare facies that was drilled, and does nothing for one that was not.
What the probabilities mean
A softmax output looks like a probability: four positive numbers that sum to one. It is a score that training produced, and whether it behaves as a frequency, so that calls made at 0.9 are right nine times in ten, is a question only wells the model never saw can answer. Plate (e) of Figure 8.5 answers it on the section. In the default state the network’s calls made at 0.8 or more average 0.99 and are right 88.7 % of the time; training longer makes them surer and no more often right, as Guo et al. (2017) found in far larger networks. The Gaussian rule, fitted to the same wells, is less sure and more honest: 93.2 % right for an average of 0.98.
Three things a probability does not say. It does not say how far a sample lies from the training data: far from every cloud the outputs still sum to one, and one facies can take nearly all of it. It cannot speak for a facies the training never held: trained on the flank wells, the network gives the gas cap a probability of gas of zero, made with near certainty. And it is not a mixture within the voxel: 0.8 for gas sand means that among voxels that look like this one, about four in five should be gas sand if the probabilities are honest, not that the voxel is four-fifths gas sand.
When machine learning beats a rule-based classifier, and when it does not
- Machine learning wins with clusters that are curved, overlapping or of several modes, with many attributes, where a Gaussian per facies becomes hard to estimate, and with abundant, well-spread training data. In Figure 8.5 the network finds each facies of the section 82.8 % of the time on four wells against the Gaussian rule’s 79.7 %, and the gap widens to 58.9 against 53.3 % at an inversion error of 8 %.
- The rule wins with few samples, with clusters close to elliptical (at log quality the two tie, 99.2 and 98.7 %), when every call must be traced to an equation, and when the rock physics is understood well enough to write down rather than learn again.
- Both together are common practice: run the rule as the default and the network beside it, trust the regions where they agree and inspect those where they differ.
In every case the number of wells matters more than the choice of model: the Gaussian rule finds each facies of the section 34.9 % of the time with one well and 84.0 % with eight.
Beyond classification: machine learning for inversion
- Learned inversion: a regressor maps seismic waveforms, near to far offsets, straight to , and density. Fast, and unreliable wherever the training data do not cover the rocks.
- Physics-informed neural networks: the wave equation’s residual joins the loss, so the result fits the data and the physics at once. Still largely research.
- Unsupervised anomaly detection: an autoencoder learns what ordinary seismic looks like and flags what it cannot reconstruct, a quick screen of large surveys.
- Generative models: networks that draw reservoir realisations consistent with the seismic, for uncertainty beyond what stochastic inversion gives.
The lessons of classification carry over unchanged: discipline in the training data and in the testing matters more than the sophistication of the architecture.
A practical checklist
- Audit the training data: samples per facies, the spread of wells over the area and the depth range, and which facies were never drilled.
- Split by well: validation and test wells fully apart from the training wells, and more than one of each, because one well is a small sample.
- Report a balanced score beside the plain accuracy, with the recall of every facies.
- Watch the validation curve and stop or regularise before the loss on held-back wells turns up.
- Test the probabilities: on blind wells, check that calls made at a given probability are right that often before handing anyone a probability cube.
- Train on the attributes you will apply it to: seismic attributes with the inversion’s error, not clean logs.
- Keep testing after deployment: compare every new well with the prediction, and retrain when they drift apart.
Machine learning is a powerful tool for quantitative interpretation and not a shortcut around its disciplines: calibrated wells, physical plausibility, honest testing and uncertainty carried through. A classifier validated on whole wells, scored facies by facies and with probabilities checked against the truth can come close to the rule-based optimum of Section 7.5 with fewer assumptions; one validated on a random split reports a confidence it has not earned. Section 8.6 closes Part 8 with CO₂ storage and monitoring, where rock physics, inversion, probabilistic facies, 4D and machine learning meet in one programme.
References
- Avseth, P., Mukerji, T., & Mavko, G. (2005). Quantitative Seismic Interpretation. Cambridge University Press.
- Bergen, K. J., Johnson, P. A., de Hoop, M. V., & Beroza, G. C. (2019). Machine learning for data-driven discovery in solid Earth geoscience. Science, 363(6433), eaau0323.
- Dramsch, J. S. (2020). 70 years of machine learning in geoscience in review. Advances in Geophysics, 61, 1-55.
- Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.
- Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 1321-1330.
- Hall, B. (2016). Facies classification using machine learning. The Leading Edge, 35(10), 906-909.
- Hall, M., & Hall, B. (2017). Distributed collaborative prediction: Results of the machine learning contest. The Leading Edge, 36(3), 267-269.
- Mavko, G., Mukerji, T., & Dvorkin, J. (2009). The Rock Physics Handbook (2nd ed.). Cambridge University Press.
- Roberts, D. R., Bahn, V., Ciuti, S., et al. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913-929.
- Wrona, T., Pan, I., Gawthorpe, R. L., & Fossen, H. (2018). Seismic facies analysis using machine learning. Geophysics, 83(5), O83-O95.