Probabilistic facies classification: uncertainty at every voxel
Learning objectives
- Explain why hard facies classification hides uncertainty and when that matters
- Write Bayes’ rule for facies classification: P(class | elastic) ∝ P(elastic | class) · P(class)
- Model per-class likelihoods as multivariate Gaussians in elastic space and compute posteriors
- Use CONFIDENCE (max posterior) to identify voxels where classification is reliable
- Turn probabilistic cubes into drilling outputs: P(gas sand) volumes, P(net pay) maps, risk-weighted STOIIP
In Section 7.4 facies came out of classification as hard calls: each voxel was assigned its single most probable facies. The map is clean, but it hides how sure each call is. A voxel whose attributes give gas sand a probability of 0.51 and oil sand 0.49 is drawn exactly like one at 0.99 gas sand, and no one should drill on the two alike.
This section keeps the probabilities. Instead of one facies per voxel the output is a set of cubes, one per facies, , , and at every voxel, summing to one. The hard map becomes a derived product, the most probable facies with its confidence beside it, and the probabilities carry the uncertainty on into volumes, chances of success and drilling risk (Mukerji et al. 2001; Avseth, Mukerji and Mavko 2005).
Bayes’ rule for facies classification
The posterior probability of facies , given the attributes measured at a voxel, is Bayes’ theorem:
Here is typically the pair from a pre-stack inversion; is the likelihood, the probability density of those attributes in facies , learned from labelled training data; and is the prior, the probability of the facies before the attributes are seen. The denominator makes the posteriors sum to one.
- The likelihood comes from calibration wells: gather the samples logged as each facies, express their attributes as the seismic would measure them, the inversion’s error included, and fit a density to them, a Gaussian or a kernel density.
- The prior encodes geology: the proportion of each facies the wells cut in the interval, or a lateral trend from mapped depositional facies. With no prior worth trusting, equal priors make the posterior the normalised likelihood.
- The posterior is what each voxel stores. From it follow the most probable facies (the maximum a posteriori, or MAP, call), its probability (the confidence), the entropy, and the probabilities of combined classes such as .
Reading Figure 7.5
The figure builds this classifier from the rock physics of Part 5. Shale and brine sand scatter about the mudrock and sandstone lines exactly as in Figure 5.5, the oil and gas sands come from the brine sands by Gassmann with 80 % saturation, and every attribute carries an inversion error, by default 4 % in and 6 % in . Sixty training samples per facies fit the four densities, and a blind well of 800 samples that took no part in the training measures the result.
- Start from the default. The headline gives the posterior at the probe and the blind well’s score: the most probable facies is right for 645 of the 800 samples (80.6 %). Half of that well is shale, so a classifier that called every sample shale would already score 50 %; the second readout averages the four rows’ share right instead, 75.6 %, so that each facies counts alike. In the confusion matrix (c) nearly every error is between neighbours in the plane, shale and brine sand, brine and oil sand, oil and gas sand; no gas sand is called shale.
- Put the probe on “the brine-oil boundary”. There the two sands’ posteriors are equal, 0.49 each, by construction. A hard map must call the point one or the other; the posterior says the attributes cannot tell.
- Set the inversion error to zero. With attributes of log quality 98.6 % of the blind well is called right and the section in (e) turns crisp. The rocks separate well; the overlap that makes the calls uncertain is the inversion’s error.
- Move the prior on shale from the wells’ 0.50 to 0.90. The shale region grows into the brine sand’s, 154 of the 200 brine sands are called shale, and the score falls to 72.3 %. The share right stays within a point of its best from 0.30 to 0.58, around the half of shale the well actually holds: the true proportion is the prior that makes the most right calls. The average over the four facies peaks lower, 77.3 % near 0.27, because a lighter prior on shale finds more of the brine sand at the cost of shale. Which score matters depends on what each error costs.
- Choose “Axes apart” with the error at zero. Ellipses held to the axes cannot follow the tilt of the clouds (within the shale and the brine sand the two attributes correlate at about −0.9), and the score falls from 98.6 to 95.8 %.
- Shade (a) by the uncertainty . It is darkest along the boundaries between clouds and pale at their centres; in the section it marks the voxels near the contacts and the edges of the sand. A common rule flags every voxel whose top posterior is below 0.6 and maps it apart; the last readout counts them. Then drag the probe to a corner of the plane, far from every cloud: one facies can still take 0.98 there, because the posteriors must sum to one, but they only rank four unlikely explanations, and the headline says so.
The Gaussian likelihood model
The most common choice for is a multivariate Gaussian:
where is the mean of facies in the attribute space of dimensions and its covariance matrix. For the pair , , both are fitted to the training samples:
- Mean: the average attributes of the samples of that facies.
- Covariance: the 2 × 2 matrix of the two variances and their covariance. It carries the tilt of each cloud: along the mudrock and sandstone lines a faster rock has a higher impedance and a lower velocity ratio, so within each facies and are anti-correlated.
With a third attribute, such as density, the matrices are 3 × 3. With many attributes a covariance needs many samples to be estimated, and kernel densities or other non-parametric classifiers take over. A kernel density, , places a small Gaussian on every training sample and follows whatever shape they have, at the price of needing more of them.
Figure 7.5 makes two warnings measurable. A fit with the covariance set to zero, the axes taken apart, misses the tilt and misclassifies samples a full Gaussian gets right. And the contours of a two-dimensional Gaussian do not hold the one-dimensional percentages: the ellipse two standard deviations out holds % of the probability, not 95 %, and the ellipse one standard deviation out holds 39 %. The figure’s contours are drawn to hold 90 % each.
Derived products from the probabilistic cubes
Once every voxel carries its posteriors, many useful quantities follow directly:
- Hard classification: , the map everyone asks for; pair it with its confidence.
- Confidence: , between and 1 for facies; low values flag ambiguous voxels.
- Entropy: , from 0 when one facies is certain to when all are equally likely. It tells a voxel torn between two facies from one spread over many, which the confidence cannot.
- Combined classes: . Most questions an asset team asks are about combinations of facies.
- Expected properties: , the porosity expected at a voxel given its facies probabilities, the honest input to a volume.
- Risk-weighted net pay: each voxel counts with the weight , multiplied by the probabilities of passing any porosity or shale cutoffs, so the net-pay map and the volume keep their uncertainty.
- Stochastic realisations: draw facies cubes from the voxel posteriors, with a spatial model for how neighbouring voxels correlate, and run each through simulation for a range of forecasts.
Where priors come from, and how to use them
The prior is what is known before the seismic is read. Common sources:
- Proportions from nearby wells: if the wells cut 50 % shale in the target interval, , and the rest is shared among the sands as the wells share it. Figure 7.5 starts from such a prior.
- Depositional setting (Part 4): a channel axis holds more sand than its levee, so priors can vary laterally with the mapped facies.
- Sequence-stratigraphic position (Section 4.2): a lowstand fan and a highstand shelf expect different proportions of sand.
- Equal priors: when no prior can be trusted. The posterior is then the normalised likelihood, driven by the data alone.
Priors matter most where the likelihoods are close, which is where the data are ambiguous. A strong prior on shale keeps the posterior on shale where the attributes somewhat favour sand: right if the prior is right, a way to suppress a real anomaly if it is not. In the figure a prior of 0.90 on shale, against the 0.50 the blind well holds, costs eight points of accuracy. Report posteriors under both equal and geological priors, and explain any large difference.
When probabilistic classification fails
- Too few training samples. A covariance fitted to a few samples is itself uncertain, and the posteriors it gives are overconfident. In the figure, eight samples per facies barely change the share called right, but the calls made with a posterior of 0.8 or more promise 0.95 on average and deliver 89 %. Mitigate by shrinking each covariance toward a pooled one, or by gathering more samples.
- Clusters that are not Gaussian. Some facies, heterolithic thin beds for one, have bimodal or curved distributions that one Gaussian misrepresents. Fit a mixture of Gaussians per facies, or a kernel density.
- Overlap by nature. If two facies occupy the same region of the attribute space, two sands of different lithology with the same and for example, no probability calculus separates them. Accept the ambiguity, or add an attribute that tells them apart.
- Priors from elsewhere. Proportions from one basin used in another bias every posterior. Calibrate priors for each basin and target interval.
- Densities from clean logs. Densities fitted to log-quality attributes and applied to attributes from seismic are too narrow, and the posteriors too confident. Build the densities from attributes that carry the inversion’s error, as Mukerji et al. (2001) do and as the figure does.
- Trusting the confidence untested. A high posterior means the model is confident, not that it is right. Test it on blind wells: of the voxels given about 0.9 for gas sand, are about 90 % gas sand? The figure’s calibration readout asks exactly this of its blind well. If the answer is no, recalibrate the posteriors.
Connecting to stochastic inversion (Section 7.3)
A complete workflow joins probabilistic facies to the stochastic pre-stack inversion of Section 7.3:
- The inversion delivers not one set of attributes per voxel but samples of their posterior, typically 50 to 500 realisations.
- The facies classifier runs on each realisation, giving the posteriors of every facies at every voxel.
- Averaging over the realisations gives the final facies probabilities, which carry both the inversion’s uncertainty and the classification’s.
In symbols, : the first factor is the classifier, the second the inversion’s posterior, and the average over realisations is a Monte Carlo estimate of the integral. Figure 7.5 takes a shortcut with the same intent: it folds a fixed inversion error into the training samples, so its densities are already those of attributes measured through seismic.
This is the standard for high-value reservoir studies, where a drilling decision turns on how uncertain the prediction is.
Probabilistic classification is hard classification with its uncertainty kept. It uses everything the inversion and the rock physics provide, says where the map can be trusted, and hands the people who decide the probabilities a risk analysis needs. Section 7.6 closes Part 7 by carrying these facies and rock-property probabilities into a reservoir model with its geological framework, volumes and their uncertainty.
References
- Mavko, G., Mukerji, T., & Dvorkin, J. (2009). The Rock Physics Handbook (2nd ed.). Cambridge University Press.
- Chopra, S., & Marfurt, K. J. (2007). Seismic Attributes for Prospect Identification and Reservoir Characterization. Society of Exploration Geophysicists.
- Hilterman, F. (2001). Seismic Amplitude Interpretation. SEG/EAGE Distinguished Instructor Short Course.
- Foster, D. J., Keys, R. G., & Lane, F. D. (2010). Interpretation of AVO anomalies. Geophysics, 75(5), 75A3-75A13.
- Mukerji, T., Jørstad, A., Avseth, P., Mavko, G., & Granli, J. R. (2001). Mapping lithofacies and pore-fluid probabilities in a North Sea reservoir: Seismic inversions and statistical rock physics. Geophysics, 66(4), 988-1001.
- Avseth, P., Mukerji, T., & Mavko, G. (2005). Quantitative Seismic Interpretation: Applying Rock Physics Tools to Reduce Interpretation Risk. Cambridge University Press.
- Silverman, B. W. (1986). Density Estimation for Statistics and Data Analysis. Chapman and Hall.