ML in processing: where it fits

Part 9, Machine Learning in Processing

Learning objectives

  • Classify processing tasks by whether ML, physics, or a hybrid approach dominates
  • Identify the data-quality and label-availability constraints that determine ML success
  • Recognise the common ML architectures used in seismic processing and the tasks they fit
  • Understand why ML complements rather than replaces physics-based processing

Machine learning has entered every corner of seismic processing over the past decade, but it has not replaced classical processing; it has augmented it. The question for a practitioner is not "should I use ML?" but "which tasks in my flow benefit from ML, which stay better with physics, and where is a hybrid approach the right answer?" This section maps that landscape before the next four sections dive into specific ML methods.

1. The three regimes

  • Physics-dominated tasks. Migration, FWI velocity updates, deghosting. These have precise physics (wave equation, reflection coefficients, ghost impulse response) and well-validated numerical methods. ML helps at the edges (starting models and low-frequency extrapolation for FWI, approximating least-squares migration) but does not lead. Do not replace RTM with a neural net: you lose the physics guarantee and gain little that the velocity model would not give you.
  • ML-dominated tasks. First-break picking, random-noise attenuation, trace reconstruction and, downstream of processing, fault and horizon labelling. These are pattern-recognition problems with no clean analytical solution. The classical methods (STA/LTA pickers, f-x deconvolution, sparse Radon reconstruction, edge attributes) are heuristic or break down on complex data; ML (U-Net, blind-spot CNNs, transformers) often beats them on both accuracy and throughput.
  • Hybrid tasks. Velocity analysis, 4D matching, residual moveout. These have decent physics-based baselines but benefit from ML refinement, typically ML-assisted auto-picking of features that a classical semblance-style method identified first.

2. The map

Put the ten tasks on two axes and the three regimes fall out of one rule. In Figure 9.1 you then change the survey (its training data, its noise, where its networks were trained, whether amplitudes feed AVO) and watch which tasks change regime.

ML scorecard for seismic modelsIoU78%Precision81%Recall74%F177%OOD drop18%Report multiple metrics + an OOD probe - a single F1 hides real failure modes

Ten tasks sit on two axes: how strong the physics-based answer is (xx) and how much ML adds to it (yy). The placements are an editorial reading of production practice around 2025, not measurements. The regime follows from the margin m=y−xm = y - x alone: physics-led at mle−1.5m \\le -1.5, ML-led at mge1.5m \\ge 1.5, hybrid between. With synthetic training labels, 10 dB of SNR and networks trained in this basin, ML leads 4 of 10 tasks, physics leads 3 and 3 are hybrid; the pickers trained on field labels keep only 70 % of their ideal gain.

Read the map along its diagonal: the further a task sits toward a large ML gain and a weak physics answer, the more ML leads. Tasks where the classical method is heuristic (first-break picking and, downstream of processing, fault and horizon labelling) are ML-led; tasks with a precise physics answer (migration, FWI, deghosting) are physics-led. The verdict is not fixed. Take the training data away and first-break picking and fault labelling fall back to hybrid, while random-noise attenuation stays ML-led because its network trains on the noisy data itself. Turn on the amplitude constraint and denoising and reconstruction become hybrid, and 4D matching becomes physics-led (windowed least squares).

3. ML architectures in seismic

  • U-Net. The workhorse. Encoder-decoder CNN with skip connections. Used for: denoising, interpolation, first-break picking, fault labelling, horizon picking. Typically from about one to a few tens of millions of parameters, trained on 2D or 3D patches.
  • DnCNN and self-supervised denoisers. DnCNN is a supervised residual CNN trained on noisy and clean pairs (in seismic, usually synthetics). Noise2Noise trains on two independent noisy copies of the same signal, and blind-spot networks (Noise2Void, StructN2V) train on a single noisy field dataset, which matters because clean ground-truth seismic is unavailable.
  • Transformers. Attention-based architectures now used for global-context tasks (fault prediction on large sections, velocity model building). Computationally expensive but superior for long-range correlations.
  • Physics-informed neural networks (PINNs). Enforce the wave equation as a loss term during training. Used in emerging FWI acceleration workflows; still at the research stage, not yet in production.
  • Diffusion models. Recent arrivals; used for generative trace-reconstruction and prior-informed FWI initialisation.

4. The training-data problem

The single largest determinant of ML success in seismic processing is training data. Key constraints:

  • Labelled seismic is scarce and expensive. Human interpreters label faults and horizons one line at a time.
  • Proprietary surveys. Most production seismic cannot be shared, limiting public training sets.
  • Domain shift. A model trained on North Sea data may underperform on Gulf of Mexico data because wavelets, noise signatures, and geology differ.
  • Synthetic training + real-data fine-tuning. Standard workaround: generate millions of synthetic examples with known ground truth, pre-train on those, then fine-tune with a small labelled real subset.
  • Self-supervised learning. Blind-spot networks (Noise2Void, StructN2V) train on a single noisy field dataset, and Noise2Noise on two independent noisy copies of the same signal, so neither needs clean ground truth. A self-supervised step does not depend on labels at all, which is why random-noise attenuation stays ML-led in Figure 9.1 when Training data is dragged to none.

5. The physics-ML spectrum in practice

A modern processing chain typically combines them:

  1. First-break picking, for refraction statics: an ML network as primary, QC by a processor.
  2. Deghosting, in marine data: a deterministic ghost operator, with ML where the sea is rough or the streamer depth uncertain.
  3. Pre-processing denoising: classical f-x deconvolution followed by a CNN denoiser for residual random noise.
  4. Trace reconstruction: sparse Radon for gap interpolation, with ML refinement for complex missing-data patterns.
  5. Velocity analysis: semblance, an ML auto-picker and human QC.
  6. Residual moveout picking: semblance or residual-curvature scans on migrated gathers, with an ML auto-picker for dense picks; with FWI and migration it forms the model-building loop.
  7. FWI model building: the adjoint-state gradient, with ML for starting models and low-frequency extrapolation.
  8. Migration: RTM or Kirchhoff PSDM (pure physics).
  9. Time-lapse: windowed least-squares matching, with ML for strongly non-stationary differences.
  10. Fault and horizon interpretation: ML as primary, human QC.

6. What ML cannot do in processing

  • Invent physics. ML extrapolates from training data; it cannot predict scattering for geology unlike anything in its training set.
  • Guarantee amplitude fidelity. Physics-based amplitude-preserving migration has well-understood amplitude behaviour under stated assumptions; ML denoisers can smooth amplitudes in subtle ways that break QI workflows.
  • Handle out-of-distribution data. A model trained on reasonable SNR data will fail on near-zero-SNR sections.
  • Provide uncertainty. Most production ML models output point predictions; Bayesian / ensemble variants provide uncertainty but at high cost.
  • Replace physical QC. An ML output always needs physics-based QC: for model-predicting tasks, forward-model the ML-predicted model through the wave equation and compare to the data; a mismatch that the modelling assumptions cannot explain means the ML output is wrong. For other steps, the physics checks are in the task cards of Figure 9.1.

7. When to invest in ML for your flow

  • A pattern-recognition task currently consuming significant interpreter time.
  • A step where the classical method gives only marginal results (denoising, fault picking).
  • Enough labelled training data or a realistic synthetic strategy.
  • Capacity to QC ML outputs with physics (so bad ML predictions are caught).

Conversely, do NOT invest in ML for tasks where the physics-based answer is precise and well validated (migration, deghosting), or where training data is genuinely impossible to generate.

The one sentence to remember

ML in seismic processing lives in the tasks where physics-based methods are heuristic (denoising, first-break picking, fault labelling) and stays out of the core of the tasks where physics is precise (migration, FWI, deghosting). The four sections that follow take denoising, interpolation, first-break picking and ML-assisted FWI in turn.

Where this goes next

Section 9.2 covers CNN-based denoising, the most widely deployed ML technique in seismic processing today. Its figure buries a synthetic section in random noise and swell bursts, denoises it with a classical method (a band-pass or f-x deconvolution) and with a small residual CNN trained on synthetics, then reads what each removed.

References

  • Yilmaz, Ö. (2001). Seismic Data Analysis (2 vols.). SEG.
  • Virieux, J., Operto, S. (2009). An overview of full-waveform inversion in exploration geophysics. Geophysics, 74, WCC1.
  • Claerbout, J. F. (1976). Fundamentals of Geophysical Data Processing. McGraw-Hill.
  • Ronneberger, O., Fischer, P., Brox, T. (2015). U-Net: convolutional networks for biomedical image segmentation. MICCAI 2015, 234-241.
  • Wu, X., Liang, L., Shi, Y., Fomel, S. (2019). FaultSeg3D: using synthetic data sets to train an end-to-end convolutional neural network for 3D seismic fault segmentation. Geophysics, 84, IM35.
  • Yu, S., Ma, J. (2021). Deep learning for geophysics: current and future trends. Reviews of Geophysics, 59, e2021RG000742.
  • Birnie, C., Ravasi, M., Liu, S., Alkhalifah, T. (2021). The potential of self-supervised networks for random noise suppression in seismic data. Artificial Intelligence in Geosciences, 2, 47-59.

This page is prerendered for SEO and accessibility. The interactive widgets above hydrate on JavaScript load.