ML-assisted first-break picking

Part 9, Machine Learning in Processing

Learning objectives

  • State the first-break picking problem and why it matters for refraction statics
  • Compare STA/LTA, AIC, and CNN pickers and their failure modes
  • Describe the U-Net training regime for first-break picking
  • Recognise why ML picking has become the production standard

The first break, the onset of the direct or refracted arrival on a shot record, is the single most-picked quantity in seismic processing. Every shot on every survey gets first-break picks used for: refraction static corrections (land), direct-wave velocity QC, near-surface tomography, overburden velocity building. Classical pickers were heuristic and required constant human oversight. Modern CNN-based pickers have taken over this role in production.

1. The three classical pickers

  • STA/LTA (short-term / long-term average). Compute the ratio of a short-window (10-40 ms) energy to a long-window (50-200 ms) energy as a function of time; trigger when the ratio crosses a threshold (typically 2-5). Simple and fast, but sensitive to noise before the arrival and to the threshold, and because the short window trails the sample being tested it picks a few milliseconds late even on clean data.
  • AIC (Akaike Information Criterion). Model the trace as two segments split at index kk: pre-arrival (white noise) + post-arrival (white noise with different variance). Compute AIC as a function of kk; the minimum marks the split. The two-segment model only holds near the arrival, so AIC is run on a window around a preliminary STA/LTA trigger (Maeda 1985). There it reads the onset to about one sample on clean data, but it assumes noise is white and stationary, and it inherits any false trigger that placed its window.
  • Correlation. Cross-correlate each trace with a template (an assumed wavelet). The correlation peak marks the first break. Good when the wavelet is known; poor when it varies shot-to-shot.

2. Why ML wins at this task

First-break picking is a pattern-recognition problem with a tight labelled training set available (any historical survey has hand-picks). A CNN sees:

  • The full waveform context of each trace (not just a moving window).
  • Adjacent traces' arrivals (moveout consistency).
  • Wavelet variations learned from training examples.
  • Noise patterns learned across many surveys.

A trained U-Net picks a trace in about a millisecond on a GPU. A careful human, even with snapping tools, spends on the order of a second per trace. Accuracy is typically 2-5 ms RMS against hand picks, which are themselves uncertain at this level.

3. The widget

In the figure below, three pickers read the same 24 traces of one shot over a slow weathering layer. Lower the signal-to-noise ratio and watch which picks leave the true onset, pick a trace to see why each picker fired where it did, then add receiver statics and replace the neighbour check with one global line.

First-break pickingFIRST-BREAK PICKINGundefinedAuto-pick the first energy on each trace - picks form the direct-wave moveout

The receivers sit every 40 m from 50 to 970 m. The direct wave, t=x/V_1t = x/V\_1, arrives first on the two nearest traces; beyond a crossover near 120 m the head wave, t=t_i+x/V_2t = t\_i + x/V\_2, overtakes it. The weathering changes thickness along the line, so every trace carries its own static delay of a few milliseconds. The wavelet is causal, so its first non-zero sample is the first break, and the arrival weakens with offset, so the far traces are lost to noise first. STA/LTA and AIC each look at one trace at a time. The third picker, a moveout-consistent picker, checks each AIC pick against a line through its neighbours and, when they disagree, re-picks that trace on its own waveform near where the neighbours point. It stands in for what a trained network learns from labelled gathers; it is not a neural network.

From 30 dB down to 16 dB on the far trace, every pick lands within a quarter period (7.1 ms at 35 Hz) of the onset; STA/LTA runs about 3 ms late on the median, the price of its trailing window, while AIC reads the onset to under a millisecond. At the figure's opening 12 dB, STA/LTA is off on 3 of the 24 traces, one trace at a time and each time early, because a patch of noise before the arrival crossed the threshold; AIC, which searches around those triggers, is off on 4, and the moveout-consistent picker on 1. At 6 dB the moveout-consistent picker still rescues 2 to 5 AIC picks on each of the four noise draws, so it is off on 4 to 9 traces where AIC is off on 7 to 11. It drags good picks toward bad ones only when most of the neighbours fail together: with an eager threshold of 1.5 at 6 dB it spoils 2 or 3 good picks and rescues none, and by 0 dB all three pickers are off on 11 to 18 of the 24 traces. Now add statics of 15 ms and switch to one global line: every single-trace pick is still right at 20 dB, but the line is off on 10 traces, because it cannot bend to the direct wave or follow the static delays that the picks exist to measure. A useful context picker lets the neighbours say where to look and the trace's own waveform say exactly when.

4. U-Net for first-break picking

Input: a 2D patch of the shot gather (traces × time). Output: a segmentation mask, 00 before the first break and 11 after it, trained with a cross-entropy loss. At inference the pick on each trace is the first time sample where the predicted probability exceeds 0.50.5. (A common alternative labels a narrow Gaussian centred on the pick and takes its per-trace maximum.) Training data: tens of thousands of shots with hand-picks, usually from multiple surveys to capture wavelet variety.

Hybrid variants use the mask output as an initial pick, then refine with AIC over a narrow window around it; this combines ML robustness with AIC's sample-level precision (sub-sample if the AIC curve is interpolated around its minimum).

5. Production deployment considerations

  • Cross-survey generalisation. A network trained on one survey (one wavelet, one noise regime) may systematically miss arrivals on another. Standard practice: pre-train on a diverse set of surveys, fine-tune per project.
  • Human QC. Every production workflow retains interpreter QC on a subset of picks. The CNN replaces the bulk labour but not the final review.
  • Uncertainty flagging. CNN outputs a confidence map; low-confidence picks are flagged for manual review. Saves time by focusing QC on ambiguous cases.
  • Outlier rejection. Post-CNN, apply a moveout-consistency filter: picks that deviate by more than some threshold from a smooth fit are flagged or rejected.

6. Failure modes

  • Refraction static errors. If the first breaks are systematically mis-picked by a few ms, near-surface velocity models become biased and refraction statics are wrong; the downstream imaging then shows residual-statics artefacts.
  • Cycle-skipping. At very low SNR the CNN may pick the wrong cycle (one period late or early). Harder to catch than STA/LTA cycle-skipping because CNN outputs look confident. Mitigation: compare predicted picks to a smoothed moveout model, flag outliers.
  • Domain shift under novel noise. Electrical hum, drone noise and unexpected infrastructure are examples: the network has not seen these in training and may produce plausible but wrong picks. Mitigation: add diverse noise to training; include "synthetic noise augmentation" during training.
The one sentence to remember

CNN first-break picking (typically U-Net on 2D gather patches) has replaced STA/LTA and AIC in production because it stays accurate where single-trace pickers fail, on weak, noisy far offsets, and so removes most of the hand QC, with the main caveat that cross-survey domain shift requires fine-tuning per project.

Where this goes next

Section 9.5 closes Part 9 with the most recent ML frontier in seismic processing: accelerating FWI. Instead of replacing the physics, ML initialises FWI with a learned prior, replaces expensive iterations with network inference, or provides a learned regulariser that keeps the solution geologically plausible.

References

  • Yilmaz, Ö. (2001). Seismic Data Analysis (2 vols.). SEG.
  • Sheriff, R. E., Geldart, L. P. (1995). Exploration Seismology (2nd ed.). Cambridge UP.
  • Claerbout, J. F. (1976). Fundamentals of Geophysical Data Processing. McGraw-Hill.
  • Allen, R. V. (1978). Automatic earthquake recognition and timing from single traces. Bulletin of the Seismological Society of America, 68(5).
  • Maeda, N. (1985). A method for reading and checking phase times in autoprocessing system of seismic wave data. Zisin, 38.
  • Sleeman, R., van Eck, T. (1999). Robust automatic P-phase picking: an on-line implementation in the analysis of broadband seismogram recordings. Physics of the Earth and Planetary Interiors, 113.
  • Yuan, S., Liu, J., Wang, S., Wang, T., Shi, P. (2018). Seismic waveform classification and first-break picking using convolution neural networks. IEEE Geoscience and Remote Sensing Letters, 15(2), 272-276.

This page is prerendered for SEO and accessibility. The interactive widgets above hydrate on JavaScript load.