Denoising with CNNs
Learning objectives
- Describe the U-Net architecture commonly used for CNN seismic denoising
- Contrast supervised (N2Clean) and self-supervised (N2N) training regimes
- Explain the amplitude-preservation trade-off CNN denoisers make
- Identify the failure modes that QC must catch
CNN-based denoising is the most widely deployed ML technique in seismic processing. A convolutional network ingests a noisy 2D patch (time × offset for a pre-stack gather, or time × trace for a stacked section) and outputs a denoised version. Trained on many patches with known ground truth, the network learns what seismic signal looks like and removes what does not fit. For random noise, and increasingly for coherent noise, published field tests often report CNN denoisers gaining a few dB of SNR over f-x deconvolution; the margin depends on the noise type and on how closely the field data resemble the training data.
1. U-Net: the workhorse architecture
Nearly all production seismic denoisers are based on U-Net: an encoder-decoder CNN with skip connections between corresponding encoder/decoder layers. The architecture:
- Encoder: 3-5 blocks, each with (conv → ReLU → conv → ReLU → downsample). Captures progressively coarser features.
- Bottleneck: 1-2 conv blocks at the coarsest resolution. Represents global context.
- Decoder: mirror of encoder, with upsample and concat from the corresponding encoder layer (skip connection). Reconstructs fine detail.
- Final layer: 1×1 conv mapping to the output channels (usually a single channel for grayscale seismic patches).
Typical parameters: 5-10 million. Patch size: 64×64 or 128×128 samples. Loss: L1 or L2 between prediction and clean ground truth.
2. The widget
Bury a synthetic section in noise, then denoise it two ways and compare: a classical method you choose, and a small network trained on synthetic data. Read the gains, then read the difference section to see what each method threw away.
The figure builds a 64-trace synthetic patch with a bright spot and a fault, adds in-band random noise and swell bursts, and denoises it with a classical method (a band-pass or f-x deconvolution) and with a small residual CNN trained on synthetic patches from the same generator. At the opening setting, 0 dB of random noise plus swell, the CNN gains +13.2 dB against +4.9 dB for f-x decon, and it keeps 98 % of the bright spot's amplitude where f-x decon keeps 64 %. Plate (e) shows what a method removed: any reflector you can see there is signal the denoiser threw away.
A band-pass removes the swell below the signal band, but the bursts reach 12 Hz, into the band of the reflections, so no low cut separates them cleanly: on swell alone the band-pass does best near a 14 Hz low cut, at +9.0 dB, while the CNN gains +15.8 dB. Random noise inside the band cannot be filtered by frequency at all. f-x deconvolution predicts each frequency across traces instead, so it smears curved and faulted events and dims a bright spot only 14 traces wide: at −10 dB it keeps 39 % of the bright spot, where the CNN keeps 89 %. The CNN removes both kinds of noise because it has learned what a reflection looks like, and the same learning is its limit: weights trained on random noise alone gain about 10 dB less on swell they never saw.
The network in the widget is deliberately tiny, 4401 parameters trained offline on 3000 random synthetic earths, but it is a real network: it never sees the clean section, and its mistakes are its own. Production networks are a thousand times larger and see far more context, which is why their failures, described below, are subtler.
3. Training regimes
- N2Clean (supervised). Train on pairs (noisy, clean). Requires clean ground truth, which is rare in real seismic. Usually clean comes from high-fold stacks, carefully processed "hero" datasets, or synthetic models. Gold standard when available.
- N2N (Noise2Noise, self-supervised). Train on pairs , two independent noisy realisations of the same underlying signal. The network cannot learn to reproduce noise (it differs between A and B) so it converges on the shared signal. Works when two realisations of the same signal exist, for example repeated shots or sweeps at one location, or two sub-stacks built from the even and odd traces of the same CMP fold. Lehtinen et al. (2018) introduced Noise2Noise for images; blind-spot self-supervised variants were adapted to seismic by Birnie et al. (2021).
- N2Self / N2Void (blind-spot). Train on noisy data alone with a blind-spot masking strategy: the network sees a masked pixel's neighbours but not the pixel itself, forcing it to predict from context. Useful when no paired data exists.
- GAN-based. An adversarial discriminator pushes the denoiser output toward the distribution of real clean seismic. Higher perceptual quality but sometimes invents plausible-looking but wrong details.
4. Amplitude fidelity: the QI caveat
CNN denoisers model a typical seismic signal distribution. An AVO anomaly (high amplitude at a reservoir) is by definition atypical. Over-aggressive CNN denoising can attenuate AVO anomalies as "outliers" from the signal manifold. Several mitigations:
- Residual-domain denoising. Apply the CNN to (noisy − classically-denoised); add the result back. Residual has limited dynamic range so amplitude regression is bounded.
- Amplitude constraint in the loss. Penalise absolute amplitude changes of reflectors during training, not just pixel-wise error.
- Post-hoc amplitude calibration. Scale CNN output by a per-trace RMS-matching factor to restore absolute amplitudes.
- Section-by-section QC. Run the constant-R test (Section 7.2) on the denoised data. If it is no longer flat, the CNN is distorting amplitudes.
5. Common failure modes
- Over-smoothing. The CNN is too aggressive and removes legitimate signal it considers anomalous. Symptoms: output looks beautiful but shallow reservoirs (bright spots) are flattened. Fix: reduce denoising strength or reduce network depth.
- Hallucination. The network synthesises plausible-looking but false signal where only noise existed. Symptoms: output shows pristine reflectors in zones where the data had no actual signal. Fix: check uncertainty metrics (Monte Carlo dropout) or switch to less aggressive architecture.
- Domain shift. Model trained on North Sea data applied to Gulf of Mexico data. Symptoms: inconsistent performance, poor QC on the new survey. Fix: fine-tune with a subset of target-survey data.
- Artifacts at patch boundaries. The network processes 64×64 patches; small inconsistencies at patch boundaries become visible as a grid pattern. Fix: overlap patches and blend in the overlap zone.
6. Typical deployment
- Training. Build a synthetic dataset (millions of synthetic traces + noise models); train U-Net for ~24-48 GPU-hours.
- Fine-tuning. Take a small labelled subset from the target survey; fine-tune the pre-trained network (~2-4 GPU-hours).
- Application. Slide the trained network over the field data in overlapping patches; blend outputs.
- QC. Amplitude check, spectral check, visual inspection of difference volume (noisy − denoised). Any remaining coherent signal in the difference is a failure mode.
- Calibration. Match output RMS to input at selected reference horizons; this preserves absolute amplitude for downstream QI.
CNN denoising (typically a U-Net) learns what seismic signal looks like, from noisy-clean pairs or from the noisy data alone, and can beat f-x deconvolution where noise shares the signal's band; the price is careful QC of the difference section and of amplitudes, to catch signal leakage, dimmed anomalies and hallucinated events.
Where this goes next
Section 9.3 covers ML-based trace reconstruction, filling in missing traces or reconstructing badly-sampled gathers using similar U-Net architectures and training strategies.
References
- Yilmaz, Ö. (2001). Seismic Data Analysis (2 vols.). SEG.
- Oppenheim, A. V., Schafer, R. W. (2009). Discrete-Time Signal Processing (3rd ed.). Prentice Hall.
- Claerbout, J. F. (1976). Fundamentals of Geophysical Data Processing. McGraw-Hill.
- Ronneberger, O., Fischer, P., Brox, T. (2015). U-Net: convolutional networks for biomedical image segmentation. MICCAI.
- Zhang, K., Zuo, W., Chen, Y., Meng, D., Zhang, L. (2017). Beyond a Gaussian denoiser: residual learning of deep CNN for image denoising. IEEE Transactions on Image Processing 26(7).
- Lehtinen, J. et al. (2018). Noise2Noise: learning image restoration without clean data. ICML.
- Krull, A., Buchholz, T.-O., Jug, F. (2019). Noise2Void: learning denoising from single noisy images. CVPR.
- Canales, L. L. (1984). Random noise reduction. SEG Expanded Abstracts.
- Yu, S., Ma, J., Wang, W. (2019). Deep learning for denoising. Geophysics 84(6).
- Birnie, C., Ravasi, M., Liu, S., Alkhalifah, T. (2021). The potential of self-supervised networks for random noise suppression in seismic data. Artificial Intelligence in Geosciences 2.