Random variables & noise

Processing Prerequisites

Learning objectives

  • Define a random variable, expectation, variance, and Gaussian distribution
  • Distinguish white from colored noise by spectrum and autocorrelation
  • Compute SNR in amplitude and dB and explain when a signal is “invisible” in noise
  • Connect noise models to denoising strategy: different statistics → different attenuators

Every seismic record is signal plus noise. Every denoising, stacking, deconvolution, and inversion algorithm assumes a model of what the noise looks like. If the model is wrong, the algorithm is wrong. So we spend some time making the assumptions explicit.

1. Random variables in one paragraph

A random variable X is a rule for generating numbers with a probability distribution. Two summary numbers describe it:

\\mathbb{E}\[X\] = \\text{mean},\\qquad \\operatorname{Var}(X) = \\mathbb{E}\[(X - \\mathbb{E}\[X\])^2\].

The square root of variance is the standard deviation sigma\\sigma; for noise with zero mean it is the rms, the typical amplitude.

2. Why Gaussian shows up everywhere

The central limit theorem says: sum many small independent contributions, and the total will look Gaussian no matter how the individuals were distributed. Ambient random noise is the sum of many tiny independent effects (thermal noise, ground motion, electronics, distant activity), so it is close to Gaussian. Spikes, bursts and bad traces are not: they give real records heavy tails, which is why robust attenuators such as median stacks and L1 fits exist.

p(x);=;frac1sigmasqrt2pi,exp!left(−frac(x−mu)22sigma2right)p(x) \\;=\\; \\frac{1}{\\sigma\\sqrt{2\\pi}}\\,\\exp\\!\\left(-\\frac{(x-\\mu)^{2}}{2\\sigma^{2}}\\right)

In Figure 0.8 below, four reflections sit on a trace buried in noise. You will stack traces until they appear, measure the noise two ways, and switch its colour between white, pink, brown and blue while you watch the histogram in (d). All four colours follow the same Gaussian bell: they are shaped from the same Gaussian samples and differ only in how neighbouring samples are correlated. Raise the sample count and the histogram settles onto the curve; brown noise, whose neighbouring samples are nearly equal, needs far more samples than white noise before it does.

Signal + noise vs clean signalSIGNAL + NOISECLEAN SIGNAL ONLYNoise attenuation recovers the buried signal - denoising is the central task

3. White, pink, brown, blue: the colour of noise

  • White noise has a flat power spectrum and a delta autocorrelation: each sample is independent of every other. This is the textbook assumption behind most deconvolution operators.
  • Pink noise has power propto1/f\\propto 1/f: more energy at low frequencies, and the same energy in every octave. Its power falls 10 dB per decade of frequency, which plate (b) measures. It appears in micro-tremor and in many electronic systems.
  • Brown (red) noise has power propto1/f2\\propto 1/f^{2}, falling 20 dB per decade. Integrating white noise gives it: a random walk, the slow drift of an instrument wandering off zero. Its neighbouring samples are nearly equal, so its autocorrelation decays slowly; in (c) it is still about 0.7 at 60 ms. Pink noise sits halfway between white and brown, and no single integrator makes it.
  • Blue noise has power proptof\\propto f, rising 10 dB per decade, so it has more energy at high frequencies. Its neighbouring samples tend to alternate in sign, so its autocorrelation dips below zero at the first lag, to about −0.4-0.4 in (c). Differencing white noise overshoots to violet noise, power proptof2\\propto f^{2}. Noise boosted at high frequencies like this appears after deconvolution with too little prewhitening.

4. Signal-to-noise ratio (SNR)

Define

\\operatorname{SNR} = \\frac{\\sum\_n s\[n\]^{2}}{\\sum\_n e\[n\]^{2}}\\qquad\\text{and}\\qquad\\operatorname{SNR}\_{\\text{dB}} = 10\\log\_{10}\\operatorname{SNR},

where s\[n\] is the signal, e\[n\] the noise, and both sums run over the same time window. The window matters. In Figure 0.8, with one trace set to 6 dB of peak SNR, the energy ratio is about −12 dB over the whole 1 s trace and about −4 dB over 100 ms around the strongest reflection, yet the picture is the same. What the eye responds to is the peak SNR, 20log_10(A_textpeak/sigma)20\\log\_{10}(A\_{\\text{peak}}/\\sigma): how far a reflection’s peak stands above the noise’s rms.

A rough guide for a single 1 s trace, in peak SNR (somewhere on such a trace the noise alone reaches 3 to 4 times its rms):

  • Above about 20 dB (a peak 10 times the rms): every reflection stands out by eye.
  • About 10 to 20 dB: the strong reflections clear the noise’s largest excursions; the weak ones may not.
  • Below about 10 dB: a reflection is no taller than the noise’s own peaks, and you cannot pick it on one trace. Only stacking, averaging many independent realisations, recovers it. This is the territory of weak reservoir anomalies and deep plays.

Try exercise 2 of the figure. With the noise-free reflections hidden, lower the peak SNR from 20 dB: at 12 dB the strongest reflection, 3.8 times the rms, still clears the noise’s largest excursion of 3.5 times, and a few dB lower it is gone. Then set 0 dB and raise the fold. Sixteen traces buy back 10log_1016approx1210\\log\_{10}16 \\approx 12 dB, and the strongest reflection returns at 4.1 times the rms of the stacked noise.

5. Stacking improves SNR by sqrtN\\sqrt{N}

Add NN traces that carry the same signal and independent noise. The signal adds coherently, so its energy scales as N2N^{2}. The noise adds incoherently, so its energy scales as NN. The energy SNR therefore grows as NN, which is +10log_10N+10\\log\_{10}N dB, and the amplitude SNR grows as sqrtN\\sqrt{N}. Figure 0.8 measures it: at 16-fold the stack gains 12.3 dB against the law’s 12.0 dB, and at 64-fold 18.1 dB against 18.1 dB.

This is why seismic is acquired with lots of fold. Sixty-fold data has sqrt60approx7.7\\sqrt{60} \\approx 7.7 times the amplitude SNR of single-fold data, +17.8 dB. That holds under two assumptions: the noise is uncorrelated between traces, and the signal is perfectly aligned on all of them after NMO and statics. Add 4 ms of residual moveout jitter at 64-fold in the figure and 3.5 dB of the gain is gone; at 8 ms, 9.5 dB. Fold is not free either: every trace in the stack was paid for in the field.

6. Autocorrelation, or: noise has a fingerprint

The autocorrelation of a signal is its convolution with its own time-reverse:

R\_{xx}\[\\tau\] = \\sum\_n x\[n\]\\, x\[n+\\tau\]

White noise has \\mathbb{E}\\{x\[n\]\\,x\[n+\\tau\]\\} = \\sigma^{2}\\delta\[\\tau\]: a spike at zero lag and, on average, nothing elsewhere (the sum above estimates N\\sigma^{2}\\delta\[\\tau\], plus fluctuations). Coloured noise has a broadened peak, as plate (c) shows for brown noise, or a negative side lobe, as it shows for blue. Coherent noise (like swell on marine data) has periodic peaks at the dominant period. Autocorrelation is the first thing you compute when characterising unknown noise.

It is also the heart of predictive deconvolution: if the noise (or the multiple train) has structure in its autocorrelation, you can predict and subtract it.

7. Spatial noise is different from temporal noise

Everything above is about noise within one trace. In 2D and 3D, we also care about noise correlation between traces:

  • Ground roll moves at a fixed low velocity across traces, coherent spatially and temporally. Attenuation happens in the f-k domain.
  • Swell noise dominates low frequencies on streamers, coherent in time, uncorrelated between streamers.
  • Random ambient is what the textbook “noise model” usually refers to, uncorrelated trace-to-trace, stacks out beautifully.

Choosing the right noise attenuator depends on knowing what kind of noise you have.

The one sentence to remember

Every denoising algorithm has a noise model in it; stacking improves amplitude SNR by N\sqrt{N} when the noise is uncorrelated between traces and the signal is aligned; autocorrelation is the diagnostic that tells you which noise model applies.

Where this goes next

Section 0.9 takes the final mathematical piece: optimization. The normal equations of Section 0.7 solve a quadratic problem exactly, but FWI and most modern inversions minimize non-linear cost functions where you can only iterate toward the answer. Gradients, step sizes, and regularization all live in the next section.

References

  • Oppenheim, A. V., Schafer, R. W. (2009). Discrete-Time Signal Processing (3rd ed.). Prentice Hall.
  • Bracewell, R. N. (1999). The Fourier Transform and Its Applications (3rd ed.). McGraw-Hill.
  • Claerbout, J. F. (1976). Fundamentals of Geophysical Data Processing. McGraw-Hill.
  • Yilmaz, Ö. (2001). Seismic Data Analysis (2 vols.). SEG.

This page is prerendered for SEO and accessibility. The interactive widgets above hydrate on JavaScript load.