ML for FWI gradient acceleration
Learning objectives
- Identify the three points where ML can accelerate FWI: initialisation, regularisation, surrogate forward
- Explain how each kind of ML acceleration changes convergence, and what each risks: a learned start can still be cycle-skipped, a learned prior can bias the model, a surrogate is approximate away from its training data
- Measure the speed-up in wave solves to a target misfit, count the training cost, and find the settings where the speed-up vanishes
- Recognise the production caveats: training-distribution match, transferability, QC
Part 9 closes with the current ML frontier in production seismic processing: accelerating FWI. Unlike the previous four sections where ML was the main algorithm, FWI acceleration uses ML alongside physics. The forward and adjoint wave-equation simulations remain the ground truth; ML reduces the number of times you have to run them or the number of iterations they must do.
1. Three entry points for ML
- Learned initial models. A CNN maps tomographic velocity models (or legacy seismic volumes) to FWI-ready starting models. Training pairs: (tomography output, converged FWI output) from historical projects. The CNN "jumps" the model part-way toward the converged state. Its lower starting misfit saves some iterations; its larger value is a start whose arrivals lie within half a period of the data at the lowest usable frequency, where the inversion converges at all.
- Learned regularisers. Instead of a hand-tuned TV or smoothness penalty, train a network to recognise "geologically plausible" velocity models and apply it as a soft constraint in the FWI loss. It can make each iteration more productive, but it changes the objective: the minimiser moves toward models like the training set, and the data fit it allows is looser.
- Surrogate forward modelling. Train a network to mimic the forward simulator. Once trained, inference can be one to two orders of magnitude cheaper than a finite-difference simulation (the widget spans 10 to 100×), but accuracy degrades for model perturbations far from the training distribution. Used for rapid Monte Carlo uncertainty estimation; not (yet) a full substitute for real FWI.
2. The widget: four routes to a target misfit
Figure 9.5 inverts one noisy zero-offset trace for the ten velocities of a layered earth, over a multiscale schedule of four bands from the lowest usable frequency up to 16 Hz, four ways: physics only from a tomography start (the smoothed truth, 6 % slow); from a learned start, a nearest-neighbour regressor over 300 simulated training earths that stands in for a network; from the learned start with a learned prior that pulls toward it for the first two bands; and with a surrogate solver, 30× cheaper but with a 3 % smooth velocity error, for the two lowest bands before the exact physics takes over. Every route finishes on the exact physics and its adjoint gradient. The score is compute, counted in wave solves (every forward model, including each line-search trial, and every adjoint), to bring the data misfit within 1.3 times a noise floor that the widget estimates from the quiet window before the first arrival. Plate (d) draws the misfit against that compute, plate (e) the model error, which only the figure knows.
At the default 3 Hz the learned start is 23 m/s from the truth (RMS) against 159 m/s for tomography, and it reaches the target in 42 solves against 83 for physics only: 2.0× less compute. The 300 training simulations are not on that axis; paid once, they pay back after 8 surveys like this one. The prior route needs 44 solves and the surrogate route 41. Now take the low frequencies away: at 6 Hz the deepest arrival of the tomography start is 0.099 s off, more than half a period (0.083 s), and physics only stalls 289 m/s from the truth while the learned start converges in 28 solves. That is the real value of a learned start: not a few iterations, but converging at all. Half a period is a rule of thumb, not a law, and the widget says so when it fails: at 1.5 Hz from a start 12 % slow, inside half a period, the lowest band itself pushes the deep arrivals out and physics only stalls; at 8 Hz from a start 9 % slow, beyond it, physics only converges, because its shallow arrivals start inside and fixing the shallow layers pulls the deep ones in.
The same controls show the costs. With a fast layer the training prior never produces, the learned start is 333 m/s wrong, worse than tomography at 324 m/s, and needs 104 solves against 96 for physics only, so it never pays back. (It matched the arrival times rather than the layer, so at 8 Hz it still converges where tomography stalls.) On that earth, leave a strong prior (weight 1) on to the end and the route stops at 28 times the noise floor, 63 m/s from the truth; on the earth like the training set the same prior does no harm. Raise the surrogate error from 3 to 10 % and the surrogate hands the physics bands a model outside their basin: it never reaches the target. The numbers come from a ten-unknown toy, so read them as comparisons, not production figures.
3. What ML cannot do for FWI
- Replace adjoint gradients. The physics-based adjoint gives the exact gradient of the misfit with respect to the model. An ML surrogate gradient is approximate; using it as the only gradient produces biased FWI. Production workflows use real adjoints for the final iterations even if ML accelerates early iterations.
- Invent geology. A CNN trained on North Sea datasets cannot produce a correct starting model for a Gulf of Mexico sub-salt project. Domain shift is a hard limit.
- Remove cycle-skipping risk. A bad ML initial model can still be cycle-skipped relative to the data. Multi-scale FWI (Section 6.2) is still required.
- Guarantee amplitude fidelity. An ML regulariser may bias amplitudes toward training-distribution statistics; QI-grade FWI still needs post-hoc amplitude calibration.
4. Production deployment patterns
- Initialiser only: ML generates the starting model; pure physics FWI from then on. Safe and widely deployed.
- Initialiser + learned prior: ML starting model plus a learned prior (in place of a hand-tuned TV penalty) for the first half of the iterations, then physics alone for the final iterations, so the prior cannot bias the answer. More aggressive; common in research projects.
- Full ML FWI: surrogate forward + learned gradient + learned prior. Highly efficient but requires strong QC and limited to training-similar projects. Emerging in time-lapse monitoring where baseline FWI has already been done carefully.
5. Quantitative expectations
Published speed-ups vary widely and depend on what is counted: iterations, wave solves or GPU-hours, and whether the training simulations are included. Read any single factor with those questions in mind. What the widget measures, in wave solves to the target and with training counted separately, over every setting its controls reach:
- A learned start inside the training distribution, where tomography also converges: usually between a small loss and about 2.3× less compute, 1.3× at the median and 2.0× at the default, with extremes from about 4× slower (8 Hz, tomography nearly exact, noisy data) to 10× less (8 Hz, tomography only just converging from beyond half a period). Where tomography is cycle-skipped, about half the settings, the difference between converging and stalling.
- A learned prior switched off for the last bands: about the same gain as the learned start alone; left on to the end, it can stop the inversion well above the noise floor when the earth is unlike its training set.
- A surrogate for the low bands: its solves cost almost nothing, so the saving is whatever the low bands used to cost, as long as its error is small (2.0× at 3 % error, 2.5× at 1 % and 100× speed); at 10 % error the route fails. The toy surrogate needs no training; a real one is trained on thousands of simulations, a cost these numbers leave out.
- Outside the training distribution, each can cost more than it saves.
6. QC for ML-accelerated FWI
- Compare final model to pure-physics FWI. On a subset of projects, run both pure and ML-accelerated FWI to verify they converge to the same model. If they disagree, diagnose whether the ML is biasing the solution.
- Well ties on final model. Compare ML-FWI output to wells; require the same tie quality as pure FWI.
- Synthetic-vs-recorded forward tests (Section 6.5). The pure-physics criterion: if the forward-modelled data from the final model matches observed data, the ML acceleration did no harm.
- Cross-survey generalisation test. Re-run the ML-FWI on a survey not seen in training; verify no systematic bias appears.
ML accelerates FWI by providing learned starting models, learned regularisers, and (experimentally) surrogate forward operators, saving compute when the learned parts are accurate and inside their training distribution, while the physics-based adjoint keeps supplying the exact gradient of the data misfit.
Part 9 closes here
You have the ML-in-seismic-processing toolkit: where ML fits (Section 9.1), CNN denoising (Section 9.2), trace interpolation (Section 9.3), first-break picking (Section 9.4), and FWI acceleration (Section 9.5). The theme across all five: ML is most valuable as a complement to physics, not a replacement: classical methods provide the exact physics and the checks, and ML provides speed and accuracy at pattern-recognition tasks. Part 10 brings all nine previous parts together in a series of end-to-end processing capstones that walk through complete projects from raw data to final deliverable.
References
- Virieux, J., Operto, S. (2009). An overview of full-waveform inversion in exploration geophysics. Geophysics, 74, WCC1.
- Pratt, R. G. (1999). Seismic waveform inversion in the frequency domain, Part 1. Geophysics, 64, 888.
- Tarantola, A. (1984). Inversion of seismic reflection data in the acoustic approximation. Geophysics, 49, 1259.
- Etgen, J., Gray, S. H., Zhang, Y. (2009). An overview of depth imaging in exploration geophysics. Geophysics, 74, WCA5.
- Bunks, C., Saleck, F. M., Zaleski, S., Chavent, G. (1995). Multiscale seismic waveform inversion. Geophysics, 60, 1457.
- Lewis, W., Vigh, D. (2017). Deep learning prior models from seismic images for full-waveform inversion. SEG Technical Program Expanded Abstracts, 1512.
- Mosser, L., Dubrule, O., Blunt, M. J. (2020). Stochastic seismic waveform inversion using generative adversarial networks as a geological prior. Mathematical Geosciences, 52, 53.
- Moseley, B., Nissen-Meyer, T., Markham, A. (2020). Deep learning for fast simulation of seismic waves in complex media. Solid Earth, 11, 1527.