Skip to main content
Reading viewAll insights →
BLOG6 min read

Closing the sim-to-real gap on synthetic seismograms

Closing the sim-to-real gap on synthetic seismograms
Tannistha Maitiby Tannistha MaitiCo-founder & Head of Engineering · 5 Oct 2026
Share

A network trained on synthetic waveforms can score near-perfectly on a synthetic test set and then read a real station as noise. The failure is not a capacity shortfall, it is a distribution mismatch, and the fix is to measure the field's own noise floor, timing spread and gain variation and reproduce them channel by channel. Realism has an optimum rather than a direction, so overshooting costs as much as undershooting, and the only instrument that can tell you which side you are on is a held-out set of real labelled records.

A model can score near-perfectly on a held-out set of synthetic waveforms and then read a real station as noise. Nothing in the network is broken. It learned the distribution it was shown, and the distribution it was shown was too clean to contain the field.

The waveform panels in Tannistha Maiti's 2018 thesis are a fair picture of what too clean means. Fig. 5.3 plots radial waveforms against backazimuth, computed from a specified subsurface model by forward physics, with the answer already attached because the model defined the structure. Every arrival sits where the physics put it. Nothing wanders, nothing drops out, no instrument drifts out of calibration halfway through a season. That is what makes such waveforms good training material, and it is the same property that ends the model's usefulness on the day it is deployed. Why synthetic waveforms exist at all is a supply argument and belongs in its own piece. This one is about what happens after you accept the supply.

The gap is a mismatch, not a shortfall

Field seismograms carry background noise, instrument response quirks, timing jitter, variable gain, and structure the synthetic model never contained. The distribution of any feature the network reads, amplitude, arrival time, spectral shape, is therefore both broader and shifted relative to the synthetic one. Where the two barely overlap, the trained model is operating outside its own support the moment it meets field data. It is not generalising. It is extrapolating, and extrapolation is where a network stays confident and stops being right.

That framing matters because it rules out the fix everyone reaches for first. The instinct on hearing that a model fails on real data is to add capacity: a deeper network, a longer schedule, a wider sweep. None of it helps. Capacity buys a closer fit to the distribution you trained on. A network that has never seen a noisy trace does not learn to read noise by getting bigger. It learns to fit clean traces more precisely, which is the opposite of what the field is asking for.

Realism is a target you measure, not a dial you turn up

The fix is to move the training distribution toward the deployment distribution, and the honest way to do that is one channel at a time. Sample background noise from real recordings made at the stations you intend to run on, rather than adding noise at whatever level feels about right. Jitter the arrival times by the spread your own picks show. Vary gain and instrument response across the range your instruments actually span. Perturb the subsurface models so the training set covers the structures you expect to meet, not the one you happened to simulate first.

Every one of those is a measurement task before it is an augmentation task. The field supplies the number and the generator reproduces it. Teams that skip the measuring step end up tuning a single global realism slider by eye, and one slider cannot be correct on more than one channel at a time.

GENERATOR REALISM, CHANNEL BY CHANNEL100.0% and 4.3 kmPOOLED OVERLAP AND REAL HELD-OUT ERROROne field record, three generated tracesthe top trace never moves: it is the deployment distribution you have to matchdirect PPs conversionfield recordmeasuredgenerated 1generated 2generated 30481216time from the direct P arrival, secondsFour channels, each against its measured field valueeach bar is the generator as a multiple of the measured field valuebackground noise floorovershot 94%field 0.080 RMS · generator 0.135 RMSquiet windows at your stationsarrival timing spreadundershot 70%field 0.19 s · generator 0.05 syour own repeat picksgain and response variationnot applied 0%field 0.080 rel · generator 0.000 relthe instruments you will run onsubsurface model spreadmatched 100%field 3.4 km · generator 3.5 kmacross your target regionthe measured field value2.5xthe summary says: pooled overlap 100.0% · synthetic test error 0.44 kmthe scoreboard says: 1 of 4 channels matched · real held-out error 4.3 kmthe pooled number is healthy because one channel's excess is paying for another channel's absencethe fieldthe generatorpast the field valuematched means overlap 0.995 or better, the field value within 15 percentField statistics are representative values, not one network's log. The algebra is not: pooled variance is a sum of squares, so one channel's excess offsets another's absence.
how realism is set:
One real field record on top, three traces from the synthetic generator underneath, and on the right the four augmentation channels with the value measured in the field marked on every track. The bench opens where the prose leaves you, on a single global realism dial parked at the setting that maximises the summary score. Read the two hero numbers against each other: pooled overlap 100.0 percent, real held-out error 4.3 kilometres. The summary is perfect and the model is still wrong, because the noise channel is overshot by enough to pay for a gain channel that was never wired, and a pooled variance is a sum of squares that cannot tell the difference. Sweep the dial anywhere you like and no position matches more than one of the four. Then switch to per channel, press set every channel to the field value, and watch the held-out error drop to its floor while the pooled score does not move at all. From there push the noise channel to the top of its range: the Ps conversion vanishes under the noise while the guide line stays where the arrival should be, and the error climbs by as much as pulling that channel down by the same factor would have cost. Realism has an optimum rather than a direction. The field statistics are representative values rather than one network's log; the algebra that lets one channel's excess hide another's absence is not.

Overshooting costs as much as undershooting

Realism has an optimum rather than a direction. Push injected noise past the field level and the arrival drowns: the network learns that onsets are unreliable and stops committing to them, so it fails on real data for the mirror image of the original reason. The target is not as messy as possible. It is as messy as the field and no messier, and both sides of that optimum degrade.

This is where the per-channel discipline earns its keep. An aggregate overlap score can look healthy while the mix underneath it is wrong, with noise overshot, timing jitter barely applied, and gain left untouched. One channel's excess quietly pays for another channel's absence and the summary number never says so. Matching each channel against a measured field statistic leaves nowhere for that compensation to hide.

The scoreboard has to be real

The last part of the discipline is the easiest to skip and the most expensive to skip. Because the failure is invisible on synthetic test sets, it cannot be verified on synthetic test sets either. A synthetic-only benchmark reports success right up to the moment the model meets a station, and then goes on reporting success afterwards.

So the work needs a held-out set of real, labelled field records, even a small one, kept out of training and used only as the scoreboard. Two numbers together tell the story. A distribution-overlap measure says whether training and deployment have converged. Error on the real held-out set says whether that convergence bought anything. Report one without the other and you have half an answer.

The precedent inside EarthScan is Tannistha Maiti's domain-adaptation work on seismic facies, where the gap ran between two real surveys rather than between simulation and field: a model trained on labelled F3 data and asked to read unlabelled Penobscot data, with an alignment loss doing the work of pulling the target features onto the source distribution [1]. Different gap, same reading. Performance on the source domain says nothing about the target until someone has measured the distance between them and closed it deliberately.

Zooming out

The sim-to-real gap is not a seismology problem. It appears wherever a simulator trains a model that a physical world then has to accept, and the treatment is the same in robotics, in autonomous systems and in geophysics. Measure the deployment distribution, match the training distribution to it channel by channel, and keep a real scoreboard. The modelling work is rarely where the fix lives. This is a data problem, and it closes when the training set stops flattering itself.

Key takeaways

  1. A model trained on synthetic seismograms fails in the field because its training distribution was too clean, not because the network was too small.
  2. Capacity does not close a distribution gap: a deeper network fits the clean distribution more precisely, which is the wrong direction entirely.
  3. Realism is measured rather than guessed, with noise floor, timing spread and gain variation sampled from the stations you intend to run on.
  4. Overshooting realism drowns the arrival and costs as much as undershooting it, so the target is the field level and not the maximum.
  5. Only a held-out set of real labelled records can verify the fix, because a synthetic benchmark cannot detect the failure it causes.

References

[1] Nasim, M. Q., Maiti, T., Srivastava, A., Singh, T., and Mei, J. Seismic Facies Analysis: A Deep Domain Adaptation Approach. IEEE Transactions on Geoscience and Remote Sensing, 60, 2022, Art. no. 4508116.

Tannistha Maiti
Tannistha Maiti

Co-founder & Head of Engineering

More from EarthScan

Related research

All insights →
Closing the Sim-to-Real Gap with Targeted Degradation Models
Research

Closing the Sim-to-Real Gap with Targeted Degradation Models

Sim-to-Real Demystified: Training on Fakes
Insight

Sim-to-Real Demystified: Training on Fakes

Designing Synthetic Data That Generalises, Not Memorises
Insight

Designing Synthetic Data That Generalises, Not Memorises

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.