A model can score near-perfectly on a held-out set of synthetic waveforms and then read a real station as noise. Nothing in the network is broken. It learned the distribution it was shown, and the distribution it was shown was too clean to contain the field.
The waveform panels in Tannistha Maiti's 2018 thesis are a fair picture of what too clean means. Fig. 5.3 plots radial waveforms against backazimuth, computed from a specified subsurface model by forward physics, with the answer already attached because the model defined the structure. Every arrival sits where the physics put it. Nothing wanders, nothing drops out, no instrument drifts out of calibration halfway through a season. That is what makes such waveforms good training material, and it is the same property that ends the model's usefulness on the day it is deployed. Why synthetic waveforms exist at all is a supply argument and belongs in its own piece. This one is about what happens after you accept the supply.
The gap is a mismatch, not a shortfall
Field seismograms carry background noise, instrument response quirks, timing jitter, variable gain, and structure the synthetic model never contained. The distribution of any feature the network reads, amplitude, arrival time, spectral shape, is therefore both broader and shifted relative to the synthetic one. Where the two barely overlap, the trained model is operating outside its own support the moment it meets field data. It is not generalising. It is extrapolating, and extrapolation is where a network stays confident and stops being right.
That framing matters because it rules out the fix everyone reaches for first. The instinct on hearing that a model fails on real data is to add capacity: a deeper network, a longer schedule, a wider sweep. None of it helps. Capacity buys a closer fit to the distribution you trained on. A network that has never seen a noisy trace does not learn to read noise by getting bigger. It learns to fit clean traces more precisely, which is the opposite of what the field is asking for.
Realism is a target you measure, not a dial you turn up
The fix is to move the training distribution toward the deployment distribution, and the honest way to do that is one channel at a time. Sample background noise from real recordings made at the stations you intend to run on, rather than adding noise at whatever level feels about right. Jitter the arrival times by the spread your own picks show. Vary gain and instrument response across the range your instruments actually span. Perturb the subsurface models so the training set covers the structures you expect to meet, not the one you happened to simulate first.
Every one of those is a measurement task before it is an augmentation task. The field supplies the number and the generator reproduces it. Teams that skip the measuring step end up tuning a single global realism slider by eye, and one slider cannot be correct on more than one channel at a time.
Overshooting costs as much as undershooting
Realism has an optimum rather than a direction. Push injected noise past the field level and the arrival drowns: the network learns that onsets are unreliable and stops committing to them, so it fails on real data for the mirror image of the original reason. The target is not as messy as possible. It is as messy as the field and no messier, and both sides of that optimum degrade.
This is where the per-channel discipline earns its keep. An aggregate overlap score can look healthy while the mix underneath it is wrong, with noise overshot, timing jitter barely applied, and gain left untouched. One channel's excess quietly pays for another channel's absence and the summary number never says so. Matching each channel against a measured field statistic leaves nowhere for that compensation to hide.
The scoreboard has to be real
The last part of the discipline is the easiest to skip and the most expensive to skip. Because the failure is invisible on synthetic test sets, it cannot be verified on synthetic test sets either. A synthetic-only benchmark reports success right up to the moment the model meets a station, and then goes on reporting success afterwards.
So the work needs a held-out set of real, labelled field records, even a small one, kept out of training and used only as the scoreboard. Two numbers together tell the story. A distribution-overlap measure says whether training and deployment have converged. Error on the real held-out set says whether that convergence bought anything. Report one without the other and you have half an answer.
The precedent inside EarthScan is Tannistha Maiti's domain-adaptation work on seismic facies, where the gap ran between two real surveys rather than between simulation and field: a model trained on labelled F3 data and asked to read unlabelled Penobscot data, with an alignment loss doing the work of pulling the target features onto the source distribution [1]. Different gap, same reading. Performance on the source domain says nothing about the target until someone has measured the distance between them and closed it deliberately.
Zooming out
The sim-to-real gap is not a seismology problem. It appears wherever a simulator trains a model that a physical world then has to accept, and the treatment is the same in robotics, in autonomous systems and in geophysics. Measure the deployment distribution, match the training distribution to it channel by channel, and keep a real scoreboard. The modelling work is rarely where the fix lives. This is a data problem, and it closes when the training set stops flattering itself.
Key takeaways
- A model trained on synthetic seismograms fails in the field because its training distribution was too clean, not because the network was too small.
- Capacity does not close a distribution gap: a deeper network fits the clean distribution more precisely, which is the wrong direction entirely.
- Realism is measured rather than guessed, with noise floor, timing spread and gain variation sampled from the stations you intend to run on.
- Overshooting realism drowns the arrival and costs as much as undershooting it, so the target is the field level and not the maximum.
- Only a held-out set of real labelled records can verify the fix, because a synthetic benchmark cannot detect the failure it causes.
References
[1] Nasim, M. Q., Maiti, T., Srivastava, A., Singh, T., and Mei, J. Seismic Facies Analysis: A Deep Domain Adaptation Approach. IEEE Transactions on Geoscience and Remote Sensing, 60, 2022, Art. no. 4508116.




