Skip to main content
Reading viewAll insights →
BLOG6 min read

32 Synthetic Models Beat 1,000 Stations

32 Synthetic Models Beat 1,000 Stations
Tannistha Maitiby Tannistha MaitiSenior AI Researcher · 28 Sep 2026
Share

A seismic station is a multi-year commitment that, once installed, only ever sees the one column of crust underneath it. A physics simulator returns a labelled receiver function for any crustal structure you care to specify, which means a network buys you recordings while a simulator buys you coverage of the structures a model actually has to learn. The two counts in the title are illustrative; the argument is about which of those quantities is the binding one.

A seismic station records the crust underneath it and nothing else. Install a thousand of them and you own a thousand fixed opinions about a thousand places you were able to get permission to build, repeated for years. That is a large quantity of data and a surprisingly small quantity of variety, and the gap between those two things decides what a model trained on the network can do.

A station is a slow way to buy a training example

Siting a station, permitting it, installing it and then waiting for enough earthquakes to arrive before it yields usable receiver functions is a multi-year commitment before the first labelled example exists. A network of a thousand stations is an enormous outlay of capital and calendar, and it still only samples the places you were able to build.

For a machine-learning model, that network is the entire training set: a few thousand labelled examples, unevenly distributed, expensive, and slow to grow. The ceiling on what the model can learn is set by data scarcity rather than by algorithm quality, which is an uncomfortable place for a research budget to sit, because it means the next increment of performance has to be bought with hardware and patience rather than with method.

What the simulator sells is not mainly speed

Give a forward simulator an earth model, a stack of layers with velocities and depths, and it computes the receiver function that a station standing on that structure would record. The answer arrives attached, because the simulator defined the structure before it computed the trace [1]. Every output is a labelled example with no annotation step and no wait for seismicity.

The obvious reading of that is cheaper examples. The more useful reading is that you now choose where in the space of crustal structures your examples come from. A station cannot be relocated to an interesting structure. A simulated station is specified rather than found, so the question shifts from where you are allowed to dig to which structures you want the model to have seen.

Count and coverage are different quantities

Two levers look alike here and behave very differently.

The first multiplies examples. One earth model can be rendered across dozens of incidence angles, noise levels and phase combinations, and each render is another labelled training pair. This is real value: it teaches invariance, and it is nearly free. What it does not do is add a crustal structure the model has never seen. Render one model ten thousand times and the set of structures in the training data is still one.

The second adds structures. Choosing a new earth model, with a different thickness or a different velocity contrast, puts a genuinely new region of the parameter space in front of the network. That is the lever a simulator hands you directly, and it is the one a field network hands you only by accident, because a new station adds a new structure only when it happens to sit on crust unlike the crust you already record. Past the first well-chosen sites, most additional stations land in terrain that resembles terrain already in the network, since both were selected under the same accessibility constraints.

STRUCTURAL COVERAGE32 vs 29STRUCTURES SEEN, OF 70A field networksited where building is possible, so it re-samples the same crustSimulated earth modelsspecified rather than found, so the structures are chosenMoho depth 26 to 44 kmVp/Vs 1.62 to 1.9Moho depth 26 to 44 kmVp/Vs 1.62 to 1.91,320 examples · 29 structures (41%)32 examples · 32 structures (46%)The accessibility pattern is a schematic, not a siting study. The structural point is not: re-sampling a distribution cannot cover the parts of it the distribution never reaches.
The same parameter space twice: Moho depth against bulk Vp/Vs, every cell one structural regime a model would have to have seen. Shading is how many labelled examples landed there. Drag the station lever and the example count climbs steeply while the occupied cells stall, because sites are chosen for accessibility and later stations keep landing on crust the network already records. Drag the model lever and coverage climbs almost one for one, because a simulated station is specified rather than found. Then drag renders per model: the example count multiplies and the coverage readout does not move by a single cell. That is the finding. Two of these three levers buy examples and exactly one buys structures, and the one that feels most productive buys none. The accessibility pattern is a schematic rather than a siting study; the structural claim under it is not, because re-sampling a distribution cannot reach the parts of it the distribution never visits.

Why the title number is deliberately small

32 is illustrative. The exact multiplier depends on how the parameter space is sampled, and anyone quoting a fixed ratio of synthetic models to real stations is quoting a modelling choice rather than a measurement. The claim underneath the number is narrower and sturdier: the count of distinct structures in a training set matters more than the count of recordings, and the count of distinct structures is precisely what a network cannot buy at any price and a simulator supplies at the cost of compute.

That reframes a capital decision. Adding stations is procurement with a multi-year lead time and a fixed geographic footprint. Adding structures is a sampling decision made in an afternoon, revisable next week when the model shows you which part of the space it handles worst.

one crustal column

What one station samples

years

Wait before first usable data

none

Manual annotation per synthetic example

illustrative

Model count in the title

What still has to come from the field

A simulated receiver function is only worth what it transfers. Real data keeps the job of telling you whether the simulator is describing the same planet, and that check is a piece of work in its own right rather than a formality. What field data loses is the job of being the only source of supply, and with it the power to set the pace of the whole programme.

The honest framing for a review board is therefore not that synthetics replace stations. It is that the network moves from being the training set to being the test set, which is a much better use of an asset that took a decade to build.

Zooming out

Other data-hungry fields have already been reshaped by this same shift. Once a credible simulator exists, synthetic data stops being the fallback used when labels run out and becomes the primary supply, with measurement retained for validation. For subsurface work the consequence is commercial rather than only technical. You can build coverage where no station exists, and you can retrain as often as you can afford to run the simulator.

The competitive question changes shape accordingly. It stops being who owns the largest network and becomes who can specify the structures worth learning, and that is a question about geological judgement and sampling design, not about procurement.

Key takeaways

  1. A seismic station is a multi-year commitment that samples one crustal column, so a large network still yields a small, uneven training set.
  2. The binding constraint on model quality in this setting is data scarcity rather than algorithm quality, which means method improvements cannot lift it.
  3. A forward simulator returns a labelled receiver function for any specified earth model, so the label costs nothing and the wait for seismicity disappears.
  4. Rendering one earth model many times multiplies examples without adding coverage; only choosing new earth models adds structures the model has not seen.
  5. The counts in the title are illustrative, and the durable claim is that a network cannot buy distinct structures at any price while a simulator supplies them at compute cost.

References

[1] Frederiksen, A. W., and Bostock, M. G. (2000). Modelling teleseismic waves in dipping anisotropic structures. Geophysical Journal International, 141, 401-412. doi:10.1046/j.1365-246X.2000.00090.x

[2] Source: 2018 PhD thesis, receiver-function imaging of the Moho/LAB, Fig 2.4 (synthetic P receiver functions). The station and model counts used here are illustrative, as the source itself states; no data-generation throughput was measured.

Tannistha Maiti
Tannistha Maiti

Senior AI Researcher

More from EarthScan

Related research

All insights →
The 21-Number Problem: Teaching a Net About Symmetry
Insight

The 21-Number Problem: Teaching a Net About Symmetry

Neural-Operator Surrogates for Receiver-Function Forward Modelling
Research

Neural-Operator Surrogates for Receiver-Function Forward Modelling

Physics-Informed Neural Networks for Seismic Noise Attenuation: When the Wave Equation Is Your Only Label
Research

Physics-Informed Neural Networks for Seismic Noise Attenuation: When the Wave Equation Is Your Only Label

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.