A seismic station records the crust underneath it and nothing else. Install a thousand of them and you own a thousand fixed opinions about a thousand places you were able to get permission to build, repeated for years. That is a large quantity of data and a surprisingly small quantity of variety, and the gap between those two things decides what a model trained on the network can do.
A station is a slow way to buy a training example
Siting a station, permitting it, installing it and then waiting for enough earthquakes to arrive before it yields usable receiver functions is a multi-year commitment before the first labelled example exists. A network of a thousand stations is an enormous outlay of capital and calendar, and it still only samples the places you were able to build.
For a machine-learning model, that network is the entire training set: a few thousand labelled examples, unevenly distributed, expensive, and slow to grow. The ceiling on what the model can learn is set by data scarcity rather than by algorithm quality, which is an uncomfortable place for a research budget to sit, because it means the next increment of performance has to be bought with hardware and patience rather than with method.
What the simulator sells is not mainly speed
Give a forward simulator an earth model, a stack of layers with velocities and depths, and it computes the receiver function that a station standing on that structure would record. The answer arrives attached, because the simulator defined the structure before it computed the trace [1]. Every output is a labelled example with no annotation step and no wait for seismicity.
The obvious reading of that is cheaper examples. The more useful reading is that you now choose where in the space of crustal structures your examples come from. A station cannot be relocated to an interesting structure. A simulated station is specified rather than found, so the question shifts from where you are allowed to dig to which structures you want the model to have seen.
Count and coverage are different quantities
Two levers look alike here and behave very differently.
The first multiplies examples. One earth model can be rendered across dozens of incidence angles, noise levels and phase combinations, and each render is another labelled training pair. This is real value: it teaches invariance, and it is nearly free. What it does not do is add a crustal structure the model has never seen. Render one model ten thousand times and the set of structures in the training data is still one.
The second adds structures. Choosing a new earth model, with a different thickness or a different velocity contrast, puts a genuinely new region of the parameter space in front of the network. That is the lever a simulator hands you directly, and it is the one a field network hands you only by accident, because a new station adds a new structure only when it happens to sit on crust unlike the crust you already record. Past the first well-chosen sites, most additional stations land in terrain that resembles terrain already in the network, since both were selected under the same accessibility constraints.
Why the title number is deliberately small
32 is illustrative. The exact multiplier depends on how the parameter space is sampled, and anyone quoting a fixed ratio of synthetic models to real stations is quoting a modelling choice rather than a measurement. The claim underneath the number is narrower and sturdier: the count of distinct structures in a training set matters more than the count of recordings, and the count of distinct structures is precisely what a network cannot buy at any price and a simulator supplies at the cost of compute.
That reframes a capital decision. Adding stations is procurement with a multi-year lead time and a fixed geographic footprint. Adding structures is a sampling decision made in an afternoon, revisable next week when the model shows you which part of the space it handles worst.
What one station samples
Wait before first usable data
Manual annotation per synthetic example
Model count in the title
What still has to come from the field
A simulated receiver function is only worth what it transfers. Real data keeps the job of telling you whether the simulator is describing the same planet, and that check is a piece of work in its own right rather than a formality. What field data loses is the job of being the only source of supply, and with it the power to set the pace of the whole programme.
The honest framing for a review board is therefore not that synthetics replace stations. It is that the network moves from being the training set to being the test set, which is a much better use of an asset that took a decade to build.
Zooming out
Other data-hungry fields have already been reshaped by this same shift. Once a credible simulator exists, synthetic data stops being the fallback used when labels run out and becomes the primary supply, with measurement retained for validation. For subsurface work the consequence is commercial rather than only technical. You can build coverage where no station exists, and you can retrain as often as you can afford to run the simulator.
The competitive question changes shape accordingly. It stops being who owns the largest network and becomes who can specify the structures worth learning, and that is a question about geological judgement and sampling design, not about procurement.
Key takeaways
- A seismic station is a multi-year commitment that samples one crustal column, so a large network still yields a small, uneven training set.
- The binding constraint on model quality in this setting is data scarcity rather than algorithm quality, which means method improvements cannot lift it.
- A forward simulator returns a labelled receiver function for any specified earth model, so the label costs nothing and the wait for seismicity disappears.
- Rendering one earth model many times multiplies examples without adding coverage; only choosing new earth models adds structures the model has not seen.
- The counts in the title are illustrative, and the durable claim is that a network cannot buy distinct structures at any price while a simulator supplies them at compute cost.
References
[1] Frederiksen, A. W., and Bostock, M. G. (2000). Modelling teleseismic waves in dipping anisotropic structures. Geophysical Journal International, 141, 401-412. doi:10.1046/j.1365-246X.2000.00090.x
[2] Source: 2018 PhD thesis, receiver-function imaging of the Moho/LAB, Fig 2.4 (synthetic P receiver functions). The station and model counts used here are illustrative, as the source itself states; no data-generation throughput was measured.




