Skip to main content
Reading viewAll insights →
RESEARCH6 min read · 9 Oct 2026

H-kappa Uncertainty, Calibrated

An H-kappa stack can be re-read as a likelihood and reported as a calibrated credible region at essentially no extra compute. The reason to bother is not tidiness. As a station's data quality falls the region inflates along the trade-off ridge while the stack maximum barely moves, so a pipeline that reports picks shows almost nothing at all for a network that has lost its constraint.

Tannistha Maitiby Tannistha Maiti · Co-founder & Head of Engineering
subsurface-aireceiver-functionsuncertainty-quantificationgeophysicsbayesian-inference

H-kappa stacking is how crustal thickness and bulk Vp/Vs get estimated from teleseismic receiver functions almost everywhere. For a grid of candidate values you sum receiver-function amplitude at the predicted arrival times of the direct Ps conversion and its two multiples, weight by phase, and read the crustal parameters off the maximum [1]. It is fast, it survives an imperfect velocity model in ways single-phase timing does not, and it has been applied to essentially every permanent broadband station on Earth.

Its weakness is epistemic rather than computational. The three phases do not constrain the two parameters independently, so the objective is a ridge in which a deeper Moho with lower Vp/Vs fits nearly as well as a shallower one with higher Vp/Vs. We have written about the shape of that posterior before, and about why a point plus two marginal error bars describes a box full of models the data has already rejected.

This piece is about something else: what the ridge does when the data gets worse, and why that is invisible to the thing most pipelines report.

The stack is already a likelihood

The upgrade is deliberately small. Treat the normalised, non-negative stack amplitude as proportional to a likelihood, with a temperature that converts amplitude into log-probability. Put a weakly informative prior over a geologically plausible box on thickness and Vp/Vs. The posterior that falls out is well approximated near its maximum by a bivariate Gaussian, so it can be summarised by a mean, a two by two covariance, and the correlation coefficient that carries the trade-off, which for typical station geometries sits near -0.82. The 95% credible region is then the ellipse enclosing 95% of the mass.

No new forward model. No retraining. The same stack the pipeline already computes, re-read.

The load-bearing step is the temperature. Set it by requiring that the posterior predictive spread of the arrival times matches the observed scatter across the event set, so a 95% region contains the truth about 95% of the time on synthetic tests. Without that calibration the region is an arbitrary contour of an arbitrary transformation and its width means nothing. With it, the interval says what it claims.

What falls apart quietly

MEAN PICK MOVEMENT SINCE FULL QUALITY0 kmACROSS THE SAME FALL IN DATA QUALITYkappa still constrainedkappa past twice a good station’s spreadA synthetic 100-station network. Good-station figures and the -0.82 correlation are from our ES-3107 note. Drag to orbit.

This is the reference network at full quality. Turn the quality down and watch which of the two readouts responds.

What each readout says as the network degrades. A good station's kappa half width is 0.04.
network qualitymean pick movedregion areakappa spreadstations past the limit
full event set, good signal to noise (shown)0 km1x1.32x0%
a thinner event set0.02 km1.43x1.58x10%
half the events, noisier0.06 km2.22x1.97x37%
a quarter of the events0.15 km4x2.64x100%
One hundred synthetic stations drawn twice. Under the pick readout the block is a plain of reported crustal thicknesses, and taking the network from a full event set to a quarter of one moves the average pick by 0.15 km. Switch to calibrated credible regions and the same stations become a range: four times the area, 2.64 times the kappa spread of a good station, and every station past the point where composition inference is supportable. Good-station figures and the H-kappa correlation of -0.82 are from our ES-3107 note; the network is synthetic.

Here is the same network twice.

Leave the readout on the stack maximum, which is what the overwhelming majority of pipelines report, and drop the network from a full event set to a quarter of one. The block stays a plain. The average reported thickness moves by 0.15 km, which is a hundred and fifty metres, and no station in the picture announces that anything has changed.

Now switch the readout to calibrated regions without touching the quality. The same hundred stations become a range. The credible regions have grown fourfold over that fall. The kappa spread is 2.64 times what a good station delivers, and every station in the network is now past twice that spread, which is the point where a composition inference, felsic against mafic, stops being supportable.

The picks behind those columns are the same numbers as before.

This is the argument for the upgrade, and it is stronger than the usual one. A pick-only pipeline does not merely under-report the loss of constraint. It has no channel for it. The maximum of a stack is a location statistic, and what degrades is the width of the thing it locates, so the reported quantity is close to stationary while the evidence behind it drains away. An analyst watching picks sees a network that looks exactly as healthy as it did last season.

What the correlation does downstream

The other reason to hand on the full covariance rather than two marginal intervals is that the correlation is the dominant term, and everything downstream inherits it.

Anything consuming Vp/Vs, crustal composition, partial melt or fluid presence, the Poisson's ratio going into a rock-physics model, inherits the trade-off whether or not the analyst acknowledges it. So does crustal thickness feeding an isostatic or thermal model. Propagating a marginal error bar treats the two as independent and produces a distribution the data never supported. The credible region is not a cosmetic improvement on the pick, it is the correct object to hand to the next model in the chain.

What this does not fix

The method inherits every assumption of the stack it re-reads. A flat, isotropic single layer is the model, so a dipping Moho or crustal anisotropy biases the entire posterior rather than merely the pick, and the credible region will then be confidently wrong rather than honestly wide. We have written separately about exactly that trap, where tilting the interface moves the answer while the stack still looks healthy, and the calibration described here does nothing about it. A bias is not a width.

Getting those cases right needs a fuller forward model and a genuine inversion, which is where amortized simulation-based inference earns its cost. For the flat-layer regime where H-kappa is already trusted and already running, calibrated credible regions are strictly better than point picks, and they are close to free.

Key takeaways

  1. An H-kappa stack can be re-read as a likelihood and summarised as a calibrated credible region using the stack the pipeline already computes.
  2. The temperature calibration is the load-bearing step: without it the region is an arbitrary contour, and with it a 95% region contains the truth about 95% of the time.
  3. As data quality falls the region inflates along the trade-off ridge while the stack maximum barely moves: 0.15 km of pick movement against a fourfold growth in region area.
  4. A pick-only pipeline therefore has no channel for the loss of constraint. It is not under-reported, it is absent.
  5. The H-kappa correlation near -0.82 is the dominant uncertainty, and anything consuming Vp/Vs or thickness downstream inherits it whether or not it is propagated.
  6. The method inherits the stack's flat isotropic assumption: a dipping Moho biases the whole posterior, and a calibrated region will then be confidently wrong rather than honestly wide.

References

  1. L. Zhu, H. Kanamori. Moho depth variation in southern California from teleseismic receiver functions. Journal of Geophysical Research, 105, 2969 to 2980, 2000. doi:10.1029/1999JB900322
  2. B. Efron, R. Tibshirani. Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy. Statistical Science, 1, 54 to 75, 1986. doi:10.1214/ss/1177013815
  3. M. Sambridge, K. Mosegaard. Monte Carlo methods in geophysical inverse problems. Reviews of Geophysics, 40, 2002. doi:10.1029/2000RG000089
  4. A. Tarantola. Inverse Problem Theory and Methods for Model Parameter Estimation. SIAM, 2005. doi:10.1137/1.9780898717921
  5. D. W. Eaton, S. Dineva, R. Mereu. Crustal structure and composition from receiver-function H-kappa analysis. Geophysical Journal International, 165, 1069 to 1083, 2006. doi:10.1111/j.1365-246X.2006.02958.x
  6. The shape of the posterior itself, and why marginals mislead: Simulation-based inference for crustal structure.
  7. The bias this calibration does not fix: The dipping-interface trap.
Tannistha Maiti
Tannistha Maiti

Co-founder & Head of Engineering

More from EarthScan

Related research

All insights →
Simulation-Based Inference for Crustal Structure
Insight

Simulation-Based Inference for Crustal Structure

The Dipping-Interface Trap: When a Flat H-κ Reads the Wrong Crust
Insight

The Dipping-Interface Trap: When a Flat H-κ Reads the Wrong Crust

Below Tuning, Thickness and Contrast Are the Same Number
Insight

Below Tuning, Thickness and Contrast Are the Same Number

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.