Skip to main content
Reading viewAll insights →
BLOG18 min read

The Error That Cancels: A Better Velocity Model Can Move the Prospect the Wrong Way

The Error That Cancels: A Better Velocity Model Can Move the Prospect the Wrong Way
Tannistha Maitiby Tannistha MaitiSenior AI Researcher · 9 Sep 2026
Share

Machine-learned velocity model building is scored the way every regression is scored, on the root-mean-square error of the residual. That scalar is a marginal statistic. What an undrilled trap depends on is not the depth error but the difference in depth error between the crest and the spill, and for a stationary field that difference has variance 2 sigma squared times one minus rho at the crest-to-spill separation. A long-correlated error is almost free and a short-correlated one costs the full square root of two, so an estimator that cuts sigma by 30% while shortening its correlation length from four closure widths to one improves its published score by 30% and makes the relief error 2.26 times worse. bp told the market in January 2019 that its full-waveform-inversion work had found a further billion barrels in place at Thunder Horse, Petrobras has published on heterogeneous salt velocity models feeding straight into gross rock volume in the Santos pre-salt, and Chevron and Shell are sanctioning multi-billion-dollar subsalt developments in the Gulf. The number to ask a vendor for is not the RMS. It is the semivariogram of the residual at the closure scale.

A vendor brings you a learned velocity model builder. The slide has one number on it, a root-mean-square error against a held-out set, and the number is better than the one your tomography produces. The question worth asking before anything else is not whether the number is real. It is whether that number, improved, moves the decision the way the slide implies. On an undrilled structure it can move it the other way, and nothing in the reported figure shows which.

The problem is on the record, from both ends

bp announced on 9 January 2019 that proprietary algorithms enhancing full-waveform inversion had let it process seismic data that would previously have taken a year to analyse in a few weeks, and that the reprocessing had identified a further 1 billion barrels of oil in place at Thunder Horse and an additional 400 million barrels in place at Atlantis. The same release described Wolfspar, an ultra-low-frequency seismic source, as letting geophysicists see deeper below salt layers and plan wells better, and gave the Atlantis Phase 3 development a 1.3 billion dollar price and a 2020 start, with peak production of about 38,000 barrels of oil equivalent per day gross. Bernard Looney, then bp's Upstream chief executive, called the Gulf business key to a strategy of growing production of advantaged high-margin oil [1].

Read that release again with an eye on the units. The headline is a volume in place, and a volume in place is an integral over a structure. The imaging work is being reported in the currency of the decision rather than in the currency of the algorithm, which is exactly right. The trouble starts when the algorithm is procured on the algorithm's own currency instead.

Petrobras and its academic collaborators have been publishing on the same problem from the other end. Maul and co-authors, writing from Petrobras, Universidade Federal Fluminense, Heriot-Watt and Emerson, set out the Santos Basin case directly: the thick and heterogeneous salt section imposes challenges for seismic imaging, signal quality and depth positioning, and most velocity models for the salt treat it as homogeneous halite when halite makes up only about 80% of the mineral content of that section. They compare gross rock volume above the oil-water contact under velocity models with and without salt stratification and report a significant effect both on the depth resolution of the events and on the volume estimate [2].

That paper is the pivot of this piece, and it is worth being precise about why. It is not a paper about velocity error in the abstract. It is a paper in which a change to the velocity model propagates into a gross rock volume, which is the number the investment decision is actually made on.

Chevron and Shell are running the same class of deepwater problem in the Gulf at a scale that fixes the stakes. Chevron started production at Anchor in August 2024, the first deepwater development to deploy 20,000 psi technology, with reservoir depths reaching about 34,000 feet below the water surface in around 5,000 feet of water and a design capacity of 75,000 barrels of oil per day and 28 million cubic feet of gas per day [3]. It started production at Ballymore in April 2025, a 1.6 billion dollar three-well subsea tieback in about 6,600 feet of water, tied back three miles to the Blind Faith facility, expected to produce up to 75,000 gross barrels of oil per day, with about 150 million barrels of oil equivalent gross recoverable over the project life [4]. Shell took a final investment decision on Sparta in December 2023, a development with about 90,000 boe per day of peak production and 244 million boe of estimated recoverable resources in more than 4,700 feet of water, Shell's first in the Gulf to produce from reservoirs at pressures up to 20,000 psi [5], and started production at Whale in January 2025, in more than 8,600 feet of water, with 15 wells, 100,000 boe per day at peak and 480 million boe estimated recoverable [6].

None of those four releases says anything about velocity model building, and nothing below attributes any velocity-model practice to Chevron or to Shell. They are here to establish the size of the decision that a depth-converted surface feeds.

The scalar everyone reports

A learned velocity model builder, like any regression, is scored on the residual. Call the residual field on the depth-converted horizon d(x), the difference between the mapped depth and the true depth at map position x. Model it as a zero-mean, second-order stationary random field with

E[d(x)]=0,Var[d(x)]=σ2,Cov[d(x1),d(x2)]=σ2ρ(h)\mathbb{E}[d(x)] = 0, \qquad \operatorname{Var}[d(x)] = \sigma^{2}, \qquad \operatorname{Cov}[d(x_1), d(x_2)] = \sigma^{2}\rho(h)

where h is the separation |x_1 - x_2| and rho is a correlation function with rho(0) = 1. The reported score is the root mean square of d over the evaluation area, which estimates sigma and nothing else. It is a marginal statistic. Two residual fields with the same sigma and completely different spatial structure produce the same score, exactly.

That is the whole problem in one sentence, and the rest of this is working out what it costs.

What an undrilled trap actually depends on

A trap is a relief integral. Gross rock volume above a bounding surface is

GRV  =  Amax ⁣(zbz(x,y),0)dA\mathrm{GRV} \;=\; \iint_{A} \max\!\bigl(z_{b} - z(x,y),\, 0\bigr)\, dA

with z the top-reservoir depth and z_b the depth of the bounding surface. There are two regimes here and they behave differently, which is worth separating because the RMS score is a fair scalar for one of them.

Where the fluid contact is known in true vertical depth, from a well penetration or a pressure gradient, z_b is fixed independently of the seismic. A bulk downward shift of the whole mapped surface then genuinely reduces the volume above that contact, and sigma is the right thing to worry about. That is the appraisal case, and it needs a discovery with wells in it.

Where there is no contact control, which is the state of every prospect before it is drilled, the column is limited by structural closure. The trap holds hydrocarbon down to its spill point and no further, so z_b is the spill depth, which is itself read off the same depth-converted surface. A bulk shift moves the crest and the spill together and leaves the relief between them untouched. The volume is driven by

H  =  z(xspill)    z(xcrest)H \;=\; z(\mathbf{x}_{\text{spill}}) \;-\; z(\mathbf{x}_{\text{crest}})

and the error in H is the error field differenced at two points a fixed distance apart, where that distance is a property of the structure rather than of the velocity model.

There is a second push in the same direction even in the appraisal case. If the depth conversion is calibrated to well markers, the residual is forced toward zero at those locations by construction. Calibration removes the long-wavelength part of d where control exists and leaves behind whatever varies between the control points. So both regimes end up caring about the two-point statistic, and only one of them also cares about the marginal one.

The two-point identity, and why it is the whole argument

Write the relief error as a difference of the field at two points separated by W, the crest-to-spill separation:

R  =  d(x1)d(x2),x1x2=WR \;=\; d(\mathbf{x}_1) - d(\mathbf{x}_2), \qquad |\mathbf{x}_1 - \mathbf{x}_2| = W

The variance of a difference of two correlated variables is the sum of the variances less twice the covariance, so under stationarity

Var[R]  =  σ2+σ22σ2ρ(W)  =  2σ2(1ρ(W))\operatorname{Var}[R] \;=\; \sigma^{2} + \sigma^{2} - 2\sigma^{2}\rho(W) \;=\; 2\sigma^{2}\bigl(1 - \rho(W)\bigr)

and the RMS relief error is

Var[R]  =  σ2(1ρ(W))\sqrt{\operatorname{Var}[R]} \;=\; \sigma\sqrt{2\bigl(1 - \rho(W)\bigr)}

Read the two limits. As W/L goes to zero the correlation goes to one and the relief error goes to zero: a perfectly long-correlated error is a bulk shift, and a bulk shift is free on an undrilled structure. As W/L grows the correlation goes to zero and the relief error goes to sqrt(2) sigma: two independent errors add in quadrature. The factor between those two extremes is arbitrarily small below and capped at 1.414 above, and sigma is identical across the whole range.

Take a Gaussian correlation rho(h) = exp(-(h/L)^2) and put numbers on it. At L/W = 4, rho is 0.9394 and the relief error is 0.348 sigma. At L/W = 1, rho is 0.3679 and the relief error is 1.124 sigma. At L/W = 0.25, rho is effectively zero and the relief error is 1.414 sigma. Same sigma, same reported score, a factor of 4.06 in the quantity that sets gross rock volume.

Where a learned estimator sits on that curve

Now the part that makes this more than an observation about variograms. A patch-wise learned estimator predicts on a tile and stitches. Whatever its accuracy, the length scale over which its errors stay coherent is bounded by the machinery that produced them, and reducing pointwise error is the objective it is trained on while error coherence is not in the loss at all. A tomographic inversion regularised toward smoothness has the opposite bias: it is built to produce long-wavelength structure and it produces long-wavelength error along with it.

That gives a trade with a sign. Suppose an estimator cuts sigma by 30%, which is a real and hard-won improvement in the reported score, while cutting the correlation length of its residual from four crest-to-spill separations to one. The relief error goes from 0.348 sigma to 0.7 times 1.124 sigma, which is 0.787 sigma, a factor of 2.26 worse on the quantity the decision uses, at the same time as the published number improves by 30%. The score and the decision move in opposite directions and the score cannot tell you.

VELOCITY MODEL SCORE · RELIEF CANCELLATION20.9 mRELIEF ERROR RMS, CREST TO SPILLdistance along section, kmdepth, mcrest to spill W = 3.00 kmpicked horizonpicked + error, one seeded realisationSCORED VERSUS DECIDINGabsolute depth error RMS, sigma60.0 mrelief error RMS, crest to spill20.9 mone scale, 0 to 300 m, common to both barsRELIEF RMS / SIGMAgaussianexponentialL / W (log)READOUTdepth error RMS, sigma60.0 mcorrelation length L12.0 kmcrest to spill W3.00 kmL / W4.00rho at W0.9394relief RMS / sigma0.3481this realisation, crest to spill5.7 mpreset ratio, learned / legacy2.26xSTIPULATED MODEL PAIRlegacy tomographicsigma 60 m, L = 4Wlearned patch-wisesigma 42 m, L = Wlegacy 20.9 m · learned 47.2 m of relief errorreported score 0.70x · relief 2.26xcorrelation length L0.1 to 40 km (log)12.0 kmdepth error RMS sigma5 to 200 m60.0 mcrest to spill W0.5 to 8 km3.00 kmBand: one seeded spectral realisation. Readouts: exact ensemble statistics. The two presets are a stipulated contrast, not measured.
A depth-error field over a picked horizon, scored two ways. The aqua bar is the absolute depth error RMS, which is sigma and carries no information about the correlation length. The amber bar is the RMS error in the crest-to-spill relief, sigma times the square root of 2 (1 - rho(W)), which is the quantity a gross rock volume depends on. Hold sigma and the crest-to-spill span W and drag the correlation length: the aqua bar does not move at all, while the amber bar runs from near nothing to a bar the square root of two times as long as the aqua one. The pills set a stipulated pair, a legacy tomographic model at sigma 60 m with L = 4W against a learned patch-wise model at sigma 42 m with L = W; those two numbers are an illustration of the trade, not a measurement of any vendor's product. Clicking from legacy to learned improves the reported RMS score by 30 per cent and makes the relief error 2.26 times worse. The inset draws the relief factor against L / W for a Gaussian correlation (solid) and an exponential one (dashed). The shaded band on the section is one seeded spectral realisation of the error field; its own crest-to-spill difference has its own readout, and being a single draw it sits wherever that draw puts it rather than on the ensemble figure the bars use.

The exhibit is that argument as one control. Open it on the legacy pill, hold the crest-to-spill span at 3 km and sigma at 60 m, then drag the correlation length across its track. The aqua bar, absolute depth error RMS, holds at 60.0 m for every position on that track, because it is sigma and the correlation length is not in it. The amber bar moves from 6.4 m at the 40 km end to 84.9 m at the 0.1 km end, which is the full sqrt(2) sigma, and the marker in the inset walks the relief curve the whole way. Then click from the legacy pill to the learned pill and read the two lines under them: the reported score improves by a factor of 0.70 and the relief error goes from 20.9 m to 47.2 m, the 2.26 the readout prints.

The two pills are a stipulated pair, not a measurement. They exist so the trade can be walked in one click, and the numbers in them are ours to illustrate the shape of the problem, not anyone's product specification.

One thing in the exhibit is deliberately not the closed form. The shaded band on the section is a single seeded realisation of the error field, synthesised from an evenly spaced frequency comb whose amplitudes come from the spectral density of the Gaussian covariance and whose phases are fixed by the seed, so the correlation length is an explicit input rather than something that emerges from a filter width. Its crest-to-spill difference is printed as its own readout, and being one draw it sits wherever that draw puts it rather than on the ensemble figure. The bars and the inset are the exact statistics.

The number to ask for instead

The fix does not need new mathematics, because the quantity is already a standard object. The semivariogram of the residual is

γ(h)  =  12E[(d(x+h)d(x))2]  =  σ2(1ρ(h))\gamma(h) \;=\; \tfrac{1}{2}\,\mathbb{E}\bigl[(d(x+h) - d(x))^{2}\bigr] \;=\; \sigma^{2}\bigl(1 - \rho(h)\bigr)

so the relief-error variance is 2 gamma(W) and the relief-error RMS is sqrt(2 gamma(W)). The RMS score is the limit of that expression as h grows past the correlation range, divided by sqrt(2). It is the square root of the sill, and the sill is the least informative point on the curve for this purpose.

So the ask is short. Do not ask for the RMS of the depth residual. Ask for gamma(h) of the depth residual, evaluated at the crest-to-spill separations of the structures in the portfolio, on the same blind set the RMS was computed on. If the vendor cannot produce it, they have the residual field and can compute it in an afternoon. If they will not produce it, the reason is worth knowing.

A useful second ask, for anyone already carrying well control: report the residual at the wells that were held out of the calibration, and report it as a function of distance from the nearest calibration well. That is the same statistic viewed through the acquisition geometry rather than the structural geometry, and it is the version an appraisal team can act on.

Reading it in velocity rather than depth

Velocity model builders are often scored on the velocity residual rather than the depth residual, and the argument survives the change of variable without modification, because it is an argument about a marginal statistic versus a two-point statistic and the RMS of any residual field is marginal.

The bridge between the two is worth having anyway. For a vertically travelling ray the depth is z = V t / 2 for a two-way time t, so a fractional velocity error produces the same fractional depth error to first order: one per cent at a 3,000 m objective is 30 m, and one per cent at 7,000 m is 70 m. That is the crude version, and a subsalt interpreter will immediately add ray bending, anisotropy and the fact that the error is concentrated in the salt rather than spread through the section. All of those change the constant. None of them changes the structure of the argument, which is that sigma and rho are separate properties of the residual field and the score only sees one of them.

Our own bench, and what it is not

We have never built a velocity model for a pre-salt or a subsalt programme. Nothing above is a measurement from our own work, and everything attributed to bp, Petrobras, Chevron and Shell is what their own releases and their own published paper say.

What we do have from our own delivery experience is the same failure mode in a smaller and better-instrumented setting, twice. On borehole image logs we found that the standard detection metrics broke on sinusoids and replaced them with a depth-tolerance confusion matrix, which counts a fracture pick correct when it lands within a physical band of the truth, scored at 3 cm and 5 cm (A Depth-Tolerance Confusion Matrix). On VeerNet, the encoder-decoder we use to lift curves off scanned logs, a headline overlap score of 0.51 was an average of a background class at 0.94 with the two curve classes at 0.26 and 0.21, so the single number read far healthier than the model was on the pixels that carry the signal (Reading IoU and Dice Scores Without Fooling Yourself).

Both are the same shape as this one. A scalar that is correct as arithmetic, computed on the right data, reported honestly, and pointed at a quantity the decision does not use. The subsurface version is more expensive because the decision on the other end of it is a platform.

Limitations

The stationarity assumption is the weakest part and it is doing real work. A subsalt depth residual is not stationary: it is structured by the salt geometry, it is larger under thick and complex salt than under thin salt, and its correlation structure inherits the shape of the overburden rather than being a property of the map. Everything above should be read as the behaviour of the leading term, not as a model of a real residual field.

The Gaussian correlation is a smooth choice and a rough one is more defensible for a geological field. With an exponential correlation rho(h) = exp(-h/L) the same numbers come out softer: L/W = 4 gives a relief error of 0.665 sigma rather than 0.348, and the legacy-to-learned move is a factor of 1.18 worse rather than 2.26. The sign of the effect does not change and the score is equally blind to it either way. The exhibit draws both families in its inset for that reason.

The two-point treatment is also a simplification of the volume integral. A real gross rock volume is an area integral over the closure and the relevant statistic is the variance of that integral, which involves rho integrated over the closure rather than evaluated at a single lag. The single-lag version is the leading behaviour and it has the same monotonicity in L, but it is not the full calculation, and anyone doing this properly for a portfolio should do the area integral.

The 30% and the four-to-one correlation length change in the pills are stipulated. They are chosen to be a plausible shape of trade, not measured from any system, and the exhibit says so on its face.

Finally, the claim about patch-wise estimators is an argument from what the loss contains, not a measurement. If a learned velocity model builder is trained with a term that penalises the spatial structure of its residual, or is post-conditioned to a long-wavelength prior, the trade described here does not apply to it. That is the interesting design question underneath all of this, and it is a better one than the RMS.

References

[1] bp. bp plans for significant growth in deepwater Gulf of Mexico. Press release, 9 January 2019. States that bp's proprietary algorithms enhance full waveform inversion and let seismic data that would previously have taken a year to analyse be processed in a few weeks; that this identified a further 1 billion barrels of oil in place at Thunder Horse and an additional 400 million barrels at Atlantis; that Wolfspar uses ultra-low frequencies to let geophysicists see deeper below salt layers; that Atlantis Phase 3 is a 1.3 billion dollar development scheduled onstream in 2020 with about 38,000 boe/d gross at peak; and quotes Bernard Looney, bp's Upstream chief executive. The bp.com URL for this release returns HTTP 403 to automated fetch, so the content was confirmed against the trade reproduction below and is cited from it. https://www.bp.com/en/global/corporate/news-and-insights/press-releases/bp-plans-for-significant-growth-in-deepwater-gulf-of-mexico.html https://www.oceannews.com/news/energy/bp-identifies-significant-additional-oil-resources-in-gulf-of-mexico

[2] Maul, A., Cetale, M., Guizan, C., Corbett, P., Underhill, J. R., Teixeira, L., Pontes, R. and Gonzalez, M. The impact of heterogeneous salt velocity models on the gross rock volume estimation: an example from the Santos Basin pre-salt, Brazil. Petroleum Geoscience 27(4), November 2021, doi 10.1144/petgeo2020-105. Author affiliations include Petrobras, Instituto de Geociencias at Universidade Federal Fluminense, Heriot-Watt University and Emerson Automation Solutions. The abstract states that halite makes up only about 80% of the mineral content of the Santos salt section, that most velocity models assume a homogeneous halite layer, and that comparing gross rock volume above the oil-water contact under models with and without salt stratification shows a significant effect on both depth resolution and volume estimation. The publisher pages at lyellcollection.org and earthdoc.org return HTTP 403 to automated fetch; the record and abstract were taken from the Crossref deposit for the DOI. https://doi.org/10.1144/petgeo2020-105

[3] Chevron. Chevron starts production at Anchor with industry-first deepwater technology, 12 August 2024, as reported by Oil & Gas Journal. States 20,000 psi subsea technology, reservoir depths reaching about 34,000 ft from the water surface, water depth about 5,000 ft, design capacity 75,000 b/d of oil and 28 MMcf/d of gas, Chevron 62.86% operator and TotalEnergies 37.14%. The chevron.com release URL returns HTTP 403 to automated fetch and is cited through the trade report. https://www.ogj.com/drilling-production/production-operations/field-start-ups/article/55132297/chevron-starts-oil-production-at-gulf-of-mexico-anchor-field

[4] Chevron. Chevron starts oil production from Ballymore project, 21 April 2025, as reproduced by BOE Report and WorkBoat. States water depth around 6,600 ft, about 160 miles southeast of New Orleans, three wells tied back three miles by one flowline to the Blind Faith facility, up to 75,000 gross barrels of oil and 50 million cubic feet of gas per day, 1.6 billion dollar project cost, 150 million boe gross recoverable over project life, Chevron 60% operator and TotalEnergies 40%. The chevron.com release URL returns HTTP 403 to automated fetch and is cited through the trade reproductions. https://boereport.com/2025/04/21/chevron-starts-oil-production-from-ballymore-project-in-gulf-of-america/ https://www.workboat.com/chevron-achieves-first-oil-at-ballymore-in-us-gulf

[5] Shell. Shell invests in the Sparta development in the Gulf of Mexico. Media release, 19 December 2023, distributed via PR Newswire. States peak production of approximately 90,000 boe/d, 244 million boe of estimated recoverable resources, water depth of more than 1,400 m (4,700 ft), reservoir pressures up to 20,000 psi, and that Sparta replicates about 95% of Whale's hull and 85% of Whale's topsides. https://www.prnewswire.com/news-releases/shell-invests-in-the-sparta-development-in-the-gulf-of-mexico-302019315.html

[6] Shell. Shell starts production at Whale in the US Gulf of Mexico. Media release, 9 January 2025, distributed via PR Newswire. States water depth of more than 8,600 ft (2,600 m), peak production of 100,000 boe/d, 15 wells tied back by subsea infrastructure, 480 million boe estimated recoverable, replication of 99% of the hull design and 80% of the topsides from Vito, Shell Offshore 60% operator and Chevron U.S.A. 40%. https://www.prnewswire.com/news-releases/shell-starts-production-at-whale-in-the-us-gulf-of-mexico-302347169.html

Tannistha Maiti
Tannistha Maiti

Senior AI Researcher

More from EarthScan

Related research

All insights →
40 Monitors, One Base: What Sets a 4D Noise Floor
Insight

40 Monitors, One Base: What Sets a 4D Noise Floor

The Subsurface Data Paradox: Why Oil & Gas Is Sitting on AI's Richest Training Ground
Insight

The Subsurface Data Paradox: Why Oil & Gas Is Sitting on AI's Richest Training Ground

Sensitivity Buys Sources, Revisit Buys Mass: What a Methane Detection Limit Is Actually For
Insight

Sensitivity Buys Sources, Revisit Buys Mass: What a Methane Detection Limit Is Actually For

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.