Skip to main content
Reading viewAll insights →
BLOG15 min read

The Saving Is Linear in the Sparsity, and the Sparsity Is the Geology

The Saving Is Linear in the Sparsity, and the Sparsity Is the Geology
Tarry Singhby Tarry SinghFounder & CEO · 15 Sep 2026
Share

ConocoPhillips says its Compressive Seismic Imaging technology has run on roughly 30 seismic surveys since 2015 for more than 250 million dollars of direct cost reductions. Shell and SparkCognition say a generative model rebuilt subsurface images from as little as one per cent of the usual shot count in completed field trials. The theory both claims stand on says something the releases do not. Herrmann's abstract puts it plainly: costs no longer grow significantly with resolution and dimensionality of the survey area, but instead depend on transform-domain sparsity only. That is not small print, it is the whole economics. Take the oversampling ratio of about five that Herrmann measured on his own spiky-deconvolution experiment and the saving is one minus five times k over N: 90% at a sparsity of 0.02, 50% at 0.10, nothing at 0.20. Transform-domain sparsity is a property of the wavefield, which is to say of the rock, so a compression ratio demonstrated in one basin is a statement about that basin.

ConocoPhillips says that its Compressive Seismic Imaging technology achieves higher seismic data quality, at lower cost, in a shorter time frame and with less exposure, that approximately 30 seismic surveys have used it since 2015, and that direct cost reductions have exceeded 250 million dollars [1]. Jie Zhang, quoted on the same page, says the technology allows the industry to produce these surveys in less time, with less shots and receivers, and with less of an environmental impact [1]. What the page does not print is a sampling-efficiency figure. There is no percentage on it, and nobody should attribute one to ConocoPhillips.

Shell prints a number. On 17 May 2023 SparkCognition and Shell announced a collaboration whose proprietary generative approach uses deep learning to generate reliable subsurface images using far fewer seismic shots, as little as one per cent in completed field trials, than traditionally necessary while preserving subsurface image quality [2]. We have cited that claim before, in a piece about why exploration AI cannot prove itself on drilling outcomes at 64 high-impact wells a year [3]. This piece is about a different question. Not whether the one per cent is real, but what kind of quantity it is.

Because the theory underneath both claims is explicit about that, and it is not the answer a procurement conversation expects.

The sentence in the abstract

Felix Herrmann's 2010 GEOPHYSICS paper on randomised sampling is the reference the whole compressive-acquisition line descends from. Its abstract does not describe a fixed compression ratio. It describes a change in what the cost depends on. The main outcome, it says, is a technology where acquisition and processing related costs are no longer determined by overly stringent sampling criteria such as Nyquist, and then the load-bearing clause: costs no longer grow significantly with resolution and dimensionality of the survey area, but instead depend on transform-domain sparsity only [4]. Later in the same abstract it repeats the point in operational form: sampling, storage and processing costs scale with transform-domain sparsity [4].

Read that as an engineer reads a spec and it says the acquisition bill has been re-parameterised. It used to be set by how big and how fine your survey is. It is now set by how compressible the wavefield is. The first quantity is a decision. The second is a property of the Earth under the survey.

What the bound actually promises

The recovery guarantee in the paper is printed as Equation 6. For a compressive sampling matrix with the right incoherence, an exactly sparse vector is recovered from nn measurements out of an ambient NN as long as it carries no more than kk nonzeros, with

Herrmann 2010, Equation 6
k    Cnlog2(N/n)k \;\le\; C\,\frac{n}{\log_{2}(N/n)}

and CC a moderately sized constant [4]. The sentence immediately after it is the one that matters for a budget, because it inverts the bound into a per-coefficient price: recovery of kk nonzeros only requires an oversampling ratio of n/kClog2Nn/k \approx C \log_{2} N, as opposed to taking all NN measurements [4].

So the theory does not hand you a compression ratio. It hands you a ratio of measurements per significant coefficient, weakly dependent on the ambient size through a base-two logarithm and multiplied by a constant the theory does not pin down. That is why the exhibit below makes n/kn/k a control rather than a number.

It does, however, leave one empirical anchor in the same paper. In the spiky-deconvolution experiment, comparing a Ricker source against a randomised one at a fixed measurement count, recovery from the randomised sampling is successful as long as the number of spikes does not exceed 11, which the paper reports as k/n0.2k/n \approx 0.2 with n=54n = 54 [4]. Invert it and the measured oversampling ratio is about five measurements per nonzero. That number is a threshold from one carefully built one-dimensional experiment, not a constant of nature. It is also the only number in the paper of the right shape to price a survey with, so it is where the exhibit opens.

The arithmetic, on one line

Write the sparsity as a fraction of the ambient dimension and the rest is division. If you need n/kn/k measurements per significant coefficient and there are kk of them out of NN, then

The survey saving, in the two quantities that set it
nN  =  nkkN,saving  =  1nkkN\frac{n}{N} \;=\; \frac{n}{k}\cdot\frac{k}{N}, \qquad \text{saving} \;=\; 1 - \frac{n}{k}\cdot\frac{k}{N}

and the saving hits zero at

The sparsity at which compressive acquisition stops paying
kN  =  1n/k\frac{k}{N} \;=\; \frac{1}{n/k}

which at five measurements per nonzero is a sparsity of one fifth. There is no curve here, no diminishing return, no regime where the method quietly keeps helping. A straight line to zero, whose slope is the measurements-per-nonzero ratio and whose root is one over it. Neither of those is a procurement decision. The ratio is set by the recovery constant, and the sparsity by the wavefield.

Two things about that line before anyone prices a survey with it. nn is a count of measurements, and in acquisition a measurement is a source position, a receiver position or both, so the same arithmetic prices a shot-sparse design and a receiver-sparse one without caring which. And kk is a property of the sampled wavefield rather than of the rock alone, so acquisition geometry and source signature move it too. Neither of those makes the quantity a knob. They make it something you have to measure on the specific survey you are about to shoot.

The exhibit

The left panel exists so that kk is a place on a decay curve before it is a symbol in a bound. The right panel prices that place. The bottom strip is the mechanism, computed rather than illustrated: a three-spike toy wavefield, sampled at the count the two sliders imply, transformed back with the positions you did not take set to zero.

COMPRESSIVE ACQUISITION, PRICED IN TRANSFORM-DOMAIN SPARSITY90.0%SURVEY SAVING AT THIS SPARSITY AND OVERSAMPLING RATIOWhere the cut sits on the wavefieldmagnitude-sorted transform coefficients, largest first10.10.010.000.050.100.150.200.250.30coefficient index, as a fraction of Ncut at k/N 0.020smallest kept 0.104 of largestWhat that cut is worthshots fired and survey saving, both as a fraction of Nyquist0%25%50%75%100%0.000.050.100.150.200.250.30saving zero at 0.200k/N 0.020transform-domain sparsity, k/NRandomised decimationbackprojected spectrum, 32 of 256 positions keptwavenumber across the bandyou set:  sparsity k/N 0.020   ·   oversampling n/k 5.0the survey:  shots 10.0% of Nyquist   ·   saving 90.0%   ·   saving reaches zero at k/N 0.200the artefact:  worst spurious peak 0.61 of the tallest true coefficient   ·   90% of the leaked energy sits in 59.7% of the spurious binssurvey savingother n/k, 2 to 12shots firedsaving reaches zerotrue coefficientleaked energySaving is one minus n/k times k/N, from Herrmann 2010. The decay curve and the three spikes are illustrative; the strip snaps to a power-of-two measurement count.
decimation rule
Left: the magnitude-sorted transform-domain coefficient spectrum on a logarithmic magnitude axis, with the sparsity cut standing on it, so k is a place on a decay curve before it is a number in a formula. Right: shots fired and survey saving against that same k/N, over a fixed fan of six faint reference lines at oversampling ratios of 2, 4, 6, 8, 10 and 12, with the live ratio drawn bold. Bottom: the backprojected wavenumber spectrum of a three-spike toy wavefield, transformed from the positions the pill kept. The exhibit opens at k/N of 0.02 and an oversampling ratio of five, which is what Herrmann's own spiky-deconvolution experiment measured, and reports a saving of 90.0%. Drag the sparsity control and the saving falls along a straight line: 50.0% at k/N of 0.10, and nothing at all at 0.20, where the amber marker stands. Drag the oversampling control instead and the live line swings across that fixed fan, landing on each faint line in turn, because its intercept is one over the ratio you set. The pill changes neither number. It changes only what the missing shots did to the record: regular decimation puts a coherent replica at the full height of a true coefficient, and, at the opening settings, randomised decimation spreads leakage of almost exactly the same total energy across the band, which the readout measures as the share of the spurious bins holding 90% of it. The decay curve and the three spikes are illustrative. The identity that prices the survey is not.

It opens at a transform-domain sparsity of 0.02 and five measurements per nonzero, and reports a survey saving of 90.0%. Drag the sparsity control right. At 0.10 the saving is 50.0%. At 0.20 it is zero, and the amber marker on the axis is standing exactly there. Push the oversampling control instead and the live line swings across a fixed fan of six reference ratios: at twelve measurements per nonzero the saving is gone by a sparsity of 0.083, and at two it survives past the right-hand edge of the axis.

Nothing in that is a claim about acquisition technique. It is a claim about the wavefield. The method is cheapest where the subsurface was already compressible, which is to say where it was already easy to image, and worth nothing where it is not.

That is the opposite of where the acquisition budget goes.

Why sparsity is geology and not a setting

The objection a geophysicist will raise here is that sparsity is not a fixed property of the rock, it is a property of the transform, and you can choose a better one. That is right, and it does not close the question, because the choice is bounded by what the wavefield is doing.

Herrmann's own account of the mechanism is about the sampling rule rather than the dictionary. Subsampling produces spectral leakage, where energy from each frequency is leaked to other frequencies, and the amount depends on the degree of subsampling. The move the paper makes is to change the character of that leakage rather than its size: the more uniformly random the sampling is, the more the leakage behaves as zero-centred Gaussian noise spread over the entire frequency spectrum [4]. The abstract calls this breaking subsampling related interferences by turning them into harmless noise, which is subsequently removed by promoting transform-domain sparsity [4].

The strip at the bottom of the exhibit is that sentence, computed. Switch the pill to regular decimation at the opening settings and the worst spurious peak stands at 1.00 of the tallest true coefficient, an exact alias replica you cannot tell from signal, with 90% of the leaked energy sitting in 7.1% of the spurious bins. Switch to randomised and the worst peak falls to 0.61 and 90% of the leaked energy is spread across 59.7% of the spurious bins. Between those two pictures the total leaked energy changes by four parts in a thousand. It does not remove leakage. It makes leakage incoherent, and incoherent is what a sparsity-promoting solve can subtract.

Which means the whole scheme is conditional on the second half of that sentence working. Removing the noise by promoting transform-domain sparsity requires the wavefield to be sparse in the transform you have. A curvelet frame is sparse on wavefields with coherent, smoothly curving events. Complex overburden, strong multiples, steep dips, scattering, a shallow gas cloud: those are the conditions that spread energy across the transform, raise kk, and walk the cut in the left panel to the right. They are also, precisely, the conditions that make a survey expensive enough to be worth shooting.

What that does to a vendor claim

None of this says the numbers the operators publish are wrong. It says they are basin statements.

ConocoPhillips's roughly 30 surveys since 2015 and more than 250 million dollars of direct cost reduction is an aggregate over a portfolio the company chose [1]. Shell and SparkCognition's one per cent came out of completed field trials [2], and a field trial is one wavefield. Neither release tells you the transform-domain sparsity of the volumes involved, which is the quantity the theory says the saving is a function of. Without it, a ratio measured in one place is not transferable, in either direction: it does not promise you the same saving, and it does not deny you a better one.

There is a second correction that runs the same way. The exhibit prices the sampling term only. Real acquisition cost also carries vessel or crew mobilisation, permitting, marine mammal observers, weather standby and the processing bill, and none of that scales with the shot count. TotalEnergies and ADNOC announced in November 2019 a pilot in which seismic sensors would be dropped by six autonomous aerial drones and later retrieved by an unmanned ground vehicle, in a 36 square kilometre desert environment in Abu Dhabi, without human intervention and therefore at a lower cost [5]. That release says nothing about compressive sensing or randomised sampling [5], and it should not: it is attacking a different term of the same bill, the cost per sensor placement rather than the number of placements. Fixed costs mean the realised saving is always less than the sampling saving the exhibit prints. The straight line is a ceiling.

We have paid for this lesson twice, on our own data

A ratio measured on one corpus is not a constant of the method. We know that because we published both times it cost us.

The first time, on a borehole-fracture detector built during a roughly twenty-month engagement with a mid-sized Middle East carbonate operator we partnered with, we ballooned the augmented training set to 92,000 overlapping patches and the model got worse, while adding three real wells lifted depth, dip and azimuth accuracy [6]. The augmentation ratio that had helped at one scale was not a property of the augmentation. It was a property of how much real distribution the corpus already held.

The second time, on the same programme, we added a fifteenth well to a training set that was working and validation F1 at a 5 cm threshold fell from about 60% to about 57. The well was not corrupt. Its ground truth had been picked with a different emphasis, with pick gaps running to 175 m, so the extra rows pulled the label distribution off the style the model had settled on across the first 14 wells [7]. The per-well gain that had held for 14 wells was not a constant of the pipeline either.

Both are the same failure as reading a compression ratio off one basin. A number measured on one corpus, under one set of conditions, gets promoted to a property of the technique, and then it is used to size the next thing. The compressive-acquisition case is cleaner than ours, because the theory tells you in advance which quantity the number is a function of, and prints it in the abstract.

Three questions worth asking

If a vendor or an internal team proposes compressive acquisition on a specific survey, the useful questions are not about the algorithm.

What is the transform-domain sparsity of a comparable volume from this basin, in the transform the recovery will actually use, and how was it measured? What oversampling ratio did the reference project use, and was it a demonstrated recovery threshold or an assumption? And of the acquisition budget for this survey, what fraction scales with the shot count at all?

The first question has an answer, because sparsity is measurable on legacy data before a single shot is fired: take a densely sampled volume from the same area, transform it, sort the coefficients by magnitude and find where the decay puts your reconstruction error inside tolerance. That is a week of work on data the operator already owns, and it converts the entire proposal from a compression ratio someone else measured into a number about this rock.

The last question is the one that most often changes the answer, and it is a procurement question rather than a geophysical one.

Key takeaways

  1. ConocoPhillips says Compressive Seismic Imaging has run on approximately 30 seismic surveys since 2015 with direct cost reductions exceeding 250 million dollars, and that the technology means less shots and receivers. The page prints no sampling-efficiency percentage, so no compression ratio should be attributed to it.
  2. Shell and SparkCognition announced on 17 May 2023 that their generative approach reconstructed subsurface images from as little as one per cent of the usual shots in completed field trials. A field trial is one wavefield, and the release does not report the transform-domain sparsity of it.
  3. Herrmann's 2010 abstract states the economics directly: costs no longer grow significantly with resolution and dimensionality of the survey area, but instead depend on transform-domain sparsity only. The recovery bound is k less than or equal to C n over log base two of N over n, and the sentence after it inverts that into an oversampling ratio of n over k.
  4. At the oversampling ratio of about five that the paper's own spiky-deconvolution experiment measured, k over n approximately 0.2 with n of 54, the survey saving is one minus five times k over N. That is 90% at a sparsity of 0.02, 50% at 0.10, and zero at 0.20.
  5. Randomised sampling does not remove leakage, it makes leakage incoherent. In the exhibit's strip at the opening settings, regular decimation puts 90% of the leaked energy in 7.1% of the spurious bins with a full-height alias replica, and randomised decimation spreads 90% of it across 59.7% of them, while the total leaked energy changes by four parts in a thousand.
  6. The exhibit prices the sampling term only. TotalEnergies and ADNOC announced in November 2019 a drone-and-rover pilot in a 36 square kilometre Abu Dhabi desert environment, aimed at the cost per sensor placement instead. The release mentions neither compressive sensing nor randomised sampling. Fixed costs make the realised saving smaller than the sampling saving, so the straight line is a ceiling.
  7. Sparsity is measurable on legacy data before a survey is shot. Transform a densely sampled volume from the same area, sort the coefficients by magnitude, and find where the decay puts reconstruction error inside tolerance. That turns someone else's compression ratio into a number about this rock.

References

[1] ConocoPhillips. Geophysicist Chengbo Li honored for seismic innovations. SPIRIT Now. https://www.conocophillips.com/spiritnow/story/geophysicist-chengbo-li-honored-for-seismic-innovations/

[2] SparkCognition and Shell. SparkCognition and Shell Announce a Technology Collaboration Aimed at Accelerating the Pace of Exploration Through the Use of Generative AI. PR Newswire, 17 May 2023. https://www.prnewswire.com/news-releases/sparkcognition-and-shell-announce-a-technology-collaboration-aimed-at-accelerating-the-pace-of-exploration-through-the-use-of-generative-ai-301826717.html

[3] EarthScan. Sixty-Four Wells a Year: Why Exploration AI Cannot Prove Itself on Outcomes. https://earthscan.io/insights/sixty-four-wells-a-year-exploration-ai-cannot-prove-itself-on-outcomes

[4] Herrmann, F. J. Randomized sampling and sparsity: getting more information from fewer samples. GEOPHYSICS 75(6), WB173 to WB187, 2010. Open preprint, University of British Columbia Technical Report TR-2010-01. https://slim.gatech.edu/Publications/Public/Journals/Geophysics/2010/herrmann2010GEOPrsg/herrmann2010GEOPrsg.pdf

[5] TotalEnergies. Abu Dhabi: ADNOC and Total Innovate in the Field of Seismic Acquisition with the Use of Unmanned Drones and Vehicle. 13 November 2019. https://totalenergies.com/media/news/press-releases/abu-dhabi-adnoc-and-total-innovate-field-seismic-acquisition-use-unmanned-drones-and-vehicle

[6] EarthScan. The 92k-Patch Trap: When More Synthetic Data Made the Model Worse. https://earthscan.io/insights/92k-patch-trap-synthetic-data-overload

[7] EarthScan. The 15th Well Made Our Model Worse: Label Consistency Beats Data Volume. https://earthscan.io/insights/the-15th-well-made-our-model-worse-label-consistency-beats-data-volume

Tarry Singh
Tarry Singh

Founder & CEO

More from EarthScan

Related research

All insights →
Four Orders of Magnitude of Archive Cost 1.9 Units of Discrimination
Insight

Four Orders of Magnitude of Archive Cost 1.9 Units of Discrimination

40 Monitors, One Base: What Sets a 4D Noise Floor
Insight

40 Monitors, One Base: What Sets a 4D Noise Floor

64 Wells a Year: Why Exploration AI Cannot Prove Itself on Outcomes
Insight

64 Wells a Year: Why Exploration AI Cannot Prove Itself on Outcomes

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.