Three operators made the same bet at three different times, and the public record of each is specific enough to argue with.
bp installed the life of field seismic system on Valhall in 2003: over 100 kilometres of cable in parallel lines buried one metre down in the seabed, more than 10,000 sensors, roughly 45 square kilometres of coverage or about 60% of the field, and surveys once to three times a year since November 2003 [1]. ConocoPhillips followed at Ekofisk, with 200 kilometres of permanent ocean bottom seismic cable installed from March 2010, a three-dimensional four-component fibre optic system using Bragg-grating sensing elements, a first acquisition scheduled for September 2010 and a second for April 2011, in a system designed for acquisition twice a year [2]. Equinor put 380 kilometres of fibre optic seismic cable and more than 6,500 acoustic sensors over more than 120 square kilometres on Johan Sverdrup, supplied by Alcatel Submarine Networks, in service of a 70% recovery ambition [3].
The argument all three make is an argument about repeat frequency. Trench the receivers once and the marginal cost of the next survey is a source vessel rather than a full deployment, so you shoot more often, and you see the reservoir move sooner. That argument is sound as far as it goes. What follows is about the part it does not cover, which is what happens to the noise when you start pooling those repeats.
What the published work actually scores
The quantity the whole field reports is NRMS. Lumley's tutorial states the reason it exists: non-repeatable noise arises because we cannot perfectly repeat the seismic imaging experiment and its environmental conditions from one survey to the next, and the standard measurement differences two datasets and compares the energy of the difference to the energy of each individual dataset [4]. He gives the calibration points a reader needs. Perfect repeatability reads zero. Two datasets that are pure random uncorrelated noise read about 1.4. Values of roughly 0.4 to 0.6 are considered good and values below 0.2 excellent [4].
Burren and Lecerf give the algebraic form, defining NRMS as the normalised RMS of the difference between base and monitor traces in their equation (1), and attributing the sensitivity literature for the metric to Landro (1999), Kragh and Christie (2002) and Eiken et al. (2003) [5]. Their paper is about a different problem, the bandwidth dependence of NRMS, but the definition is the one everybody uses.
Estimators of two kinds are now pointed at these streams. Alali, Kazei, Sun and Alkhalifah train a recurrent network on the non-reservoir part of the data to match time-lapse surveys, in place of a conventional cross-equalisation matching filter, and measure repeatability using normalised root mean square and predictability metrics [6]. Romero, Luiken and Ravasi attack the Sleipner 4D dataset with a joint inversion and segmentation algorithm built on convex optimisation with Total-Variation and segmentation priors, which is a regularised inversion rather than a learned method, and they state that the approach mitigates non-repeatable noise by exploiting similarities across multiple surveys [7].
That last clause is the interesting one, because it names the aggregation step. Pooling several surveys to suppress non-repeatable noise is the move. The question this piece asks is what the covariance of that pool actually looks like when the surveys are acquired the way a permanent array acquires them.
The algebra of a shared base
Take the conventional programme. One base survey , then monitor surveys , and every monitor differenced against that same base. Write the non-repeatable noise on the base as and on monitor as , take them zero-mean, mutually independent and of equal variance , and set the time-lapse signal aside so the arithmetic is about noise alone.
The th difference carries noise , so its variance is and its RMS is . That is the same factor of root two behind the value of about 1.4 that Lumley quotes for two datasets of pure uncorrelated noise [4], which is a useful check that the model is describing the metric people report and not a private abstraction.
Now take any two different monitors. Their differences share exactly, with the same sign:
One half, and it does not depend on . Average the differences and the monitor terms average down while the base term does not move at all:
As grows this falls to and stops. The stacked difference noise goes from to , a total improvement of dB, for any repeat count whatsoever. The same statement in effective-sample terms uses the standard correlated-average formula with :
Forty surveys against one base are worth 1.95 independent repeats. Not 40, and not the 6.3-fold noise reduction a root-N reading of that count promises, but slightly under two repeats.
The distribution of that 3.01 dB across the programme is the part that decides a budget. The first difference is the reference point at 0 dB. The second monitor brings 1.25 dB. The fourth has taken 2.04 dB, which is 68% of everything on offer. Going from the fourth survey to the fortieth, 36 more acquisitions, moves the number from 2.04 dB to 2.90 dB. That is 0.86 dB for 36 surveys.
The lever is the reference, not the repeat count
Generalise the base. Let the reference be a stack of independent base acquisitions rather than one, so its noise variance is . The stacked difference noise, normalised so that the single-monitor single-base case reads 1.00, is
Two things follow immediately. The floor moves as , so it is set by the reference volume and by nothing else. And the naive expectation, , is not a wrong law at all: it is exactly , the case where the reference is re-shot as often as the monitors. Anyone quoting for a fixed-base programme has silently assumed an acquisition design that programme does not have.
Set the reference depth to 1 and drag the monitor count from 4 to 40. The gain readout climbs from 2.04 dB to 2.90 dB and the aqua curve never reaches the amber floor at 0.707, while the steel curve beside it has fallen to 0.158 and is promising 16.02 dB. Now leave the monitor count at 40 and drag the reference depth instead. At the same programme reads 0.274 and 11.25 dB, and the effective repeat count moves from 1.95 to 7.5. The control that moved the answer is the one nobody buys.
The toggle makes the identity explicit. Turn on a re-shot reference and the aqua curve lands on the steel one, because the reference depth is then tracking the monitor count and . Note what the toggle does not do: the steel curve is and never moves under any control in the instrument, and the monitor count slider never reshapes either curve, it only moves the marker along them and drives the readouts.
What the array actually buys
A geophysicist will object at this point, correctly, that the whole purpose of trenching receivers is to make small in the first place. That objection is right and it is worth stating precisely, because it is a different quantity from the one above.
A permanent array removes receiver positioning error and receiver coupling variation between vintages, which are among the largest contributors to the acquisition-geometry sensitivity the NRMS literature documents [5]. It lowers per survey. It does not touch , because is a property of the differencing architecture and not of the hardware. And it leaves the sources moving: an ocean bottom installation fixes the receivers, while source positioning, source signature and sea state stay non-repeatable, which is the general condition Lumley describes when he says the imaging experiment and its environmental conditions cannot be perfectly repeated [4].
So the two facts sit side by side and point in different directions. Permanent instrumentation gives you a lower noise level per survey and a cadence that resolves faster reservoir dynamics. It does not give you an unbounded stacking gain, and any claim that a survey count of 40 is worth in noise reduction against a fixed base is off by a factor of 4.5 in amplitude.
Where we have hit this ourselves
We have never delivered a 4D monitoring programme, and nothing here is drawn from one. Our exposure to this algebra is adjacent, from aggregation steps in our own segmentation work, and it has the same shape.
In our segmentation ensembling work on raster well-log curves we published the constraint plainly: every member of our ensemble sits in the same encoder-decoder lineage, our diversity comes from the loss rather than from the architecture, and the operating rule we arrived at was to stop adding members the moment a new one stops moving the hardest curve [8]. That rule is the empirical face of the same covariance. Members that share a lineage share errors, the aggregate improves at the rate the shared component allows, and the member count runs on long after the improvement has stopped.
Test-time augmentation was the sharper version, because there the correlation is structural rather than incidental. Eight augmented forward passes go through one frozen set of weights, so they are eight looks at the same estimator, not eight estimators. We wrote at the time that a vote where every variant agrees on the wrong answer recovers nothing [9]. Replace "one frozen set of weights" with "one base survey" and you have the 4D case exactly: the shared component is the thing the average cannot reach.
What to ask a vendor, and what to ask an acquisition team
For anyone buying a learned or regularised 4D estimator, the question is not what NRMS it reaches on one monitor. It is what its aggregation step assumes about the covariance between difference volumes. A method that pools surveys and reports an improvement should say whether its error model treats those differences as independent. If it does, and the programme differences against a fixed base, its uncertainty estimate is optimistic by a factor that grows with and it will quietly overstate the confidence on small time-lapse anomalies.
For the acquisition side, the question is whether the reference can be deepened. Re-processing several early vintages into a single stacked reference costs processing time rather than vessel time, and it moves the floor as where more repeats move it not at all. On a field that has been shooting since November 2003 at up to three surveys a year [1], the volumes needed to make are already on disk. The binding constraint is not the archive, it is the reservoir. The model behind that floor takes the base volumes to differ from one another only in non-repeatable noise, and vintages separated by real production do not: stack across them and the reference sits at a time-averaged reservoir state, which shifts every 4D difference by the same amount instead of adding noise that averages away. So the reference has to be stacked from a window over which the field moved by less than the anomaly being hunted, and even at the fastest cadence [1] reports, eight volumes span more than two years of production. Where the quiet window holds eight, that is the cheapest 8.3 dB in the programme and it is an architecture decision rather than a modelling one. Where it holds three, the same forty-monitor programme reads 7.47 dB instead of 2.90 dB, which is four and a half decibels no further monitor can buy.
The inversion is worth stating flatly. The permanent array buys cadence, and cadence is the right thing to buy for a fast-moving reservoir. Cadence alone cannot buy signal to noise past 3 dB against a fixed reference.
Takeaways
- bp's Valhall installation (2003, over 100 km of cable buried one metre into the seabed, more than 10,000 sensors, roughly 45 square kilometres, one to three surveys a year since November 2003), ConocoPhillips at Ekofisk (installation from March 2010, 200 km of permanent ocean bottom cable, designed to be shot twice a year) and Equinor on Johan Sverdrup (380 km of fibre, more than 6,500 sensors, more than 120 square kilometres) all make the same case: repeat frequency.
- When every monitor is differenced against the same base, the base survey's non-repeatable error appears in every difference with the same sign. Under equal-variance independent survey noise the correlation between any two 4D differences is exactly one half, independent of the repeat count.
- Stacking N monitor differences against one base takes the noise from root-two sigma to sigma and no further: 3.01 dB in total for any N, and an effective independent repeat count that caps at two. Forty surveys against one base are worth 1.95 independent repeats.
- Roughly two thirds of that 3.01 dB is already spent by the fourth survey. Going from four monitors to forty, 36 more acquisitions, is worth 0.86 dB.
- The floor is set by the reference volume and moves as one over the square root of the reference stack depth. At N = 40 the difference between a single base and an eight-deep reference is 2.90 dB against 11.25 dB, and on a field shooting since 2003 the volumes are already acquired, provided the eight are drawn from a window over which the reservoir has not moved.
- The naive one-over-root-N curve is not wrong, it is the special case where the reference is re-shot as often as the monitors. A permanent array lowers the per-survey noise level and does not change the correlation, because the correlation is a property of the differencing architecture.
- Our own record here is adjacent rather than direct: we have never delivered a 4D monitoring programme, but the same covariance capped our segmentation ensembles, whose members share one encoder-decoder lineage, and our test-time-augmentation voting, where eight augmented passes share one frozen set of weights.
Limitations
The model behind every number in this piece is deliberately plain: zero-mean, mutually independent non-repeatable noise of equal variance on the base and on each monitor. Real survey noise breaks all three assumptions, and mostly in the direction that makes the result stronger rather than weaker. Seasonal near-surface velocity variation, which is one of the effects Alali and colleagues set out to match away [6], correlates monitors acquired in the same season with each other, which adds a second shared component on top of the base term and lowers the effective repeat count further. Survey noise levels are not equal across vintages, so the equal-variance normalisation is a convenience. Cross-equalisation and matching filters reduce the shared component before any stacking happens, and the residual after matching is what this arithmetic should be applied to, not the raw difference.
The time-lapse signal was set aside throughout. Averaging N monitor differences also averages the reservoir change across those monitors, so the stack described here estimates a time-averaged change or a noise floor, not a per-vintage change. The same objection lands on the reference, and it lands harder there, because the reference is the side this piece recommends acting on. Stacking k base vintages assumes those k volumes differ from one another only in non-repeatable noise. If the reservoir moved between them, the stack is a reference at a time-averaged state, and every 4D difference then measures change from that average rather than from a single date, which is a bias shared by all N differences rather than noise that averages away. That is why the k worth buying is bounded by the length of the window over which the field was quiet, and not by the number of volumes on disk.
The three field sources differ in provenance and none of them is an audited disclosure. [1] is a documentation project of the Norwegian Petroleum Museum, with its text credited to Gunleiv Hadland of that museum and its own claims footnoted to trade press and to a 2003 Offshore Europe paper, so it is not a bp publication. [2] is an article in the trade magazine GEO ExPro, bylined Per Gunnar Folstad of ConocoPhillips and published on 3 July 2010, so it is an operator's account of a system that was then still being installed, and its twice a year figure is a design intention rather than a delivered record. Only [3] is an operator's own release, and it is a contract announcement. None of the three reports an NRMS value for the field it describes, and no NRMS value in this piece is attributed to any of them. The dB figures are arithmetic consequences of the stated model, not measurements from any field. Our own ensembling and test-time-augmentation material is from raster well-log digitisation, so it generalises here as a shape of argument, not as a number.
References
[1] Norwegian Petroleum Museum. Life of field seismic system on Valhall. Industrial heritage Valhall, a documentation project of the museum; text credited to Gunleiv Hadland, published 25 June 2019 and updated 10 August 2020. valhall.industriminne.no
[2] Folstad, P. G. (ConocoPhillips). Monitoring of the Ekofisk Field. GEO ExPro, 3 July 2010. geoexpro.com
[3] Equinor (17 January 2018). Johan Sverdrup: contract for fibre optic seismic cable with Alcatel Submarine Networks. equinor.com
[4] Lumley, D. (October 2009). 4D seismic monitoring of subsurface fluid flow. CSEG RECORDER, Vol. 34 No. 08. csegrecorder.com
[5] Burren, J. and Lecerf, D. (Petroleum Geo-Services). Repeatability Measure for Broadband 4D seismic. Second EAGE/SBGf Workshop on Broadband Seismic, Rio de Janeiro, 4 to 5 November 2014. tgs.com
[6] Alali, A., Kazei, V., Sun, B. and Alkhalifah, T. (2022). Time-lapse data matching using a recurrent neural network approach. arXiv:2204.00941. arxiv.org
[7] Romero, J., Luiken, N. and Ravasi, M. (2023). Seeing through the CO2 plume: joint inversion-segmentation of the Sleipner 4D Seismic Dataset. arXiv:2303.11662. arxiv.org
[8] EarthScan. Voting Models: Ensembling Segmenters for Cheaper, Steadier Labels. earthscan.io
[9] EarthScan. Augmenting at Inference: Test-Time Augmentation for Noisy Field Scans. earthscan.io




