Skip to main content
Reading viewAll insights →
BLOG18 min read

More Wells Is Not More Signal: The Partner Share Where Pooled Exploration Data Starts to Pay

More Wells Is Not More Signal: The Partner Share Where Pooled Exploration Data Starts to Pay
Tarry Singhby Tarry SinghFounder & CEO · 25 Sep 2026
Share

On 24 August 2026 Equinor, Aker BP and Vår Energi said they would combine expertise, data, technology and exploration capacity on the Norwegian continental shelf, across 20 to 25 opportunities in four to five years. The release says nothing about AI; the question here is ours. ExxonMobil's second-quarter prepared remarks, out on 31 July, say a model trained on its Guyana discovery data identified nearly 90% of known discoveries in validation and has since flagged four new Guyana prospects. PETRONAS said on 3 September that agentic AI will be added to its myPROdata platform, with nothing yet live. Pooling raises the well count and it also raises the number of picking conventions in the labels, and our own ablation measured what that costs: class error fell from 93.115% at 3 wells to a floor of 0.817% at 11, then rose to 2.536% at 14 when three wells picked in a second hand came in. Decompose that one reversal and the pooled set behaves like a clean set the size of its majority convention plus 8.02 points of class error per unit minority share. The instrument computes the surface that follows, and it has a hard rule in it: below a partner share of one half the pooled error is your own error plus a strictly positive term, so pooling cannot win until the partner brings more wells than you have. At our measured pool size that crossing sits at 52.0%, and re-picking every partner well to the shared standard moves it to 21.4%.

On 24 August 2026 Equinor, Aker BP and Vår Energi announced a strategic exploration collaboration on the Norwegian continental shelf, and all three companies published it on the same day [1][2][3]. The companies will "combine expertise, data, technology and exploration capacity" across approximately 20 to 25 opportunities over the next four to five years, with an ambition of around five high-impact exploration wells a year [1]. Kjetil Hove, Equinor's executive vice president for Exploration and Production Norway, says that "by combining our expertise and exploration portfolios, we can high-grade the best opportunities, move faster and do more" [1]. Karl Johnny Hersvik, Aker BP's chief executive, says that "by bringing together three strong exploration organisations, we can pursue more of the most promising opportunities and increase the chances of significant new discoveries" [2]. Nick Walker, Vår Energi's chief executive, calls it "a transformative step for the future exploration of high-impact opportunities on the NCS" [3].

The release says nothing about artificial intelligence, machine learning, digital tools or seismic. Those words are not on the page. The question in this piece is ours and not theirs, and it starts from the one word that is on the page: data.

Two days later Equinor's own ONS programme listed a stand talk called "Industrial AI turning data into barrels", 10:30 to 11:00 on Wednesday 26 August [4]. That is a title on an event page with no speaker, abstract or figure attached, and it is cited here as a title and nothing more.

Two other operators did attach numbers to AI in the same season, and they attached very different ones. ExxonMobil's second-quarter 2026 prepared remarks, the document that accompanied its results release of 31 July 2026 [6], say the company has "created sophisticated exploration models powered by both AI and the world's largest seismic database", that it "trained one model using our Guyana discovery data, and validation testing showed it could identify nearly 90% of known discoveries", and that "to date, this model identified four new Guyana prospects" [5]. On 3 September PETRONAS said its myPROdata upstream data platform will integrate geological, geophysical and engineering data with agentic AI that "can autonomously analyse and connect multiple data sources", with Iraya Energies as Consortium Lead Integrator [8]. Nothing in that release is described as running.

Five operators, three announcements, and one training-set question none of them is obliged to answer. When three exploration organisations pool their data, the well count goes up. So does the number of picking conventions inside the labels. Those two facts pull in opposite directions, and we have measured the second one on our own work.

Three announcements, read exactly

The alliance is an exploration collaboration, not a data platform. What the three releases commit to is a joint portfolio and a drilling ambition. The word "data" arrives inside a list of four things being combined, and no model, system or method is named anywhere on the page [1][2][3]. Read it as intent about acreage and wells, because that is what it says.

ExxonMobil's is the only first-party document among these with a validation figure in it, and its scope is narrow in ways worth stating. The nearly 90% is a validation-testing result on a model trained on the company's own Guyana discovery data [5]. The four are prospects, "ones previously not identified through traditional methods" [5], which is a statement about leads on a map and not about barrels in a tank. The document is careful about that: "It's already delivering promising results" is the strongest verb attached to them [5]. Two paragraphs matter for scope. The approximately 2 billion dollars of net present value quoted immediately above the AI passage belongs to 4D seismic and high-performance computing applied to reservoir insight, not to the exploration model [5], and the earnings release itself, published the same day, contains no mention of the model at all [6]. When the item reached the trade press on 13 August it arrived without the validation figure and attributed to the company's vice president of exploration rather than to the prepared remarks [7]. That is trade coverage, not the operator's own document, and every ExxonMobil figure here is read off the document.

PETRONAS names the substrate and says the least about deployment. myPROdata is a data platform for exploration blocks, discovered resource opportunities and producing fields in Malaysia; the agreement is that it will integrate subsurface and engineering data with agentic AI, and Iraya Energies holds the Consortium Lead Integrator role [8]. The RM50 billion to RM60 billion of annual upstream investment and the ambition to shorten discovery to first hydrocarbon from 100 months to 50 months are quoted in the release as PETRONAS' ambition, in a statement by Malaysia Petroleum Management's senior vice president, and they are not outcomes of this agreement [8].

An alliance that combines exploration portfolios combines the labels attached to them. That is the part none of these releases prices, and it is the part we have a measurement for.

One training set, more than one hand

For a mid-sized Middle East carbonate operator we ran a well-count ablation on a Detection Transformer trained to pick fractures and beddings from borehole image logs. Same architecture, same augmentation, same optimiser; the only thing that moved was how many wells went into the training pool. Validation class error fell from 93.115% at 3 wells to 18.370% at 6, to 1.055% at 9, to a floor of 0.817% at 11, and then rose to 2.536% at 14 [9].

The reversal is the useful part, and it is not noise. The three wells added between the 11-well and 14-well runs carried less consistent sinusoid picks than the core set. In a set-prediction model trained with Hungarian matching the optimal assignment is computed against the labels, so a sinusoid that one interpreter picked and another did not does not add a hard example. It teaches the matcher a contradictory target. The signature is unmistakable in the same table: the geometry losses kept falling through the reversal, from 0.083 to 0.059 on the L1 parameter loss, while classification regressed [9]. Geometry better, classification worse, is what a labelling problem looks like and not what a capacity problem looks like.

The counts underneath that reversal are the ones this piece runs on. The combined fractures-and-beddings set was consistently picked across 11 wells; the fractures-only set spans all 14 [9]. So the 14-well run is a pool of 11 wells in one picking convention and 3 in another, and the 11-well run is the same pool with the second convention removed. Two runs, one difference, and the difference is a mixture.

Two other pieces of our record set the boundaries of what follows. We validated the transfer mechanism for well-to-well correlation on the open FORCE 2020 set, 118 Norwegian Sea wells, before touching the operator's 14, on the principle that a public proxy can certify the machinery and must never be asked to certify the basin [10]. And when the same fracture detector was read blind on horizontal wells it had never trained near, its azimuth output reached 65.4 degrees of mean error against roughly 10 degrees on verticals, which we published and wrote into the deployment protocol rather than averaging away [11]. We say where our models stop. The model below is built the same way.

The arithmetic

Take a pool of NN wells, a share hh of them contributed by a partner and picked to a second convention. Let C(n)C(n) be our measured class error on nn consistently picked wells. A set-prediction model trained against these labels learns the convention that dominates them, so the wells in the other hand do not add usable examples; they supply targets that contradict the ones the matcher is optimising against. The pool then behaves like a clean set the size of its majority convention, plus a penalty that grows with the share sitting in the minority.

Wells of the dominant picking convention
M(N,h)  =  N max⁡ ⁣(h,  1−h)M(N,h) \;=\; N\,\max\!\left(h,\; 1-h\right)
The pooled class error: the clean curve at the majority count, plus what the minority contradicts
E(N,h)  =  C ⁣(M(N,h))  +  A min⁡ ⁣(h,  1−h)E(N,h) \;=\; C\!\bigl(M(N,h)\bigr) \;+\; A\,\min\!\left(h,\; 1-h\right)

That model has exactly one free parameter and exactly one calibration point, and the two happen to line up. At N=14N=14 and h=3/14h = 3/14 the majority is 14×11/14=1114 \times 11/14 = 11 wells, whose clean error we measured at 0.817%. The mixed run scored 2.536%. So

Points of class error per unit minority share, from the one mixed run we measured
A  =  2.536−0.8173/14  =  1.719×143  =  8.022A \;=\; \frac{2.536 - 0.817}{3/14} \;=\; \frac{1.719 \times 14}{3} \;=\; 8.022

Two consequences follow without any further assumption. The first is that E(N,h)=E(N,1−h)E(N,h) = E(N,1-h): a pool that is entirely the partner's wells is as internally consistent as one that is entirely yours, so the damage lives in the mixture rather than in the partner. The second is the one worth arguing about. Below a partner share of one half your own wells are the majority, so M(N,h)=N(1−h)M(N,h) = N(1-h), which is precisely the count you would have had if you had trained on your own wells alone. Subtract the two:

Below one half, pooling adds a term and removes nothing
E(N,h)  −  C ⁣(N(1−h))  =  A h  >  0for all 0<h≤12E(N,h) \;-\; C\!\bigl(N(1-h)\bigr) \;=\; A\,h \;>\; 0 \qquad \text{for all } 0 < h \le \tfrac{1}{2}

Pooling adds a strictly positive term and takes none away, for every contradiction cost above zero. That is not a consequence of our value of AA. It survives any value the reader puts on the slider.

The bench

The exhibit below computes that surface rather than drawing it. Left to right is the partner share of the pooled wells, 0 to 100%. Back to front is the pool, 3 to 30 wells. Height is validation class error on a log scale, from 0.5% at the bottom of the block to 100% at the top, so the floor sits low and the mixture rises. The amber sheet is our measured 11-well floor at 0.817%. The white curtain is your partner share and the bright line is your pool size, and the labels follow both.

Four controls. The partner share and the pool size move the cursor across the surface, and neither reshapes it. The re-picking control does reshape it: it is the share of the partner's wells re-picked to the shared standard, and it enters the surface through the effective partner share, h(1−u)h(1-u). It does not enter the comparison against your own wells, because re-picking the partner's picks does not change how many wells you brought. The fourth control is the contradiction cost, and its derivation is printed on the plate.

POOLED PICK HETEROGENEITY2.54%VALIDATION CLASS ERROR, 14 WELLS POOLED, PARTNER SHARE 21.4%pooling starts to paypartner share 52.0%partner share 21.4%, 14 wells pooled2.54% class erroryour own wells alone11 wells, 0.82%partner share of the pooled wells, 0% left to 100% rightwells pooled, 3 at the back to 30 at the frontvalidation class error, 0.5% low to 100% high, log scaleYOUR OWN WELLS ALONE0.82% on 11 wellsPOOLING AT THIS SHAREcosts 1.72 pointsPOOLING PAYS FROM SHARE52.0%RE-PICKS TO STOP LOSING100.0%class error 0.5% to 100%, logour 11-well floor, 0.82%your partner shareyour pool sizeClean curve: our measured class error at 3, 6, 9, 11 consistently picked wells (93.115, 18.370, 1.055, 0.817%), log-interpolated, flat outside.Cost: our 14-well run, 11 wells one convention and 3 another, scored 2.536%. (2.536 less 0.817) over 3/14 is 8.02 points. Ours; the slider is yours.
The surface is the validation class error of a fracture model trained on a pooled well set, over every partner share of the pool (left to right) and every pool size from 3 to 30 wells (back to front), on a log height scale. Its height is computed from our own measured ablation plus one modelled response: the pool behaves like a consistently picked set of the size of its majority picking convention, plus a penalty proportional to the share of the pool sitting in the minority convention. The amber sheet is our measured 11-well floor. The white curtain is your partner share and the bright line is your pool size, and the labels follow both. The clean curve is measured; the mixture response is our model of it. With no wells re-picked the surface is symmetric about a partner share of one half, because a pool that is entirely the partner's wells is as internally consistent as one that is entirely yours. Drag the partner share up from the opening settings and watch the pooling readout: below one half your own wells are the majority, so the pooled error is your own error plus a strictly positive term and pooling cannot win. Drag the re-picking control to 100% and the break-even share falls. Drag or use the arrow keys to orbit, Home to reset.

What the bench shows

The exhibit opens on the run we measured: 14 wells pooled, a partner share of 21.4%, nothing re-picked, the contradiction cost at our derived 8.02. The hero reads 2.54%, which is our 14-well ablation result recovered as an output of the model rather than typed into it. Your own wells alone reads 0.82% on 11 wells. Pooling at this share costs 1.72 points, which is a factor of 3.1 on the floor you were already standing on.

Now drag the pool size to 30 and watch what does not happen. The margin still reads 1.72 points and the hero still reads 2.54%. Taking the pool from 14 wells to 30 at the same partner share changes neither, because our clean curve has been flat since 11 wells and the penalty is a function of the share rather than of the count. Above the floor, extra volume at a fixed heterogeneity buys nothing at all. That is the sentence to hold when a partner offers wells.

Now drag the partner share up. At 60%, with the pool back at 14, the hero reads 5.08% and your own wells alone reads 22.81% on 5.6 wells, so pooling saves 17.73 points. The same partner data, the same picking mismatch, and the verdict has flipped, because you no longer have enough wells of your own to be on the flat part of the curve. The readout that names the crossing says pooling pays from a partner share of 52.0%. Below that the partner is not bringing enough to outweigh what the mixture costs; above it they are.

The last control is the one an alliance can actually pull. Set re-picking to 100% and the break-even share falls from 52.0% to 21.4%, which is the share at which the partner's volume first takes your own count below 11 wells. The re-picking readout at the opening settings is blunter still: it says 100.0%, meaning every one of the partner's three wells has to be re-picked before the pooled model stops being worse than the 11 you started with. It never becomes better. It stops being worse. Our clean curve has no room above 11 wells, so re-picking buys back the floor and nothing beyond it.

That is the shape of the finding, and it is uncomfortable in a specific way. Below a partner share of one half, pooling on our numbers has no upside at all; it has only a smaller or larger downside. The upside lives entirely in the case where the partner brings more wells than you have, which is exactly the case in which the model is going to learn the partner's picking convention and not yours.

What an alliance actually buys a model

The three companies said they would combine expertise, data, technology and exploration capacity [1]. The word that is not on that list is convention, and on our measured curve it is the one that decides whether the combination is worth anything to a model.

Read the two disclosed AI programmes against that. ExxonMobil trained on its own Guyana discovery data [5], which is one operator, one basin and one interpretation practice: the easy case, and the one where extra volume is not fighting extra heterogeneity. PETRONAS is proposing to put an agentic layer over a national data platform that aggregates across operators [8], which is the hard case, and the release does not say what the picking conventions underneath it look like because nothing obliges it to. The alliance sits between the two: three organisations, three exploration histories, and no public statement about whose picking standard wins.

The honest asymmetry is that this is easier to say about a borehole image log than about a seismic interpretation. A sinusoid pick is a discrete annotation with an interpreter's name on it. A horizon pick across a shelf is the same kind of object with the same kind of variance, and the release that says three organisations will combine their data is the release that has not yet said which of the three hands the labels will be in.

Where these numbers come from, and what they are not

The measured half of this bench is one engagement: 14 vertical wells in a confidential Middle East carbonate field, logged with two different microresistivity imaging tools, on a fracture and bedding detection task [9]. It is not seismic interpretation, it is not the Norwegian continental shelf, and it is not an exploration prospect ranking. The clean curve is measured; the mixture response is our model of it, calibrated on one point, and it is labelled that way on the plate.

The contradiction cost in particular is the softest number here. It is a single ratio derived from a single reversal, so the slider is not decoration. What survives any value you put on it is the structure: the flat region above 11 wells where volume stops paying, the symmetry about a partner share of one half, and the rule that below one half the pooled error is your own error plus a positive term. Those three are properties of the decomposition, not of our constants.

One more boundary. This prices one use of a partner's wells: supervised training against their picks. It says nothing about pretraining on their images without their labels, which is a different question with a different answer, and nothing about what a partner's wells are worth to a geologist rather than to a matcher.

Three questions for a pooled exploration data set

How many picking conventions are in the labels, and how many wells sit in each? Not how many wells, and not how many terabytes. The split. On our data that one count is what separates a 0.817% model from a 2.536% one, on the same 14 wells.

What does the model score on the largest single-convention subset alone? That is the number the pooled score has to beat, and it is cheap to produce because the subset is already in the pool. If the pooled model does not beat it, the pool is not a training set, it is two training sets in a bag.

What would it cost to re-pick the minority to the shared standard, and what does the model score afterwards? On our curve the answer to the second half is bounded by the floor, so the re-picking budget should be sized against being no worse rather than against being better. An alliance that agrees a picking standard before it agrees a data-sharing schedule has bought something a model can use. One that agrees the schedule first has bought volume, and volume is the part our ablation says stops paying.

Key takeaways

  1. Equinor, Aker BP and Vår Energi will combine expertise, data, technology and exploration capacity across 20 to 25 opportunities in four to five years on the NCS. The release names no AI, no model and no method; the training-set question here is ours and is attributed as ours throughout.
  2. ExxonMobil's second-quarter 2026 prepared remarks, the document accompanying its 31 July results release, say a model trained on its Guyana discovery data identified nearly 90% of known discoveries in validation and has since flagged four new Guyana prospects. The approximately 2 billion dollars of net present value in the paragraph above belongs to 4D seismic and high-performance computing, not to that model, and the results release itself does not mention it.
  3. PETRONAS says agentic AI will be integrated into the myPROdata upstream data platform with Iraya Energies as Consortium Lead Integrator. Nothing is described as running, and the RM50 billion to RM60 billion and the 100 months to 50 months figures are quoted in the release as PETRONAS' ambition, not as outcomes of the agreement.
  4. Our own well-count ablation measured class error at 93.115% on 3 wells, 18.370% on 6, 1.055% on 9, a floor of 0.817% on 11, and 2.536% on 14 once three wells picked in a second hand entered. Decomposed, the 14-well run is the 11-well floor plus 1.719 points at a minority share of 3 in 14, which is 8.02 points per unit minority share.
  5. The instrument computes the surface that follows from that decomposition. It is symmetric about a partner share of one half, and below one half the pooled error equals your own error plus a strictly positive term, so pooling cannot strictly beat your own wells until the partner brings more wells than you have. That holds for every contradiction cost above zero, not only ours.
  6. At the opening settings pooling costs 1.72 points and pays only from a partner share of 52.0%; taking the pool from 14 wells to 30 at the same share moves neither the hero nor the margin, because our clean curve is flat above 11 wells. Re-picking every partner well to the shared standard moves the crossing to 21.4% and returns the model to the floor it started on, which is the whole of what harmonisation buys.

References

[1] Equinor. Equinor, Aker BP and Vår Energi join forces to search for Norway's next major discoveries. 24 August 2026. https://www.equinor.com/news/20260824-equinor-aker-bp-and-var-energi-join-forces

[2] Aker BP ASA. Equinor, Aker BP and Vår Energi join forces to search for Norway's next major discoveries. Stock exchange announcement, 24 August 2026. https://akerbp.com/en/borsmelding/equinor-aker-bp-and-var-energi-join-forces-to-search-for-norways-next-major-discoveries-2/

[3] Vår Energi ASA. Equinor, Aker BP and Vår Energi join forces to search for Norway's next major discoveries. Stock exchange announcement, 24 August 2026. https://varenergi.no/en/newsroom/stock-exchange-announcements/?release=47ACA6E7A9D6823F

[4] Equinor. ONS 2026 programme, Equinor Talks. Talk title "Industrial AI turning data into barrels", 10:30 to 11:00, Wednesday 26 August 2026. Title only; no speaker, abstract or figure is published on the page. https://www.equinor.com/about-us/ons

[5] ExxonMobil. Second Quarter 2026 Prepared Remarks. Published with the second-quarter results, 31 July 2026. https://d1io3yog0oux5.cloudfront.net/_ffcbd7a566640bfe3df0b1e14b50bc22/exxonmobil/db/2404/22710/pdf/2Q26+Prepared+Remarks.pdf

[6] ExxonMobil Holdings Corporation. ExxonMobil Announces Second-Quarter 2026 Results. 31 July 2026. Cited for the date of the disclosure; the results release itself does not mention the exploration model. https://investor.exxonmobil.com/company-information/press-releases/detail/1208/exxonmobil-announces-second-quarter-2026-results

[7] World Oil. ExxonMobil identifies four Guyana exploration opportunities using AI. 13 August 2026. Trade coverage, not the operator's own document: it carries the four prospects without the validation figure and attributes the account to the company's vice president of exploration. https://www.worldoil.com/news/2026/8/13/exxonmobil-identifies-four-guyana-exploration-opportunities-using-ai/

[8] PETRONAS. PETRONAS Advances Malaysia's Upstream Data Platform with Agentic AI to Accelerate E&P Investment. 3 September 2026. https://www.petronas.com/media/media-releases/petronas-advances-malaysias-upstream-data-platform-agentic-ai-accelerate-ep

[9] EarthScan. How Many Wells Is Enough? A Well-Count Ablation for Fracture Detection. https://earthscan.io/case-studies/geobfdt-well-count-ablation-how-many-wells-fracture-detection

[10] EarthScan. Validating Well-to-Well Transfer on 118 Public Wells Before Touching a Single Client Log. https://earthscan.io/case-studies/force-2020-proxy-validation-well-to-well-transfer

[11] EarthScan. The Generalization Cliff: Turning a Broken Azimuth Metric Into a Deployment Protocol. https://earthscan.io/case-studies/the-generalization-cliff-what-horizontal-wells-did-to-our-fracture-model

Tarry Singh
Tarry Singh

Founder & CEO

More from EarthScan

Related research

All insights →
Fully Deployed Is Not Machine-Readable: The Raster Share Behind Four NOC Data Platforms
Insight

Fully Deployed Is Not Machine-Readable: The Raster Share Behind Four NOC Data Platforms

64 Wells a Year: Why Exploration AI Cannot Prove Itself on Outcomes
Insight

64 Wells a Year: Why Exploration AI Cannot Prove Itself on Outcomes

Two to Three Times More Rigs per Engineer Is a Claim About Alert Precision
Insight

Two to Three Times More Rigs per Engineer Is a Claim About Alert Precision

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.