On 13 August 2026 ExxonMobil said that artificial intelligence had identified four new exploration opportunities on the Stabroek Block offshore Guyana, applying deep learning, machine learning, high performance computing and advanced seismic imaging to its own historical discoveries, drilling results and subsurface data [1]. Seven months earlier, on 7 January 2026, Equinor reported that AI created 130 million dollars of value in 2025, that its interpretation capacity is up tenfold, that it processed two million square kilometres of seismic across the year, and that AI-driven planning on Johan Sverdrup phase 3 found a solution nobody had considered and saved the partnership 12 million dollars [2]. Shell and SparkCognition have described generative reconstruction of subsurface images from as little as one percent of the usual shot count [3].
Every one of those statements is checkable. Each names a thing that happened, to a stated asset, in a stated period, and an operator's own disclosure carries it. The statement the market actually wants is a different kind of sentence: that prospects selected with machine learning drill better than prospects selected without it. Nobody has said that, and the reason is not modesty. It is that the industry does not drill enough wells to find out.
The sample the industry has
Westwood recorded 64 high-impact wells worldwide in 2025, the lowest annual total since 2008, delivering 4.8 billion barrels of oil equivalent at a 31% commercial success rate against 28 to 29% in the two prior years [4]. Rystad puts 2025 high-impact wildcat success at 38%, up from 23% in 2024, on discovered volumes of about 2.3 billion barrels of oil equivalent, and identifies 42 high-impact wells for 2026 [5]. Those two volume figures sit roughly a factor of two apart for the same calendar year, because the two houses count different things under the same label. Name the gap rather than split the difference. Neither definition moves the well count off dozens a year. Equinor plans about 26 exploration and appraisal wells offshore Norway in 2026 [6], on a shelf where every exploration well is published individually and the counting is not in dispute [7].
Those are the sample sizes on offer. Tens per operator per year. Dozens per year across the entire high-impact programme of the whole industry. That is the denominator any outcome-based claim about exploration AI has to live inside.
What it costs to see a lift
Treat the question as a comparison of two success rates, one arm screened with the model and one without. At significance level and power , the wells needed per arm to resolve a difference against a control rate are
At 5% significance and 80% power, with a control rate of 30% and a claimed move to 40%, that is about 330 wells per arm, so roughly 660 wells in total. At Equinor's planned 26 exploration and appraisal wells a year it is a 25 year experiment. Give the trial the entire industry's high-impact programme of 64 wells a year and it still takes a decade. Nothing in this is specific to Equinor: TotalEnergies and Saudi Aramco run their exploration campaigns against the same arithmetic, because the arithmetic is a property of the well count and not of the operator.
Read the same relationship the other way and it gives the smallest lift a programme of a given size could ever see:
That is the surface in the exhibit below. Its height over the plane of well rate and programme length is the smallest lift the programme could detect, and the translucent plane is the lift somebody is claiming. Where the surface stands above the plane, the programme cannot see the claim at all.
The base rate is a distraction and the well count is not
Arguments about exploration AI tend to fix on the base rate. Which wells count as high impact, whether appraisal belongs in the denominator, whether a 31% commercial success rate is the right comparator or a flattering one. Drag the base rate across the whole plausible range, 20 to 50%, and the smallest detectable lift moves by about a quarter, because only runs from 0.57 to 0.71 over that span. The argument everyone is having changes the answer by a fifth to a quarter.
Now drag the well count. Going from one operator's 26 wells a year to the industry's entire 64 high-impact wells is a factor of 2.5 in sample size, and it improves the detectable lift only from roughly 23 percentage points to roughly 14, a factor of 1.6, because resolution scales as the inverse square root of the sample. A ten point lift, which would be commercially decisive and would reprice an exploration portfolio, sits far under both numbers. It stays invisible in drilling outcomes for a quarter of a century at operator scale and for a decade at industry scale, and no amount of arguing about the correct base rate rescues it.
The design toggle in the exhibit carries the sharper version of the same result. A historical-baseline comparison, where every well goes to the screened arm and the archive supplies the control, is more sensitive than a concurrent split early on, because it does not spend half the programme on wells it already knows how to drill. Then it stops improving. The archive's own variance is a floor, and past a certain point more drilling cannot buy resolution the control arm does not have. At 26 wells a year a ten point lift sits permanently under that floor.
Fourteen wells taught us this first
We have argued this in public about our own work, at a scale two orders of magnitude smaller and with the same shape.
On a Middle East carbonate field we had 14 vertical wells of borehole-image data, 11 of them with consistent bedding picks. The textbook evaluation is to hold out whole wells so no depth-adjacent near-duplicate can sit on both sides of the split. We could not afford it. One well is about seven percent of the corpus, and because every well in a carbonate field carries its own structural character, the well you happen to reserve swings the test distribution as hard as the training distribution. A one-well test set at that scale is a coin flip you paid a training well to perform. So we shuffled at patch level, defended it with validation and test parity and a continuous blind zone, and published the constraint rather than the flattering number [8].
Separately, a speed figure of under 30 seconds per metre for vug detection left a draft because the scope behind it would not survive a stranger reading the sentence. We retired it instead of defending it on a sample that could not carry it [9].
Both decisions come from the same place the exploration arithmetic comes from. When the sample is small, the honest move is to say what it cannot support and then design an experiment that fits inside it.
What process evidence looks like
The evidence that is actually available for exploration AI is process evidence, and it is stronger than most buyers assume.
Take prospects that have already been drilled, strip the outcome, and hand the model the pre-drill data package. Score its ranking against what the wells later found. The sample here is the archive rather than the forward programme, which is hundreds of prospects rather than dozens, and the outcome is already known so the experiment costs nothing but discipline. The discipline is the hard part: the outcome has to be withheld from everyone who touches the model, the pre-drill package has to be reconstructed as it stood on the decision date rather than as it stands now, and the split has to be by prospect and by basin rather than by seismic patch, for exactly the leakage reason that forced the argument at 14 wells.
That design is checkable, repeatable, and available today. It is also the one thing a vendor with a real capability can produce and a vendor without one cannot.
What to ask, and what to stop asking
Stop asking an exploration AI vendor for a success-rate lift. The number cannot be substantiated on any programme that exists, and a vendor who offers one is either quoting a simulation or quoting a sample too small to carry the claim. Ask instead for three things that can be produced. A blind re-screening on already-drilled prospects, with the withholding protocol written down. A count of interpretation work done, on the model of Equinor's two million square kilometres, which is an operational fact rather than a causal one. And a statement of what the vendor's own evaluation cannot resolve, expressed as a minimum detectable difference at the well count the buyer actually drills.
The last one is the tell. A team that has done the arithmetic knows its own detection floor and will say it out loud. A team that has not will keep arguing about the base rate.
Takeaways
- ExxonMobil, Equinor and Shell have each disclosed specific, checkable AI results in exploration and subsurface imaging; none of them claims a drilling success-rate lift, and the arithmetic explains why.
- Resolving a move from a 30% to a 40% commercial success rate needs roughly 330 wells per arm at 5% significance and 80% power, about 660 wells in total.
- At Equinor's planned 26 exploration and appraisal wells a year that is a 25 year experiment; at the industry's entire 64 high-impact wells a year it is still a decade.
- The base rate argument changes the answer by about a quarter across the whole plausible range; the well count changes it far more, and resolution improves only as the inverse square root of the sample.
- A historical-baseline comparison is sharper early and then floors out, because the archive's own variance is a limit more drilling cannot get under.
- The evidence a buyer can actually get is blind re-screening of already-drilled prospects with the outcome withheld, split by prospect and basin, which is the same discipline that forced us to publish the split constraint on a 14-well borehole-image corpus.
Limitations
The power calculation assumes independent wells and a normal approximation to the binomial, and exploration wells are not independent: they cluster by basin, by play and by seismic vintage, so the effective sample is smaller than the raw count and the horizons in the exhibit are optimistic rather than conservative. The variance inflation applied to the historical-baseline design, a factor of 2.5 over ten prior years, is a stated assumption chosen to represent epoch-to-epoch drift in portfolio mix and rig fleet. It is not a measurement, so the position of that floor is illustrative while its existence is not. The operator disclosures cited here are the operators' own accounts of their own programmes and have not been independently audited. Our 14-well material is a single carbonate field in the Middle East and a single log family, so the split constraint it describes generalises as a method rather than as a number.
References
[1] World Oil (13 August 2026). ExxonMobil identifies four Guyana exploration opportunities using AI. worldoil.com
[2] Equinor (7 January 2026). Artificial intelligence saved Equinor USD 130 million. equinor.com
[3] Journal of Petroleum Technology. SparkCognition and Shell team up to push generative AI. jpt.spe.org
[4] Westwood Global Energy Group (25 June 2026). Westwood Insight: the state of exploration 2026. westwoodenergy.com
[5] Rystad Energy (28 January 2026). High-impact wells: Africa will continue to drive global drilling activity in 2026. rystadenergy.com
[6] World Oil (25 November 2025). Equinor to drill 26 exploration and appraisal wells offshore Norway in 2026. worldoil.com
[7] Norwegian Petroleum. Exploration activity on the Norwegian shelf. norskpetroleum.no
[8] EarthScan. Patch splits or well splits? Evaluating subsurface AI honestly when you only have 14 wells. earthscan.io
[9] EarthScan. The self-scrub: two gates every number should pass before a reviewer ever sees it. earthscan.io




