Four disclosures, all inside two years, all selling the same thing. On 10 March 2025 AIQ announced a 340 million dollar contract, over three years, to deploy ENERGYai and related solutions across ADNOC's upstream operations, built on 70 years of proprietary data and knowledge, with five fully operational AI agents covering subsurface tasks and a stated plan to scale to thousands of additional wells across more than 28 producing fields [1]. Two months earlier ADNOC reported the completed trial: over 50 years of its knowledge and petabytes of its proprietary data, data from over 15% of its onshore and offshore wells, and a 70% improvement in accuracy in major seismic interpretation aspects [2]. At LEAP in Riyadh on 4 March 2024, Saudi Aramco presented Metabrain, 250 billion parameters trained on 7 trillion data points spanning more than 90 years of company history, aimed among other things at analysing drilling plans and geological data [3]. Equinor, having reported its 2025 AI results in January 2026, says it has identified more than 100 further AI use cases [4]. And ExxonMobil, which does not lead with an archive figure at all, is putting 4,032 NVIDIA Grace Hopper superchips behind elastic full wavefield inversion to take 4D seismic processing from months to weeks [5].
Notice which number appears in three of those five statements and which does not. Years of archive: 70, over 50, more than 90. Documents in the archive: nowhere. Not in any of them. That asymmetry is not a presentational accident, and it is the subject of this piece, because the size of the archive is the one term in the retrieval problem that behaves benignly and the term that decides everything is the one nobody prints.
The archive is the denominator, not the asset
Strip the marketing off an agentic subsurface assistant and the load-bearing component is a retriever. Somebody asks a question. The system ranks the corpus, takes what clears a threshold, and hands the survivors to a language model or to a person. Everything downstream is conditioned on that step, and the step is a plain signal-detection problem.
Write it in the equal-variance Gaussian form, which is the one that makes the scaling visible. Let the archive hold documents, of which actually answer the question that was asked. The retriever assigns a score. Irrelevant documents score ; answer-bearing ones score . The separation is the discrimination index, and it is the only property of the retriever that enters. The reader accepts everything above a threshold .
Now the assumption everyone skips. is set by the question, not by the archive. Adding four decades of end-of-well reports to a corpus does not add four decades of answers to "what mud weight did we run through the Shuaiba in this block" or "which wells in this field logged a resistivity image over the reservoir interval". There is a small number of documents that answer any specific question, and it is roughly the same small number whether the archive holds a hundred thousand documents or a billion. So the base rate
falls as the archive grows. The archive is the denominator. It is not, by itself, the asset.
Two growth laws, one of them brutal
Pin the threshold at , the mean of the answer-bearing scores. That is the median-recall operating point, where exactly half the answer-bearing documents clear the bar, and it has the pleasant property that the threshold and the discrimination index are the same number, so one axis carries both. Nothing that follows depends on that choice: at any fixed recall the threshold is , every in this piece shifts by the same constant , and every difference between two of them is unchanged.
With the upper tail of the standard normal, the irrelevant documents a reader must wade through is
which is linear in . A better ranker changes , which is the slope. It does not change the fact that the relationship is a straight line in , crossing zero at . On a log-log plot the line has slope one at any whatsoever, and dragging the retriever slides it up and down without ever bending it. The slope is not exactly one, because the subtracted still counts at the left of the axis, but the departure bends the drawn line by under a hundredth of a pixel.
Ask the other question instead. Fix a reading budget: irrelevant documents is all anyone will open before giving up. What discrimination holds the pile at as the archive grows? Solve :
The right-hand form is the leading-order Gaussian tail asymptotic, and it is the shape rather than the number. The shape is the whole argument. Reading grows like . The discrimination that would hold reading constant grows like .
Put the two together at figures a buyer would recognise. Set , ten documents in the archive that answer the question, and a reading budget of ten. At the retriever that exactly meets the budget is , and it puts 10.00 irrelevant documents in front of the reader. Hold that retriever fixed and grow the archive to , and the same retriever now puts 100,007 in front of them. The exact ratio of those two counts is , a factor of ten thousand. To hold the pile at ten instead, the required discrimination goes from 3.719 to 5.612. Four decades of archive cost either a factor of ten thousand in reading or 1.893 units of discrimination, and there is nothing in between.
What the instrument is for
Drag the archive slider and watch one readout explode while the other creeps. That is the finding, and it is worth reaching by hand rather than by reading a table, because the asymmetry is much larger than intuition allows for.
Then drag the other two. The reading budget is the lever most procurement conversations end up pulling, usually dressed as a larger context window: we will just show the model more documents. At a billion-document archive, running the budget from 1 to 100 walks the required discrimination from 5.998 down to 5.199, a span of under 0.8. Four decades of archive demand 1.893. The budget lever is worth less than half of what the archive growth costs, because enters only inside a logarithm while enters linearly. Buying context is the expensive lever, and it is the one everybody pulls.
The score strip along the bottom is where the whole effect lives, and it is drawn to make one point that a table cannot. The amber rule is the threshold that admits exactly irrelevant documents, and it slides right as the archive grows. Everything the reader must wade through comes from the region beyond it. Measure the height of the irrelevant bell at that rule and you get the point in one number: about one part in 15 at the shallowest setting the controls allow, one part in a thousand at the default, one part in 65 million at the deepest. So across almost the whole control space the region that carries the entire reading cost has no height you could draw, which is why the band on the floor marks the region and the count sits in the panel. The effect that dominates the economics of the system is not visible at any scale you would draw the distributions on.
One number in that frame is measured rather than modelled, and it is ours.
The base rate is the number nobody measures
We indexed a real archive for a Texas onshore operator: 136,771 scanned TIF rasters and 7,781 machine-readable LAS files, deduplicated and catalogued into something a model could ingest (Indexing a 136,771-Scan Raster Archive [6], and the ratio argument that came out of it in What 136,771 TIFs and 7,781 LAS Files Teach You About Real Data [7]). That is the single marker on the archive axis in the instrument above, at 144,552 files, and it is alone there for a reason worth stating plainly: not one of the operator disclosures at the top of this piece states an archive scale in documents. Petabytes, data points, years, square kilometres. Not documents. So there is nothing else honest to plot.
What that indexing job taught us is not the count. It is what the count did not tell us. We knew, exactly, how many files there were. What we did not have, and what nobody on either side of the engagement asked for, was : how many documents in that corpus actually answer a given question. Counting the archive is a data-engineering task and it finishes. Estimating the base rate needs a labelled question set, somebody who knows the field well enough to say which documents genuinely answer each question, and the discipline to hold that set back from whatever is being tuned. Nobody budgets for it, because the archive count feels like the same kind of number and is a thousand times cheaper to get.
It is not the same kind of number. Look again at the readout column. At 100,000 documents with ten answers and a retriever at 3.719, precision is 33.3%: five answers among 15 documents opened, and a geologist will happily read 15. At a billion documents the same retriever gives 0.0050%. Same retriever, same question, same ten answers. What changed was only the denominator, and the denominator is the thing being sold.
This is also why the ADNOC trial's headline figure is hard to act on. A 70% improvement in accuracy in major seismic interpretation aspects is a real, specific claim, and the release states no baseline for it [2]. An accuracy improvement without a denominator is structurally the same gap as an archive without a base rate: a ratio quoted with one side missing. That is not a criticism of the result, which may well be excellent. It is an observation about what a press release can carry and a note on what a buyer has to ask for separately.
What to ask for instead
Three things, none of which is hard for a vendor who has done the work.
First, the base rate on a real question set. Not the corpus size. Take 50 questions the asset team actually asks, have somebody qualified mark which documents answer each one, and report the median count. That number, divided by the corpus size, is , and it is the number that decides whether the system is usable. It is also the number that tells you whether the corpus should be partitioned before it is indexed, which is usually the cheapest real fix available: retrieval over one field's end-of-well reports is a different problem from retrieval over the whole company.
Second, discrimination rather than recall. Recall at some unspecified operating point is close to meaningless, because it can always be bought by lowering the threshold and paying in reading. Ask for the pair: the recall and the number of irrelevant documents that came with it, at the archive size the system will actually run against. Two numbers, one operating point. From those you can back out and put it on the axis above.
Third, the scaling test the vendor has already run, or an admission that they have not run one. Index a tenth of the corpus, then the whole corpus, and report the reading pile at fixed recall for both. The prediction from the algebra is a factor of ten. If the measurement comes in materially under that, something good is happening that the vendor should be able to explain, most often a partition or a metadata filter doing work the ranker is being credited with. If it comes in at ten, the system is behaving exactly as this model says it will, and the buyer now knows what the next order of magnitude of archive is going to cost.
None of this argues against large archives. 70 years of proprietary well data is a genuine asset, and Equinor's published interpretation and value figures, which we worked through separately in Sixty-Four Wells a Year [8], are genuine results. The argument here is narrower. The archive is what makes the answer available at all, and it is simultaneously what makes the answer hard to find, and only one of those two effects appears in the announcement. A buyer who reads "70 years of proprietary data" as a capability claim has read the numerator and skipped the denominator.
Takeaways
- AIQ, ADNOC and Saudi Aramco all lead with the age of the archive, at 70, over 50 and more than 90 years; not one of the five disclosures cited here, theirs or Equinor's or ExxonMobil's, states an archive scale in documents, which is the unit the retrieval problem is posed in.
- At a fixed retriever, the irrelevant documents a reader must wade through is (N minus R) times Q(d), exactly linear in the archive: on a log-log frame that reads as a straight line of slope one, and a better ranker moves its height without ever bending it.
- Holding that pile at a fixed reading budget needs discrimination growing like the square root of twice the log of the archive, so four decades cost 1.893 units of d against a factor of ten thousand in reading.
- The reading budget, which is what a larger context window buys, is the weak lever: at a billion documents, running it from 1 to 100 walks the requirement from 5.998 to 5.199, under 0.8, against the 1.893 that four decades of archive demand.
- The base rate of an answer-bearing document is the term that decides precision, and it falls as 1/N because a bigger archive holds more of everything except answers to your question.
- We counted a real archive exactly, 136,771 TIF rasters and 7,781 LAS files, and still did not have the base rate; counting the corpus is a data-engineering job that finishes, while estimating the base rate needs a labelled question set nobody budgets for.
Limitations
This is exact algebra for an idealised detector, not a measured retrieval benchmark, and four things about it are worth holding at arm's length.
The equal-variance Gaussian model is a convenience. Real retrieval scores are not Gaussian and the two populations rarely share a variance, so here is a summary of separation rather than a quantity you would measure directly off a production system. The scaling result survives that in two parts. Reading is linear in at any fixed operating point whatever the score distribution is, because it is a count of documents multiplied by a fixed probability, and no tail assumption enters. The requirement grows sublinearly for any tail decaying at least as fast as a power law with exponent above one: as a power of for a power-law tail, logarithmically for an exponential one, as the square root of the log for a Gaussian one. What does not survive is the figure 1.893, which belongs to the Gaussian.
The form is the leading-order asymptotic and runs high at the tail depths in play. Measured against the exact quantile, it overstates the requirement by 0.573 at and by 0.458 at . The instrument and every number quoted here use the exact inverse tail; the closed form is in the prose for its shape only.
Holding fixed as the archive grows is the strong assumption in the piece, and it is a modelling choice rather than a measurement. It is right for a specific factual question about a specific well or field, which is most of what an asset team asks. It is wrong for a question whose answer set genuinely scales with the corpus, such as a survey of every wellbore that encountered a given formation, and for those the base rate does not fall and this argument does not apply.
Finally, the instrument's fixed value of ten answer-bearing documents is chosen so the base-rate readout is exactly ten over the archive size and can be checked by eye. Neither growth law depends on it while stays small against , but the precision readout does, in direct proportion.
References
[1] AIQ (10 March 2025). AIQ announces 340 million dollar contract for large-scale deployment of agentic AI across ADNOC operations. aiq.ae
[2] ADNOC (16 January 2025). ADNOC and AIQ successfully complete trial phase of agentic AI solution. adnoc.ae
[3] Offshore Technology (7 March 2024). Saudi Aramco unveils industry-first generative AI model, presented at LEAP in Riyadh on 4 March 2024. offshore-technology.com
[4] Journal of Petroleum Technology (7 January 2026). Equinor says AI saved it 130 million dollars in 2025. jpt.spe.org
[5] ExxonMobil. The future of seismic imaging and technology: Discovery 6 and elastic full wavefield inversion. corporate.exxonmobil.com
[6] EarthScan. Indexing a 136,771-scan raster archive for ML ingestion. earthscan.io
[7] EarthScan. What 136,771 TIFs and 7,781 LAS files teach you about real data. earthscan.io
[8] EarthScan. Sixty-four wells a year: why exploration AI cannot prove itself on outcomes. earthscan.io




