Skip to main content
Reading viewAll insights →
BLOG19 min read

Foundations First Is a Schedule Claim: What Time-to-Data Does to an Industrial AI Programme

Foundations First Is a Schedule Claim: What Time-to-Data Does to an Industrial AI Programme
Tarry Singhby Tarry SinghFounder & CEO · 1 Oct 2026
Share

Inside 16 days of August 2026, Aramco Digital said its agreement with SANAD would build the digital foundations for Industrial AI, Saudi Aramco signed a memorandum of understanding with Aramco Digital framing potential collaboration in industrial AI and virtual twin technologies, Baker Hughes said Kuwait Oil Company had awarded it a multi-year contract to build a research and technology development centre in the Ahmadi Innovation Valley, and CNOOC Limited said it had fully deployed the Haineng-Zhiqing platform and strengthened the digital foundation for intelligent oil and gas fields. Foundations first is a sensible order of work, and it is also a schedule claim, and not one of the four releases states a date, a term or a go-live. On our own three-phase carbonate programme in Oman the ledger read 5 asked and 0 delivered, then 10 and 8, then 25 and 8, and the quantity that moved the model milestone was how long one data request took to come back. Put that queue in front of a fixed-budget phase and the whole distance in mean lead time between a phase on plan and a phase over its budget is the headroom priced in days by the queue itself, which narrows as one over the number of requests the phase depends on. At the recorded Phase 3 ask of 25 requests that band is 17.5 days wide. At 400 requests it is 1.10 days, and the instrument prints what it takes to widen it.

Four announcements inside 16 days, and all four are about foundations. On 13 August 2026, by the date its own news listing carries, Aramco Digital said it had signed an agreement to give SANAD a cloud-based enterprise resource planning platform, infrastructure modernisation and managed technology services, and its chief executive Dr Ashraf Al-Tahini framed the work with the sentence "Industrial AI begins with secure, connected, and trusted digital foundations" [1]. On 24 August Saudi Aramco announced three instruments under the heading of collaboration with French companies, the third of them a memorandum of understanding with Aramco Digital that "establishes a framework for potential collaboration in industrial AI, virtual twin/digital twin technologies, and related technologies, including potential applications in the oil and gas sector" [2]. On 11 August Baker Hughes announced that Kuwait Oil Company had awarded it a multi-year contract for the Ahmadi Innovation Valley project, under which Baker Hughes "will build a dedicated research and technology development center" [3]. And on 26 August CNOOC Limited reported that it had "fully deployed the 'Haineng-Zhiqing' digital platform, and strengthened the digital foundation for intelligent oil and gas fields" [4].

Building the foundations before the models is the right order of work. Nobody who has run a subsurface-AI programme on an operator's own data would argue otherwise. But foundations first is not only an architecture statement. It is a claim about a schedule, because the AI work that sits on top of the foundation starts when the data underneath it can be asked for and got. None of the four releases says when that is. Three give no date at all, and the fourth says a platform is deployed without saying what a request against it costs in days.

That gap is the subject of this piece, and the reason to write it is that on the one programme where we tracked the number, it was the number that decided the phase.

Four releases, read exactly

What each release does and does not say is the whole of the evidence, so it is worth reading to the letter.

The Aramco Digital and SANAD agreement is an ERP, infrastructure and managed-services engagement, and the release describes the AI as future tense on both sides. Al-Tahini says Aramco Digital will support SANAD "in building the integrated technology environment required to enhance today's operations", and the Industrial AI opportunities that environment is meant to open are put, in the same sentence, in the future tense [1]. SANAD's chief executive Miguel Sanchez says modernising the platforms and infrastructure is "preparing the organization to benefit from emerging digital and Industrial AI technologies" [1]. There is no contract value, no term, no go-live date and no headcount anywhere on the page, and the page names no joint-venture partner for SANAD, so neither does this piece.

The Aramco release of 24 August is the one most likely to be misread, and the misreading is a money one. Its headline says the "Agreements and MoU have potential combined value of more than $3.7 billion" [2]. That is the combined potential value of all three instruments, and the release allocates none of it to the memorandum. The memorandum is a framework for potential collaboration, which is the softest of the three, and its named counterparty is Aramco Digital rather than any French company [2]. Anyone attaching the $3.7 billion to industrial AI is attaching a number the company did not attach.

The Kuwait Oil Company item is not Kuwait Oil Company speaking. It is Baker Hughes announcing a contract it has won, and every quotation in it is Baker Hughes's own: the release carries a statement from chief executive Lorenzo Simonelli, that Baker Hughes "is committed to deeply understanding KOC's development aspirations and providing the solutions needed to help achieve them", and no statement from KOC [3]. What the vendor says it will deliver is a research and technology development centre in the Ahmadi Innovation Valley, described as KOC's flagship initiative for an in-country research and innovation hub, together with a portfolio described as "digital and artificial intelligence (AI) automation solutions" [3]. Multi-year is the only duration given. No value, no term in years, no headcount. Read it as the service major's account of the operator's programme, because that is what it is.

CNOOC's is the strongest deployment claim of the four and it is still a claim about a platform. The sentence sits inside an interim-results release and says the company formulated the scenario blueprint for its digital and intelligent initiative, fully deployed the platform, and strengthened the digital foundation [4]. There is no count of assets, users, wells or workloads attached to it anywhere in the release.

Four companies, four foundations, and not one date. That is not a complaint about the releases. Three are results or partnership notices and one is a vendor award announcement, and none of them claims to be a project plan. It is the reason the arithmetic below is worth doing, because the operator running each of those programmes has the date, and the date is set by something the release does not mention.

What a data request actually costs

On a three-phase carbonate borehole-imaging programme with a major operator in Oman, the artifact we managed the engagement by was a two-column ledger: wells asked for against wells actually in our hands, one row per phase, updated every month. Phase 1 recorded 5 asked and 0 delivered, all five rejected at intake because none carried the apparent dip and azimuth picks the detector needed as ground truth. Phase 2 recorded 10 asked and 8 delivered, two of the ten excluded at QC. Phase 3 recorded 25 asked and 8 delivered by the final report [5].

The word request is doing real work in those rows. A request for an annotated well is a request for a completed human interpretation: a geologist sitting with an image log and picking the dip, the azimuth and the feature class for every sinusoid, then exporting those picks out of the operator's confidential systems under their own review cadence and competing priorities [5]. It is not a file copy. It is a queue with people in it, and the people are not ours.

The tracking of that queue became its own piece of software. The data-management dashboard went from V1 to V11 across the middle of Phase 2, and each version number was a defect class we had learned to expect: received-and-missing columns per well, manifest depth ranges reconciled against the depth index inside the file, sampling intervals interrogated inside the raw pick sheets, static-range checks per logical file [6]. At V9 it stopped being a QC crutch. It could put the delivered count, the exclusions and the next phase's minimum in one view: all 10 wells delivered, two excluded for abnormal static ranges, 8 usable, against a Phase 3 minimum on the order of 10 to 15 wells and a phase boundary two months out. The sentence that reached the board deck was that data collection was moving at a really slow pace and timelines would shift [6].

That sentence is the thing this piece is about. It reads as a management observation and it was a dashboard reading. The middle of that phase is recorded in a running log of 288 dated blocks between February and December 2022 [7], and what those blocks record is not a modelling problem. Two of ten delivered wells were unusable for reasons that had nothing to do with any model [7]. The models came later and they came out fine. The phase was decided by the queue.

The arithmetic

Take a model phase that depends on nn usable data deliveries. Requests that come back clear the intake gate at a yield yy, so the programme has to issue n/yn/y of them to end with nn. The operator's data owners keep kk of them open at once, each taking a mean lead time of LL days, so the last one lands after

Days until the last needed request has landed
T(L,n)  =  L ny kT(L,n) \;=\; \frac{L\,n}{y\,k}

The phase also carries a tail τ\tau, the work that can only start once the last needed delivery is in hand: retrain on the complete set, validate, report. Everything else in the phase overlaps the queue. So against a phase window of WW months, with DD days in a month, the months by which the first supervised model misses that window are

Slip of the first supervised model, in months
slip(L,n)  =  max⁡ ⁣(0,  T(L,n)D+τ−W)\mathrm{slip}(L,n) \;=\; \max\!\left(0,\;\frac{T(L,n)}{D} + \tau - W\right)

Two regimes are already in that expression and no further assumption is needed to find them. Where the queue drains inside the phase window the maximum returns zero, and on that part of the surface neither the lead time nor the request count nor the in-flight count nor the yield appears at all: the phase's own work sets the date. Past that fold, every extra day of lead time is (n/y)/k(n/y)/k days of slip and the access queue sets the date alone.

The fold is a hyperbola, and so is every contour above it. Setting the slip to a level ss and solving for the lead time,

The mean lead time at which the slip reaches s months
L(s,n)  =  (s+W−τ) D k ynL(s,n) \;=\; \frac{(s + W - \tau)\,D\,k\,y}{n}

so the lead time at which the phase first slips at all is L(0,n)L(0,n), the lead time at which it passes a schedule headroom of BB months is L(B,n)L(B,n), and the whole distance between them is

The band between a phase on plan and a phase over its budget, in days of mean lead time
band(n)  =  L(B,n)−L(0,n)  =  B D k yn\mathrm{band}(n) \;=\; L(B,n) - L(0,n) \;=\; \frac{B\,D\,k\,y}{n}

The phase window and the tail cancel out of that difference. What is left is the schedule headroom priced in days by the queue's own service rate, and it narrows as one over the number of data requests the phase depends on. That is the finding, and the bench below is where it is read.

The bench

The exhibit computes that surface rather than drawing it. Left to right is the mean lead time per data request, 2 to 60 days. Back to front is the number of data requests the model phase depends on, 10 to 400. Height is the slip of the first supervised model in months, capped at 24 so the far corner does not leave the block. The amber sheet is your schedule headroom and the amber line on the surface is where the two meet, which is where the programme stops being late and becomes cancelled. The bright line is the fold, where the phase first slips at all. The white curtain is your request count and the labels follow it.

Three of the opening values are ours and recorded. The request count opens at the Phase 3 ask of 25 [5]. The intake yield opens at the Phase 2 gate, 8 of the 10 that arrived [5]. The phase window is 10 months, the span of the recorded running log from February to December 2022 [7], and the plate derives it from those two months rather than quoting a contract. The other three are stated assumptions and the plate says so: a tail of four weeks, six requests open at once, and three months of schedule headroom. None of the four releases at the top of this piece gives a lead time per data request, and neither did we record one, which is exactly why the lead time is an axis and not a default.

PROGRAMME SLIP SURFACE17.5 daysOF MEAN LEAD TIME BETWEEN ON PLAN AND OVER BUDGET, AT 25 REQUESTSon plan up to53.1 days10 days of lead time, 25 requestsslip nonemean lead time per data request, 2 days left to 60 rightdata requests the phase depends on, 10 at the back to 400 at the frontslip of the first supervised model, 0 to 24 monthsON PLAN UP TO53.1 daysOVER BUDGET PAST70.6 daysIN FLIGHT FOR A 20-DAY BAND6.84DATE AT 10 DAYS SET BYthe work itselfslip, 0 to 24 monthsbudget planeyour request countphase starts to slipover budgetRecorded, ours: an Oman programme asked 5, 10 and 25 wells and kept 0, 8 and 8. The cursor opens at that 25; the yield at that 8 of 10.Phase window 10 months, the recorded Feb to Dec 2022 log span (288 blocks). Tail 4 weeks, 6 in flight and the headroom: assumed. Height caps at 24 months.
The surface is the number of months by which the first supervised model misses its phase window, over every mean lead time per data request (left to right) and every count of data requests the phase depends on (back to front). Its height is computed from one queue: to end with the requests the phase needs, the programme issues more of them at the intake yield, the operator's data owners serve them a few at a time at the mean lead time, and a fixed tail of work can only start once the last one has landed. The amber sheet is your schedule headroom, and the amber line on the surface is where the two meet. The white curtain is your request count, and the labels follow it. The bright line is where the phase first slips at all; to the left of it the surface is exactly flat, because that branch of the model returns the phase window and contains neither the lead time nor any control. The request count, the intake yield and the phase window open at figures recorded on our own Oman programme; the tail, the in-flight count and the headroom are stated assumptions, and every value is yours to move. Drag or use the arrow keys to orbit, Home to reset.

What the bench shows

At the opening settings the phase is comfortable. On plan up to 53.1 days of mean lead time, over budget only past 70.6 days, a band 17.5 days wide between the two, and at a readout lead time of ten days the slip is none and the date is set by the work itself. Six requests open at once is very nearly what a 20-day band would take: that readout says 6.84 in flight. A programme that depends on 25 data requests can absorb a mean lead time of seven weeks per request and still hit its date.

Now drag the request count forward. At 400 requests the same phase, with every other control untouched, is on plan up to 3.32 days, over budget past 4.41 days, and the band between them is 1.10 days. The slip at a ten-day lead time is 18.3 months, and the readout for a 20-day band says 110 requests in flight against the six the programme has. One point one days of mean lead time is the entire distance between a phase that hits its date and a phase that has spent its schedule headroom. No enterprise access process is specified to that tolerance. Most are not measured to it.

The flip is worth finding by hand, because it is sharper than any summary of it. Drag the request count slowly. At 132 requests the readout says the date is set by the work itself and the slip at ten days is none. At 133 it says the access queue, and the slip reads 0.0 months. One request either side of that boundary and the programme has changed which department owns its schedule, while the slip readout still shows nothing at all. The number that then grows is not gradual: by 200 requests the slip at the same ten-day lead time is 4.6 months, and by 400 it is 18.3.

The other three controls say which levers are worth buying. Raising the schedule headroom from three months to twelve, at 400 requests, widens the band from 1.10 days to 4.38 days: four times the headroom buys four times the band, and four and a bit days is still not a procurement cycle. Raising the requests open at once from six to 24 widens the band by exactly the same factor, to 4.38 days, and it does something the headroom cannot: it moves the fold, so the on-plan ceiling goes from 3.32 days to 13.3 days and the slip at a ten-day lead time falls from 18.3 months to none. Buying schedule is buying tolerance for a queue you still have. Buying concurrency is buying the queue away.

The intake yield is the one that is hardest to move and easiest to underestimate. Take it from the recorded 8 of 10 all the way to a perfect gate at 400 requests and the slip at ten days falls from 18.3 months to 12.8, which is an improvement and is nowhere near a rescue. Drop it to the low end and the band collapses to 0.27 days. On our own programme the yield was not a policy choice; it was what arrived, and one phase of it was zero out of five [5].

That is the shape of the finding. A programme with a handful of data requests is in a work-bound regime where the access queue is genuinely not the risk, and the platform milestone is a fair thing to talk about. A programme with hundreds of them is in a queue-bound regime where the platform milestone is not the risk and the access queue behind it is, and the tolerance is not weeks. The four releases at the top of this piece are all about the foundation. Not one of them says which regime the programme on top of it is in.

Where the ten months comes from, and what it is not

The one recorded quantity that carries the most weight on the plate is the phase window, so it is worth being exact about what it measured. It is the span of a running log: 288 dated blocks between February and December 2022 on one confidential carbonate engagement [7]. It is a work diary, not a contract, and the dates order the milestones reliably while the boundaries of the phase are softer than a single number implies. It is on the plate because it is the one phase-length figure we have published from our own record, and because the derivation from two month numbers to one window is short enough to print.

What the model does not have is our own lead time per data request. We never recorded one, and the plate does not pretend we did. The lead time is the left-to-right axis for exactly that reason: it is the reader's number, and the whole point of the surface is that the answer depends on where their programme sits along it. The tail of four weeks, the six requests in flight and the three months of headroom are stated assumptions in the same spirit, and every one of them is a control.

What survives any of those choices is the structure. The flat floor where the lead time cannot hurt, the fold where it starts to, and a band whose width is the headroom divided by the request count. Those are properties of a queue in front of a fixed window, not of our figures.

Three questions for a programme building foundations

How many distinct data requests does the first model milestone depend on? Not terabytes and not systems integrated: the count of separate asks that have to be answered by somebody before the model can be trained. On our record that count ran 5, then 10, then 25 [5], and it is the coordinate that decides which regime the programme is in.

How many of them can be open at once, and with whom? The bench shows this is the only control that moves the fold rather than the band. If the answer is a single custodian working requests in series, the concurrency is one and the arithmetic is unforgiving at any request count above a handful.

What is the mean lead time per request today, and who measures it? Access, trust and time-to-data are the three questions a stalled programme never answers, and time-to-data is the one that is a distribution rather than a headline, so a programme is only as fast as the request that blocks the decision actually waiting [8]. An ERP go-live does not answer it. A knowledge-graph ingest that turns a request into a query does, and that is the difference between a foundation and a schedule.

Key takeaways

  1. Aramco Digital says its SANAD agreement builds the digital foundations for Industrial AI, Saudi Aramco has signed a memorandum of understanding with Aramco Digital framing potential collaboration in industrial AI and virtual twin technologies, Baker Hughes says Kuwait Oil Company has awarded it a multi-year contract to build a research and technology development centre, and CNOOC Limited says it fully deployed the Haineng-Zhiqing platform. All four are about foundations and none states a date, a term or a go-live.
  2. The Kuwait Oil Company leg is the vendor speaking: every quotation in that release is Baker Hughes's and there is none from KOC. The $3.7 billion in the Aramco release is the potential combined value of a set of agreements and one memorandum of understanding, and the release allocates none of it to the AI memorandum, whose named counterparty is Aramco Digital.
  3. On our own three-phase carbonate programme in Oman the ledger read 5 asked and 0 delivered, then 10 and 8, then 25 and 8. A request for an annotated well is a completed human interpretation exported from an operator's confidential systems, so the queue has people in it and the people are not ours.
  4. Against a fixed phase window the slip is zero wherever the request queue drains inside it, and in that branch the lead time, the request count, the concurrency and the intake yield are all absent from the date. Past the fold the access queue sets it alone.
  5. The whole distance in mean lead time between a phase on plan and a phase over its budget is the schedule headroom priced in days by the queue, B times the days in a month times the concurrency times the yield, over the request count. The phase window and the tail cancel out of it, and it narrows as one over the number of requests.
  6. At the recorded ask of 25 requests that band is 17.5 days and the phase is on plan up to 53.1 days of lead time. At 400 requests the band is 1.10 days, the on-plan ceiling is 3.32 days, and a ten-day lead time costs 18.3 months. The readout flips from the work itself to the access queue between 132 and 133 requests.
  7. Four times the schedule headroom widens the band from 1.10 days to 4.38 at 400 requests and moves nothing else. Four times the concurrency widens it by the same factor and also moves the fold, taking the on-plan ceiling from 3.32 days to 13.3 and the slip at ten days from 18.3 months to none. Concurrency is the lever that buys the queue away.

References

[1] Aramco Digital. Aramco Digital and SANAD Sign Digital Transformation Agreement. 13 August 2026, per the date on the aramcodigital.com news listing; the article page itself carries no date. https://aramcodigital.com/news/news-article-2

[2] Saudi Arabian Oil Company (Aramco). Aramco enhances its global partnership ecosystem through collaboration with French companies. 24 August 2026. https://www.aramco.com/en/news-media/news/2026/aramco-enhances-its-global-partnership-ecosystem-through-collaboration-with-french-companies

[3] Baker Hughes. Baker Hughes Awarded Multi-Year Contract by Kuwait Oil Company for Ahmadi Innovation Valley Project. GlobeNewswire, 11 August 2026. The same text is on the Baker Hughes investor newsroom. This is the vendor's release, not Kuwait Oil Company's, and it carries no KOC quotation. https://www.globenewswire.com/news-release/2026/08/11/3342544/0/en/baker-hughes-awarded-multi-year-contract-by-kuwait-oil-company-for-ahmadi-innovation-valley-project.html

[4] CNOOC Limited. CNOOC Limited Focuses on Value Creation, Production and Profit Hit New Highs in H1 2026. 26 August 2026. https://www.cnoocltd.com/english/presscenter/pressreleases/2026/202608/t20260826_122439.html

[5] EarthScan. The Attrition Ledger: How We Ran a Subsurface-AI Program on One Number, Asked-Versus-Delivered. https://www.earthscan.io/case-studies/25-wells-asked-for-8-delivered-the-data-scarcity-gap

[6] EarthScan. The Data-Management Dashboard That Grew From V1 to V11 Across a Phase. https://www.earthscan.io/case-studies/data-management-dashboard-v1-to-v11

[7] EarthScan. The Messy Middle of Phase 2: From Data-Quality Chaos to a Working Supervised Model. https://www.earthscan.io/case-studies/the-messy-middle-of-phase-2-from-data-quality-chaos-to-a-working-supervised-model

[8] EarthScan. Access, Trust, Time-to-Data: Why Enterprise AI Initiatives Fizzle. https://www.earthscan.io/whitepapers/access-trust-time-to-data-why-enterprise-ai-initiatives-fizzle

Tarry Singh
Tarry Singh

Founder & CEO

More from EarthScan

Related research

All insights →
Progressively Autonomous Still Waits: The Approval Round Three Operators Do Not Price
Insight

Progressively Autonomous Still Waits: The Approval Round Three Operators Do Not Price

More Wells Is Not More Signal: The Partner Share Where Pooled Exploration Data Starts to Pay
Insight

More Wells Is Not More Signal: The Partner Share Where Pooled Exploration Data Starts to Pay

Two to Three Times More Rigs per Engineer Is a Claim About Alert Precision
Insight

Two to Three Times More Rigs per Engineer Is a Claim About Alert Precision

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.