CNOOC Limited's interim results of 26 August 2026 say the company "fully deployed" the Haineng-Zhiqing digital platform and "strengthened the digital foundation for intelligent oil and gas fields" [1]. Four days later PetroChina said it had "vigorously implemented" its Artificial Intelligence Plus special initiative [2]. Sinopec's interim release of 23 August reports the launch of what it calls "the industry's first digital expert", the Fenghuo industrial AI agent, and says the capabilities of its Great Wall large model "further improved" [3]. On 3 September PETRONAS, through Malaysia Petroleum Management, announced an agreement with Iraya Energies to add agentic AI to the myPROdata upstream data platform, with Iraya as Consortium Lead Integrator [4].
Four national oil companies, four verbs, and not one denominator. None of the four says how many well records sit behind the platform, what share of them are scanned paper, or how many a model can read today. That is not a complaint about the releases: three are results announcements and one is a partnership notice, and none of them claims to be a data inventory. It is the reason the sentence in the title is worth saying. "Fully deployed" is a statement about a platform. A platform can be fully deployed over an archive most of which no model can read, and the platform's own release would look exactly the same either way.
Four verbs, read exactly
What each release does and does not say is the whole of the evidence, so it is worth reading to the letter.
CNOOC's is the strongest deployment claim of the four, and it is a claim about a platform, not about a model reading anything. The only item on the page that is specifically artificial intelligence is that the intelligent injection-production interaction scenario for offshore oilfield production was selected as a high-value scenario at the 2026 World Artificial Intelligence Conference [1]. There is no count of assets, users or workloads on the platform anywhere in the release.
PetroChina's sentence on the initiative is one sentence, with no system named and no number attached [2]. The same release reports that newly granted invention patents rose 174% year on year in the first half and that the company led the development of four international and seven national standards [2]. Those are company-wide figures. The release does not attribute them to the AI initiative and neither should anyone else.
Sinopec names two things, the Fenghuo agent and the Great Wall model, and quantifies neither [3]. The capital expenditure line that sits nearest to them, RMB 1.2 billion for corporate and others in the first half, is described as "mainly for R&D and digital intelligence projects" [3]. That is a blended line. It is not an AI spend, and the second-half figure of RMB 3.6 billion for the same segment is blended the same way.
PETRONAS is the one release that names the substrate, and it is the one that says least about deployment. myPROdata is a data platform; the agreement is that the platform "will integrate" geological, geophysical, engineering and other data with agentic AI that "can autonomously analyse and connect multiple data sources" [4]. Nothing is described as live. The RM50 billion to RM60 billion of annual upstream investment and the shortening of discovery to first hydrocarbon from 100 months to 50 months are quoted in the release as PETRONAS' ambition, in a statement by Malaysia Petroleum Management's senior vice president, and they are not outcomes of this agreement [4].
An agent that connects data sources can connect a LAS file. It cannot read a TIF. That distinction is the subject of this piece, because on the one archive we have counted, the TIFs were nearly all of it.
What a well archive is made of
For a Texas onshore operator we worked with, the raw material was the public Texas Railroad Commission record: 136,771 scanned TIF rasters against 7,781 digital LAS files, a ratio of 17.6 to 1 [5]. A TIF is an image of a printed log, a picture of ink on paper. A LAS file is the same log already digitised into depth-indexed curves. One is what a person reads and the other is what a machine reads, and the archive did not know which of its files was which. There was no index, no deduplication, and no way to tell an image from its digital twin without opening it [5].
So the first deliverable of that engagement was not a network. It was a catalog: every file opened, its real format confirmed, its dimensions recorded, a stable identifier assigned, and the duplicates found by content and metadata rather than by name, because the same log rescanned at a different resolution is bit-for-bit a different file [5]. Two consequences of that order matter for the model below. Deduplication needs the whole census, since a record is not known to be unique until every other record has been seen. And the digitiser that turns a raster into curves cannot be validated honestly until the deduplication is done, because a duplicate that lands on both sides of the train and validation split lets the model study the exam before it sits it [5].
That is why, in what follows, the digitiser waits for the index. It is not a modelling convenience. It is the order the work ran in, for the reason it ran in that order.
The arithmetic
Take an archive of well records, a share of them raster. Records are made addressable at a week, so the index is complete at
A digital record is readable the moment it is indexed. A raster record is readable only after a digitiser has reached it, at records a week, starting when the index is complete. The share of the archive a model can read at week is then
Two regimes are already in that expression, and no further assumption is needed to find them. For the second term is zero at every raster share, so does not appear: raising the digitiser rate changes nothing on that part of the surface. For the second term has slope in time, and the digitiser rate is the whole story.
The date at which a share of the archive is readable makes the crease explicit. The digital records alone reach inside the indexing phase exactly when , so
Below a raster share of the digitiser rate is not in the expression at all. Above it, the digitiser rate enters and, as approaches one, dominates. For half the archive, the crease sits at a raster share of exactly one half.
The bench
The exhibit below computes that surface rather than drawing it. Left to right is the raster share of the legacy logs, from 0 to 100%. Back to front is months since the platform went live, 0 to 36. Height is the share of the archive a model can read. The amber sheet is half of the archive; the white curtain is the raster share you set, and the labels follow it. The bright line across the surface is the month the index completes.
The archive opens at our own recorded counts, 136,771 rasters and 7,781 LAS files, so the raster share starts at 94.6% [5]. The digitiser default is derived on the plate: 604,800 seconds in a week over 30 seconds a metre times 1,000 metres, which is 20.16 records a week for one instance running round the clock. The 30 seconds a metre is our recorded machine interpretation rate on borehole image logs [6]; the 1,000 metre interval and the single instance are stated assumptions. A raster curve digitiser is a different model from the one that rate was measured on, and the digitiser slider is where your own rate goes. The indexing default of 5,000 records a week is an assumption too: we did not record our own indexing rate, and the plate says so.
What the bench shows
At the opening settings the index completes at month 6.6 and 5.7% of the archive is readable at month 12. Half of it is readable in year 61.9. The readout for what it would take to have half the archive readable by month 36 says 505 records a week, which is 25 times the derived default. At a raster share this high and one digitiser instance, the AI layer on the platform is waiting on the digitiser, and it will be waiting for a working lifetime.
Now drag the raster share down through one half and watch the half-readable readout. At 49% it says month 6.5, and it says month 6.5 at every digitiser rate the slider reaches, because the branch that computes it does not contain the digitiser rate. At 51% it says month 23.1 at the default digitiser and month 7.3 at 500 a week. Two points of raster share on either side of one half, and the question of which rate to buy has a different answer. Whether the digitiser matters to the half-readable date is decided by the raster share and by nothing else, which is why it is the number to ask for first.
The indexing rate is the other lever, and the bench shows where it is strong and where it is not. Set it to 500 records a week and the index completes in year 5.5, past the frame: the month-12 readout is then 1.0% at our recorded share and it does not move when the digitiser slider is dragged across its whole range, because nothing raster is readable before the crease. Set it to 50,000 a week and the index completes at month 0.7, but at 94.6% raster the half-readable date only comes forward from year 61.9 to year 61.4, because after the crease the digitiser owns the slope. A tenfold faster index buys six months off 62 years. A digitiser at 5,000 a week, with the index back at its default, brings the same date to month 9.6.
That is the shape of the finding. A raster share below one half puts the platform in an indexing problem, and it is the kind of problem that finishes: a census is a data-engineering job with an end. A raster share above one half puts it in a digitising problem, and the arithmetic is unforgiving at the scale of a real archive, because every raster record has to be read once by something that is not a person. The four releases at the top of this piece do not say which problem their platforms are in. Our archive was in the second one, by a margin of 17.6 to 1.
Where the 30 seconds a metre comes from, and what it is not
The one machine rate on the plate is a measured one, and it is worth being exact about what it measured. For a mid-sized Middle East carbonate operator, manual sinusoid picking on high-resolution borehole image logs ran to days of expert time per well, and the productised pipeline interpreted at roughly 30 seconds per metre, which the operator's own figures put at about five times the manual baseline across a 16-well programme [6]. The same operator held more than 80 processed image logs against a fixed expert headcount, and the scoping made plain that no realistic headcount clears a backlog that grows with every well drilled [7].
That is an interpretation rate on image logs, not a digitising rate on scanned curves, and the bench does not pretend otherwise. It is on the plate because it is the one per-metre machine rate we have published, and because the derivation from it to a per-week figure is short enough to print. What survives any digitiser rate you put on the slider is the structure: the plateau before the index completes, where the digitiser cannot help, and the crease at one half, where the digitiser starts to matter. Those two features are properties of the pipeline, not of our numbers.
Three questions for a platform that is fully deployed
What share of the well records behind the platform are raster? Not petabytes and not years of history: the count of scanned records against the count of digital ones. On our archive that ratio, once measured, reframed the entire engagement [5], and it is the coordinate that decides which regime the platform is in.
Is the index complete, and when was the census taken? Until it is, the digitiser is not the constraint and buying a faster one changes nothing, which the bench shows by leaving the month-12 readout where it was.
What is the digitiser rate, in records a week, and does the AI layer read its output or the scans? An agent that can connect multiple data sources [4] connects what is addressable. If the answer to the first question is above one half, the digitiser rate is the number that sets the date, and the platform's release will not carry it, because none of these four did.
Key takeaways
- CNOOC Limited says it fully deployed the Haineng-Zhiqing digital platform, PetroChina that it vigorously implemented its Artificial Intelligence Plus initiative, Sinopec that it launched the Fenghuo industrial AI agent, and PETRONAS that Iraya Energies will add agentic AI to myPROdata. All four are statements about a platform or an initiative, and none states a count of well records, a raster share, or a rate.
- On the one archive we counted, 136,771 of 144,552 records were scanned TIF rasters, 17.6 to 1 against the LAS files. The first deliverable was a catalog, because deduplication needs the whole census and an honest train and validation split for the digitiser needs the deduplication. The digitiser waited for the index.
- In a two-stage pipeline where raster records are readable only after indexing and then digitising, the digitiser rate does not appear on the surface before the index completes, and the date at which half the archive is readable has a crease at a raster share of exactly one half: below it the indexing rate alone sets the date, above it the digitiser rate takes over.
- At the opening settings, our recorded raster share and one digitiser instance at 20.16 records a week derived from our recorded 30 seconds a metre, 5.7% of the archive is readable at month 12 and half of it in year 61.9; half by month 36 needs 505 records a week, 25 times the default.
- Two points either side of one half give different answers: at 49% raster the half-readable date is month 6.5 at every digitiser rate, at 51% it is month 23.1 at the default digitiser and month 7.3 at 500 a week.
- A tenfold faster index moves the half-readable date at our recorded share from year 61.9 to year 61.4; a digitiser at 5,000 a week moves it to month 9.6. Which lever matters is set by the raster share, which is the number to ask an NOC data platform for first.
References
[1] CNOOC Limited. CNOOC Limited Focuses on Value Creation, Production and Profit Hit New Highs in H1 2026. 26 August 2026. https://www.cnoocltd.com/english/presscenter/pressreleases/2026/202608/t20260826_122439.html
[2] PetroChina Company Limited. PetroChina Achieves a Strong Start for "the 15th Five-Year Plan": Interim Operating Results for the First Half of 2026 Hit New Record Highs. PR Newswire APAC, 30 August 2026. https://www.prnewswire.com/apac/news-releases/petrochina-achieves-a-strong-start-for-the-15th-five-year-plan-interim-operating-results-for-the-first-half-of-2026-hit-new-record-highs-302864395.html The company's own site returned HTTP 412 to our request, so the wire copy of the company's release is cited.
[3] China Petroleum and Chemical Corporation (Sinopec Corp). Press Release: Sinopec FY2026 Interim Results. EQS Newswire, 23 August 2026. https://www.eqs-news.com/news/corporate-news/en-press-release-sinopec-fy2026-interim-results/514ef82e-238a-4aa2-9f2d-32f1161d68b0_en
[4] PETRONAS. PETRONAS Advances Malaysia's Upstream Data Platform with Agentic AI to Accelerate E&P Investment. 3 September 2026. https://www.petronas.com/media/media-releases/petronas-advances-malaysias-upstream-data-platform-agentic-ai-accelerate-ep
[5] EarthScan. Indexing a 136,771-Scan Raster Archive for ML Ingestion. https://earthscan.io/case-studies/indexing-136771-scan-raster-archive
[6] EarthScan. From weeks to hours per well: the throughput dividend of an automated interpretation pipeline. https://earthscan.io/case-studies/weeks-to-hours-per-well-interpretation-throughput-roi
[7] EarthScan. Closing the Interpretation Capacity Gap: 80+ Wells of Image-Log Backlog Against a Fixed Expert Headcount. https://earthscan.io/case-studies/interpretation-capacity-gap-80-wells-backlog-roi




