Skip to main content
Reading viewAll insights →
BLOG17 min read

The Digital Expert Reads the Label: What a Large Model Can and Cannot Take From a Well File

The Digital Expert Reads the Label: What a Large Model Can and Cannot Take From a Well File
Tarry Singhby Tarry SinghFounder & CEO · 8 Oct 2026
Share

In one fortnight Sinopec announced what it calls the industry's first digital expert, the Fenghuo industrial AI agent, alongside its Great Wall large model; PetroChina said it had vigorously implemented an Artificial Intelligence Plus initiative; and YPF said its renewed Corva agreement now aims at digital copilots and agents for well construction. Three operators, three announcements, and not one word about what the model reads. On the scans we digitised, the answer splits in two: the model resolves the header taxonomy and the depth scale reliably, and it recovered only 0.37 and 0.32 of the two curve masks against 0.97 on the background. So a digital expert answers a question about the label far better than a question about a value, and this piece computes by how much. The instrument prints the field count at which the two kinds of question are equally reliable, 38.4 on our opening settings, and that number does not move when you drag the stain or the skew: on an unrestored sheet both answers carry the same damage exponent, so the ordering is settled on the clean sheet. Deskew and stain normalisation are the only controls that move it, because they are the only ones that treat printed header glyphs and faint curve ink differently.

Sinopec's interim results of 23 August 2026 report the launch of what the company calls "the industry's first digital expert", the "Fenghuo" industrial AI agent, and say the capabilities of its Great Wall large model "further improved" [1]. A week later PetroChina said it had "vigorously implemented" its "Artificial Intelligence Plus" special initiative and had "promoted the deep integration of digital and intelligent technologies" with the energy and chemical industry [2]. On 1 September YPF said the renewal of its agreement with Corva enters a stage aimed at "operaciones progresivamente autónomas", progressively autonomous operations, by way of predictive analytics, intelligent alerts, automation, and "copilotos digitales y agentes", digital copilots and agents able to assist operational processes of growing complexity [3].

All three are the operator's own words. Two of them reached us on a wire rather than from the company's own domain, because Sinopec's own sites refuse HTTPS and PetroChina's returned HTTP 412 to our request, so the wire copy of each company's release is what is cited below. None of the three is a vendor writing about a customer, and none of the three is trade press.

What none of the three says is what the model reads. A digital expert, a copilot, an agent: each of those words describes a thing that answers questions. The question that decides whether it can is whether the file it is pointed at is a LAS curve or a photograph of ink on paper, and if it is the second, which parts of that photograph a model has actually recovered. That is not a criticism of the releases. Two are results announcements and one is a partnership notice; none claims to be a data inventory. It is the reason this piece exists, because on the archives we have digitised the answer is not one number. It is two, and they are very far apart.

The two questions a digital expert is asked

Put a language model over a well archive and the questions arriving at it fall into two shapes.

A LABEL question is answered from the header. Which curves are in this file, in which track, over which depth interval, on which scale. The header is a block of printed text at the top of the log, and printed text is exactly what optical character recognition is good at [5]. Answering a label question means reading some number of header fields and getting all of them right.

A VALUE question is answered from a curve. What is the resistivity at 2,412 metres. Answering it means two things must have worked: the depth scale must have been read, because without it a mask is a shape in image coordinates with no physical meaning [5], and the curve itself must have been recovered from the ink. Those are different problems solved by different components, and only the first is text.

The gap between the two is the whole product. A model that knows this file is a resistivity log over 2,300 to 2,600 metres, and cannot tell you the resistivity at 2,412 metres, has answered the label question and failed the value question. Every one of the three announcements above is a statement about a system that answers questions. Not one of them says which of the two shapes it answers.

What we measured, and on what

Our own record on this splits the same way, and it splits sharply.

The header reader we built resolves the taxonomy of a 3-track log before the segmenter runs: Track 1 carrying gamma ray, spontaneous potential and caliper, Track 2 carrying the shallow, medium and deep resistivity triple, and Track 3 carrying neutron porosity and bulk density, so 8 named curves over 3 tracks, with the depth scale as the load-bearing field that is read first and that every depth-indexed value hangs from [5]. That is 9 fields for a full inventory answer, and it is the default the instrument below opens at.

The curve segmenter's per-class recall is the other half, and it is the half nobody quotes in a launch release. On the first production run over eight real field scans, the background mask recalled at 0.97 and the two curve masks recalled at 0.37 and 0.32 [4]. The model almost never missed the empty page and found roughly a third of the curve pixels it should have found. Those three numbers were used at the time as a routing instruction rather than as a score: the background was trusted and auto-accepted, and all 16 curve regions across the eight scans went to a human reviewer, whose corrected curves landed at a mean absolute error of 0.11 and 0.12 against the operator's own log [4].

The third recorded thing is the ceiling, and it belongs to the sheet rather than to the model. Those scans were paper folded in a drawer for three decades, photographed on a flatbed at whatever angle the technician dropped them, and in more than one case stamped with the ring of a coffee cup. The clean-image ceiling the restoration stage exists to reach was an intersection over union of 0.51 with recall 0.97 and F1 0.55, and the three repairs ran in one order for a reason: deskew first, because a rotated sheet puts every later step in the wrong coordinate frame; fold repair second; stain and shadow normalisation last [6].

Three numbers, one taxonomy, one ceiling. Everything below is built from those and says so where it is not.

The arithmetic

Write the per-field header read accuracy on a clean square scan as hh, and a class recall on the same clean scan as RR. A label question needing FF header fields is a conjunction over those fields, and a value question is the depth scale times one curve:

The two answers, as shares of the questions asked
label(F)  =  h FEh,value(R)  =  h Eh R Ec\mathrm{label}(F) \;=\; h^{\,F E_h}, \qquad \mathrm{value}(R) \;=\; h^{\,E_h}\, R^{\,E_c}

The exponents are where damage enters, and the form is not arbitrary. Read a class mask as a per-pixel decision against a fixed threshold with the evidence per pixel exponentially distributed. A clean recall RR then fixes that threshold at t=−ln⁡Rt = -\ln R. Let damage divide every pixel's evidence by E≥1E \ge 1, and the recall becomes e−tE=REe^{-tE} = R^{E}. One exponent acts on every class and no per-class fragility constant is invented: the whole spread between the classes under damage is carried by the three recalls we recorded. For the fold and stain intensity ss from 0 to 1 and the skew angle aa in degrees,

The damage exponent on an unrestored sheet, both channels
E  =  1  +  s  +  a10E \;=\; 1 \;+\; s \;+\; \frac{a}{10}

so the clean square corner sits at E=1E = 1 and reproduces the recorded numbers exactly, and the worst corner of the frame sits at E=3E = 3.

Now the part that decides the piece. On an unrestored sheet the header channel and the curve channel carry the same exponent, so the ratio of the two answers is

The ratio between the two answers, unrestored
label(F)value(R)  =  (h F−1R) ⁣E\frac{\mathrm{label}(F)}{\mathrm{value}(R)} \;=\; \left(\frac{h^{\,F-1}}{R}\right)^{\!E}

which is monotone in EE but never changes sign. The two surfaces cannot cross. The field count at which they are equally reliable,

The field count at which a label question and a value question are equally reliable
F∗  =  1  +  EcEh ln⁡Rln⁡hF^{*} \;=\; 1 \;+\; \frac{E_c}{E_h}\,\frac{\ln R}{\ln h}

collapses to 1+ln⁡R/ln⁡h1 + \ln R / \ln h when the two exponents are equal, which is to say it is the same number at every stain level and every skew angle on the sheet. Damage moves both answers down together and decides nothing about which kind of question is safer to ask.

Restoration is the only thing in the model that breaks that, because it is the only thing that treats printed header glyphs and faint curve ink differently, and the asymmetry is the restoration stage's own stated limit. Deskew serves both channels and is the most consequential single repair, so it leaves a residual 0.15 of the skew term. Stain normalisation divides out a slowly varying background, which is most of what stands between OCR and printed header text, so it leaves 0.25 of the stain term there. It does far less for curve ink, because a dark sharp-edged stain that overlaps a curve at the curve's own spatial frequency cannot be divided out without taking the curve with it [6], so it leaves 0.60. Those three residuals are our model of the repair, not a measurement of it, and they are the only place the two channels are given different numbers.

The bench

The exhibit below computes the two answers rather than drawing them. Left to right is the fold and stain intensity of the incoming sheet, 0 to 100%. Back to front is the skew angle it was scanned at, 0 to 10 degrees. Height is the share of questions answered. The aqua surface is the label question at the field count you set; the steel and dark surfaces are the value question at the bold and the faint curve. The white curtain is your own stain level and the labels follow it, and where a label question and a faint-curve value question are exactly as reliable as each other, a pin marks the stain at which they meet.

The clean corner is anchored to the recorded numbers and nothing else. Everything about how the surfaces fall away from it is our model, stated on the plate: the exponent, the two unit weights in it, and the three restoration residuals. The per-field header accuracy is a control, because we never measured ours, and its opening value of 97% is a starting point rather than a figure.

WHAT THE FILE ANSWERS4.38xLABEL ANSWER OVER FAINT CURVE VALUE ANSWER, AT THIS SCANlabel answer, 9 header fields63.6%faint curve value answer14.5%fold and stain intensity, 0% left to 100% rightskew angle, 0 deg at the back to 10 deg at the frontshare of questions answered, 0 to 1LABEL ANSWER63.6%FAINT CURVE VALUE ANSWER14.5%BOLD CURVE VALUE ANSWER18.4%THE TWO SWAP AT38.4 fieldsAHEAD HEREthe label answerlabel answerbold curve valuefaint curve valueyour stainyour skewRecorded, ours: background recall 0.97, bold curve 0.37, faint curve 0.32; header 3 tracks, 8 curves, 1 depth scale.MODELLED, not measured: recall = clean ^ (1 + stain + skew/10); repair keeps 0.15 skew, 0.25 stain on text, 0.60 on ink.
Three surfaces over the same plane of scan conditions: fold and stain intensity left to right, skew angle back to front. Height is the share of questions answered. The aqua surface is a label question, answered from the header alone and needing every one of the fields you set to be read correctly. The steel and dark surfaces are value questions, answered from the depth scale plus one curve, at the bold and the faint curve recalls we recorded. The clean corner uses exactly those recorded recalls; the response to damage is our model of it, stated on the plate. Leave restoration off and drag the stain and skew controls anywhere: the swap readout does not move, because without restoration both answers carry the same exponent and their ratio does not depend on the scan. Now switch restoration on and set the field count near the swap readout: it stops being a constant, and a pin marked the two answers meet gives the stain at which they are equally reliable. Drag or use the arrow keys to orbit, Home to reset.

What the bench shows

The plate opens on a moderately damaged sheet, 35% stain at 3.0 degrees of skew, asking the 9-field inventory question against a 97% per-field header. It reads 63.6% on the label answer and 14.5% on the faint curve, a ratio of 4.38 to 1. Drag both controls back to the clean square corner and the two answers are 76.0% and 31.0%, a ratio of 2.45 to 1. Push them to the far corner, 100% stain at 10 degrees, and they are 43.9% and 3.0%, a ratio of 14.7 to 1. Underneath those two readouts the class recalls themselves, which the plate does not print, move by very different amounts over the same span: the background class from 0.970 to 0.913 and the faint curve from 0.320 to 0.033. One exponent acts on both, and that is exactly why they separate: the strongest class loses under 6 points of recall while the weakest loses almost everything it had. The strongest class is also the one nobody queries.

Now the finding, and it is in the readout marked THE TWO SWAP AT. With restoration off, drag the stain anywhere from 0 to 100% and the skew anywhere from 0 to 10 degrees. The readout says 38.4 fields and stays at 38.4 fields, in every one of those states. Not approximately: the branch that computes it multiplies by the ratio of the two exponents, and with restoration off that ratio is exactly 1 at every point of the plane. So on an unrestored sheet the question of which kind of question is safer to ask has already been settled before the sheet is scanned, by the field count and the header accuracy alone. Damage is not a tiebreaker.

Set the field count to 44, which is past that boundary, and the aqua surface drops below the dark one everywhere at once: 26.2% against 31.0% at the clean corner, 1.8% against 3.0% at the worst. The whole surface flips together and the two never touch, which is what parallel exponents look like.

Then switch restoration on, at 44 fields. The swap readout stops being a constant. It reads 38.4 fields at the clean corner and 48.9 fields at 100% stain with the sheet square, and the meeting pin appears: at 0 degrees of skew the two answers are equal at a stain of 47.8%, and at 10 degrees at 55.0%. On the clean side of that boundary a value question at the faint curve is the more reliable one, and on the damaged side the 44-field label question is. It exists only because deskew and stain normalisation do not help the header and the curve by the same amount.

That is the shape of it. Scan damage sets how much a digital expert can answer. Restoration sets what it can answer. They are different questions with different budgets, and the second one is the one the three releases at the top of this piece are silent about.

Where our numbers stop and our model starts

Being exact about that boundary matters more here than usual, because the finding is a structural claim and a reader is entitled to ask which part of it is load-bearing.

Measured, ours: the three recalls, 0.97, 0.37 and 0.32, on eight real field scans against the operator's own logs [4]. The 3-track, 3 + 3 + 2 taxonomy and the depth scale the header reader resolves, which is where the 9-field default comes from [5]. The clean-image ceiling of 0.51 intersection over union with recall 0.97 and F1 0.55, and the deskew-first ordering of the repairs [6]. Nothing else on the plate is a measurement.

Modelled, ours, and labelled that way on the plate: the exponential-evidence response, the two unit weights in the exponent, and the three restoration residuals. The case study that describes the repair stage says plainly that its per-stage increments are a teaching model of which damage costs what rather than measured per-stage scores [6], and nothing here upgrades them.

What survives whatever you do to those assumptions is the structure, and it is worth naming precisely what is structural. That both answers carry an exponent is the modelling choice. That the label answer is a conjunction over fields and the value answer is a product with one weak factor is not a choice: it is what those two questions are. And given both, the flat boundary follows from the two channels sharing an exponent, which is true of any damage that does not distinguish printed text from curve ink. Give the two channels different responses to raw damage and the boundary tilts even before restoration. The instrument gives them the same one, on the honest ground that we have no measurement that separates them on an unrepaired sheet.

The other honest limit is the one the header reader carries. It is only as good as the header it is given, and a sheet whose header is torn, stamped over or missing pushes the whole label answer back onto defaults and human review [5]. That is what the header accuracy control is for. Drag it down and watch the swap readout fall with it, which is the same finding read from the other end: a poor header does not make a value question easier, it makes a label question harder.

Three questions for a digital expert

What share of the archive behind it is scanned paper rather than curves, and for the paper share, what is the per-class recall? Not an average and not an intersection over union on a clean validation set. The number that decides what a model can answer is the recall on the class the question is about, and on our scans the spread across classes was 0.97 against 0.32 [4].

Which of the two questions is the system being sold on? A digital expert [1] that resolves what is in a file is answering from the header, and that is real and worth having. An agent that assists an operational process [3] is doing something else again. Neither is the same as returning a value at a depth, and the vocabulary of these releases does not distinguish them.

Has the restoration stage run, and what did it leave behind on each channel? On our archive the fold, the skew and the stain arrived before the model did, and the repairs were where the recoverable share was actually decided [6]. The bench says that repair is the only lever that changes which kind of question is safer to ask. A platform release will not carry that number, because none of these three did.

Key takeaways

  1. Sinopec reports launching what it calls the industry's first digital expert, the Fenghuo industrial AI agent, alongside its Great Wall large model; PetroChina says it vigorously implemented an Artificial Intelligence Plus initiative; YPF says its renewed Corva agreement now aims at digital copilots and agents for well construction. All three are the operator's own words, two carried on a wire because the companies' own domains refused our request, and not one describes what the system reads.
  2. Questions to a model over a well archive split in two. A label question is answered from the header and needs every field it asks for to be read correctly. A value question needs the depth scale plus the one curve it asks about. On the eight scans we digitised, the background mask recalled at 0.97 and the two curve masks at 0.37 and 0.32, so the two shapes of question are not close to equally answerable.
  3. At the opening settings, a moderately damaged sheet at 35% stain and 3.0 degrees of skew asking the recorded 9-field inventory question, the label answer is 63.6% and the faint-curve value answer 14.5%, a ratio of 4.38 to 1. At the clean corner it is 76.0% against 31.0%; at 100% stain and 10 degrees it is 43.9% against 3.0%, a ratio of 14.7 to 1.
  4. The finding: on an unrestored sheet both answers carry the same damage exponent, so the field count at which they are equally reliable is 38.4 in every state of the plane, and the two surfaces never cross. No amount of fold, stain or skew changes which kind of question is safer to ask.
  5. Restoration is the only control that tilts that boundary, because deskew and stain normalisation do not help printed header glyphs and faint curve ink by the same amount. Switch it on at 44 fields and the swap readout runs from 38.4 fields at the clean corner to 48.9 at full stain, and the two answers become equal at 47.8% stain when the sheet is square and 55.0% at 10 degrees of skew.
  6. The clean corner is our recorded per-class recall and our recorded header taxonomy. The response to damage and the three restoration residuals are our model of it, printed on the plate as such, and the per-field header accuracy is the reader's own control because we never measured ours.

References

[1] China Petroleum and Chemical Corporation (Sinopec Corp). Press Release: Sinopec FY2026 Interim Results. EQS Newswire, 23 August 2026. https://www.eqs-news.com/news/corporate-news/en-press-release-sinopec-fy2026-interim-results/514ef82e-238a-4aa2-9f2d-32f1161d68b0_en Sinopec's own domains refuse HTTPS, so the wire copy of the company's release is cited.

[2] PetroChina Company Limited. PetroChina Achieves a Strong Start for "the 15th Five-Year Plan": Interim Operating Results for the First Half of 2026 Hit New Record Highs. PR Newswire APAC, 30 August 2026. https://www.prnewswire.com/apac/news-releases/petrochina-achieves-a-strong-start-for-the-15th-five-year-plan-interim-operating-results-for-the-first-half-of-2026-hit-new-record-highs-302864395.html The company's own site returned HTTP 412 to our request, so the wire copy is cited.

[3] YPF. YPF y Corva renuevan un acuerdo clave para el RTIC de Perforación y Workover. 1 September 2026. https://novedades.ypf.com/ypf-corva-renovacaciondelacuerdo.html

[4] EarthScan. Eight Real Scans, One Reviewer: Standing Up Human-in-the-Loop Validation. https://earthscan.io/case-studies/human-in-the-loop-review-eight-real-scans

[5] EarthScan. Adding an OCR Header Reader to Auto-Detect Tracks and Depth Scales. https://earthscan.io/case-studies/ocr-header-reader-auto-detect-tracks

[6] EarthScan. The Coffee Ring and the Crease: Recovering Folded, Stained, Skewed Field Scans. https://earthscan.io/case-studies/deskew-and-fold-stain-recovery-on-real-field-scans

Tarry Singh
Tarry Singh

Founder & CEO

More from EarthScan

Related research

All insights →
Fully Deployed Is Not Machine-Readable: The Raster Share Behind Four NOC Data Platforms
Insight

Fully Deployed Is Not Machine-Readable: The Raster Share Behind Four NOC Data Platforms

12.2% Against What? The Missing Baseline Behind an AI Production Uplift
Insight

12.2% Against What? The Missing Baseline Behind an AI Production Uplift

More Data, A Narrower Band, Less Truth Inside It: Ensemble Collapse Under a Reservoir Twin
Insight

More Data, A Narrower Band, Less Truth Inside It: Ensemble Collapse Under a Reservoir Twin

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.