Skip to main content
Case studya major operator in Oman

The Generalization Cliff: Turning a Broken Azimuth Metric Into a Deployment Protocol

The blind horizontal azimuth metric broke: 65.4 degrees of mean error, close to reading a compass at random, against about 10 degrees on vertical wells. We measured that gap and published it in a companion research piece. This case-study is about the decision that number forced on the ground: instead of shipping one blended score, we and the operator's expert interpreter wrote a per-output hand-off protocol, agreed permissible-error tolerances output by output, and deployed the detector as a pre-processing aid that clears the vertical backlog and flags horizontal fractures for a human rather than pretending to resolve their azimuth.

10 Sep 20269 min read
The Generalization Cliff: Turning a Broken Azimuth Metric Into a Deployment Protocol

Most write-ups of a model that partly fails stop at the measurement. We had the measurement. Read blind on horizontal wells it never trained near, our vertical-trained fracture detector held on depth and detection but blew its azimuth: a mean error of 65.4 degrees, against roughly 10 degrees on vertical wells, which is close to reading a compass at random. We accounted for that gap and its physics in a companion piece, From 82% on Paper to 63% Blind. This case-study is about the harder, less-told half: what a major operator in Oman and their expert interpreter actually did with a number that says one of the model's four outputs is untrustworthy on a whole class of wells. The answer was not to retrain until the number looked better. It was to write a deployment protocol around the boundary and ship the model with the boundary drawn on it.

The one line of measurement this case rests on

The full accounting lives in the companion research piece, so one paragraph here is enough. On a blind July 2023 run at a 0.55 probability threshold and the 6.0 cm depth operating point, the detector scored F1 60.87%, with depth MAE 2.11 cm and dip MAE 11.12 deg holding at workable levels, while azimuth MAE reached 65.4 deg. Three of four outputs transferred with a haircut; one collapsed. The reason is geometric and we do not re-derive it here: a horizontal borehole flattens the fracture sinusoid whose phase carries azimuth, so azimuth is the output most coupled to the exact thing that changed. That physics, and why a flattened curve punishes the angle before the depth, is set out in Keypoint Detection vs Direct Dip/Azimuth Regression. What matters for this piece is only the shape of the result: not a uniform drop, but a clean split between outputs that held and one that broke.

VERTICAL-TRAINED DETECTOR ON BLIND HORIZONTAL WELLS65.4 degblind azimuth MAE, up from ~10 deg verticalDepth and F1 held on blind horizontal data. Azimuth fell off a cliff.Same detector, no retrain: teal is vertical wells (the training regime), the lower bar is the blind horizontal split.vertical wellsblind horizontalthe azimuth cliff0204060Depth MAElower is better1.1 cm2.1139 cmF1 at operating pointhigher is better82%60.87%Dip MAElower is better1.6 deg11.1181 degAzimuth MAElower is better10 deg65.4072 degazimuth is near-uninformative here: ~78% accuracy only near a 90 deg toleranceabsoluteH/V ratioblind horizontal, 6 cm op point:P 56.25% / R 66.32% / F1 60.87%
The vertical-trained fracture detector read on blind horizontal wells with no retrain. Teal is the vertical-well regime the model was built on; the lower bar in each pair is the blind horizontal split. Depth localisation and F1 at the operating point held (depth MAE 2.11 cm, F1 60.87% at the 6 cm point, precision 56.25% / recall 66.32%), and dip degraded only modestly (11.12 deg). The one orange arm is azimuth MAE, which climbed from about 10 degrees on vertical wells to 65.4 degrees on the blind horizontal data - near-uninformative, since azimuth reaches about 78% accuracy only near a 90 degree tolerance. Toggle to H/V ratio to see the roughly 6.5x azimuth blow-up against sub-2x moves on every other metric. Run: ResNet-18 backbone, L1+focal loss, LR 0.0007736, 200 epochs, inference threshold 0.55, July 2023, trained on a five-well horizontal ceiling. The horizontal metrics and the azimuth cliff are the sourced same-run figures; the vertical depth, dip and F1 comparators are drawn from the combined vertical model as reference points, not the same run.

The instrument above is that split in one frame. It is not the argument of this case-study; it is the input to it. The argument starts one step later, at the question the operator asked when they saw it: given that the model is right about depth and wrong about azimuth on these wells, how do we put it to work without either throwing away the good outputs or trusting the bad one?

The protocol: agree the tolerances output by output

The first thing the azimuth number changed was how we defined "correct" at all. A single accuracy figure assumes one bar for the whole model. That assumption is what hides a broken output inside a decent average, and it was the first thing we and the operator's interpreter threw out.

Instead we sat with the interpreter and fixed a permissible-error tolerance per output, tied to what a human picker's own hand-to-hand variability actually is on this log set. A predicted fracture counts as a hit only if its depth lands within tolerance of the interpreter's pick, and dip and azimuth error are then scored on matched hits alone. The agreed relaxations, output by output, were 2, 4 and 6 cm on depth, 1, 3 and 5 degrees on dip, and 5, 10 and 15 degrees on azimuth. Those azimuth bands are the tell. On vertical wells the model cleared the 5-to-15-degree azimuth window comfortably. On horizontal wells it reached usable accuracy only when the azimuth tolerance was relaxed toward 90 degrees, roughly six times looser than the band the interpreter would accept. A tolerance the expert would not sign off on is not a tolerance you deploy behind. So azimuth on horizontal wells did not get a looser bar. It got taken out of the automated path.

How the azimuth cliff changed the acceptance criteria on the ground

Before

One blended score, one bar for all outputs

The leaderboard default: a single F1 that averages a survivable depth error and a catastrophic azimuth error into one unremarkable number

After

Per-output tolerances, agreed with the interpreter

Depth 2/4/6 cm, dip 1/3/5 deg, azimuth 5/10/15 deg; azimuth on horizontal wells cannot meet its band and is routed to the human instead

The broken output is isolated, not buried

The hand-off: a pre-processing aid, not a replacement

With the tolerances agreed, the deployment decision wrote itself, and it is a decision worth stating plainly because it is the opposite of how these models are usually sold. The detector was deployed as a pre-processing aid that routes candidate fractures to the interpreter, not as a replacement for the interpreter's eye. On vertical wells, where all four outputs met their agreed bands, that aid runs close to end to end: it clears the high-F1 backlog and the human spot-checks it. On horizontal wells, the same tool runs deliberately short of its outputs. It surfaces where a fracture is and how deep, because depth held, and it explicitly declines to assert which way the fracture faces, because azimuth did not. The interpreter picks azimuth by hand on exactly those wells, and the tool's job there is to narrow the search, not to answer.

That division is the whole product. The model earns its keep on the wells and outputs where it meets a bar a human expert set, and it hands back the one output, on the one geometry, where it cannot. An interpreter who is told "trust the depth, distrust the horizontal azimuth, here are the candidate locations" can work faster and still be right. An interpreter handed a single confident-looking score has no way to know that a horizontal azimuth reading is noise, and will eventually feed that noise into a stress or permeability-anisotropy call and discover the error the expensive way.

Why the disclosure was the commercial move, not the risk

The tempting alternative was to describe the horizontal work as "extended to deviated wells," report the outputs that held, and let the reader assume azimuth came along. We put the azimuth failure in front of the operator instead, and it is worth being precise about why that was the stronger commercial position and not a confession.

A model that ships with a stated boundary is a model an interpreter can build a workflow around. The boundary is what makes the good outputs usable: because the tool is explicit that horizontal azimuth is out of scope, the interpreter keeps trusting it for depth and fracture presence without hedging everything. A model that hides the boundary buys one demo and loses the account on the first bad azimuth that reaches a decision. This exact axis had already been the sharp end of peer review on the underlying method: one reviewer pressed hard that a model tested only on its own kind of well overstates its reach, a fight recounted in The Reviewer Said Unpublishable, and a second reviewer's 33 comments forced every implicit claim to be stated as a measured fact, told in 33 Comments That Turned a Rejected Manuscript Into a Journal Paper. The deployment protocol here is the same discipline pushed past the paper and onto the rig: a transfer limit you have measured, named, and written into the acceptance criteria is a stronger claim than a limit you have papered over, because it tells the buyer precisely where the model's competence ends and where their interpreter still has to carry the well.

What transfers beyond one reservoir

The reusable artifact is not the azimuth number. It is the sequence that turned a broken output into a shippable product: decompose the model's score by output rather than reporting one blend, set the acceptance tolerance for each output with the human expert who will consume it, deploy the model only inside the outputs and geometries that clear their agreed bands, and hand the rest back explicitly. A blended metric would have hidden the whole problem and shipped a tool that was quietly wrong a quarter of the time. The per-output split, plus the interpreter-in-the-loop hand-off it justified, is what turned "trust this model" into the far more defensible "trust these three outputs on these wells, and this is the one we route to you." That protocol, not the model weights, is what we carry to the next transfer question.

Limitations

The protocol described here rests on figures from a single confidential carbonate engagement in Oman and a blind evaluation on five horizontal wells, which is a small sample and is itself part of why azimuth degraded so far; the underlying metrics and their physics are accounted for in the companion research piece and are only summarised here. The per-output tolerance bands (2/4/6 cm depth, 1/3/5 deg dip, 5/10/15 deg azimuth) were agreed with this operator's expert interpreter for this log set and are not a universal standard; another interpreter or basin would set their own. The vertical comparators for depth, dip and F1 are drawn from the combined vertical model as reference points and are not the same training run as the horizontal split, so the paired bars should be read as regime-versus-regime context, not a controlled ablation. The 6.0 cm depth operating point is one choice on a precision-recall curve; a different tolerance moves the detection numbers, though not the azimuth conclusion or the hand-off decision it forced. All figures are method-level results on this operator's logs under confidentiality, not a benchmark to quote against a different tool stack or basin.

References

  1. Blind horizontal-well fracture-detection metrics (F1 60.87% at 6 cm; MAE depth 2.11 cm, dip 11.12 deg, azimuth 65.4 deg; inference threshold 0.55, July 2023; five-well horizontal training ceiling) and the vertical comparators, accounted for in full in the companion research piece From 82% on Paper to 63% Blind and derived from internal blind-evaluation figure sets on a 14-vertical-well, five-horizontal-well confidential carbonate dataset in Oman; data and code withheld under operator confidentiality.
  2. Per-output permissible-error tolerances (depth 2/4/6 cm, dip 1/3/5 deg, azimuth 5/10/15 deg), agreed with the operator's expert interpreter and used as the acceptance criteria behind the deployment protocol; drawn from the engagement's evaluation-design records, anonymised per house style.
  3. The depth-versus-angle recoverability argument that explains why azimuth breaks first, treated in Keypoint Detection vs Direct Dip/Azimuth Regression.
  4. The peer-review pressure on transfer claims that the boundary-disclosure answers, recounted in The Reviewer Said Unpublishable and 33 Comments That Turned a Rejected Manuscript Into a Journal Paper.
Narendra Patwardhan
Narendra Patwardhan

Research Collaborator

Discuss a programme
Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.