Most write-ups of a model that partly fails stop at the measurement. We had the measurement. Read blind on horizontal wells it never trained near, our vertical-trained fracture detector held on depth and detection but blew its azimuth: a mean error of 65.4 degrees, against roughly 10 degrees on vertical wells, which is close to reading a compass at random. We accounted for that gap and its physics in a companion piece, From 82% on Paper to 63% Blind. This case-study is about the harder, less-told half: what a major operator in Oman and their expert interpreter actually did with a number that says one of the model's four outputs is untrustworthy on a whole class of wells. The answer was not to retrain until the number looked better. It was to write a deployment protocol around the boundary and ship the model with the boundary drawn on it.
The one line of measurement this case rests on
The full accounting lives in the companion research piece, so one paragraph here is enough. On a blind July 2023 run at a 0.55 probability threshold and the 6.0 cm depth operating point, the detector scored F1 60.87%, with depth MAE 2.11 cm and dip MAE 11.12 deg holding at workable levels, while azimuth MAE reached 65.4 deg. Three of four outputs transferred with a haircut; one collapsed. The reason is geometric and we do not re-derive it here: a horizontal borehole flattens the fracture sinusoid whose phase carries azimuth, so azimuth is the output most coupled to the exact thing that changed. That physics, and why a flattened curve punishes the angle before the depth, is set out in Keypoint Detection vs Direct Dip/Azimuth Regression. What matters for this piece is only the shape of the result: not a uniform drop, but a clean split between outputs that held and one that broke.
The instrument above is that split in one frame. It is not the argument of this case-study; it is the input to it. The argument starts one step later, at the question the operator asked when they saw it: given that the model is right about depth and wrong about azimuth on these wells, how do we put it to work without either throwing away the good outputs or trusting the bad one?
The protocol: agree the tolerances output by output
The first thing the azimuth number changed was how we defined "correct" at all. A single accuracy figure assumes one bar for the whole model. That assumption is what hides a broken output inside a decent average, and it was the first thing we and the operator's interpreter threw out.
Instead we sat with the interpreter and fixed a permissible-error tolerance per output, tied to what a human picker's own hand-to-hand variability actually is on this log set. A predicted fracture counts as a hit only if its depth lands within tolerance of the interpreter's pick, and dip and azimuth error are then scored on matched hits alone. The agreed relaxations, output by output, were 2, 4 and 6 cm on depth, 1, 3 and 5 degrees on dip, and 5, 10 and 15 degrees on azimuth. Those azimuth bands are the tell. On vertical wells the model cleared the 5-to-15-degree azimuth window comfortably. On horizontal wells it reached usable accuracy only when the azimuth tolerance was relaxed toward 90 degrees, roughly six times looser than the band the interpreter would accept. A tolerance the expert would not sign off on is not a tolerance you deploy behind. So azimuth on horizontal wells did not get a looser bar. It got taken out of the automated path.
Before
One blended score, one bar for all outputs
The leaderboard default: a single F1 that averages a survivable depth error and a catastrophic azimuth error into one unremarkable number
After
Per-output tolerances, agreed with the interpreter
Depth 2/4/6 cm, dip 1/3/5 deg, azimuth 5/10/15 deg; azimuth on horizontal wells cannot meet its band and is routed to the human instead
The broken output is isolated, not buried
The hand-off: a pre-processing aid, not a replacement
With the tolerances agreed, the deployment decision wrote itself, and it is a decision worth stating plainly because it is the opposite of how these models are usually sold. The detector was deployed as a pre-processing aid that routes candidate fractures to the interpreter, not as a replacement for the interpreter's eye. On vertical wells, where all four outputs met their agreed bands, that aid runs close to end to end: it clears the high-F1 backlog and the human spot-checks it. On horizontal wells, the same tool runs deliberately short of its outputs. It surfaces where a fracture is and how deep, because depth held, and it explicitly declines to assert which way the fracture faces, because azimuth did not. The interpreter picks azimuth by hand on exactly those wells, and the tool's job there is to narrow the search, not to answer.
That division is the whole product. The model earns its keep on the wells and outputs where it meets a bar a human expert set, and it hands back the one output, on the one geometry, where it cannot. An interpreter who is told "trust the depth, distrust the horizontal azimuth, here are the candidate locations" can work faster and still be right. An interpreter handed a single confident-looking score has no way to know that a horizontal azimuth reading is noise, and will eventually feed that noise into a stress or permeability-anisotropy call and discover the error the expensive way.
Why the disclosure was the commercial move, not the risk
The tempting alternative was to describe the horizontal work as "extended to deviated wells," report the outputs that held, and let the reader assume azimuth came along. We put the azimuth failure in front of the operator instead, and it is worth being precise about why that was the stronger commercial position and not a confession.
A model that ships with a stated boundary is a model an interpreter can build a workflow around. The boundary is what makes the good outputs usable: because the tool is explicit that horizontal azimuth is out of scope, the interpreter keeps trusting it for depth and fracture presence without hedging everything. A model that hides the boundary buys one demo and loses the account on the first bad azimuth that reaches a decision. This exact axis had already been the sharp end of peer review on the underlying method: one reviewer pressed hard that a model tested only on its own kind of well overstates its reach, a fight recounted in The Reviewer Said Unpublishable, and a second reviewer's 33 comments forced every implicit claim to be stated as a measured fact, told in 33 Comments That Turned a Rejected Manuscript Into a Journal Paper. The deployment protocol here is the same discipline pushed past the paper and onto the rig: a transfer limit you have measured, named, and written into the acceptance criteria is a stronger claim than a limit you have papered over, because it tells the buyer precisely where the model's competence ends and where their interpreter still has to carry the well.
What transfers beyond one reservoir
The reusable artifact is not the azimuth number. It is the sequence that turned a broken output into a shippable product: decompose the model's score by output rather than reporting one blend, set the acceptance tolerance for each output with the human expert who will consume it, deploy the model only inside the outputs and geometries that clear their agreed bands, and hand the rest back explicitly. A blended metric would have hidden the whole problem and shipped a tool that was quietly wrong a quarter of the time. The per-output split, plus the interpreter-in-the-loop hand-off it justified, is what turned "trust this model" into the far more defensible "trust these three outputs on these wells, and this is the one we route to you." That protocol, not the model weights, is what we carry to the next transfer question.
Limitations
The protocol described here rests on figures from a single confidential carbonate engagement in Oman and a blind evaluation on five horizontal wells, which is a small sample and is itself part of why azimuth degraded so far; the underlying metrics and their physics are accounted for in the companion research piece and are only summarised here. The per-output tolerance bands (2/4/6 cm depth, 1/3/5 deg dip, 5/10/15 deg azimuth) were agreed with this operator's expert interpreter for this log set and are not a universal standard; another interpreter or basin would set their own. The vertical comparators for depth, dip and F1 are drawn from the combined vertical model as reference points and are not the same training run as the horizontal split, so the paired bars should be read as regime-versus-regime context, not a controlled ablation. The 6.0 cm depth operating point is one choice on a precision-recall curve; a different tolerance moves the detection numbers, though not the azimuth conclusion or the hand-off decision it forced. All figures are method-level results on this operator's logs under confidentiality, not a benchmark to quote against a different tool stack or basin.
References
- Blind horizontal-well fracture-detection metrics (F1 60.87% at 6 cm; MAE depth 2.11 cm, dip 11.12 deg, azimuth 65.4 deg; inference threshold 0.55, July 2023; five-well horizontal training ceiling) and the vertical comparators, accounted for in full in the companion research piece From 82% on Paper to 63% Blind and derived from internal blind-evaluation figure sets on a 14-vertical-well, five-horizontal-well confidential carbonate dataset in Oman; data and code withheld under operator confidentiality.
- Per-output permissible-error tolerances (depth 2/4/6 cm, dip 1/3/5 deg, azimuth 5/10/15 deg), agreed with the operator's expert interpreter and used as the acceptance criteria behind the deployment protocol; drawn from the engagement's evaluation-design records, anonymised per house style.
- The depth-versus-angle recoverability argument that explains why azimuth breaks first, treated in Keypoint Detection vs Direct Dip/Azimuth Regression.
- The peer-review pressure on transfer claims that the boundary-disclosure answers, recounted in The Reviewer Said Unpublishable and 33 Comments That Turned a Rejected Manuscript Into a Journal Paper.




