Skip to main content
Reading viewAll insights →
BLOG8 min read

The Mid-Project Data Intervention: Running a Workshop When Data Collection Stalls

The Mid-Project Data Intervention: Running a Workshop When Data Collection Stalls
Quamer Nasimby Quamer NasimML Research Engineer · 23 Jul 2026
Share

Mid-way through a roughly twenty-month applied-AI engagement with a major operator in Oman, the model was not the problem. Well deliveries had slowed enough to threaten the next phase, which needed a minimum of 10-15 wells to start in November 2022. Instead of escalating by email, the team ran a 39-slide interactive workshop that renamed one vague complaint - data collection is slow - as three separate, separately-owned problems: I can't access the data, I don't trust the data in the report, it takes too long to get to it. This is a template for any vendor blocked on the client's data supply chain: name the problems, show the honest infrastructure gap, and ask for a specific resourcing decision.

There is a point in most applied-AI engagements where the burndown chart stops being about the model. On a roughly twenty-month program with a major operator in Oman, building a fracture and bedding detector on borehole image logs, we hit that point halfway through. The code was moving. The wells were not. Deliveries of annotated well data had slowed to the point where the next phase was at risk, and that phase had a hard floor: it needed a minimum of somewhere around 10 to 15 wells to be worth starting, and it was expected to begin in November 2022. The obvious move, the one every delivery lead reaches for first, is an email. Data collection is behind, timelines will shift, please expedite. We did something different, and it is the part of the engagement I would most want another vendor to copy.

We ran a workshop. Not a status meeting with a slide that said "data is a risk," but a 39-slide interactive session built around the data supply chain, whose single job was to convert one blurred complaint into a set of problems the client could act on.

Why an email was the wrong instrument

An escalation email has a structural flaw that has nothing to do with tone. It collapses everything into one channel. "Data collection is moving at a really slow pace" reads as a single problem with a single implied owner, usually the person the email is addressed to, and a single implied fix: try harder. But a stalled data supply chain is almost never one problem. When we sat with the client's own framing, three distinct complaints fell out, and they had different owners, different causes, and different fixes.

The first was access: I can't get to the data. The second was trust: I don't trust the numbers in the report I did get. The third was latency: even when I can get the data and I trust it, it takes too long to reach me. Those are not three symptoms of one issue. Access is a permissions and infrastructure problem. Trust is a lineage and validation problem. Time-to-data is a pipeline and tooling problem. An email that says "please expedite" answers none of them, because it never names them. The workshop's first act was to write all three on the wall as separate lanes.

WHEN THE DATA SUPPLY CHAIN, NOT THE MODEL, IS THE BINDING CONSTRAINTName the stall as three distinct problems, then schedule itNov 2022 phase needed 10-15 wells; deliveries stalled and timelines would shift.Escalating by email keeps one blurred channel; a workshop splits it into owned lanes.THREE NAMED DATA PROBLEMS4-DAY FORMAT x 3 WORKSHOP TYPES1ACCESSI can't access the data2TRUSTI don't trust the data in the report3TIME-TO-DATAIt takes too long to get to the dataDefine requirementsRequirements + data discoveryD1D2D3D4Design practicesData practices + RACID1D2D3D4Architect platformReference architectureD1D2D3D4RESPONSE TO THE STALLEscalate by emailone blurred channelStructured workshopthree lanes, scheduledTHE ASK THAT COMES OUT OF ITClient-side annotation + labelling resourcingunblocks the supply chain and aids knowledge retentionSame stalled supply chain, 8-dimension MLOps maturity gap on the tablenamed + owned + scheduledEach problem gets a lane and a slot in the 4-day format; the resourcing ask is a decision the client can act on.Sourced: three-problem framing, 10-15-well minimum, 4-day x 3-type format. Lane weighting is illustrative.
Two ways a vendor blocked on the client's data supply chain can respond to a stall. The one orange lever flips the response. Escalate-by-email collapses everything into a single undifferentiated thread: the three lanes go grey, nothing routes, and the four-day workshop grid stays dark. Structured-workshop decomposes the same slowdown into the three distinct data problems the deck named - access, trust, and time-to-data - lights each lane teal, and routes them into a four-day format run across three workshop types, ending in a concrete client-side annotation and labelling resourcing ask. The Phase-3 minimum of 10-15 wells expected to start November 2022, the three-problem framing, the eight-dimension MLOps maturity model, and the four-day-by-three-type workshop format are sourced from the mid-engagement data-management workshop deck; the lane routing and the qualitative lane weighting are an illustrative reading of that framing, not measured quantities.

The instrument above is the argument in one frame. Flip the response from escalate-by-email to structured-workshop and watch what happens: the single grey channel splits into three teal lanes, each of which now routes somewhere. That routing is the whole point. Once "the data is slow" becomes three named lanes, each lane can be assigned, scheduled, and resourced. Left as one email thread, it stays un-named, un-owned, and un-schedulable, and it will still be un-scheduled the following month.

Showing the greyed-out slide

A workshop that only names problems the client already feels is a therapy session. The one we ran also had to put a hard, slightly uncomfortable fact on the table, and we chose to do it visually rather than in prose. One architecture slide showed the client's data estate with a region deliberately greyed out, labelled as the area the project team was not aware of when the project commenced.

That is an admission against interest, and it was deliberate. It said, in effect, some of this stall is on us: we did not have full visibility into your infrastructure until we were deep into the heavy engineering, and there is a part of your data landscape we are still partly blind to. Naming your own blind spot on a slide does two things at once. It makes the trust lane concrete, because the client can see exactly where the report numbers they distrust are coming from. And it earns the right to make the ask that comes next, because you have shown you are not simply routing blame back across the table.

The ask: resource the annotation, on the client side

Every intervention workshop has to end in a decision the other party can make, or it was theatre. Ours ended in a specific one: add client-side technical and domain resourcing for data collection and annotation. Not "please expedite," but "assign named people, on your side, to the labelling and data-preparation work, because that is the lane on the critical path."

We framed it as more than an unblocking measure, because it genuinely is. Client-side annotators move wells through faster, and they also build durable knowledge inside the client organisation, so the interpretation capacity outlives the engagement instead of leaving with the vendor. Framed that way, the resourcing ask stops being a favour the client does the vendor and becomes an investment the client makes in its own retained capability. That reframe is what moves a budget line.

What the workshop was measuring against

The session did not float free of a standard. Behind the three-lane map sat an eight-dimension before-and-after model of MLOps maturity, adapted from McKinsey's "Scaling AI like a tech native," covering data management, AI development, deployment, live model operations, the technology stack, governance and security, ML-asset management, and people and skills. Each dimension carried a concrete contrast between where the program was and where a mature operation would be, so the stall read as a known gap on a known ladder rather than a crisis. The point is not to grade the client. It is to show that the slow-data problem is a predictable stage, not a failure, which lowers the defensiveness that kills these conversations.

The format was productized, not improvised

The reason we could run this as a workshop rather than a difficult meeting is that the format was pre-built. The data-practice method was packaged as a four-day workshop, run across three workshop types: one to define requirements and discover the data, one to design the data practices, and one to architect the data platform, each with day-by-day activity, deliverable, and participant grids. That structure is doing quiet work. A blocked vendor who calls a meeting looks like a vendor with a complaint. A vendor who arrives with a four-day, three-type format and a day-one agenda looks like a vendor running a known process. The productized format is what let a moment of project risk read as competence rather than escalation.

For the underlying reason the well count mattered so much in the first place - the accuracy curve that tracked labelled-well supply, not algorithm changes - see The Data Bottleneck Is the Real Bottleneck in Subsurface AI. This piece is the process companion to that one: given that data supply is the binding constraint, this is how you intervene when it stalls mid-flight.

The template, for any vendor blocked on a client's data

Strip out the borehole logs and the pattern generalises to any engagement where you sit downstream of a client's data supply chain and it dries up. Do not escalate by email, because email flattens three problems into one. Run a session that names the distinct problems separately, with the client's own words on the wall: access, trust, and time-to-data are a reliable starting decomposition. Put your own blind spot on a slide, because the admission is what earns the ask. Measure the gap against a maturity model the client can see themselves on, so the stall reads as a stage rather than a verdict. And close on one concrete resourcing decision, framed as the client's own capability investment. The workshop is not a soft alternative to escalation. It is the more demanding one, and the one that actually gets the wells moving.

Limitations

This is a single engagement, and the decomposition into access, trust, and time-to-data came from one client's framing of its own data estate; other organisations will surface different problem sets, and the three-lane map is a method for finding the problems, not a fixed taxonomy. The workshop reduced a specific stall on a specific program; it is not a controlled comparison against a matched engagement that escalated by email instead, so the claim is about mechanism and structure, not measured effect size. The 39-slide count, the 10-15 well minimum, the November 2022 phase start, the eight-dimension maturity model, and the four-day-by-three-type format are drawn from the workshop deck; the lane weighting in the instrument is an illustrative reading of the three-problem framing, not a measured quantity. Finally, a workshop only moves things when the client can act on the resourcing ask; where the constraint is external, such as partner sign-off or a producing asset's own priorities, naming the problem well is necessary but not sufficient.

References

[1] McKinsey & Company, "Scaling AI like a tech native: The CEO's role." McKinsey Digital. https://www.mckinsey.com/

Quamer Nasim
Quamer Nasim

ML Research Engineer

More from EarthScan

Related research

All insights →
The 288-Slide Running Log: Weekly Evidence Beats Monthly Polish in Applied-AI Delivery
Insight

The 288-Slide Running Log: Weekly Evidence Beats Monthly Polish in Applied-AI Delivery

Effort Obligations and a Liability Cap: Contracting AI R&D So Both Sides Survive Failure
Insight

Effort Obligations and a Liability Cap: Contracting AI R&D So Both Sides Survive Failure

What a EUR 86k Subsurface-AI Pilot Actually Buys: Line-Item Anatomy of a 2020 Engagement
Insight

What a EUR 86k Subsurface-AI Pilot Actually Buys: Line-Item Anatomy of a 2020 Engagement

Stay ahead

EarthScan insights, in your inbox.

Field-tested research on subsurface and energy-transition AI. About twice a month. No noise.

We use your email only for this newsletter. Unsubscribe anytime Privacy.