Most plants do not have an AI problem. They have an authority problem: nobody has written down what an agent is allowed to do to a machine, and under what evidence.
Engineering organizations pushed AI coding adoption up 65% and got roughly 8% faster software delivery. Teams with heavy AI use closed 21% more tasks while code review time climbed 91%.
Those two findings come from DX and Faros, and they describe the same event from opposite ends of the pipeline.1,2 Generation got cheap. Verification did not. Writing code was only about 16% of a developer's day to begin with, so making that part faster moves the delivery number by single digits while loading the review queue.3
Then there is the controlled result. METR ran a randomized trial with 16 experienced open-source developers across 246 real tasks in repositories they had worked in for about five years. Going in, the developers forecast a 24% time saving. Coming out, they reported a 20% saving. The stopwatch said tasks took 19% longer.4
Set aside the headline. The durable finding is the 39-point spread between what practitioners believed about their own speed and what was measured. Skilled people, on their own code, were wrong about the direction of the effect.
That is the reason this matters more in operations than in software. An agent on a plant floor does not emit diffs. It emits actions: a work order, a dispatch, a speed change, a valve command, a hold on a lot. Same stochastic engine, same failure signature of output that looks right and is not, and a blast radius measured in scrap, downtime, and occasionally people.
Practitioners already sense this. In one survey, 96% of developers said they do not fully trust AI-generated code and 61% said it often looks correct but proves unreliable.5 Useful instinct, useless control. "Do not fully trust" cannot be implemented. Thresholds, envelopes, expiry timers, and audit trails can.
Gartner expects 60% of organizations to run smaller software engineering teams by 2029, and warns that cutting junior roles in response to AI hollows out the talent pipeline.6 Hold that thought until section 02, failure F6. The people who handle the rare, ugly cases are made by handling the ordinary ones.
Steve Bennett of Wireless Logic framed the endpoint of this argument in EE Times: the winners will be the organizations that can say exactly where they draw the line and hold it.7 This paper is about how to draw that line as an engineering artifact rather than a policy statement, and how to move it deliberately as evidence accumulates.
None of these are model quality problems. Every one of them is an architecture decision that nobody made on purpose.
Every one of these is addressable. None of them are addressed by a better model, a longer prompt, or a vendor demo. They are addressed by where you put the boundaries and what evidence moves them.
The workable pattern separates the stochastic part of the system from everything that touches a machine, and puts a deterministic checkpoint between them.
Filtering, normalization, and provenance. This is unglamorous and it is where the return hides. Most mid-market plants are running first-generation IoT on top of equipment older than the network it sits on, which means duplicate tags, drifted calibrations, and units that changed during a retrofit nobody documented.
One rule holds the plane together: nothing reaches the reasoning layer without a source, a unit, and a timestamp you would be willing to produce in a warranty dispute. Sparkplug B and OPC UA give you the vocabulary. Enforcement is yours to build.
This is the only stochastic component in the design. NVIDIA's research group argued the case for small models here better than most: agentic systems mostly invoke a language model to do a narrow, repetitive job, and for those jobs small models are sufficiently capable, operationally better suited, and far cheaper.12 A 3B-class model on a gateway, fine-tuned on your incident history and your manuals, will beat a frontier model reached over a flaky cellular link on every axis that matters in a plant.
Structure it hierarchically. Run the small local model first and escalate only what it cannot handle confidently. Measured across five devices and three datasets, hierarchical inference held a target accuracy while cutting latency up to 73% and device energy up to 77% against on-device-only inference; adding early exit cut both further.13 The commercial argument arrives with the engineering one: escalation costs money and bandwidth, so the routing policy is also the cost model.
The routing rule that sends a case to a bigger model is the same rule that sends a case to a human. Build one policy, express it once, and apply it in both directions. Two separate escalation mechanisms drift apart within a quarter.
Control engineering solved a version of this problem before the current AI wave started. The Simplex architecture, formalized by Lui Sha, pairs a high-performance controller of uncertain reliability with a verified baseline controller and a monitor that hands authority to the baseline the moment safety conditions come under threat.14 That pattern has since been extended to machine-learned controllers, where design-time formal verification is not available and runtime assurance takes its place.15
Apply it directly. The language model is the high-performance controller. Your PID loops, interlocks, and standard operating procedures are the verified baseline. The assurance plane is the monitor, and it does five jobs: check the proposed action against the physical and process envelope, rate-limit it, classify its reversibility, execute it in a sandbox with scoped credentials, and write an immutable record of what was proposed, what was allowed, and what happened next.
Actuation is whatever writes to the physical world, including the ones people forget to count: a CMMS work order that dispatches a technician at 2am, or a hold that idles a line. Governance runs alongside all four planes rather than sitting at the top of a diagram, because identity, evidence, and authority have to be present at every hop or they are present at none.
"Is this agent autonomous?" is the wrong question. The answerable version: for this specific action, on this specific asset, at this measured confidence, what is the agent allowed to do without a human?
That question has two independent inputs. How reliable is the judgment, and how bad is it if the judgment is wrong. Plot them against each other and the policy writes itself.
| Tier | The agent | The human | Evidence to promote out |
|---|---|---|---|
| T0 observe | Annotates and correlates. Emits no recommendation. | Works exactly as before, with better context. | 200+ annotated cases graded for usefulness by the crew. |
| T1 recommend | Proposes a ranked action with its evidence and its confidence. | Decides and executes. Records agreement or override with a reason. | Agreement rate and override reasons stable across 60 days and two crews. |
| T2 stage and confirm | Prepares the action fully, then waits. Expires to a no-op on timeout. | Approves inside a time box. Silence means no action. | Zero envelope violations; approval latency inside the process window. |
| T3 act and report | Executes inside a pre-approved envelope, rate-limited, then reports. | Reviews after the fact, on a schedule, with sampling. | Measured error at the operating point below the agreed risk budget. |
| T4 act within envelope | Runs the loop. Reports exceptions and drift only. | Owns the envelope, the risk budget, and the quarterly recertification. | Nothing. T4 is the ceiling, and it is reserved for reversible actions. |
Demotion is automatic; promotion is manual. Drift detection, an envelope violation, or a sensor calibration event drops the affected action to T1 without a meeting. Moving back up takes a human signature and fresh evidence.
Silence is not consent. A staged action that nobody approves must expire into a no-op. Systems that default to execute after a timeout have quietly promoted themselves to T3 while the org chart still says T2.
Vendors sell tier 4 across the board because tier 4 demos well. In brownfield environments it is rarely the right answer, and the honest version of the pitch is narrower: tier 4 for reversible, high-frequency, well-instrumented actions, tier 2 for anything that costs real money to undo, and a documented path between them.
A confidence score from a language model is a number the model produced. Calling it a measurement is a category error that has quietly authorized a lot of bad decisions.
Recent work on selective prediction makes the point precisely: token entropy on its own is not enough to decide when a model should answer and when it should stand down.16 The number correlates with correctness under favorable conditions and stops correlating exactly when conditions get unfavorable, which is when you needed it.
The fix is old and well understood outside the AI industry. Give the system an explicit third option beyond right and wrong: abstain. Chow formalized the reject option in the 1970s, and modern selective prediction turned it into a risk-coverage problem where the objective is low error on the cases you accept, with everything else deferred.17
Conformal prediction supplies the missing guarantee. Rather than trusting the model's self-reported confidence, you calibrate on held-out data and get prediction sets with finite-sample coverage that holds without assuming anything about the model's internals.18 In clinical triage work, pairing conformal sets with cost-aware deferral cut error on retained cases by roughly 47% to 50% at fixed coverage, including on out-of-distribution splits.19 The same shape of problem appears on a plant floor: high volume, rare positives, asymmetric costs, and a human expert who is expensive but available.
Equipment failures are rare events, and conformal coverage computed across the whole population will happily meet its target while systematically under-covering the class you actually care about. Class-conditional (Mondrian) conformal prediction fixes that. In a benchmark across 15 imbalanced real-world datasets, class-conditional calibration restored minority-class coverage by an average of 61.7 percentage points over the marginal version.20 Calibrate per failure mode, not per fleet.
Conformal guarantees rest on exchangeability between calibration and deployment data. A plant violates that assumption every time it changes product mix, replaces a bearing supplier, or crosses into summer. Treat empirical coverage as a live production metric, recalibrate on a rolling window, and set drift detection to trip an automatic tier demotion.
Accuracy is not on this list. Accuracy without a coverage figure describes nothing, because a system that answers 4% of cases can post any accuracy it likes.
| Metric | Why it earns its place on the wall |
|---|---|
| Coverage at target risk | The share of decisions the agent takes while staying inside the risk budget. This is the actual productivity number. |
| Selective risk | Error measured only on the cases the agent accepted. The number your operations lead cares about. |
| Envelope violation rate | Proposals the assurance plane blocked. Should be non-zero in testing and zero in production. |
| Override rate and reasons | A rising override rate is drift. A falling one with no accuracy change is complacency. |
| Calibration error over time | Panel B as a monthly trend line rather than a one-time acceptance test. |
| Mean time to human takeover | How long the plant runs on a bad decision before a person intervenes. Tests the alerting path, not the model. |
| Cost-weighted error | A false dispatch and a missed bearing failure are not the same event. Weight them and the threshold moves on its own. |
Two different things get called a harness in this field, and an industrial deployment needs both: the runtime scaffolding that surrounds the model in production, and the test rig that decides whether it is allowed to run at all.
Everything around the model that governs what it sees, what it can call, and what happens when it is wrong. The empirical case for taking it seriously keeps strengthening. One recent framework separated task execution from state management and independent progress verification, giving each execution round a fresh context, and lifted pass rate on a long-horizon benchmark from 51.8% to 80.7% using the same underlying models.21 Researchers have started arguing that comparing agents without disclosing their harness is close to meaningless, because that is where reliability lives.22
Translated to a plant, the harness carries four responsibilities.
Generic benchmarks tell you nothing about your gearboxes. What you need is a plant-specific regression suite that runs on every change and produces the evidence that moves a cell in Figure 2.
Version bumps, prompt edits, tool schema changes, and gateway firmware updates all invalidate your calibration. Run the suite. Manufacturing already knows how to do this; it is the same discipline that governs a change to a control program, applied to an artifact that happens to be probabilistic.
Governance in agentic ops comes down to one structural question: which identity proposes, which approves, which executes, and which audits. Collapse any two of those into the same identity and the controls become decoration.
| Function | Held by | Failure when it collapses |
|---|---|---|
| Propose | The agent, with its evidence and confidence attached. | Proposals without evidence cannot be reviewed, only believed. |
| Approve | A named human at T1 and T2; a pre-signed envelope at T3 and T4. | An agent that approves its own actions has one control, not two. |
| Execute | A scoped service identity with parameterized, typed tool access. | Shared credentials make attribution impossible after an incident. |
| Audit | A separate reader with no write path, on a schedule with sampling. | Self-reported logs are testimony from an interested party. |
Three controls do most of the work in an OT environment. Agent identities stay distinct from human identities, so the log answers who without ambiguity. Network egress is denied by default, because an agent that cannot reach an unexpected endpoint cannot be talked into using one. Tools are parameterized rather than general, so the agent calls set_speed(line=3, rpm=<bounded>) instead of holding a shell on the historian.
Very little of this needs inventing. Manufacturing has spent forty years building the frameworks; the work is mapping agent behavior onto them.
| Framework | What it governs in an agentic deployment |
|---|---|
| NIST AI RMF + AI 600-1 | The lifecycle wrapper: govern, map, measure, manage. The generative AI profile adds the risk categories specific to language models.24 |
| ISA/IEC 62443 | Zones and conduits. The reasoning plane belongs in its own zone with an inspected conduit to the control zone, and the gateway is a new asset in your inventory. |
| IEC 61508 / ISO 13849 | Safety functions. An agent stays outside the safety function entirely. Anything with a safety integrity level keeps its existing certified path and the agent gets no authority over it. |
| OWASP Top 10 for Agentic Applications | Agent-specific threat modeling: goal hijack, tool misuse, identity abuse, memory poisoning, rogue agents. Use it for red-team scoping before go-live.10 |
| EU Cyber Resilience Act | Product security obligations for anything connected that you place on the EU market, including the vulnerability reporting clock now running.23 |
Every agentic deployment has one on the slide. Ask three questions of yours. Who is authorized to pull it, by name and at 2am on a Sunday. How long does the plant take to reach a known-good state after it is pulled, measured with a stopwatch rather than estimated. When was it last tested end to end, in production, with the crew that would actually use it. A switch tested once at commissioning is a documented intention.
Mid-market plants have an advantage over enterprise here that rarely gets named: one meeting can hold everyone whose approval matters. Use it. The sequence below assumes a single named workflow and refuses to widen until the gates pass.
The step most often skipped, and the one that decides whether savings reach the P&L. Decide in writing, before deployment, what the recovered hours are for: a maintenance backlog, a second shift avoided, a certification that has been waiting eighteen months. Attach a name. Capacity that is not pre-committed gets quietly reabsorbed into the same queue it came from, and six months later the CFO asks a fair question that nobody can answer.
Bainbridge's warning has an operational answer. Rotate a share of agent-handled cases back to humans on purpose, even when the agent would have been right. Run quarterly manual-recovery drills on the workflows the agent now owns. Budget for it, schedule it, and treat the cost as insurance against the day the agent hits a case nobody has seen since 2019.
These are not gotchas. A good vendor answers all eight in twenty minutes and will be relieved someone finally asked.
Every architectural choice in this paper reduces to one artifact: a written statement of what an agent may do to a machine, under what evidence, with whose signature, and what automatically takes that permission away.
Most plants running agentic pilots today do not have that artifact. They have a vendor default, a Slack thread, and an assumption held slightly differently by the maintenance lead and the IT director. That works until the first case where the agent was confident and wrong, which is also the first case where somebody asks who approved it.
So a question worth taking into your next operations review. Pick the single action your agents would take most often. Can you say, in writing, what confidence and what consequence class would let that action run without a person in the loop, and who signs when that boundary moves? If the sentence does not exist yet, that is the artifact worth building first, ahead of any model selection.
Stragentech is a fractional CTO practice building Agentic Operational Intelligence for industrial manufacturers, managed service providers, and the private equity firms that own them. The work draws on 25 years of production systems: AWS IoT Analytics and Project Kuiper enterprise engineering at Amazon, legacy modernization and agentic workflow automation at Boeing / Jeppesen ForeFlight, and the CTO seat at a PE-backed industrial computing company.
Bilal A. Khan is based in Seattle and works remotely nationwide. Engagements start with an Operational Intelligence Assessment and a prioritized roadmap.
stragentech.com · info@stragentech.com · linkedin.com/in/bilalakhan
© 2026 Stragentech LLC · Agentic Technology Partners · Seattle, WA. Shared for the use of operations and engineering leaders evaluating agentic deployments. Figures may be reproduced with attribution.