11% of annual revenue lost to unplanned downtime, on average, across global manufacturers — Siemens/Senseye "True Cost of Downtime" 60% average OEE in manufacturing vs. 85% "world-class" benchmark — Vorne / OEE.com 98% of organizations say a single hour of downtime costs over $100K — ITIC

Operational Incident Cost Glossary

The reliability and incident-response vocabulary, translated for the plant floor — from classic IT-ops terms like MTTR and RTO to manufacturing-specific ones like OEE, MTBF, and andon.

30+ terms Key distinctions table Plant-floor context throughout

Key Distinctions

The pairs of terms people mix up most often.

Term ATerm BThe Difference
MTTDMTTAMTTD is how long until the fault is detected (by a sensor, operator, or system). MTTA is how long until someone acknowledges it once an alert has already fired.
MTTR (Repair)MTTR (Recover)Same acronym, two meanings — "time to repair" the failed component vs. "time to recover" the full process back to normal output. Always clarify which one a report means.
Unplanned downtimePlanned downtimeUnplanned is the failure you didn't schedule — the subject of this whole suite. Planned downtime (changeovers, scheduled PM, holidays) is a capacity-planning problem, not an incident-cost problem.
OEETEEPOEE measures effectiveness against scheduled production time. TEEP measures it against all calendar time, including hours you never scheduled to run — useful for spotting unused capacity.
Preventive maintenancePredictive maintenancePreventive is maintenance on a fixed schedule (every 500 hours, every quarter). Predictive uses condition data (vibration, temperature, oil analysis) to maintain equipment only when it actually needs it.
RTORPORTO is the maximum acceptable time a line or system can be down. RPO is the maximum acceptable data loss, measured in time — more relevant for MES/ERP systems than for the machines themselves.

Full Glossary

Annualized Loss Expectancy (ALE) risk quantification
ALE = Single Loss Expectancy (SLE) × Annual Rate of Occurrence (ARO). Originally a cybersecurity risk formula, but directly reusable for operational risk: multiply a typical incident's cost by how often it happens per year to get your expected annual exposure — exactly what the Calculator's "annualized exposure" figure does.
Andon
A visual (and often audible) signal system — a light, a cord pull, a digital board — that lets any operator flag a problem and, in many lean systems, stop the line. Makes small problems visible before they become large ones.
Blast Radius
How many lines, cells, shifts, or downstream processes are affected by a given incident. A single workstation jam has a small blast radius; a plant-wide power outage has a large one.
CMMS (Computerized Maintenance Management System)
Software that tracks maintenance schedules, work orders, asset history, and downtime by equipment — the foundation for calculating MTBF and justifying a predictive-maintenance investment with real data instead of a hunch.
Changeover Time
Time spent reconfiguring a line between product runs — different from an incident, since it's planned, but it competes for the same production hours and is worth tracking separately from unplanned downtime.
Event
Any observable occurrence in your operation — a sensor reading, an alarm, a state change. Not every event is an incident; most are routine. An incident is an event (or combination of events) that actually disrupts production, quality, or safety.
HMI (Human-Machine Interface)
The screen or panel an operator uses to monitor and control equipment — often the first thing to fail or freeze during a controls/software incident, even when the underlying machine is fine.
Incident (Operational)
Any unplanned event that disrupts production, quality, or safety on the plant floor — equipment failure, a software/OT/controls failure, a power or utility outage, a quality defect, a supply chain disruption, or a safety stoppage. See the full breakdown on By Incident Type.
Incident Management
The ongoing program that defines severity levels, escalation paths, and response processes before anything goes wrong — distinct from incident response, which is what happens during a specific incident.
Incident Response
The reactive process of handling one specific incident, from detection through containment, repair, and recovery to full output.
MTBF (Mean Time Between Failures)
Average operating time between failures for a repairable piece of equipment. The core metric for deciding whether an asset needs a predictive-maintenance program or critical spares on hand.
MTTF (Mean Time To Failure)
Like MTBF, but for components that are replaced rather than repaired when they fail (e.g., a bearing or a belt) — average lifespan rather than average time between repairable failures.
MTTA (Mean Time to Acknowledge)
Time between an alert or alarm firing (on a SCADA system, andon board, or CMMS) and a human acknowledging it. A long MTTA often means alert fatigue — too many low-value alarms drowning out the ones that matter.
MTTD (Mean Time to Detect)
Time from when a fault actually begins to when it's detected — by a sensor, an operator, or a quality check. Shorter MTTD generally means a smaller blast radius and lower cost, since the problem is caught before it cascades.
MTTR (Mean Time to Repair / Recover)
Ambiguous on purpose in common usage — "time to repair" (fix the failed component) and "time to recover" (get the full process back to normal output, including restart/ramp-up) are different numbers. Always clarify which one a report is using; the Calculator separates them into "full-stoppage duration" and "ramp-up time."
OEE (Overall Equipment Effectiveness)
Availability × Performance × Quality, multiplied together into a single percentage. Industry average is roughly 60%; "world-class" is considered 85% (Vorne/OEE.com). The single most common manufacturing reliability metric.
OT/IT Convergence
The ongoing merging of operational technology (PLCs, SCADA, the physical plant floor) with information technology (networks, cloud, MES/ERP). Explains why software and cybersecurity incidents increasingly have real, physical production consequences instead of staying contained to an office network.
P1 / P2 / P3 / P4 (Severity Tiers)
P1: full line or plant stoppage, immediate response. P2: partial or single-cell stoppage, urgent response. P3/P4: degraded performance or non-production-affecting issues, handled during normal business hours or on a scheduled basis.
Planned vs. Unplanned Downtime
Planned downtime (changeovers, scheduled maintenance, holidays) is a capacity-planning line item. Unplanned downtime — the subject of this entire suite — is what an incident actually costs you beyond what was already budgeted for.
Playbook / Runbook
A documented, pre-built procedure for responding to a specific type of incident — who does what, in what order, with which tools. Plants with tested runbooks for their top failure modes consistently recover faster than those improvising each time.
PLC (Programmable Logic Controller)
The industrial computer that directly controls machinery and processes on the plant floor — the thing an HMI displays and a SCADA system aggregates data from. A PLC crash or firmware bug is a classic "software/OT failure."
Postmortem / Post-Incident Review
Structured analysis after an incident is resolved — what happened, why, what the response did well or poorly, and what changes would prevent a recurrence. The step most plants skip when they're busy, and the one that compounds the most value over time.
Predictive Maintenance
Maintenance triggered by actual equipment condition data — vibration, temperature, oil analysis, current draw — rather than a fixed schedule. Catches developing failures before they become unplanned downtime, at the cost of needing sensors and some analysis capability.
Preventive Maintenance
Maintenance performed on a fixed schedule (every 500 operating hours, every quarter) regardless of actual equipment condition. Simpler to run than predictive maintenance, but can mean servicing healthy equipment or missing a failure that develops faster than the schedule anticipated.
Root Cause
The fundamental reason an incident occurred, as distinct from its symptoms. A conveyor jam might be the symptom; a misaligned sensor or a worn guide rail might be the root cause. Fixing symptoms without finding root cause is why the same incident keeps recurring.
RPO (Recovery Point Objective)
The maximum acceptable amount of data loss, measured in time — e.g., "we can afford to lose up to 15 minutes of MES production data." More relevant to your IT/MES systems than to the physical machines themselves.
RTO (Recovery Time Objective)
The maximum acceptable time a line, cell, or system can be down before the business impact becomes unacceptable. Setting an explicit RTO for your most critical lines is what tells you whether a given backup/redundancy investment is actually justified.
SCADA (Supervisory Control and Data Acquisition)
The system that aggregates data from PLCs and other plant-floor devices into dashboards and historical records for monitoring and control at a higher level than any single machine's HMI.
SLA (Service Level Agreement)
A contractual commitment — either one you've made to a customer (on-time delivery, quality specs) or one a supplier or utility has made to you. SLA breaches are a common secondary cost driver, especially for supply chain and power/utility incidents.
Takt Time
The pace of production needed to meet customer demand (available production time ÷ customer demand). Useful context when estimating lost-production cost — it tells you how much output a given stoppage actually should have produced.
TEEP (Total Effective Equipment Performance)
Like OEE, but measured against all calendar time rather than just scheduled production time — surfaces unused capacity (unscheduled shifts, weekends) that OEE doesn't account for.

Frequently Asked Questions

Why does this glossary include IT-operations terms like MTTR and RTO?

Because the plant floor increasingly runs on them too. As OT and IT converge — PLCs talking to MES systems, cloud-based analytics on production data — the reliability vocabulary that used to live only in a data center now shows up in maintenance meetings. We've kept the IT-native terms that generalize well and reframed each one for a production context.

What did you leave out from a typical IT/security glossary, and why?

Pure cybersecurity terms — EDR, XDR, SIEM, SOAR, breach notification, dwell time — don't apply to most operational incidents and would clutter a glossary meant for plant-floor reliability. If your incident genuinely is a cybersecurity event that happens to hit production (increasingly common with OT/IT convergence), your IT or security team's own glossary is the better reference for that half of it.

Which of these terms actually matters most for a small plant?

If you only track three things, track OEE (how effectively you're using scheduled time), MTBF by critical asset (so you know what's about to fail), and a clear RTO for your most important line (so you know how fast "fast enough" actually is).

// Put these terms to work

Run your numbers on the Calculator.

Now that the vocabulary is covered, see what an actual incident — using these terms — costs your plant.

The Complete Suite

Overview
Operational Incident & Outage Cost

The hub page — headline benchmarks, cost components, and links into the full suite.

Open Overview →
Calculator
Incident Cost Calculator

Model a real incident against your own revenue, labor, and recovery costs.

Open the Calculator →
Data
Cost Index

Benchmark tables by incident type, duration, and company size for SMB manufacturing.

View the Index →
Reference
By Incident Type

Deep dive on each incident category — cost drivers, typical duration, and mitigations.

Browse Types →