- Annualized Loss Expectancy (ALE) risk quantification
- ALE = Single Loss Expectancy (SLE) × Annual Rate of Occurrence (ARO). Originally a cybersecurity risk formula, but directly reusable for operational risk: multiply a typical incident's cost by how often it happens per year to get your expected annual exposure — exactly what the Calculator's "annualized exposure" figure does.
- Andon
- A visual (and often audible) signal system — a light, a cord pull, a digital board — that lets any operator flag a problem and, in many lean systems, stop the line. Makes small problems visible before they become large ones.
- Blast Radius
- How many lines, cells, shifts, or downstream processes are affected by a given incident. A single workstation jam has a small blast radius; a plant-wide power outage has a large one.
- CMMS (Computerized Maintenance Management System)
- Software that tracks maintenance schedules, work orders, asset history, and downtime by equipment — the foundation for calculating MTBF and justifying a predictive-maintenance investment with real data instead of a hunch.
- Changeover Time
- Time spent reconfiguring a line between product runs — different from an incident, since it's planned, but it competes for the same production hours and is worth tracking separately from unplanned downtime.
- Event
- Any observable occurrence in your operation — a sensor reading, an alarm, a state change. Not every event is an incident; most are routine. An incident is an event (or combination of events) that actually disrupts production, quality, or safety.
- HMI (Human-Machine Interface)
- The screen or panel an operator uses to monitor and control equipment — often the first thing to fail or freeze during a controls/software incident, even when the underlying machine is fine.
- Incident (Operational)
- Any unplanned event that disrupts production, quality, or safety on the plant floor — equipment failure, a software/OT/controls failure, a power or utility outage, a quality defect, a supply chain disruption, or a safety stoppage. See the full breakdown on By Incident Type.
- Incident Management
- The ongoing program that defines severity levels, escalation paths, and response processes before anything goes wrong — distinct from incident response, which is what happens during a specific incident.
- Incident Response
- The reactive process of handling one specific incident, from detection through containment, repair, and recovery to full output.
- MTBF (Mean Time Between Failures)
- Average operating time between failures for a repairable piece of equipment. The core metric for deciding whether an asset needs a predictive-maintenance program or critical spares on hand.
- MTTF (Mean Time To Failure)
- Like MTBF, but for components that are replaced rather than repaired when they fail (e.g., a bearing or a belt) — average lifespan rather than average time between repairable failures.
- MTTA (Mean Time to Acknowledge)
- Time between an alert or alarm firing (on a SCADA system, andon board, or CMMS) and a human acknowledging it. A long MTTA often means alert fatigue — too many low-value alarms drowning out the ones that matter.
- MTTD (Mean Time to Detect)
- Time from when a fault actually begins to when it's detected — by a sensor, an operator, or a quality check. Shorter MTTD generally means a smaller blast radius and lower cost, since the problem is caught before it cascades.
- MTTR (Mean Time to Repair / Recover)
- Ambiguous on purpose in common usage — "time to repair" (fix the failed component) and "time to recover" (get the full process back to normal output, including restart/ramp-up) are different numbers. Always clarify which one a report is using; the Calculator separates them into "full-stoppage duration" and "ramp-up time."
- OEE (Overall Equipment Effectiveness)
- Availability × Performance × Quality, multiplied together into a single percentage. Industry average is roughly 60%; "world-class" is considered 85% (Vorne/OEE.com). The single most common manufacturing reliability metric.
- OT/IT Convergence
- The ongoing merging of operational technology (PLCs, SCADA, the physical plant floor) with information technology (networks, cloud, MES/ERP). Explains why software and cybersecurity incidents increasingly have real, physical production consequences instead of staying contained to an office network.
- P1 / P2 / P3 / P4 (Severity Tiers)
- P1: full line or plant stoppage, immediate response. P2: partial or single-cell stoppage, urgent response. P3/P4: degraded performance or non-production-affecting issues, handled during normal business hours or on a scheduled basis.
- Planned vs. Unplanned Downtime
- Planned downtime (changeovers, scheduled maintenance, holidays) is a capacity-planning line item. Unplanned downtime — the subject of this entire suite — is what an incident actually costs you beyond what was already budgeted for.
- Playbook / Runbook
- A documented, pre-built procedure for responding to a specific type of incident — who does what, in what order, with which tools. Plants with tested runbooks for their top failure modes consistently recover faster than those improvising each time.
- PLC (Programmable Logic Controller)
- The industrial computer that directly controls machinery and processes on the plant floor — the thing an HMI displays and a SCADA system aggregates data from. A PLC crash or firmware bug is a classic "software/OT failure."
- Postmortem / Post-Incident Review
- Structured analysis after an incident is resolved — what happened, why, what the response did well or poorly, and what changes would prevent a recurrence. The step most plants skip when they're busy, and the one that compounds the most value over time.
- Predictive Maintenance
- Maintenance triggered by actual equipment condition data — vibration, temperature, oil analysis, current draw — rather than a fixed schedule. Catches developing failures before they become unplanned downtime, at the cost of needing sensors and some analysis capability.
- Preventive Maintenance
- Maintenance performed on a fixed schedule (every 500 operating hours, every quarter) regardless of actual equipment condition. Simpler to run than predictive maintenance, but can mean servicing healthy equipment or missing a failure that develops faster than the schedule anticipated.
- Root Cause
- The fundamental reason an incident occurred, as distinct from its symptoms. A conveyor jam might be the symptom; a misaligned sensor or a worn guide rail might be the root cause. Fixing symptoms without finding root cause is why the same incident keeps recurring.
- RPO (Recovery Point Objective)
- The maximum acceptable amount of data loss, measured in time — e.g., "we can afford to lose up to 15 minutes of MES production data." More relevant to your IT/MES systems than to the physical machines themselves.
- RTO (Recovery Time Objective)
- The maximum acceptable time a line, cell, or system can be down before the business impact becomes unacceptable. Setting an explicit RTO for your most critical lines is what tells you whether a given backup/redundancy investment is actually justified.
- SCADA (Supervisory Control and Data Acquisition)
- The system that aggregates data from PLCs and other plant-floor devices into dashboards and historical records for monitoring and control at a higher level than any single machine's HMI.
- SLA (Service Level Agreement)
- A contractual commitment — either one you've made to a customer (on-time delivery, quality specs) or one a supplier or utility has made to you. SLA breaches are a common secondary cost driver, especially for supply chain and power/utility incidents.
- Takt Time
- The pace of production needed to meet customer demand (available production time ÷ customer demand). Useful context when estimating lost-production cost — it tells you how much output a given stoppage actually should have produced.
- TEEP (Total Effective Equipment Performance)
- Like OEE, but measured against all calendar time rather than just scheduled production time — surfaces unused capacity (unscheduled shifts, weekends) that OEE doesn't account for.