What is the system supposed to do?
Define success first.
Failure cannot be understood without a clear intended function and acceptable performance range.
Side 25
A study of how systems fail, how weak signals become incidents, and how design can limit damage. Reliability is not the absence of failure; it is the disciplined management of failure probability, detection, containment and recovery.
Incidents emerge when hazards, component weaknesses, operating conditions and failed defenses align.
Define success first.
Failure cannot be understood without a clear intended function and acceptable performance range.
Energy, pressure, data, authority, delay?
A hazard can exist without an incident if effective controls keep it contained.
Omission, commission, degradation, timing?
Failure modes describe specific ways a component or process can cease to perform as intended.
Local or cascading?
Coupling determines whether one failure remains contained or spreads into dependent systems.
Safety, cost, delay, trust, data?
Severity and probability are separate dimensions of risk.
No single analysis technique captures every pathway. Use the method that matches the system and the uncertainty.
Ask how each element can fail, what the effect would be, how severe it is and how easily it would be detected.
Combine lower-level causes through AND/OR logic to understand pathways into a defined failure.
Trace which barriers succeed or fail and how those branches produce different consequences.
Ask why the system allowed the local error to become consequential instead of stopping at “operator mistake.”
Preventive barriers sit before the central event; mitigative barriers sit after it.
Series and parallel structures reveal how component reliability combines into system reliability.
People make mistakes, but reliability analysis asks why the environment made that mistake likely, invisible or consequential.
Execution deviates through attention, interface or motor error.
Memory load, interruption and poor cueing can make intended steps disappear.
Knowledge, diagnosis or rule selection fails before execution begins.
Local incentives, unrealistic procedures or normalized workarounds can make deviation rational in context.
Automation and interfaces can hide which control mode is active.
If bad outcomes fail to appear immediately, risky workarounds can gradually become normal practice.
When the same class of error can recur, redesign the task, interface, constraint or detection mechanism rather than relying only on vigilance.
Barriers reduce the probability that one error becomes a catastrophic outcome.
The strongest control is often to redesign the system so the dangerous condition cannot arise.
Interlocks, permissions, physical constraints and validation can stop a failure before it propagates.
Alarms, checksums, sensors, reconciliation and peer review shorten time-to-detection.
Segmentation, isolation, circuit breakers and compartmentalization keep local failures local.
Backups, rollback, failover and emergency procedures reduce duration and consequence.
Multiple components can improve reliability when failures are independent enough and switching actually works.
Redundancy fails when supposedly separate backups share the same power source, software defect, supplier or operating assumption.
An incident that almost happened reveals a pathway before the full consequence arrives.
A value, behavior or event falls outside the expected pattern.
The anomaly is small enough to dismiss but may indicate a degrading condition.
A failure pathway activates but a barrier, chance event or late intervention prevents full consequence.
The system experiences loss, harm or disruption but remains within recoverable bounds.
Multiple defenses fail and consequences exceed ordinary recovery capacity.
A near miss with a dangerous pathway may deserve more attention than a low-impact incident caused by an isolated, easily contained error.
Complex systems cannot eliminate every failure mode, so adaptive capacity matters alongside prevention.
Observability makes degradation visible before operators are forced to infer it from consequences.
Inventory, time, compute, staffing or financial reserves absorb shocks that optimization might otherwise remove.
When control is lost, the system should move toward a state with lower expected harm.
Partial service can be preferable to total collapse when components fail.
Recovery procedures should be tested before they are needed under pressure.
Post-incident learning should change architecture, procedure or training rather than merely document what happened.
Operational impact may be small, yet the underlying failure path still deserves analysis. The successful recovery barrier reduced consequence; it did not erase the initiating defect.
Look for common-mode dependence: shared power, network, credentials, code, supplier or maintenance process.
Study why. If the formal procedure makes normal work impractical, the violation may be evidence of system design conflict rather than individual carelessness.