Skip to content

Side 25

Failure, Error
& Reliability

A study of how systems fail, how weak signals become incidents, and how design can limit damage. Reliability is not the absence of failure; it is the disciplined management of failure probability, detection, containment and recovery.

hazard→failure mode→barrier→detection→recovery
06failure lenses
05barrier types
06resilience habits
25Side

Failure is usually a path, not a point.

Incidents emerge when hazards, component weaknesses, operating conditions and failed defenses align.

01 · Function

What is the system supposed to do?

Define success first.

Failure cannot be understood without a clear intended function and acceptable performance range.

02 · Hazard

What could cause harm?

Energy, pressure, data, authority, delay?

A hazard can exist without an incident if effective controls keep it contained.

03 · Failure mode

How can function break?

Omission, commission, degradation, timing?

Failure modes describe specific ways a component or process can cease to perform as intended.

04 · Propagation

What happens next?

Local or cascading?

Coupling determines whether one failure remains contained or spreads into dependent systems.

05 · Consequence

What is the damage?

Safety, cost, delay, trust, data?

Severity and probability are separate dimensions of risk.

Reliability habitfailure probability × exposure × consequence ≠ one universal risk score, but forces the dimensions apart

Different methods ask different failure questions.

No single analysis technique captures every pathway. Use the method that matches the system and the uncertainty.

FMEA

Start from component failure modes.

Ask how each element can fail, what the effect would be, how severe it is and how easily it would be detected.

Fault tree

Work backward from a top event.

Combine lower-level causes through AND/OR logic to understand pathways into a defined failure.

Event tree

Work forward from an initiating event.

Trace which barriers succeed or fail and how those branches produce different consequences.

Root cause

Go beyond the immediate trigger.

Ask why the system allowed the local error to become consequential instead of stopping at “operator mistake.”

Bow-tie

Connect causes, event and consequences.

Preventive barriers sit before the central event; mitigative barriers sit after it.

Reliability block

Model component dependence.

Series and parallel structures reveal how component reliability combines into system reliability.

Human error is often a system output.

People make mistakes, but reliability analysis asks why the environment made that mistake likely, invisible or consequential.

Slip

Right intention, wrong action.

Execution deviates through attention, interface or motor error.

Lapse

Something is forgotten.

Memory load, interruption and poor cueing can make intended steps disappear.

Mistake

The plan itself is wrong.

Knowledge, diagnosis or rule selection fails before execution begins.

Violation

A rule is knowingly bypassed.

Local incentives, unrealistic procedures or normalized workarounds can make deviation rational in context.

Mode confusion

The system is in a different state than believed.

Automation and interfaces can hide which control mode is active.

Normalization

Repeated deviation stops feeling exceptional.

If bad outcomes fail to appear immediately, risky workarounds can gradually become normal practice.

“Be more careful” is weak engineering.

When the same class of error can recur, redesign the task, interface, constraint or detection mechanism rather than relying only on vigilance.

Reliable systems assume some things will fail.

Barriers reduce the probability that one error becomes a catastrophic outcome.

Eliminate

Remove the hazard.

The strongest control is often to redesign the system so the dangerous condition cannot arise.

Prevent

Block the initiating event.

Interlocks, permissions, physical constraints and validation can stop a failure before it propagates.

Detect

Notice the failure quickly.

Alarms, checksums, sensors, reconciliation and peer review shorten time-to-detection.

Contain

Limit blast radius.

Segmentation, isolation, circuit breakers and compartmentalization keep local failures local.

Recover

Restore safe function.

Backups, rollback, failover and emergency procedures reduce duration and consequence.

Redundancy

Multiple components can improve reliability when failures are independent enough and switching actually works.

Common-mode failure

Redundancy fails when supposedly separate backups share the same power source, software defect, supplier or operating assumption.

Near misses are free information — if the system records them.

An incident that almost happened reveals a pathway before the full consequence arrives.

Anomaly

A value, behavior or event falls outside the expected pattern.

Weak signal

The anomaly is small enough to dismiss but may indicate a degrading condition.

Near miss

A failure pathway activates but a barrier, chance event or late intervention prevents full consequence.

Incident

The system experiences loss, harm or disruption but remains within recoverable bounds.

Catastrophe

Multiple defenses fail and consequences exceed ordinary recovery capacity.

Outcome severity can hide process severity.

A near miss with a dangerous pathway may deserve more attention than a low-impact incident caused by an isolated, easily contained error.

Reliability tries to prevent failure. Resilience prepares to survive it.

Complex systems cannot eliminate every failure mode, so adaptive capacity matters alongside prevention.

Monitor

Know current system state.

Observability makes degradation visible before operators are forced to infer it from consequences.

Buffer

Carry spare capacity.

Inventory, time, compute, staffing or financial reserves absorb shocks that optimization might otherwise remove.

Fail safe

Choose the safer default.

When control is lost, the system should move toward a state with lower expected harm.

Degrade gracefully

Lose function progressively.

Partial service can be preferable to total collapse when components fail.

Recover

Restore from known good state.

Recovery procedures should be tested before they are needed under pressure.

Learn

Feed failure back into design.

Post-incident learning should change architecture, procedure or training rather than merely document what happened.

A deployment fails but rollback works immediately.

Operational impact may be small, yet the underlying failure path still deserves analysis. The successful recovery barrier reduced consequence; it did not erase the initiating defect.

Two backups fail during the same outage.

Look for common-mode dependence: shared power, network, credentials, code, supplier or maintenance process.

An operator repeatedly bypasses a safety step.

Study why. If the formal procedure makes normal work impractical, the violation may be evidence of system design conflict rather than individual carelessness.

Normal AccidentsCharles Perrow · complexity and coupling
Human ErrorJames Reason · error and system defenses
Engineering a Safer WorldNancy Leveson · systems safety
Resilience EngineeringHollnagel, Woods & Leveson · adaptive capacity