Independent computer or process.
Own memory + own clock.
Nodes do not share instantaneous state.
Side 58
A study of computation spread across unreliable machines. Distributed systems replace shared memory with messages, introducing delay, partial failure, uncertain ordering and the need for explicit coordination.
One process can no longer directly know the state of another; it only receives messages that may be delayed, duplicated or lost.
Own memory + own clock.
Nodes do not share instantaneous state.
Delay is variable.
A missing reply does not reveal whether the peer failed or the network is slow.
Who knows the latest value?
Replication creates availability and coordination problems simultaneously.
Partial failure.
Distributed systems must operate without assuming all components share fate.
Replay, reconcile, elect?
Recovery logic is part of the protocol, not an afterthought.
Distributed systems often need to reason about order without relying on synchronized wall time.
Synchronization reduces but does not eliminate uncertainty.
If one event can influence another, the system can order them causally.
Lamport clocks track a partial temporal relation among events.
Vector clocks can distinguish concurrent updates from causally ordered ones.
Timeouts turn uncertainty into a decision point without proving failure.
Leases depend on carefully managed clock assumptions.
Copies increase resilience and read capacity, but writes must be propagated and reconciled.
Followers replicate the leader’s log or state.
Useful across regions but creates conflict-resolution work.
Quorums and repair mechanisms reconcile divergent copies.
Read behavior depends on whether stale results are acceptable.
Systems need deterministic merge or application-level resolution rules.
Protocols such as Raft and Paxos coordinate nodes on an ordered sequence of decisions.
Nodes must avoid accepting conflicting leaders for the same logical term.
Overlap prevents two independent decisions from both appearing committed.
State machines can replay the same committed commands deterministically.
Consensus separates proposed work from work known to be committed.
Safety may be preserved even when progress temporarily stops.
Progress depends on enough functioning nodes and communication.
Different guarantees make different promises about ordering and freshness.
| Model | Guarantee | Cost |
|---|---|---|
| Linearizable | Operations appear to occur atomically in real-time order | Coordination and latency |
| Sequential | All clients see one order consistent with program order | Still coordinated, weaker than real-time |
| Read-your-writes | A client sees its own completed updates | Session routing/state |
| Eventual | Replicas converge if updates stop | Temporary divergence allowed |
| Causal | Causally related operations preserve order | Metadata and dependency tracking |
Distributed systems need idempotence, retries, redundancy and clear ownership of uncertain work.
Repeat transiently failed operations, but make repeated execution safe where possible.
Ensure duplicate requests do not create duplicate effects.
Slow retries to avoid amplifying an overloaded dependency.
Stop sending work to a dependency that is repeatedly failing.
Preserve work that cannot be processed automatically for later inspection.
Correlate logs, metrics and traces across nodes to reconstruct failure paths.