Failure Modes

Short analytical pieces on the mechanics of failure in distributed systems: what the received wisdom gets wrong, and the conditions under which it stops being true.

Failure Modes are written from published postmortems, incident write-ups and engineering documentation, which are listed at the foot of every piece. Nothing in this section has been run, measured or operated here — where a number, a threshold or a result appears, it belongs to the source it is credited to.

The Second Region Illusion

7 min read

The standard architectural argument for multi-region deployment rests on a simple premise: if one location fails, the other remains functional. This logic treats geographic distance as a sufficient condition for fault isolation. However, redundancy only…

The Timeout That Saves No One

5 min read

The standard advice for building resilient clients is straightforward: always set a timeout on every remote call. The gRPC documentation explicitly states that without an explicit deadline, a client may wait for a response effectively forever. This guidance…

The Cache Is Not a Load Shedder

6 min read

The standard justification for placing a cache in front of a database is straightforward: it absorbs repeated reads, reducing the number of queries that reach the primary storage layer. On a normal day, this works as advertised. The application layer checks…

The scheduler is not a spare node

5 min read

The common assumption is that when a node fails, the Kubernetes control plane automatically moves the affected workload to a healthy node, rendering the incident self-healing. This view conflates the creation of a replacement object with the successful…

The Health Check That Kills the Fleet

6 min read

A liveness probe is a binary question: is this process alive and responsive? It is not a question about correctness, nor about the state of the database, nor about the availability of the message queue. The Kubernetes documentation defines the liveness probe…