Failure Testing

Lesson 2 · System Design (medior → senior) · mission · glossary

Back in the calibration, this was your one real miss. Part 5 asked "how do you prove it holds up before going live?" and you answered with load testing against projected users — "set the pod config, be confident it'll be okay." That answer proved the happy path scales. But your original design didn't die from traffic — it died from failure: a process crash silently dropped in-flight retries, provider throttling broke latency, a single AZ took everything down.

Rule 1 — load tests prove capacity; failure tests prove guarantees. They are different suites for different questions. You need both.

The method: guarantees first, failures second

Never start from "what can break" — you'll think of a million things and test the wrong ones. Start from your guarantees, then ask the only question that matters: what failure would break this guarantee? Every answer is a test you must run.

Your guaranteeFailure that breaks itTest asserts
Never lostWorker killed mid-process; Kafka leader dies; provider outageEvery event still delivered exactly once, even across crashes
<5s p95Provider throttling; slow consumer; poison messageLatency holds under the tail; queue drains
Per-user orderingWorker crash mid-partition; rebalanceOrder survives handoff to a new consumer
No SPOFAZ loss; LB death; DB failureStill serving, no data loss

Notice the first row reuses everything from Lesson 1: the state-per-partition store, the changelog replay on takeover, and your trx_id idempotency key are precisely the machinery that makes "kill a worker mid-process" survivable. Failure testing is testing Lesson 1.

Fault injection & chaos engineering

Fault injection is the act of deliberately making things fail in a controlled way: kill a process, blackhole network traffic, inject latency, break a dependency, fail a disk. A chaos experiment adds discipline on top:

  1. State a hypothesis — "if a worker dies, no alert is lost because the queue redelivers and dedup skips it."
  2. Define the blast radius — start with one worker, one broker, one AZ. Not all of them.
  3. Define an abort condition — the moment the system breaks its SLO or loses data, the experiment stops and you learn.

Chaos is a test discipline, not "break production for fun." Netflix's Chaos Monkey is the origin — randomly kill instances so engineers are forced to build systems that survive. Tools you'll meet: Toxiproxy (inject network faults into a service you control), k6 fault injection, Gremlin, and your own shell scripts that kill a pod and watch what happens.

The code that makes failures graceful

Fault injection only works if the code has the defense in depth that turns a fault into a blip instead of an outage:

ControlWhy
Timeouts everywhereA dead dependency hangs forever without one; timeouts bound the damage.
Retries + backoff + jitterRetry with exponential backoff; add random jitter so retries don't hit all at once (the retry storm).
Circuit breakerAfter N failures, stop calling the dependency for a while — fail fast instead of hammering a dead provider.
Queue as the bufferA provider outage should fill the queue, not fail the request. That's what the queue is for.

The "never lost" test that never stops

The single most powerful proof of "no loss" is one you already know how to build: reconciliation as a test. A continuous job compares the source of truth (the transaction events) against the outcome (notification records in a delivered state). Any event without a matching delivered record is a bug — found automatically, not in an incident. You ran this at Jenius for money; run it for your guarantees. This is the test that runs forever, in production, quietly.

Exercise — write the test plan for your notification system

  1. Pick your four guarantees. For each, name the single failure you'd inject first, and what your assertion would be.
  2. Describe the provider-outage drill: MoEngage/APNs is down for 30 minutes on payday at 9 AM. What do you expect, what do you monitor, and when do you declare it recovered?
  3. What's your abort condition if the chaos experiment starts losing alerts?

Check yourself

Q1. The original design died from:

Q2. A chaos test asserts that:

Q3. The best continuous "no loss" check:

Q4. When the push provider is down:

Go deeper

Primary sources: Chaos Engineering — Rosenthal & Jones (O'Reilly), and principlesofchaos.org. Netflix's chaos engineering blog posts are where it started.