Back in the calibration, this was your one real miss. Part 5 asked "how do you prove it holds up before going live?" and you answered with load testing against projected users — "set the pod config, be confident it'll be okay." That answer proved the happy path scales. But your original design didn't die from traffic — it died from failure: a process crash silently dropped in-flight retries, provider throttling broke latency, a single AZ took everything down.
Rule 1 — load tests prove capacity; failure tests prove guarantees. They are different suites for different questions. You need both.
Never start from "what can break" — you'll think of a million things and test the wrong ones. Start from your guarantees, then ask the only question that matters: what failure would break this guarantee? Every answer is a test you must run.
| Your guarantee | Failure that breaks it | Test asserts |
|---|---|---|
| Never lost | Worker killed mid-process; Kafka leader dies; provider outage | Every event still delivered exactly once, even across crashes |
| <5s p95 | Provider throttling; slow consumer; poison message | Latency holds under the tail; queue drains |
| Per-user ordering | Worker crash mid-partition; rebalance | Order survives handoff to a new consumer |
| No SPOF | AZ loss; LB death; DB failure | Still serving, no data loss |
Notice the first row reuses everything from Lesson 1: the state-per-partition store, the changelog replay on takeover, and your trx_id idempotency key are precisely the machinery that makes "kill a worker mid-process" survivable. Failure testing is testing Lesson 1.
Fault injection is the act of deliberately making things fail in a controlled way: kill a process, blackhole network traffic, inject latency, break a dependency, fail a disk. A chaos experiment adds discipline on top:
Chaos is a test discipline, not "break production for fun." Netflix's Chaos Monkey is the origin — randomly kill instances so engineers are forced to build systems that survive. Tools you'll meet: Toxiproxy (inject network faults into a service you control), k6 fault injection, Gremlin, and your own shell scripts that kill a pod and watch what happens.
Fault injection only works if the code has the defense in depth that turns a fault into a blip instead of an outage:
| Control | Why |
|---|---|
| Timeouts everywhere | A dead dependency hangs forever without one; timeouts bound the damage. |
| Retries + backoff + jitter | Retry with exponential backoff; add random jitter so retries don't hit all at once (the retry storm). |
| Circuit breaker | After N failures, stop calling the dependency for a while — fail fast instead of hammering a dead provider. |
| Queue as the buffer | A provider outage should fill the queue, not fail the request. That's what the queue is for. |
The single most powerful proof of "no loss" is one you already know how to build: reconciliation as a test. A continuous job compares the source of truth (the transaction events) against the outcome (notification records in a delivered state). Any event without a matching delivered record is a bug — found automatically, not in an incident. You ran this at Jenius for money; run it for your guarantees. This is the test that runs forever, in production, quietly.
Q1. The original design died from:
Q2. A chaos test asserts that:
Q3. The best continuous "no loss" check:
Q4. When the push provider is down:
Primary sources: Chaos Engineering — Rosenthal & Jones (O'Reilly), and principlesofchaos.org. Netflix's chaos engineering blog posts are where it started.