A worked example running the whole senior discipline on one scenario: payday, Friday 09:00 WIB. Two rounds of sizing → load-test plan → chaos plan → degradation policy. Every claim is a number, written to be checked.
1M users. Payday — salary credits land, and marketing fires a sweep at the same moment:
| Flow | Volume | Shape |
|---|---|---|
| Transaction alerts | 400k events (200k users × 2) | 08:00–10:00, spiky in the first 5 min |
| Marketing push | 500k pushes, scheduled 09:00 | one-shot burst (the trap) |
| Baseline background | ~700k events | spread across the day |
| Peak day total | ≈1.6M events (vs ~821k avg) | ≈2× a normal day |
Firing marketing flat at 09:00:
The naive math then forces an unnatural design: Kafka sized for a 9,000/s spike you see once a month, 100+ partitions to absorb it, autoscaling that panics. The trap is reframing a rate problem as a scale problem while the volume (500k users) was always the same.
Same volume, different rate, same users served:
Reframing rule: spreading the same work over time changes the rate, never the volume.
| Component | Value | Why |
|---|---|---|
| Transaction topic | 16 partitions, key user_id |
dedup + per-user ordering + co-location (Lesson 1) |
| Marketing topic | 8 partitions, salted keys | ordering irrelevant; pace-limited and priority-queued |
| Consumers | 16 txn workers (fixed) + marketing pool (autoscale) | txn ordering needs stable ownership; marketing drains as capacity allows |
| Storage | ~1.5 kB/record → ≈2.4 GB/day → 17 GB @ 7-day retention (+changelog ≈26 GB) | not a design driver |
| Bandwidth | not a constraint at 840/s | the constraint is ingest QoS, not pipe |
k6, the "payday simulation": run the exact payday mix at 840 events/s for 2 hours — a soak, not a 60-second spike. 70% txn-shaped (spiky arrivals), 30% marketing-paced.
Assertions: 1. Ingest p95 < 5 s, p99 < 8 s — the transport guarantee 2. Zero dropped messages — reconciliation + DLQ counters flat 3. State-store replay healthy across forced rebalances 4. Worker memory flat over the full 2 h (catches the leak) 5. Autoscale by queue depth scales consumers up within 60 s
This plan answers only "does it hold when healthy?" — the chaos plan answers the rest.
Guarantee-first: for each promise, find the failure that would break it, and assert.
| Guarantee | Failure injected | Hypothesis | Blast radius → abort |
|---|---|---|---|
| Never lost | kill a worker mid-processing during peak | queue redelivers; dedup skips the duplicate; nothing lost | 1 worker → abort if any delivered record missing (reconciliation) |
| Never lost | kill a Kafka partition leader | failover promotes a replica; committed offsets intact | 1 leader → abort if committed offsets lost / gap in delivered |
| <5 s p95 ingest | gate push calls +12 s (Toxiproxy) | ingest unaffected (queue absorbs); txn alerts drain ahead; delivery SLA holds | 1 provider channel → abort if ingest p95 > 5 s for 2 min |
| Per-user ordering | kill a consumer mid-partition | new owner replays changelog, resumes same-key order | 1 partition → abort if users see out-of-order |
| No SPOF | kill an entire AZ | surviving AZ absorbs; Kafka 3× replica serves; zero alerts lost | 1 AZ → abort on missing records or drain > SLA |
| Never blocked | publish a poison message to the txn topic | DLQ after N retries; partition never blocks; dedup skips | 1 message → abort if partition p95 > 5 s |
Every guarantee ships with its fallback. The policy is the contract the chaos plan asserts.
| Guarantee | Failed by | Accepted degradation | Time bound | Watch |
|---|---|---|---|---|
| Never lost | anything | none — zero loss, always | forever | reconciliation every 5 min; alert on any mismatch |
| <5 s p95 ingest | provider outage | ingest stays <5 s; delivery falls back to 99% within 15 min, 100% eventually | page at 10 min | queue depth; delivery-age p95 |
| Provider down 30+ min | outage beyond backoff | 100% eventually delivered; txn alerts prioritized, marketing deferred | page at 10 min | delivery age, DLQ depth |
| No SPOF | AZ loss | surviving AZ serves; brief p95 blip < 5 min; txn ordering intact | < 5 min | cross-AZ traffic, ingest p95 |
| Marketing arrival | any failure | marketing delays until txn backlog drains (priority queue) | indefinite | marketing queue lag |