0002-payday-gauntlet

A worked example running the whole senior discipline on one scenario: payday, Friday 09:00 WIB. Two rounds of sizing → load-test plan → chaos plan → degradation policy. Every claim is a number, written to be checked.

The scenario

1M users. Payday — salary credits land, and marketing fires a sweep at the same moment:

Flow Volume Shape
Transaction alerts 400k events (200k users × 2) 08:00–10:00, spiky in the first 5 min
Marketing push 500k pushes, scheduled 09:00 one-shot burst (the trap)
Baseline background ~700k events spread across the day
Peak day total ≈1.6M events (vs ~821k avg) ≈2× a normal day

Round 1 — sizing, the naive pass (the trap)

Firing marketing flat at 09:00:

The naive math then forces an unnatural design: Kafka sized for a 9,000/s spike you see once a month, 100+ partitions to absorb it, autoscaling that panics. The trap is reframing a rate problem as a scale problem while the volume (500k users) was always the same.

Round 2 — sizing, the correct pass (reframe the rate, not the volume)

Same volume, different rate, same users served:

Reframing rule: spreading the same work over time changes the rate, never the volume.

Infrastructure numbers from the sizing

Component Value Why
Transaction topic 16 partitions, key user_id dedup + per-user ordering + co-location (Lesson 1)
Marketing topic 8 partitions, salted keys ordering irrelevant; pace-limited and priority-queued
Consumers 16 txn workers (fixed) + marketing pool (autoscale) txn ordering needs stable ownership; marketing drains as capacity allows
Storage ~1.5 kB/record → ≈2.4 GB/day → 17 GB @ 7-day retention (+changelog ≈26 GB) not a design driver
Bandwidth not a constraint at 840/s the constraint is ingest QoS, not pipe

Load-test plan (proves the happy path)

k6, the "payday simulation": run the exact payday mix at 840 events/s for 2 hours — a soak, not a 60-second spike. 70% txn-shaped (spiky arrivals), 30% marketing-paced.

Assertions: 1. Ingest p95 < 5 s, p99 < 8 s — the transport guarantee 2. Zero dropped messages — reconciliation + DLQ counters flat 3. State-store replay healthy across forced rebalances 4. Worker memory flat over the full 2 h (catches the leak) 5. Autoscale by queue depth scales consumers up within 60 s

This plan answers only "does it hold when healthy?" — the chaos plan answers the rest.


Chaos plan (proves the guarantees)

Guarantee-first: for each promise, find the failure that would break it, and assert.

Guarantee Failure injected Hypothesis Blast radius → abort
Never lost kill a worker mid-processing during peak queue redelivers; dedup skips the duplicate; nothing lost 1 worker → abort if any delivered record missing (reconciliation)
Never lost kill a Kafka partition leader failover promotes a replica; committed offsets intact 1 leader → abort if committed offsets lost / gap in delivered
<5 s p95 ingest gate push calls +12 s (Toxiproxy) ingest unaffected (queue absorbs); txn alerts drain ahead; delivery SLA holds 1 provider channel → abort if ingest p95 > 5 s for 2 min
Per-user ordering kill a consumer mid-partition new owner replays changelog, resumes same-key order 1 partition → abort if users see out-of-order
No SPOF kill an entire AZ surviving AZ absorbs; Kafka 3× replica serves; zero alerts lost 1 AZ → abort on missing records or drain > SLA
Never blocked publish a poison message to the txn topic DLQ after N retries; partition never blocks; dedup skips 1 message → abort if partition p95 > 5 s

Degradation policy (what "OK" means when it breaks)

Every guarantee ships with its fallback. The policy is the contract the chaos plan asserts.

Guarantee Failed by Accepted degradation Time bound Watch
Never lost anything none — zero loss, always forever reconciliation every 5 min; alert on any mismatch
<5 s p95 ingest provider outage ingest stays <5 s; delivery falls back to 99% within 15 min, 100% eventually page at 10 min queue depth; delivery-age p95
Provider down 30+ min outage beyond backoff 100% eventually delivered; txn alerts prioritized, marketing deferred page at 10 min delivery age, DLQ depth
No SPOF AZ loss surviving AZ serves; brief p95 blip < 5 min; txn ordering intact < 5 min cross-AZ traffic, ingest p95
Marketing arrival any failure marketing delays until txn backlog drains (priority queue) indefinite marketing queue lag

What this audit forces you to decide, before launch

  1. The <5 s p95 guarantee is scoped to ingest — it says nothing about delivery, which is its own SLA.
  2. Marketing is pace-limited and depth-queued — it is never fired flat, and it waits behind a txn backlog.
  3. Reconciliation is the always-on never-lost test — not an experiment, a 24/7 job.
  4. Sizing gives you margin, not precision — that's why you scale on queue depth, not on Monday-morning hope.