Memahami circuit breaker (Hystrix/Resilience4j), bulkhead isolation, retry with exponential backoff & jitter, graceful degradation, fallback strategies, dan chaos engineering (Netflix Chaos Monkey) untuk sistem yang bisa bertahan dari kegagalan

Setelah di episode 13 kita memahami sharding & partitioning, pada episode ini kita masuk ke topik yang menentukan apakah sistem bisa bertahan saat komponen gagal: fault tolerance dan resilience. Di production, kegagalan bukan kemungkinan — ia kepastian. Server mati, network terputus, dependency down. Pertanyaannya bukan "apakah kegagalan akan terjadi" tapi "bagaimana sistem merespons saat kegagalan terjadi."
Netflix mempopulerkan istilah "chaos engineering" — secara sengaja menyuntikkan kegagalan ke production untuk memastikan sistem bisa bertahan. Ini bukan gila — ini realistis. Sistem yang belum pernah gagal adalah sistem yang belum pernah diuji.
Circuit breaker memutus sementara request ke service yang sedang gagal, menghindari cascade failure.
CLOSED → (failure threshold tercapai) → OPEN → (timeout) → HALF-OPEN
↑ ↓
└─────────────── (success) ────────────────────────────────┘| State | Penjelasan |
|---|---|
| CLOSED | Normal — request diteruskan ke downstream |
| OPEN | Gagal — request langsung ditolak/tidak diteruskan |
| HALF-OPEN | Test — beberapa request diteruskan untuk cek recovery |
100 requests terakhir:
- 60% gagal (60 dari 100) → circuit opens
- Semua request berikutnya langsung ditolak selama 30 detik
- Setelah 30 detik → HALF-OPEN: 10 request test
- Jika 90% berhasil → circuit closes (normal)
- Jika masih gagal → circuit opens lagiBulkhead membatasi jumlah concurrent request ke setiap dependency — mencegah satu dependency yang lambat mempengaruhi seluruh sistem.
Service A:
├── Connection pool ke DB: max 20 connections
├── Connection pool ke Payment API: max 10 connections
└── Connection pool ke Notification API: max 5 connections
Jika Notification API lambat → hanya 5 request terpengaruh
→ Service A tetap bisa handle DB dan Payment APIThread pool (bulkhead) per dependency:
- DB pool: 20 threads, timeout 5s
- Payment pool: 10 threads, timeout 10s
- Notification pool: 5 threads, timeout 3s
Jika pool penuh → reject request dengan ThreadPoolRejectionException
→ client handle gracefullyRetry 1: tunggu 1s
Retry 2: tunggu 2s
Retry 3: tunggu 4s
Retry 4: tunggu 8s
Max retries: 3-5Tanpa jitter: semua client retry pada waktu yang sama → thundering herd
Dengan jitter: tambahkan randomness (0-1s) ke setiap retry
Retry 1: tunggu 1.2s (1 + random 0.2)
Retry 2: tunggu 2.8s (2 + random 0.8)
Retry 3: tunggu 5.1s (4 + random 1.1)Max retries: 3 per request
Max retry rate: 10% dari total traffic
Timeout per retry: 5s
Jika retry budget tercapai → stop retry, fail fastKetika dependency gagal, sistem tetap berfungsi dengan fitur yang berkurang.
Dependency: Product recommendation service DOWN
→ Tampilkan default/popular products (fallback)
→ User experience berkurang tapi tidak broken
Dependency: Payment gateway DOWN
→ Simpan order ke queue
→ Proses payment nanti (async recovery)
→ User dapat "order confirmed, processing"| Strategy | Contoh |
|---|---|
| Cache stale | Tampilkan data cache meskipun sudah expired |
| Default value | Rating = 4.5 jika rating service down |
| Feature flag | Disable fitur premium jika billing service down |
| Queue for later | Simpan ke queue, proses saat service recovery |
Chaos Monkey secara acak mematikan instance di production untuk memastikan sistem bisa bertahan.
1. Define "steady state" (baseline metrics)
2. Hypothesis: "killing instance X tidak mempengaruhi user"
3. Inject failure (kill instance, latency injection, network partition)
4. Observe: apakah hypothesis benar?
5. Fix weakness yang ditemukan
6. Repeat| Fault Type | Tool | Efek |
|---|---|---|
| Pod kill | Chaos Mesh, Litmus | Simulate server crash |
| Network latency | tc (traffic control) | Simulate slow network |
| Network partition | iptables | Simulate network split |
| CPU stress | stress-ng | Simulate high CPU |
| Disk failure | fault injection frameworks | Simulate disk I/O error |
Tip
Mulai chaos engineering dari environment staging/development. Jika staging bisa bertahan dari Chaos Monkey, production kemungkinan besar juga bisa. Netflix menguji di production karena mereka sudah mature — jangan contoh mereka tanpa fondasi yang kuat.
Service A → CircuitBreaker → Service B
Config:
- failureThreshold: 5 failures dalam 60 detik
- resetTimeout: 30 detik
- halfOpenMax: 10 requests
Flow:
1. Normal: request diteruskan ke B
2. B mulai gagal: 5 failures terdeteksi
3. Circuit OPENS: semua request langsung return fallback
4. Setelah 30 detik: HALF-OPEN, 10 request test
5. Jika 90% berhasil: circuit CLOSES (normal)
6. Jika masih gagal: circuit OPENS lagiInti yang harus dibawa pulang:
Di episode 15 selanjutnya kita akan membahas microservices architecture — service boundary, synchronous vs asynchronous communication, service mesh, trade-off complexity, dan kapan memilih microservices vs monolith. Arsitektur microservices adalah pilihan paling populer untuk scalable systems — tapi bukan tanpa trade-off!