Mempelajari cyber resilience, business continuity planning, disaster recovery architecture, dan designing for failure dalam security architecture

Setelah di episode 19 kita mempelajari secure AI/LLM architecture, pada episode ini kita dalami resilience & BCP security — bagaimana mendesain sistem yang bisa bertahan dan pulih dari serangan. Resilience bukan hanya availability — ia mencakup kemampuan mendeteksi, contain, recover, dan belajar dari insiden.
Mengapa resilience penting? Karena breach sudah tidak bisa dicegah sepenuhnya — yang bisa dilakukan adalah membatasi dampak dan mempercepat recovery. Sistem yang resilient tidak hanya survive serangan, tapi juga belajar dari itu.
| Phase | Activities |
|---|---|
| Prepare | Planning, training, backup |
| Protect | Defense in depth, redundancy |
| Detect | Monitoring, anomaly detection |
| Respond | Incident response, containment |
| Recover | Restoration, validation |
| Learn | Post-mortem, improvement |
| Metric | Target |
|---|---|
| RTO (Recovery Time Objective) | < 4 hours |
| RPO (Recovery Point Objective) | < 1 hour |
| MTBF (Mean Time Between Failures) | > 30 days |
| MTTR (Mean Time To Recover) | < 2 hours |
| Component | Detail |
|---|---|
| Business impact analysis | Identify critical functions |
| Recovery strategies | How to recover |
| Plan development | Detailed procedures |
| Testing & exercises | Validate plan |
| Maintenance | Keep plan current |
| Level | Description | RTO |
|---|---|---|
| Critical | Revenue-generating, regulatory | < 1 hour |
| Important | Business operations | < 4 hours |
| Normal | Non-critical | < 24 hours |
| Low | Can wait | < 72 hours |
Tip
BCP harus diuji secara berkala — minimal annually, idealnya quarterly. Tabletop exercises adalah cara termurah untuk menguji rencana. Jangan menunggu insiden nyata untuk mengetahui apakah rencana berfungsi.
| Pattern | Description | RTO | Cost |
|---|---|---|---|
| Backup & restore | Periodic backup | Hours | Low |
| Pilot light | Minimal standby | Minutes | Medium |
| Warm standby | Partial standby | Minutes | High |
| Multi-site active | Full active-active | Seconds | Very High |
| Component | Resilience Strategy |
|---|---|
| Compute | Auto-scaling, multi-AZ |
| Storage | Replication, versioning |
| Network | Redundant paths, failover |
| Database | Read replicas, failover |
| DNS | Multi-provider, health checks |
Note
Disaster recovery architecture harus dirancang untuk failure mode yang berbeda: component failure, AZ failure, region failure, dan account compromise. Setiap mode membutuhkan strategi yang berbeda.
| Threat | Resilience Response |
|---|---|
| Ransomware | Backup integrity, immutable storage |
| DDoS | CDN, rate limiting, failover |
| Data breach | Data minimization, encryption |
| Supply chain attack | SBOM, vendor diversity |
| Insider threat | Least privilege, monitoring |
Inti yang harus dibawa pulang:
Di episode 21 selanjutnya kita akan membahas enterprise security patterns — microservices security, API gateway patterns, dan service mesh patterns.