Reliability
Keep critical user journeys available and recover within agreed limits. Set an SLO for service quality, an RTO for restoration time, and an RPO for acceptable data loss. Analyze dependency failures, remove single points of failure, and test recovery.
- Example
- Run the portal across availability zones, handle transient dependency failures with bounded retries, and rehearse database restoration. Check that the whole service meets its recovery targets.
- Tradeoff
- Extra replicas and standby regions increase cost and operational complexity; use the business impact of downtime to justify them.