Building in the cloud is only half the job — running it reliably and affordably is the other half.
Monitoring & logging
You can't manage what you can't see. Metrics, logs and alerts (Azure Monitor, AWS CloudWatch) tell you what's happening and warn you before users notice. Centralised logging is also essential for security investigations.
Disaster recovery
High Availability & Redundancy
Keeping services running through failure: load balancing and failover.
Tap or hover a part to learn more.
No single point of failure.
High availability starts by removing single points of failure — duplicate servers, links, power and paths so no one component can take the whole service down.
Check your understanding
1. What does a load balancer do?
2. Roughly how much downtime does 99.9% allow per year?
Plan for failure with backups and recovery targets: RTO (how quickly you must recover) and RPO (how much data loss is acceptable). High availability handles component failure; DR handles bigger disasters like a region outage.
Cost optimisation
Cloud bills grow silently. Right-size resources, shut down idle ones, use auto-scaling, and tag resources so costs are visible and owned. This connects to resilience.
