Labs ICT
⭐ Pro Login

High Availability and Fault Tolerance

Designing systems that stay up even when things fail

2 min read | Cloud Computing
⭐

Want the full learning experience?

Get structured courses, certificates, projects, and instructor support with LabsICT Pro.

Explore Pro Courses

High Availability and Fault Tolerance

High Availability (HA) means your system remains operational for a high percentage of time. Fault Tolerance means the system continues operating even when components fail. Together, they ensure your applications stay up when things go wrong.

Key Concepts


  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚  Availability = Uptime / Total Time Γ— 100            β”‚
  β”‚                                                      β”‚
  β”‚  99.9%  = 8.76 hours downtime/year  (Three 9s)      β”‚
  β”‚  99.99% = 52.6 minutes downtime/year (Four 9s)      β”‚
  β”‚  99.999%= 5.26 minutes downtime/year (Five 9s)      β”‚
  β”‚                                                      β”‚
  β”‚  Each "9" is 10x more expensive to achieve           β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

AWS High Availability Architecture


  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚                                                      β”‚
  β”‚           Region: us-east-1                          β”‚
  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”‚
  β”‚  β”‚  AZ-a           β”‚    β”‚  AZ-b           β”‚         β”‚
  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚    β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚         β”‚
  β”‚  β”‚  β”‚    ALB    β”‚  β”‚    β”‚  β”‚    ALB    β”‚  β”‚         β”‚
  β”‚  β”‚  β”‚  (active) β”‚  β”‚    β”‚  β”‚ (standby) β”‚  β”‚         β”‚
  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β”‚    β”‚  β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β”‚         β”‚
  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”  β”‚    β”‚  β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”  β”‚         β”‚
  β”‚  β”‚  β”‚  EC2      β”‚  β”‚    β”‚  β”‚  EC2      β”‚  β”‚         β”‚
  β”‚  β”‚  β”‚  Instance β”‚  β”‚    β”‚  β”‚  Instance β”‚  β”‚         β”‚
  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚    β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚         β”‚
  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚    β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚         β”‚
  β”‚  β”‚  β”‚   RDS     β”‚  │◄───│─▢│   RDS     β”‚  β”‚         β”‚
  β”‚  β”‚  β”‚  Primary  β”‚  β”‚syncβ”‚  β”‚  Standby  β”‚  β”‚         β”‚
  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚    β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚         β”‚
  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β”‚
  β”‚                                                      β”‚
  β”‚  If AZ-a fails β†’ ALB routes to AZ-b                 β”‚
  β”‚  RDS fails over to standby automatically             β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Strategies for High Availability

Redundancy: Run multiple instances of everything. If one fails, another takes over.

Multi-AZ: Spread resources across availability zones within a region.

Multi-Region: Deploy across regions for disaster recovery and global performance.

Health Checks: Automatically detect and replace unhealthy instances.

Load Balancing: Distribute traffic so no single instance is a bottleneck.

Common Anti-Patterns

Running everything in a single AZ (single point of failure). Not testing failover procedures. Assuming a single database instance is enough. Not monitoring system health. Ignoring DNS TTL for faster failover.

πŸ§ͺ Quick Quiz

What is high availability in cloud architecture?