Availability Patterns
What is Availability?
Section titled “What is Availability?”Availability measures how often your system is up and working. It’s the percentage of time users can successfully use your service.
The Nines of Availability
Section titled “The Nines of Availability”The industry measures availability in “nines” — each additional nine dramatically reduces allowed downtime:
Downtime by Availability Level
Section titled “Downtime by Availability Level”| Availability | Downtime/Year | Downtime/Month | Downtime/Week | Use Case |
|---|---|---|---|---|
| 99% | 3.65 days | 7.2 hours | 1.68 hours | Development/Test |
| 99.9% | 8.76 hours | 43.8 min | 10.1 min | Standard apps |
| 99.95% | 4.38 hours | 21.9 min | 5 min | E-commerce |
| 99.99% | 52.6 min | 4.38 min | 1 min | Financial services |
| 99.999% | 5.26 min | 26.3 sec | 6 sec | Life-critical systems |
SLI, SLO, and SLA: The Availability Triangle
Section titled “SLI, SLO, and SLA: The Availability Triangle”Understanding these three terms is crucial for any engineer:
The Relationship
Section titled “The Relationship”| Term | Definition | Owner | Consequence of Miss |
|---|---|---|---|
| SLI | The metric you measure | Engineering | Investigation triggered |
| SLO | Internal target (stricter than SLA) | Engineering + Product | Team prioritizes fixes |
| SLA | External promise to customers | Business | Financial penalties, lost trust |
Common SLIs to Track
Section titled “Common SLIs to Track”| Category | SLI | What It Measures |
|---|---|---|
| Availability | Success rate | % of requests that succeed |
| Latency | P50, P95, P99 response time | How fast responses are |
| Throughput | Requests per second | System capacity |
| Error Rate | 5xx errors / total requests | Failure frequency |
| Saturation | CPU, memory, queue depth | How “full” the system is |
High Availability Patterns
Section titled “High Availability Patterns”Pattern 1: Redundancy (Eliminate Single Points of Failure)
Section titled “Pattern 1: Redundancy (Eliminate Single Points of Failure)”A Single Point of Failure (SPOF) is any component whose failure brings down the entire system. HA systems eliminate SPOFs through redundancy at every layer.
Redundancy Levels
Section titled “Redundancy Levels”| Level | Description | Example |
|---|---|---|
| Active-Passive | Backup sits idle until needed | Secondary DB that takes over on primary failure |
| Active-Active | All copies handle traffic | Multiple servers behind a load balancer |
| N+1 | Run one extra node for safety | Need 3 servers? Run 4 |
| N+2 | Extra buffer for maintenance + failure | Need 3 servers? Run 5 |
Pattern 2: Failover (Automatic Recovery)
Section titled “Pattern 2: Failover (Automatic Recovery)”When the primary fails, traffic automatically switches to the backup.
Key Failover Metrics:
| Metric | Description | Typical Target |
|---|---|---|
| Detection Time | How quickly we notice the failure | < 10 seconds |
| Failover Time | How long to switch to backup | < 30 seconds |
| Recovery Time | Total time until service restored | < 1 minute |
Pattern 3: Graceful Degradation
Section titled “Pattern 3: Graceful Degradation”When parts of your system fail, continue serving users with reduced functionality rather than complete failure.
The Principle: Identify which features are core vs nice-to-have, and ensure core features work even when nice-to-haves fail.
| E-commerce Example | Category | On Failure |
|---|---|---|
| Product info | Core | Must work — show error page if down |
| Recommendations | Nice-to-have | Hide section, show empty |
| Reviews | Nice-to-have | Hide section, show cached |
| Real-time inventory | Nice-to-have | Show “In Stock” (cached) |
| Checkout | Core | Must work — queue if payment down |
Real-World Examples
Section titled “Real-World Examples”Example 1: AWS S3 Availability (99.999999999% Durability)
Section titled “Example 1: AWS S3 Availability (99.999999999% Durability)”Company: Amazon Web Services (AWS)
Scenario: AWS S3 (Simple Storage Service) stores trillions of objects for millions of customers. The service must guarantee that data is never lost, even with hardware failures, data center outages, or natural disasters.
Implementation: Uses multi-region replication and erasure coding:
Why This Works:
- Erasure Coding: Can lose up to 4 fragments and still reconstruct the object
- Multi-AZ: Fragments stored across multiple availability zones
- Multi-Region: Complete copies in different geographic regions
- Result: 99.999999999% (11 nines) durability
Real-World Impact:
- Scale: Stores over 100 trillion objects
- Uptime: 99.99% availability SLA
- Durability: Designed for 99.999999999% (losing 1 object per 10,000 years)
Example 2: Google Search Availability (99.999% Uptime)
Section titled “Example 2: Google Search Availability (99.999% Uptime)”Company: Google
Scenario: Google Search handles billions of queries daily. Even a few minutes of downtime would impact millions of users worldwide.
Implementation: Uses massive redundancy and automatic failover:
Why This Works:
- Geographic Redundancy: Multiple data centers per region
- N+2 Redundancy: Always run 2 extra servers beyond capacity needs
- Automatic Failover: Failed servers removed in seconds
- Result: 99.999% (five nines) availability
Real-World Impact:
- Queries: 8.5+ billion searches per day
- Uptime: Less than 5 minutes downtime per year
- Failover Time: < 30 seconds for automatic recovery
Example 3: Netflix Streaming (99.99% Availability)
Section titled “Example 3: Netflix Streaming (99.99% Availability)”Company: Netflix
Scenario: Netflix streams content to 200+ million subscribers. Service interruptions directly impact user experience and subscription retention.
Implementation: Uses CDN distribution and graceful degradation:
Why This Works:
- CDN Distribution: Content cached at edge locations worldwide
- Multiple CDN Providers: Fallback to different CDN if one fails
- Quality Degradation: Lower bitrate if bandwidth limited
- Result: 99.99% availability even during peak hours
Real-World Impact:
- Peak Traffic: 15% of global internet bandwidth during peak hours
- Streaming Quality: Automatic quality adjustment based on network conditions
- Availability: 99.99% uptime despite massive scale
Example 4: GitHub Availability (99.95% SLA)
Section titled “Example 4: GitHub Availability (99.95% SLA)”Company: GitHub (Microsoft)
Scenario: GitHub hosts millions of repositories and serves millions of developers. Downtime impacts productivity and developer workflows globally.
Implementation: Uses active-active replication and read replicas:
Why This Works:
- Active-Active: Multiple primary databases can handle writes
- Read Replicas: Distribute read load across multiple regions
- Automatic Failover: Primary failure triggers automatic promotion of replica
- Result: 99.95% availability SLA
Real-World Impact:
- Repositories: 100+ million repositories
- Users: 100+ million developers
- Uptime: 99.95% SLA with credits for downtime
- Failover: < 60 seconds for automatic failover
Example 5: Stripe Payment Processing (99.99% Availability)
Section titled “Example 5: Stripe Payment Processing (99.99% Availability)”Company: Stripe
Scenario: Stripe processes billions of dollars in payments. Payment failures directly impact merchant revenue and customer trust.
Implementation: Uses multi-region active-active architecture:
Why This Works:
- Multi-Region Active-Active: Both regions can process payments
- Synchronous Replication: Critical payment data replicated synchronously
- Automatic Failover: < 30 seconds failover time
- Result: 99.99% availability with financial guarantees
Real-World Impact:
- Transaction Volume: Billions of dollars processed monthly
- Uptime: 99.99% availability SLA
- Failover: < 30 seconds automatic failover
- Financial Guarantees: Credits for downtime exceeding SLA
LLD ↔ HLD Connection
Section titled “LLD ↔ HLD Connection”How availability concepts affect your class design:
Failover Pattern Implementation
Section titled “Failover Pattern Implementation”When a primary service fails, automatically switch to a backup:
from typing import List, Optionalfrom abc import ABC, abstractmethod
class ServiceClient(ABC): @abstractmethod def call(self, request: str) -> str: pass
@abstractmethod def is_healthy(self) -> bool: pass
class FailoverService: def __init__(self, clients: List[ServiceClient]): self.clients = clients self.current_index = 0
def call_with_failover(self, request: str) -> Optional[str]: attempts = 0 max_attempts = len(self.clients)
while attempts < max_attempts: client = self.clients[self.current_index]
if not client.is_healthy(): # Move to next client self.current_index = (self.current_index + 1) % len(self.clients) attempts += 1 continue
try: return client.call(request) except Exception as e: # Failover to next client self.current_index = (self.current_index + 1) % len(self.clients) attempts += 1
return None # All clients failedimport java.util.*;
public interface ServiceClient { String call(String request) throws Exception; boolean isHealthy();}
public class FailoverService { private List<ServiceClient> clients; private int currentIndex = 0;
public FailoverService(List<ServiceClient> clients) { this.clients = clients; }
public String callWithFailover(String request) { int attempts = 0; int maxAttempts = clients.size();
while (attempts < maxAttempts) { ServiceClient client = clients.get(currentIndex);
if (!client.isHealthy()) { currentIndex = (currentIndex + 1) % clients.size(); attempts++; continue; }
try { return client.call(request); } catch (Exception e) { currentIndex = (currentIndex + 1) % clients.size(); attempts++; } }
return null; // All clients failed }}interface ServiceClient { call(request: string): Promise<string>; isHealthy(): boolean;}
class FailoverService { private clients: ServiceClient[]; private currentIndex: number = 0;
constructor(clients: ServiceClient[]) { this.clients = clients; }
async callWithFailover(request: string): Promise<string | null> { let attempts = 0; const maxAttempts = this.clients.length;
while (attempts < maxAttempts) { const client = this.clients[this.currentIndex];
if (!client.isHealthy()) { this.currentIndex = (this.currentIndex + 1) % this.clients.length; attempts++; continue; }
try { return await client.call(request); } catch (error) { this.currentIndex = (this.currentIndex + 1) % this.clients.length; attempts++; } }
return null; // All clients failed }}#include <vector>#include <string>#include <memory>
class ServiceClient {public: virtual ~ServiceClient() = default; virtual std::string call(const std::string& request) = 0; virtual bool isHealthy() const = 0;};
class FailoverService {private: std::vector<std::shared_ptr<ServiceClient>> clients; size_t currentIndex = 0;
public: FailoverService(std::vector<std::shared_ptr<ServiceClient>> clients) : clients(std::move(clients)) {}
std::string callWithFailover(const std::string& request) { size_t attempts = 0; size_t maxAttempts = clients.size();
while (attempts < maxAttempts) { auto client = clients[currentIndex];
if (!client->isHealthy()) { currentIndex = (currentIndex + 1) % clients.size(); attempts++; continue; }
try { return client->call(request); } catch (...) { currentIndex = (currentIndex + 1) % clients.size(); attempts++; } }
return ""; // All clients failed }};using System;using System.Collections.Generic;
public interface IServiceClient { string Call(string request); bool IsHealthy();}
public class FailoverService { private List<IServiceClient> clients; private int currentIndex = 0;
public FailoverService(List<IServiceClient> clients) { this.clients = clients; }
public string CallWithFailover(string request) { int attempts = 0; int maxAttempts = clients.Count;
while (attempts < maxAttempts) { var client = clients[currentIndex];
if (!client.IsHealthy()) { currentIndex = (currentIndex + 1) % clients.Count; attempts++; continue; }
try { return client.Call(request); } catch (Exception e) { currentIndex = (currentIndex + 1) % clients.Count; attempts++; } }
return null; // All clients failed }}Graceful Degradation Implementation
Section titled “Graceful Degradation Implementation”Continue serving users with reduced functionality when dependencies fail:
from typing import Optionalfrom enum import Enum
class ServiceStatus(Enum): AVAILABLE = "available" DEGRADED = "degraded" UNAVAILABLE = "unavailable"
class ProductService: def __init__(self): self.recommendation_service = None self.review_service = None
def get_product_page(self, product_id: str) -> dict: # Core feature - must work product = self._get_product_details(product_id)
# Nice-to-have features - degrade gracefully recommendations = self._get_recommendations(product_id) reviews = self._get_reviews(product_id)
return { "product": product, # Core - always present "recommendations": recommendations, # Optional "reviews": reviews # Optional }
def _get_recommendations(self, product_id: str) -> Optional[list]: try: if self.recommendation_service and self.recommendation_service.is_healthy(): return self.recommendation_service.get_recommendations(product_id) except Exception: pass return None # Gracefully degrade - hide recommendations
def _get_reviews(self, product_id: str) -> Optional[list]: try: if self.review_service and self.review_service.is_healthy(): return self.review_service.get_reviews(product_id) except Exception: # Return cached reviews if available return self._get_cached_reviews(product_id) return None # Gracefully degrade - hide reviewsimport java.util.*;
public class ProductService { private RecommendationService recommendationService; private ReviewService reviewService;
public ProductPage getProductPage(String productId) { // Core feature - must work Product product = getProductDetails(productId);
// Nice-to-have features - degrade gracefully List<Recommendation> recommendations = getRecommendations(productId); List<Review> reviews = getReviews(productId);
return new ProductPage(product, recommendations, reviews); }
private List<Recommendation> getRecommendations(String productId) { try { if (recommendationService != null && recommendationService.isHealthy()) { return recommendationService.getRecommendations(productId); } } catch (Exception e) { // Log error } return null; // Gracefully degrade - hide recommendations }
private List<Review> getReviews(String productId) { try { if (reviewService != null && reviewService.isHealthy()) { return reviewService.getReviews(productId); } } catch (Exception e) { // Return cached reviews if available return getCachedReviews(productId); } return null; // Gracefully degrade - hide reviews }}interface ProductPage { product: Product; recommendations?: Recommendation[]; reviews?: Review[];}
class ProductService { private recommendationService?: RecommendationService; private reviewService?: ReviewService;
getProductPage(productId: string): ProductPage { // Core feature - must work const product = this.getProductDetails(productId);
// Nice-to-have features - degrade gracefully const recommendations = this.getRecommendations(productId); const reviews = this.getReviews(productId);
return { product, // Core - always present recommendations, // Optional reviews // Optional }; }
private getRecommendations(productId: string): Recommendation[] | undefined { try { if (this.recommendationService?.isHealthy()) { return this.recommendationService.getRecommendations(productId); } } catch (error) { // Log error } return undefined; // Gracefully degrade - hide recommendations }
private getReviews(productId: string): Review[] | undefined { try { if (this.reviewService?.isHealthy()) { return this.reviewService.getReviews(productId); } } catch (error) { // Return cached reviews if available return this.getCachedReviews(productId); } return undefined; // Gracefully degrade - hide reviews }}#include <optional>#include <vector>#include <memory>
struct ProductPage { Product product; // Core - always present std::optional<std::vector<Recommendation>> recommendations; // Optional std::optional<std::vector<Review>> reviews; // Optional};
class ProductService {private: std::shared_ptr<RecommendationService> recommendationService; std::shared_ptr<ReviewService> reviewService;
public: ProductPage getProductPage(const std::string& productId) { // Core feature - must work Product product = getProductDetails(productId);
// Nice-to-have features - degrade gracefully auto recommendations = getRecommendations(productId); auto reviews = getReviews(productId);
return {product, recommendations, reviews}; }
private: std::optional<std::vector<Recommendation>> getRecommendations(const std::string& productId) { try { if (recommendationService && recommendationService->isHealthy()) { return recommendationService->getRecommendations(productId); } } catch (...) { // Log error } return std::nullopt; // Gracefully degrade - hide recommendations }
std::optional<std::vector<Review>> getReviews(const std::string& productId) { try { if (reviewService && reviewService->isHealthy()) { return reviewService->getReviews(productId); } } catch (...) { // Return cached reviews if available return getCachedReviews(productId); } return std::nullopt; // Gracefully degrade - hide reviews }};using System;using System.Collections.Generic;
public class ProductPage { public Product Product { get; set; } // Core - always present public List<Recommendation> Recommendations { get; set; } // Optional public List<Review> Reviews { get; set; } // Optional}
public class ProductService { private RecommendationService recommendationService; private ReviewService reviewService;
public ProductPage GetProductPage(string productId) { // Core feature - must work var product = GetProductDetails(productId);
// Nice-to-have features - degrade gracefully var recommendations = GetRecommendations(productId); var reviews = GetReviews(productId);
return new ProductPage { Product = product, // Core - always present Recommendations = recommendations, // Optional Reviews = reviews // Optional }; }
private List<Recommendation> GetRecommendations(string productId) { try { if (recommendationService != null && recommendationService.IsHealthy()) { return recommendationService.GetRecommendations(productId); } } catch (Exception e) { // Log error } return null; // Gracefully degrade - hide recommendations }
private List<Review> GetReviews(string productId) { try { if (reviewService != null && reviewService.IsHealthy()) { return reviewService.GetReviews(productId); } } catch (Exception e) { // Return cached reviews if available return GetCachedReviews(productId); } return null; // Gracefully degrade - hide reviews }}Key Takeaways
Section titled “Key Takeaways”What’s Next?
Section titled “What’s Next?”Now that you understand availability concepts, let’s dive into how systems stay consistent across replicas:
Next up: Replication Strategies — Learn how data is replicated for both availability and performance.