Why System Design Matters
The Journey from Class to System
Section titled “The Journey from Class to System”You’ve written a beautiful class. It’s well-designed, follows SOLID principles, and has great test coverage. But software doesn’t run in isolation—it runs on servers, handles thousands of users, and must work 24/7.
What is System Design?
Section titled “What is System Design?”System design is the process of defining the architecture, components, and data flow of a system to meet specific requirements. It’s about making decisions that affect:
- How your code runs - On one server or thousands?
- How data flows - Synchronous or asynchronous?
- How failures are handled - What happens when things break?
- How the system scales - Can it handle 10x more users?
The Two Levels of Design
Section titled “The Two Levels of Design”| Aspect | High-Level Design (HLD) | Low-Level Design (LLD) |
|---|---|---|
| Focus | System architecture | Class structure |
| Scope | Multiple services | Single service/module |
| Artifacts | Architecture diagrams | Class diagrams |
| Decisions | Which database? How many servers? | Which pattern? What interface? |
| Scale | Millions of users | Thousands of objects |
Why LLD Engineers Need System Design
Section titled “Why LLD Engineers Need System Design”1. Your Code Doesn’t Run in Isolation
Section titled “1. Your Code Doesn’t Run in Isolation”Every class you write will eventually run in a system with:
class OrderService: """Looks simple, but consider the system context..."""
def __init__(self, db: Database, payment: PaymentGateway, inventory: InventoryService): self.db = db # Which database? Replicated? Sharded? self.payment = payment # External API - what if it's slow? self.inventory = inventory # Another service - what if it's down?
def place_order(self, order: Order) -> OrderResult: # What if this takes 30 seconds? # What if 1000 users call this simultaneously? # What if the database is in another data center?
self.inventory.reserve(order.items) # Network call #1 payment_result = self.payment.charge(order.total) # Network call #2 self.db.save(order) # Network call #3
return OrderResult(success=True, order_id=order.id)public class OrderService { // Looks simple, but consider the system context...
private final Database db; // Which database? Replicated? Sharded? private final PaymentGateway payment; // External API - what if it's slow? private final InventoryService inventory; // Another service - what if it's down?
public OrderService(Database db, PaymentGateway payment, InventoryService inventory) { this.db = db; this.payment = payment; this.inventory = inventory; }
public OrderResult placeOrder(Order order) { // What if this takes 30 seconds? // What if 1000 users call this simultaneously? // What if the database is in another data center?
inventory.reserve(order.getItems()); // Network call #1 PaymentResult paymentResult = payment.charge(order.getTotal()); // Network call #2 db.save(order); // Network call #3
return new OrderResult(true, order.getId()); }}class OrderService { // Looks simple, but consider the system context...
constructor( private db: Database, // Which database? Replicated? Sharded? private payment: PaymentGateway, // External API - what if it's slow? private inventory: InventoryService // Another service - what if it's down? ) {}
async placeOrder(order: Order): Promise<OrderResult> { // What if this takes 30 seconds? // What if 1000 users call this simultaneously? // What if the database is in another data center?
await this.inventory.reserve(order.items); // Network call #1 const paymentResult = await this.payment.charge(order.total); // Network call #2 await this.db.save(order); // Network call #3
return new OrderResult(true, order.id); }}class OrderService {private: Database* db; // Which database? Replicated? Sharded? PaymentGateway* payment; // External API - what if it's slow? InventoryService* inventory; // Another service - what if it's down?
public: OrderService(Database* db, PaymentGateway* payment, InventoryService* inventory) : db(db), payment(payment), inventory(inventory) {}
OrderResult placeOrder(const Order& order) { // What if this takes 30 seconds? // What if 1000 users call this simultaneously? // What if the database is in another data center?
inventory->reserve(order.getItems()); // Network call #1 PaymentResult paymentResult = payment->charge(order.getTotal()); // Network call #2 db->save(order); // Network call #3
return OrderResult(true, order.getId()); }};public class OrderService { // Looks simple, but consider the system context...
private readonly Database db; // Which database? Replicated? Sharded? private readonly PaymentGateway payment; // External API - what if it's slow? private readonly InventoryService inventory; // Another service - what if it's down?
public OrderService(Database db, PaymentGateway payment, InventoryService inventory) { this.db = db; this.payment = payment; this.inventory = inventory; }
public OrderResult PlaceOrder(Order order) { // What if this takes 30 seconds? // What if 1000 users call this simultaneously? // What if the database is in another data center?
inventory.Reserve(order.Items); // Network call #1 var paymentResult = payment.Charge(order.Total); // Network call #2 db.Save(order); // Network call #3
return new OrderResult(true, order.Id); }}2. Design Decisions Have System Implications
Section titled “2. Design Decisions Have System Implications”Every LLD decision affects the system:
| LLD Decision | System Implication |
|---|---|
| Using Singleton pattern | Won’t work across multiple servers |
| Storing state in instance variables | Can’t scale horizontally |
| Synchronous method calls | Creates coupling, blocks resources |
| In-memory caching | Each server has different cache |
| Auto-increment IDs | Conflicts in distributed databases |
3. Interviews Test Both Levels
Section titled “3. Interviews Test Both Levels”In senior engineering interviews, expect questions like:
The Five Pillars of System Design
Section titled “The Five Pillars of System Design”Every system design discussion involves these key concerns:
1. Scalability
Section titled “1. Scalability”Can the system handle growth?
LLD Impact: Design classes that can work in a distributed environment. Avoid global state, use dependency injection, make components stateless where possible.
2. Reliability
Section titled “2. Reliability”Does the system work correctly, even when things fail?
- Hardware fails (servers crash, disks die)
- Software has bugs
- Networks are unreliable
- Users make mistakes
LLD Impact: Implement proper error handling, use retry patterns, design for idempotency.
3. Availability
Section titled “3. Availability”Is the system accessible when users need it?
- 99.9% uptime = 8.76 hours downtime/year
- 99.99% uptime = 52.6 minutes downtime/year
- 99.999% uptime = 5.26 minutes downtime/year
LLD Impact: Design classes with fallback behaviors, implement circuit breakers, handle graceful degradation.
4. Maintainability
Section titled “4. Maintainability”Can the system be easily modified and operated?
- New features can be added
- Bugs can be fixed quickly
- Operations are simple
- System is observable
LLD Impact: Follow SOLID principles, write clean code, use design patterns appropriately.
5. Performance
Section titled “5. Performance”Does the system respond quickly and efficiently?
- Low latency (fast responses)
- High throughput (many requests)
- Efficient resource usage
LLD Impact: Choose appropriate data structures, optimize algorithms, minimize unnecessary operations.
Real-World Examples
Section titled “Real-World Examples”Example 1: Twitter’s Tweet Counter Evolution
Section titled “Example 1: Twitter’s Tweet Counter Evolution”Company: Twitter (now X)
Scenario: Twitter needs to display view counts, like counts, and retweet counts for billions of tweets. Initially, they used in-memory counters, but this failed at scale.
Implementation: Evolved from naive to distributed design:
Why This Matters:
- Scale: Billions of tweets, millions of interactions per second
- Consistency: Users expect accurate counts
- Performance: Counts must load instantly
- Result: Redis-based distributed counters handle millions of increments per second
Real-World Impact:
- Throughput: Millions of counter increments per second
- Latency: Sub-millisecond counter updates
- Consistency: All users see same counts globally
Example 2: Instagram’s Photo View Counter
Section titled “Example 2: Instagram’s Photo View Counter”Company: Instagram (Meta)
Scenario: Instagram displays view counts on photos and videos. With billions of photos and millions of views per second, they need a scalable counting system.
Implementation: Uses distributed counters with sharding:
Why Sharding?
- Scale: Distributes load across multiple Redis instances
- Capacity: Each shard handles subset of photos
- Performance: Parallel processing increases throughput
- Result: Handles billions of views with low latency
Real-World Impact:
- Scale: Billions of photos, trillions of views
- Performance: < 1ms counter increment latency
- Availability: 99.99% uptime despite massive scale
Example 3: YouTube’s View Counter System
Section titled “Example 3: YouTube’s View Counter System”Company: Google (YouTube)
Scenario: YouTube tracks view counts for billions of videos. The system must handle massive spikes during viral videos while maintaining accuracy.
Implementation: Uses hybrid approach with batching:
Why Batching?
- Efficiency: Reduces database writes by 100x
- Performance: Handles traffic spikes gracefully
- Accuracy: Eventually consistent, acceptable for views
- Result: Handles viral video traffic spikes
Real-World Impact:
- Scale: Billions of videos, trillions of views
- Spike Handling: Handles 10x traffic spikes during viral events
- Efficiency: 100x reduction in database writes through batching
Real-World Example: A Simple Counter
Section titled “Real-World Example: A Simple Counter”Let’s see how system thinking changes a simple class design:
Version 1: The Naive Approach
Section titled “Version 1: The Naive Approach”class PageViewCounter: """Simple counter - works perfectly on one server"""
def __init__(self): self.counts = {} # page_id -> count
def increment(self, page_id: str) -> int: if page_id not in self.counts: self.counts[page_id] = 0 self.counts[page_id] += 1 return self.counts[page_id]
def get_count(self, page_id: str) -> int: return self.counts.get(page_id, 0)
# Usagecounter = PageViewCounter()counter.increment("homepage") # 1counter.increment("homepage") # 2import java.util.HashMap;import java.util.Map;
public class PageViewCounter { // Simple counter - works perfectly on one server private Map<String, Integer> counts = new HashMap<>();
public synchronized int increment(String pageId) { int count = counts.getOrDefault(pageId, 0) + 1; counts.put(pageId, count); return count; }
public int getCount(String pageId) { return counts.getOrDefault(pageId, 0); }}
// Usagepublic class Main { public static void main(String[] args) { PageViewCounter counter = new PageViewCounter(); counter.increment("homepage"); // 1 counter.increment("homepage"); // 2 }}class PageViewCounter { // Simple counter - works perfectly on one server private counts: Map<string, number> = new Map();
increment(pageId: string): number { const current = this.counts.get(pageId) || 0; const newCount = current + 1; this.counts.set(pageId, newCount); return newCount; }
getCount(pageId: string): number { return this.counts.get(pageId) || 0; }}
// Usageconst counter = new PageViewCounter();counter.increment("homepage"); // 1counter.increment("homepage"); // 2#include <unordered_map>#include <string>
class PageViewCounter {private: std::unordered_map<std::string, int> counts;
public: int increment(const std::string& pageId) { counts[pageId]++; return counts[pageId]; }
int getCount(const std::string& pageId) const { auto it = counts.find(pageId); return it != counts.end() ? it->second : 0; }};
// UsagePageViewCounter counter;counter.increment("homepage"); // 1counter.increment("homepage"); // 2using System;using System.Collections.Generic;
public class PageViewCounter { private Dictionary<string, int> counts = new Dictionary<string, int>();
public int Increment(string pageId) { if (!counts.ContainsKey(pageId)) { counts[pageId] = 0; } counts[pageId]++; return counts[pageId]; }
public int GetCount(string pageId) { return counts.ContainsKey(pageId) ? counts[pageId] : 0; }}
// Usagevar counter = new PageViewCounter();counter.Increment("homepage"); // 1counter.Increment("homepage"); // 2Problems with this design:
- Data lost if server restarts
- Different counts on each server
- No persistence
- Memory grows unbounded
Version 2: System-Aware Design
Section titled “Version 2: System-Aware Design”The key insight is to externalize state to a shared store that all servers can access. This requires:
- Abstraction - Define an interface for storage (Dependency Inversion Principle)
- Shared State - Use Redis, a database, or similar shared storage
- Atomic Operations - Use Redis’s
INCRcommand which is atomic
import redis
class PageViewCounter: """Counter that works in distributed systems"""
def __init__(self, redis_client): self.redis = redis_client # External shared state
def increment(self, page_id: str) -> int: return self.redis.incr(f"pageview:{page_id}") # Atomic operation
def get_count(self, page_id: str) -> int: return int(self.redis.get(f"pageview:{page_id}") or 0)
# Now works across all servers!counter = PageViewCounter(redis.Redis(host='redis-cluster'))counter.increment("homepage")import redis.clients.jedis.Jedis;
public class PageViewCounter { private final Jedis redis; // External shared state
public PageViewCounter(Jedis redis) { this.redis = redis; }
public long increment(String pageId) { return redis.incr("pageview:" + pageId); // Atomic operation }
public long getCount(String pageId) { String count = redis.get("pageview:" + pageId); return count != null ? Long.parseLong(count) : 0; }}import Redis from 'ioredis';
class PageViewCounter { private redis: Redis;
constructor(redis: Redis) { this.redis = redis; // External shared state }
async increment(pageId: string): Promise<number> { return await this.redis.incr(`pageview:${pageId}`); // Atomic operation }
async getCount(pageId: string): Promise<number> { const count = await this.redis.get(`pageview:${pageId}`); return count ? parseInt(count, 10) : 0; }}
// Now works across all servers!const counter = new PageViewCounter(new Redis('redis-cluster'));await counter.increment("homepage");#include <hiredis/hiredis.h>#include <string>
class PageViewCounter {private: redisContext* redis; // External shared state
public: PageViewCounter(redisContext* redis) : redis(redis) {}
long increment(const std::string& pageId) { std::string key = "pageview:" + pageId; redisReply* reply = (redisReply*)redisCommand(redis, "INCR %s", key.c_str()); long result = reply->integer; freeReplyObject(reply); return result; // Atomic operation }
long getCount(const std::string& pageId) { std::string key = "pageview:" + pageId; redisReply* reply = (redisReply*)redisCommand(redis, "GET %s", key.c_str()); long result = reply ? std::stol(reply->str) : 0; freeReplyObject(reply); return result; }};using StackExchange.Redis;
public class PageViewCounter { private IDatabase redis; // External shared state
public PageViewCounter(IDatabase redis) { this.redis = redis; }
public long Increment(string pageId) { return redis.StringIncrement($"pageview:{pageId}"); // Atomic operation }
public long GetCount(string pageId) { var count = redis.StringGet($"pageview:{pageId}"); return count.HasValue ? (long)count : 0; }}
// Usagevar redis = ConnectionMultiplexer.Connect("redis-cluster").GetDatabase();var counter = new PageViewCounter(redis);counter.Increment("homepage");What changed and why:
| Change | System Design Reason |
|---|---|
Added CounterStorage interface | Decouples from specific storage (DIP) |
| Used Redis instead of in-memory | Shared state across servers |
| Dependency injection | Testable, flexible, swappable |
Atomic operations (INCR) | Handles concurrent requests |
Key Takeaways
Section titled “Key Takeaways”What’s Next?
Section titled “What’s Next?”Now that you understand why system design matters, let’s dive into the first fundamental concept:
Next up: Scalability Fundamentals - Learn how systems grow and the strategies to handle that growth.