Skip to content
LowLevelDesign Mastery

Understanding Bottlenecks

Find the weakest link before it breaks

A bottleneck is the component that limits your system’s overall performance. No matter how fast other parts are, the system can only go as fast as its slowest component.

Bottleneck in a pipeline: components A and C handle 1000 requests per second but component B handles only 100, capping system throughput at 100

Symptoms: High CPU usage, slow computations

💡 Tip: Click dropdown to switch between languages
cpu_bottleneck.py
# CPU-bound operation blocking the event loop
def calculate_fibonacci(n: int) -> int:
if n <= 1:
return n
return calculate_fibonacci(n-1) + calculate_fibonacci(n-2)
# Solution: Use efficient algorithm or offload to worker
from functools import lru_cache
@lru_cache(maxsize=1000)
def calculate_fibonacci_cached(n: int) -> int:
if n <= 1:
return n
return calculate_fibonacci_cached(n-1) + calculate_fibonacci_cached(n-2)

Symptoms: High memory usage, OOM errors, GC pauses

💡 Tip: Click dropdown to switch between languages
memory_bottleneck.py
# Loading everything into memory
def process_large_file(filename: str) -> list:
with open(filename) as f:
data = f.readlines() # Loads entire file into memory!
return [process(line) for line in data]
# Solution: Stream processing
def process_large_file_streaming(filename: str):
with open(filename) as f:
for line in f: # Reads one line at a time
yield process(line)

Symptoms: Slow queries, connection pool exhaustion, high DB CPU

💡 Tip: Click dropdown to switch between languages
database_bottleneck.py
# N+1 query problem
def get_orders_with_items(user_id: str) -> list:
orders = db.query("SELECT * FROM orders WHERE user_id = ?", user_id)
for order in orders:
# This runs a query for EACH order!
order.items = db.query("SELECT * FROM items WHERE order_id = ?", order.id)
return orders
# Solution: Use JOIN or batch query
def get_orders_with_items_optimized(user_id: str) -> list:
return db.query("""
SELECT o.*, i.*
FROM orders o
LEFT JOIN items i ON o.id = i.order_id
WHERE o.user_id = ?
""", user_id)

Symptoms: High network latency, waiting on external services

Network and I/O bottleneck: service calls an external API with 500ms latency, mitigated with caching, async calls, circuit breakers and timeouts
ResourceToolWarning Signs
CPUtop, htop, metrics>80% sustained
Memoryfree, vmstat>90%, frequent GC
Diskiostat, iotopHigh wait times
Networknetstat, ssPacket loss, high latency
💡 Tip: Click dropdown to switch between languages
profiling.py
import cProfile
import pstats
def profile_function(func):
"""Decorator to profile a function"""
def wrapper(*args, **kwargs):
profiler = cProfile.Profile()
profiler.enable()
result = func(*args, **kwargs)
profiler.disable()
stats = pstats.Stats(profiler)
stats.sort_stats('cumulative')
stats.print_stats(10) # Top 10 slowest
return result
return wrapper
@profile_function
def my_slow_function():
# Your code here
pass
End-to-end request trace: a 500ms request breaks down into auth, validation, a 400ms DB query and serialization, making the DB query the bottleneck
Bottleneck TypeSolutions
CPUOptimize algorithms, caching, horizontal scaling
MemoryStreaming, pagination, efficient data structures
DatabaseIndexing, query optimization, caching, read replicas
NetworkCaching, compression, connection pooling
External APIsCaching, async calls, circuit breakers, timeouts

Advanced: Latency vs Throughput Bottlenecks

Section titled “Advanced: Latency vs Throughput Bottlenecks”

Understanding the difference is crucial for senior engineers:

Latency bottleneck from slow queries versus throughput bottleneck from a 100-connection limit, and how having both causes a performance crisis
TypeSymptomDiagnosisSolution
LatencyHigh response timesProfile shows slow operationsOptimize the slow code
ThroughputRequests queue upResources saturatedAdd capacity or optimize resource usage
BothSlow AND queueingEverything is redTriage: fix biggest impact first

Deep Dive: Production Bottleneck Investigation

Section titled “Deep Dive: Production Bottleneck Investigation”

Here’s how senior engineers approach bottleneck investigation in production:

Before optimizing, you need to know what “normal” looks like. Track these key metrics:

CategoryMetrics to Track
Request LatencyP50, P95, P99 response times
DatabaseQuery times by operation, connection pool usage
External APIsCall durations by service, error rates
ResourcesCPU, memory, disk I/O, network

Tools:

  • APM (Application Performance Monitoring): Datadog, New Relic, Dynatrace
  • Metrics: Prometheus + Grafana
  • Distributed Tracing: Jaeger, Zipkin, AWS X-Ray

The hot path is the code that runs most frequently or consumes most resources:

Hot path versus cold path: about 5 percent of code (handlers, DB queries, serialization) takes 80 percent of execution time, so focus there first

Example 1: Facebook News Feed Database Bottleneck

Section titled “Example 1: Facebook News Feed Database Bottleneck”

Company: Meta (Facebook)

Scenario: Facebook News Feed was loading slowly for users. Investigation revealed a database bottleneck causing high latency.

Implementation: Identified and fixed N+1 query problem:

Facebook News Feed N+1 fix: 81 separate queries taking 2.4 seconds replaced by one batched JOIN query, cutting latency to 50ms

Why This Matters:

  • Scale: Billions of users, trillions of posts
  • Impact: 48x latency reduction
  • User Experience: Faster feed loading increases engagement
  • Result: Reduced database load by 98%

Real-World Impact:

  • Queries: Reduced from 81 to 1 query per feed load
  • Latency: 2.4 seconds → 50ms (48x improvement)
  • Database Load: 98% reduction in queries

Example 2: Twitter Timeline CPU Bottleneck

Section titled “Example 2: Twitter Timeline CPU Bottleneck”

Company: Twitter (now X)

Scenario: Timeline generation was slow during peak hours. CPU usage was maxed out, causing high latency.

Implementation: Identified CPU bottleneck in ranking algorithm:

Twitter timeline CPU bottleneck: O(n squared) ranking at 100 percent CPU optimized to O(n log n) with caching, dropping latency from 2-5s to 100-200ms

Why This Matters:

  • Scale: Millions of timeline requests per second
  • Impact: 20-50x latency reduction
  • Resource Usage: CPU reduced from 100% to 30%
  • Result: Handles 10x more traffic with same hardware

Real-World Impact:

  • Latency: 2-5 seconds → 100-200ms (20-50x improvement)
  • CPU Usage: 100% → 30% (70% reduction)
  • Capacity: 10x more traffic handled

Example 3: Netflix Streaming Memory Bottleneck

Section titled “Example 3: Netflix Streaming Memory Bottleneck”

Company: Netflix

Scenario: Video encoding service was running out of memory when processing large video files. OOM errors caused encoding failures.

Implementation: Fixed memory bottleneck with streaming:

Netflix encoding memory bottleneck: loading a 50GB video into RAM caused OOM errors, fixed by streaming 100MB chunks using about 2GB of memory

Why This Matters:

  • Scale: Thousands of videos encoded daily
  • Impact: Eliminated OOM errors, 2x throughput increase
  • Reliability: Encoding no longer fails due to memory
  • Result: Can process larger videos with same hardware

Real-World Impact:

  • Memory Usage: 100% → 3% (97% reduction)
  • Reliability: OOM errors eliminated
  • Throughput: 2x increase in encoding speed

Example 4: Amazon Checkout Network Bottleneck

Section titled “Example 4: Amazon Checkout Network Bottleneck”

Company: Amazon

Scenario: Checkout page was slow during Prime Day. Investigation revealed external payment API was the bottleneck.

Implementation: Fixed network bottleneck with caching and async processing:

Amazon checkout network bottleneck: a 500ms external payment API fixed with cached payment methods, async processing and a circuit breaker, to 50ms

Why This Matters:

  • Scale: 100K checkout requests per second during Prime Day
  • Impact: 10x latency reduction, 15% conversion increase
  • User Experience: Faster checkout increases conversions
  • Result: Handles traffic spikes without degradation

Real-World Impact:

  • Latency: 500ms → 50ms (10x improvement)
  • Conversion Rate: +15% increase
  • Revenue: Millions in additional revenue during Prime Day

Real-World Case Study: E-Commerce Checkout Bottleneck

Section titled “Real-World Case Study: E-Commerce Checkout Bottleneck”

Situation: Checkout page takes 8 seconds to load during sales events.

Step 1: Add timing instrumentation to each component

Checkout timing breakdown during a sale: a 6000ms inventory check is 75 percent of total time, traced to an N+1 query problem

Step 2: Diagnose the root cause

The inventory service was making one database query per item:

  • Cart with 100 items = 100 database queries
  • Each query ~60ms = 6 seconds total

Step 3: Fix with batch query

-- BEFORE: N+1 queries (100 queries for 100 items)
SELECT * FROM inventory WHERE sku = 'SKU001';
SELECT * FROM inventory WHERE sku = 'SKU002';
-- ... 98 more queries
-- AFTER: Single batch query
SELECT * FROM inventory WHERE sku IN ('SKU001', 'SKU002', ...);
MetricBeforeAfterImprovement
P50 Latency6.2s0.4s93% reduction
P99 Latency12s0.8s93% reduction
DB Queries103595% reduction
Conversion Rate2.1%3.8%81% increase


You’ve completed the Foundations section! You now understand:

  • Why system design matters for LLD
  • Scalability fundamentals
  • Latency and throughput metrics
  • How to find and fix bottlenecks

Continue your journey: Explore other HLD Concepts sections to deepen your understanding of distributed systems.