Skip to content
LowLevelDesign Mastery

Latency and Throughput

Measuring what matters in system performance

When measuring system performance, two metrics matter most:

Latency versus throughput: latency is how fast one request completes in ms, throughput is how many requests per second, and systems need both

Latency is the time between sending a request and receiving a response.

Every request goes through multiple stages:

Request latency breakdown: client to network, server processing, database and back over the network, totalling roughly 26ms to 360ms
TypeDescriptionTypical Values
Network LatencyTime for data to travel over network1-100ms
Processing LatencyTime for server to process request1-50ms
Database LatencyTime for database queries1-10ms
Queue LatencyTime spent waiting in queues0-1000ms+

Average latency can be misleading. Percentiles give a better picture:

Percentile latencies for 1000 requests: P50 of 50ms as typical, P95 of 200ms for most users and P99 of 500ms as the worst case for most

Real-World Example: E-Commerce Checkout Latency

Section titled “Real-World Example: E-Commerce Checkout Latency”

Company: Amazon, eBay, Shopify

Scenario: Checkout pages must load quickly to prevent cart abandonment. Even small latency increases can significantly impact conversion rates.

Implementation: Uses parallel API calls and caching:

Average latency is misleading: a 100ms average hides a real distribution of P50 50ms, P95 150ms and P99 2000ms where 1 percent wait 2 seconds

Throughput is the amount of work done per unit of time.

MetricDescriptionExample
RPSRequests per second10,000 RPS
TPSTransactions per second5,000 TPS
QPSQueries per second50,000 QPS
BandwidthData transferred per second1 Gbps

Throughput is calculated as:

Throughput = Total Requests / Time Period

Example: If your server handles 10,000 requests in 10 seconds, your throughput is 1,000 RPS.

Key considerations:

  • Measure over time - instantaneous measurements fluctuate
  • Track success rate - failed requests count against effective throughput
  • Monitor under load - throughput often drops when the system is stressed

They’re related but independent:

Latency and throughput scenarios: API server with low latency and high throughput, batch job with high latency and high throughput, overloaded system

A fundamental relationship:

Average Concurrent Requests = Throughput × Average Latency

Example:

  • Throughput: 1,000 RPS
  • Average Latency: 100ms = 0.1 seconds
  • Concurrent Requests: 1,000 × 0.1 = 100 requests in flight

Your code decisions directly affect latency and throughput:

💡 Tip: Click dropdown to switch between languages
sequential_bad.py
class OrderService:
"""Sequential calls - high latency"""
def get_order_details(self, order_id: str) -> OrderDetails:
# Each call waits for the previous one
order = self.db.get_order(order_id) # 10ms
user = self.user_service.get(order.user_id) # 20ms
items = self.inventory.get_items(order.items) # 15ms
shipping = self.shipping.get_status(order_id) # 25ms
# Total: 10 + 20 + 15 + 25 = 70ms
return OrderDetails(order, user, items, shipping)
💡 Tip: Click dropdown to switch between languages
parallel_good.py
import asyncio
class OrderService:
"""Parallel calls - lower latency"""
async def get_order_details(self, order_id: str) -> OrderDetails:
# First, get the order (we need it for user_id)
order = await self.db.get_order(order_id) # 10ms
# Then fetch everything else in parallel
user, items, shipping = await asyncio.gather(
self.user_service.get(order.user_id), # 20ms
self.inventory.get_items(order.items), # 15ms } All run
self.shipping.get_status(order_id) # 25ms } in parallel
)
# Total: 10 + max(20, 15, 25) = 10 + 25 = 35ms
# Saved 35ms (50% reduction!)
return OrderDetails(order, user, items, shipping)
Sequential versus parallel calls: DB then user, items and shipping in sequence takes 70ms, running the last three in parallel cuts it to 35ms

Caching is the most effective way to reduce latency. The idea is simple: store frequently accessed data closer to where it’s needed.

Multi-level cache: request checks an L1 in-memory cache at 0.01ms, then L2 Redis at 1-2ms, and only about 1 percent reach the 10-50ms database

With this caching strategy:

LevelLatencyHit Rate
L1 (Local)0.01ms90%
L2 (Redis)2ms9%
Database30ms1%

Average latency = 0.9(0.01) + 0.09(2) + 0.01(30) = 0.49ms

That’s a 60x improvement over hitting the database every time!


MetricWhat It Tells YouAction Threshold
P50Typical user experienceBaseline for “normal”
P95Most users’ worst caseWatch for drift
P99Outlier experienceInvestigate if > 3x P50
Error RateSystem healthAlert if > 1%

In Production:

  • APM Tools: Datadog, New Relic, Dynatrace
  • Metrics: Prometheus + Grafana
  • Distributed Tracing: Jaeger, Zipkin

In Development:

  • Profilers: cProfile (Python), JProfiler (Java)
  • Benchmarking: pytest-benchmark, JMH

Example 1: Google Search Latency Optimization

Section titled “Example 1: Google Search Latency Optimization”

Company: Google

Scenario: Google Search must return results in milliseconds. Even 100ms delay can reduce user satisfaction and search volume.

Implementation: Uses parallel processing and caching:

Google Search latency: a query fans out in parallel to web, image, news and maps search, backed by a multi-level cache to meet a sub-100ms target

Why This Matters:

  • Scale: Billions of queries per day
  • Latency Impact: 100ms delay = 0.2% search volume reduction
  • Performance: Parallel processing reduces latency by 60%
  • Result: Sub-100ms average response time

Real-World Impact:

  • Queries: 8.5+ billion searches per day
  • Latency: Average 50ms, P99 200ms
  • Revenue Impact: Every 100ms delay costs millions in ad revenue

Company: Amazon

Scenario: Product pages must load quickly. Amazon found that every 100ms delay reduces sales by 1%.

Implementation: Uses parallel API calls and edge caching:

Amazon product page latency: parallel API calls for details, reviews, recommendations, inventory and pricing plus CDN-cached static content in 200ms

Why Parallel Processing?

  • User Experience: Fast page loads increase conversions
  • Revenue Impact: 100ms delay = 1% sales reduction
  • Performance: Parallel calls reduce latency by 70%
  • Result: 200ms average page load time

Real-World Impact:

  • Scale: Billions of product page views daily
  • Latency: 200ms average, P99 500ms
  • Revenue: Every 100ms optimization worth millions annually

Example 3: Netflix Video Streaming Latency

Section titled “Example 3: Netflix Video Streaming Latency”

Company: Netflix

Scenario: Video playback must start quickly. Users expect playback to begin within 2 seconds of clicking play.

Implementation: Uses CDN distribution and adaptive bitrate streaming:

Netflix playback latency: play request served by a nearby CDN edge with adaptive streaming starting at low quality, first frame under 2s

Why Low Latency Matters:

  • User Experience: Fast playback start increases engagement
  • Retention: Slow start causes users to abandon
  • Performance: CDN reduces latency by 90%
  • Result: < 2 seconds time to first frame

Real-World Impact:

  • Scale: 200+ million subscribers, billions of plays daily
  • Latency: < 2 seconds time to first frame
  • Engagement: Fast playback increases watch time by 20%


Now that you understand performance metrics, let’s learn how to identify and fix performance issues:

Next up: Understanding Bottlenecks - Learn to find and eliminate performance bottlenecks.