The Art of Estimation
Good estimates are not about being exact — they are about being useful. An estimate of “somewhere between 1,000 and 100,000 QPS” is nearly useless because those two numbers imply completely different architectures. But “roughly 10,000 QPS with peaks around 30,000” is actionable because it tells you: a single database can handle this, but you need a cache layer and probably 20-30 application servers. The art is in:- Showing structured thinking — break complex estimates into simple multiplications, then combine them. Interviewers care more about your process than your final number.
- Understanding orders of magnitude — the difference between 1 GB and 1 TB is the difference between “fits on a laptop” and “needs a cluster.” Getting within 2-3x of the right answer is success; being off by 10x means your architecture is wrong.
- Identifying bottlenecks early — the estimate often reveals the hardest part of the design. If your storage estimate says 500 PB, storage is your primary constraint, not compute.
- Making reasonable assumptions — and stating them clearly. “I will assume 200 million DAU with 5 requests per user per day” gives the interviewer a chance to course-correct early.
🧠 Numbers You Must Memorize
Time Constants
Latency Numbers (Jeff Dean’s Famous List)
Data Size Reference
Capacity Numbers
Estimation Formulas
The formulas below are your toolkit for converting business requirements (“100 million users, 10 actions per day”) into infrastructure numbers (“10K QPS, 3TB storage, 20 servers”). The key insight: you almost never need all four estimates (QPS, storage, bandwidth, memory). Identify the binding constraint early — the one number that dictates your architecture — and spend your interview time on that. For a read-heavy social feed, QPS is the binding constraint. For a video platform, bandwidth and storage are. For a messaging app, concurrent connections are.Query Per Second (QPS)
Storage Estimation
Bandwidth Estimation
🎯 Step-by-Step Estimation Framework
Step 1: Identify Key Metrics
Always estimate reads and writes separately — they have different scaling characteristics and often different bottlenecks. A read-heavy system (100:1 read/write) calls for aggressive caching and read replicas. A write-heavy system (e.g., logging, analytics) calls for append-optimized storage (LSM trees) and message queues for buffering.Step 2: Make Assumptions Clear
Step 3: Round Smartly
🔥 Practice Problems
Problem 1: Design a URL Shortener
Work through the estimation
Work through the estimation
Given assumptions:Read QPS:Storage (5 years):Short URL Length:
- 100M new URLs per month
- 10:1 read-to-write ratio
- 5 years of data retention
- Average URL: 500 bytes
Problem 2: Design Twitter
Work through the estimation
Work through the estimation
Given assumptions:Timeline Read QPS:Storage (5 years):Timeline Fanout:
- 500M total users
- 200M DAU
- Average user follows 200 people
- 10% of users tweet daily
- Average tweet: 500 bytes
- Read:Write ratio: 50:1
Problem 3: Design YouTube
Work through the estimation
Work through the estimation
Given assumptions:Video Upload QPS:Storage (per year):Bandwidth:
- 2B monthly active users
- 500M DAU watching videos
- Average video watch: 5 min, 5 videos/day
- 500K creators uploading daily
- Average video: 10 min, 100 MB (compressed)
Problem 4: Design WhatsApp
Work through the estimation
Work through the estimation
Given assumptions:Storage (5 years):Concurrent Connections:
- 2B users total
- 500M DAU
- Average: 50 messages/day/user
- Average message: 200 bytes
- 5% with images (100 KB avg)
Interview Tips
The goal of estimation in an interview is not to arrive at a number — it is to arrive at an architectural insight. Every estimate should end with a sentence that connects the number to a design decision.Techniques That Impress
Common Mistakes to Avoid
📊 Quick Reference Tables
Scale Levels
Common Patterns by QPS
Storage Growth Rates
Estimation Interview Questions
These are the exact style of back-of-envelope estimation questions you will face in system design interviews. Each includes a full worked solution showing the thought process, the math, and the insights the interviewer expects you to extract from the numbers.Question 1: “Estimate the storage requirements for Netflix’s entire video library.”
Difficulty: Foundational What the interviewer is really testing: Can you break a large unknown into smaller estimable pieces? Do you understand video encoding and multi-resolution storage?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Let me break this down step by step.Step 1 — How many titles does Netflix have?
Netflix has approximately 15,000 titles (movies and TV series). A TV series averages 3 seasons at 10 episodes each = 30 episodes. Let’s say 5,000 movies and 10,000 series. Total episodes/movies:
- 5,000 movies
- 10,000 series x 30 episodes = 300,000 episodes
- Total content items: ~305,000
- Movies: ~100 minutes
- Episodes: ~40 minutes
- Total hours: (5,000 x 100 + 300,000 x 40) / 60 = (500K + 12M) / 60 = ~208,000 hours
- 240p: ~0.3 GB/hour
- 480p: ~0.7 GB/hour
- 720p: ~1.5 GB/hour
- 1080p: ~3 GB/hour
- 4K HDR: ~7 GB/hour
- “Netflix adds 1,000 new titles per month. How does this change their annual storage growth rate?”
- “Netflix pre-positions popular content on ISP servers (Open Connect). How much storage does a single ISP appliance need?”
- “If Netflix wanted to support 8K streaming, how would that change your storage estimate?”
- “What is the cost difference between storing this in S3 vs in a CDN edge cache?”
Question 2: “How many servers does Google need to handle Search?”
Difficulty: Intermediate What the interviewer is really testing: Can you chain multiple estimates together (users -> queries -> QPS -> servers)? Do you account for the difference between peak and average?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Let me work backwards from the query volume.Step 1 — Query volume:
Google processes approximately 8.5 billion searches per day (public data).Step 2 — QPS:
QPS = 8.5B / 86,400 ≈ 8.5B / 100K = 85,000 QPS average
Peak QPS (3x): ~255,000 QPSStep 3 — How much work is one search query?
This is the critical estimation. A single Google search query:
- Hits the index (likely distributed across 1,000+ shards)
- Each shard scans its portion of the index and returns top results
- Results are merged, ranked, and personalized
- Total compute: roughly equivalent to scanning 1 million documents in ~200ms
- Index storage: Google’s index is estimated at 100+ PB. At 10 TB per server (SSD), that is 10,000 servers just for index storage.
- Crawling and indexing: Continuous pipeline processing billions of pages.
- Ads serving: A separate system with its own servers.
- Caching layer: Frequently searched queries (head queries like ‘weather’, ‘facebook’) are cached. This reduces the 255K QPS hitting the full search pipeline to perhaps 100K QPS.
- Query serving: ~20,000-50,000 servers
- Index storage: ~10,000-20,000 servers
- Crawling/indexing pipeline: ~5,000-10,000 servers
- Ads, caching, infrastructure: ~10,000-20,000 servers
- Total for Search alone: ~50,000-100,000 servers
- “What is the cost of running these servers? Estimate the annual infrastructure cost for Google Search.”
- “Google serves results in 200ms. Break down the latency budget — where does each millisecond go?”
- “How would the server count change if Google moved to a fully GPU-based search ranking pipeline?”
- “How does Google handle a 10x traffic spike during a major world event?”
Question 3: “Estimate the bandwidth required for Spotify’s streaming service.”
Difficulty: Intermediate What the interviewer is really testing: Do you understand audio streaming bitrates? Can you distinguish between concurrent users and DAU? Do you think about CDN and edge caching?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Let me build this up from user behavior.Step 1 — User base and concurrency:
- Spotify has ~600M total users, ~220M paying subscribers
- DAU: roughly 200M users
- Average listening time: 30 minutes/day across all users (some listen for hours, many just a few minutes)
- Concurrent listeners at peak: ~10% of DAU = 20M concurrent streams
- Average concurrency: ~5% of DAU = 10M concurrent streams
- Free tier (normal quality): 160 kbps
- Premium (high quality): 320 kbps
- Let’s assume weighted average of 200 kbps (mix of free and premium, mobile and desktop)
- Origin bandwidth: ~200-400 Gbps
- CDN edge bandwidth: ~3.6 Tbps
- “Spotify is launching in 20 new countries in Africa and Southeast Asia. How does this change your bandwidth estimate and infrastructure requirements?”
- “What happens to your bandwidth when Spotify releases a new Taylor Swift album and 50M users try to stream it simultaneously?”
- “Compare the infrastructure cost of streaming audio vs streaming video. Why is Spotify profitable while most video streaming services are not?”
- “How would Spotify reduce bandwidth costs by 30% without degrading user experience?”
Question 4: “How much storage does Uber need for trip data over 5 years?”
Difficulty: Foundational / Intermediate What the interviewer is really testing: Can you estimate data per event, multiply by event frequency, and account for different data types (structured metadata vs GPS traces)?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Uber’s trip data has multiple components, each with different storage characteristics.Step 1 — Trip volume:
- Uber reports ~25 million trips per day globally
- Per year: 25M x 365 = ~9 billion trips/year
- Over 5 years: ~45 billion trips
- Trip ID, rider ID, driver ID: 3 x 8 bytes = 24 bytes
- Pickup/dropoff coordinates: 4 x 8 bytes = 32 bytes
- Start/end timestamps: 2 x 8 bytes = 16 bytes
- Distance, duration, fare, surge multiplier: 4 x 8 bytes = 32 bytes
- Payment details, promo codes, ratings: ~100 bytes
- Status history, cancellation info: ~50 bytes
- Total per trip record: ~250 bytes, round to 500 bytes with indexes and overhead
- Average trip duration: 15 minutes = 900 seconds
- GPS points per trip: 900 / 4 = 225 points
- Each point: latitude (8 bytes) + longitude (8 bytes) + timestamp (8 bytes) + speed (4 bytes) + heading (4 bytes) = 32 bytes
- GPS data per trip: 225 x 32 bytes = 7,200 bytes ≈ 7 KB
- 5M active drivers, ~10 hours/day active
- Updates: 5M x (10 x 3600 / 4) = 45 billion location pings/day
- Each ping: 32 bytes
- Daily: 45B x 32 bytes = 1.44 TB/day
- Yearly: 1.44 TB x 365 = 526 TB
- 5 years with replication: ~8 PB
- Trip records: ~70 TB
- GPS traces: ~1 PB
- Driver location pings: ~8 PB
- Total: ~9-10 PB over 5 years
- “Uber wants to retain GPS traces for only 90 days to save storage but keep trip records forever. How does this change the architecture?”
- “How would you estimate the QPS for Uber’s location service, given 5M drivers sending updates every 4 seconds?”
- “Uber Eats adds 10M food delivery orders per day. Each order has additional data (restaurant, menu items, delivery instructions). How does this change your estimate?”
- “At 10 PB, what is Uber’s approximate annual storage cost? Compare cloud vs on-premise.”
Question 5: “Estimate the number of chat messages WhatsApp processes per second.”
Difficulty: Foundational What the interviewer is really testing: Basic QPS estimation. Can you make reasonable assumptions about user behavior and convert daily volume to per-second rate?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Let me start from users and work toward QPS.Step 1 — User base:
- WhatsApp has ~2.5 billion monthly active users
- DAU: roughly 70% of MAU = 1.75 billion DAU (WhatsApp has unusually high daily engagement)
- This varies wildly by market. In India and Brazil, power users send 100+ messages/day. In the US, maybe 20-30.
- Weighted global average: ~50 messages/day per active user
- But messages are bidirectional — for every message sent, someone receives it. The system processes both.
- Messages processed (sent): 1.75B x 50 = 87.5 billion messages/day
- Text messages (~80%): 700K/sec
- Image/media messages (~15%): 131K/sec
- Voice messages (~5%): 44K/sec
- Connections: 1.75B DAU with ~30% concurrent at peak = 525M WebSocket connections
- Bandwidth for text: 875K/sec x 200 bytes = 175 MB/sec = 1.4 Gbps (trivial)
- Bandwidth for media: 131K/sec x 200 KB average = 26 GB/sec = 208 Gbps (significant!)
- Status updates (online/offline, typing, read receipts): probably 5-10x the message rate = 5-9 million events/sec
- “WhatsApp must deliver messages in order within a conversation. How does this constraint affect their server architecture at 875K messages/sec?”
- “During New Year’s Eve, messaging volume spikes 10-20x. How does WhatsApp prepare for this?”
- “WhatsApp stores messages on the server only until delivered. Estimate the peak undelivered message storage when 30% of users are offline.”
- “Compare WhatsApp’s infrastructure requirements to Slack’s. Why can WhatsApp serve 10x more users with a fraction of the engineering team?”
Question 6: “Estimate the QPS and storage for a system like Instagram.”
Difficulty: Intermediate What the interviewer is really testing: Can you separately estimate reads and writes, account for media storage vs metadata, and identify the read:write ratio as a key architectural signal?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Instagram is interesting because it has asymmetric read/write patterns and media-heavy storage. Let me break it down.Step 1 — User activity:
- 2B monthly active users, ~500M DAU
- Content creation: ~5% of DAU posts daily = 25M posts/day
- Feed consumption: average user views feed 7 times/day, 20 posts per view = 140 posts viewed/day
- Likes: average user likes 10 posts/day
- Comments: average user comments on 1 post/day
- Photo/video uploads: 25M / 86,400 ≈ 290/sec (peak: 870/sec)
- Likes: 500M x 10 / 86,400 ≈ 58,000/sec (peak: 174K/sec)
- Comments: 500M x 1 / 86,400 ≈ 5,800/sec (peak: 17K/sec)
- Total write QPS: ~64,000/sec average, ~192K/sec peak
- Feed views: 500M x 7 / 86,400 ≈ 40,000 feed requests/sec
- But each feed request loads 20 posts, each with images: 40K x 20 = 800K image loads/sec
- Profile views, search, explore page: ~20,000 requests/sec
- Total read QPS: ~860,000/sec peak
- Read:Write ratio: roughly 13:1 — read-heavy, which means caching is critical
- 25M photos/day, average 2 MB after compression
- Instagram stores 4-5 resolution variants per photo: thumbnail (10 KB), small (100 KB), medium (500 KB), large (1 MB), original (2 MB) = ~3.6 MB total per photo
- Daily: 25M x 3.6 MB = 90 TB/day
- Yearly: 90 TB x 365 = 32.8 PB/year
- 5 years with 3x replication: ~500 PB
- Perhaps 10% of posts are Reels, average 30 seconds at ~5 MB after compression, with 4 quality variants: ~15 MB total
- 2.5M Reels/day x 15 MB = 37.5 TB/day
- Yearly: 13.7 PB
- Post metadata: 25M x 1 KB = 25 GB/day
- User profiles, follows, likes: ~50 GB/day
- Yearly: ~27 TB
- “Instagram Explore page needs to serve personalized content to 500M users. Estimate the recommendation system’s QPS.”
- “What percentage of Instagram’s total stored photos are never viewed after the first week? How would you use this to optimize storage costs?”
- “Instagram Stories disappear after 24 hours. Estimate the daily storage churn from Stories creation and deletion.”
- “Compare Instagram’s infrastructure profile to YouTube’s. Which is more compute-intensive vs storage-intensive, and why?”
Question 7: “How many machines would you need to serve a real-time global leaderboard for a game with 100 million players?”
Difficulty: Senior What the interviewer is really testing: Can you reason about data structure operations per second, memory sizing, and the trade-off between exactness and scalability? Do you distinguish between writes (score updates) and reads (rank lookups)?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“Let me model both the data size and the operation rate.Step 1 — Data size:
- 100M players, each with: player_id (8 bytes) + score (8 bytes) + metadata (username, avatar_url) = ~50 bytes
- Total data: 100M x 50 bytes = 5 GB
- This fits comfortably in a single Redis instance’s memory (modern servers have 128-512 GB RAM)
- Not all 100M players are active simultaneously. Maybe 10M concurrent.
- Active players score every ~30 seconds (game action completes)
- Write QPS: 10M / 30 = 333,000 score updates/sec
- Peak (tournament event, 2x): 666,000 writes/sec
- Every player checks their rank after each game action, plus periodic checks
- Read QPS: ~500,000 rank lookups/sec peak
- Redis sorted set operations (ZADD for write, ZREVRANK for rank): each is O(log N) = O(log 100M) = ~27 operations
- A single Redis instance can handle ~100K-200K sorted set operations/sec
- We need ~1.2M ops/sec (666K writes + 500K reads)
- Single instance: No. Need to partition.
- Divide the score range into buckets. Partition 1 handles scores 0-999, Partition 2 handles 1000-1999, etc.
- Problem: hot partitions. If most players have scores in the 500-2000 range, those partitions are overloaded.
- Advantage: global rank can be computed by summing player counts in higher partitions + rank within partition.
- Hash player_id to one of N partitions (N = 10).
- Each partition maintains its own sorted set of ~10M players.
- Write: O(1) routing, O(log 10M) per partition = fast.
- Global rank: Cannot be determined from a single partition. Need a separate aggregation service.
- Aggregation: maintain a coarse-grained global histogram (score buckets of size 100). Each partition publishes its bucket counts. Global rank = sum of global counts above your score + rank within your bucket.
- 10 Redis instances, each handling ~120K ops/sec (well within limits)
- 1 aggregation service updating the global histogram every 5-10 seconds
- Approximate global rank (accurate to within ~100 positions for mid-ranked players, exact for top 1000)
- Top-1000 leaderboard maintained separately with exact rankings
- 10 Redis instances (32 GB RAM each, 8 cores): 10 servers
- 10 Redis replicas for read scaling and failover: 10 servers
- 2-3 aggregation service instances: 3 servers
- 2-3 API gateway servers: 3 servers
- Total: ~26 servers
ORDER BY queries.
Follow-ups:
- “The game runs a 24-hour tournament with 50M concurrent players and score updates every 5 seconds. Re-estimate the infrastructure.”
- “How do you display ‘You are ranked #4,847,329 out of 100M players’ when your rank is only approximate? What is the user experience trade-off?”
- “The game wants friend-group leaderboards — your rank among your 200 friends. How do you efficiently compute this without N separate sorted sets?”
- “During a server failure, some score updates are lost. How do you reconcile scores after recovery?”
Question 8: “Estimate the cost of sending 1 billion push notifications per day.”
Difficulty: Senior What the interviewer is really testing: Do you understand the full cost stack (infrastructure + third-party services)? Can you identify that the dominant cost is not always what you expect? Do you think about failure rates and retries as cost multipliers?Strong Answer (Worked Solution)
Strong Answer (Worked Solution)
“This is a cost estimation, not just a capacity estimation. I need to think about every layer.Step 1 — Volume breakdown:
1 billion push notifications/day:
- QPS average: 1B / 86,400 ≈ 11,600/sec
- Peak QPS (3x): ~35,000/sec
- Platform split (approximate): 70% Android (FCM), 25% iOS (APNs), 5% Web Push
- FCM: 700M/day, APNs: 250M/day, Web: 50M/day
- FCM (Firebase Cloud Messaging): Free for unlimited push notifications. Google does not charge per message.
- APNs (Apple Push Notification service): Free. Apple does not charge per message.
- Web Push: Free (standards-based).
- At 35K/sec peak, assuming each notification takes 5ms of server time (template rendering, user preference check, dedup, API call to FCM/APNs):
- Concurrent work: 35,000 x 0.005 = 175 concurrent tasks
- With 8-core servers handling 50 concurrent tasks each: 4 servers
- With 3x headroom for retries and bursts: 12 servers
- Cost: 12 x
c5.2xlargeat ~3,000/month**
- 1B messages/day at 500 bytes each = 500 GB/day throughput
- 3-day retention: 1.5 TB storage
- 6-node Kafka cluster with replication
- Cost: 6 x
i3.xlargeat ~1,800/month**
- Need to check each user’s notification preferences, device tokens, opt-out status
- 500M users, ~200 bytes per user = 100 GB
- Redis cluster or DynamoDB
- Cost: DynamoDB at 35K reads/sec = ~2,000-5,000/month**
- 1B notifications x 500 bytes = 500 GB/day egress to FCM/APNs
- AWS egress: 45/day = $1,350/month
- Compute: $3,000
- Kafka: $1,800
- User preference store: $3,500
- Network egress: $1,350
- Analytics pipeline: $2,000
- Monitoring and infrastructure overhead: $1,000
- Total: ~150,000-180,000/year
- “If you switch from FCM free tier to a paid notification service like OneSignal or Airship, how does the cost change?”
- “Your notification system has a 5% delivery failure rate. The product team says it should be under 1%. What do you do, and what does it cost?”
- “Compare the cost of push notifications vs SMS notifications vs email. At what volume does each channel make economic sense?”
- “A bug sends 10x the intended notifications over 2 hours. What is the cost of this incident beyond the extra infrastructure — in terms of user opt-outs and app uninstalls?”