Skip to main content
Interview Critical: Every system design interview starts with capacity estimation. Master these calculations to ace the first 5 minutes and build credibility.

The Art of Estimation

Good estimates are not about being exact — they are about being useful. An estimate of “somewhere between 1,000 and 100,000 QPS” is nearly useless because those two numbers imply completely different architectures. But “roughly 10,000 QPS with peaks around 30,000” is actionable because it tells you: a single database can handle this, but you need a cache layer and probably 20-30 application servers. The art is in:
  1. Showing structured thinking — break complex estimates into simple multiplications, then combine them. Interviewers care more about your process than your final number.
  2. Understanding orders of magnitude — the difference between 1 GB and 1 TB is the difference between “fits on a laptop” and “needs a cluster.” Getting within 2-3x of the right answer is success; being off by 10x means your architecture is wrong.
  3. Identifying bottlenecks early — the estimate often reveals the hardest part of the design. If your storage estimate says 500 PB, storage is your primary constraint, not compute.
  4. Making reasonable assumptions — and stating them clearly. “I will assume 200 million DAU with 5 requests per user per day” gives the interviewer a chance to course-correct early.

🧠 Numbers You Must Memorize

Time Constants

Latency Numbers (Jeff Dean’s Famous List)

Data Size Reference

Capacity Numbers

Estimation Formulas

The formulas below are your toolkit for converting business requirements (“100 million users, 10 actions per day”) into infrastructure numbers (“10K QPS, 3TB storage, 20 servers”). The key insight: you almost never need all four estimates (QPS, storage, bandwidth, memory). Identify the binding constraint early — the one number that dictates your architecture — and spend your interview time on that. For a read-heavy social feed, QPS is the binding constraint. For a video platform, bandwidth and storage are. For a messaging app, concurrent connections are.

Query Per Second (QPS)

Storage Estimation

Bandwidth Estimation

🎯 Step-by-Step Estimation Framework

Step 1: Identify Key Metrics

Always estimate reads and writes separately — they have different scaling characteristics and often different bottlenecks. A read-heavy system (100:1 read/write) calls for aggressive caching and read replicas. A write-heavy system (e.g., logging, analytics) calls for append-optimized storage (LSM trees) and message queues for buffering.

Step 2: Make Assumptions Clear

Step 3: Round Smartly

🔥 Practice Problems

Problem 1: Design a URL Shortener

Given assumptions:
  • 100M new URLs per month
  • 10:1 read-to-write ratio
  • 5 years of data retention
  • Average URL: 500 bytes
Write QPS:
Read QPS:
Storage (5 years):
Short URL Length:

Problem 2: Design Twitter

Given assumptions:
  • 500M total users
  • 200M DAU
  • Average user follows 200 people
  • 10% of users tweet daily
  • Average tweet: 500 bytes
  • Read:Write ratio: 50:1
Tweet Write QPS:
Timeline Read QPS:
Storage (5 years):
Timeline Fanout:

Problem 3: Design YouTube

Given assumptions:
  • 2B monthly active users
  • 500M DAU watching videos
  • Average video watch: 5 min, 5 videos/day
  • 500K creators uploading daily
  • Average video: 10 min, 100 MB (compressed)
Video Watch QPS:
Video Upload QPS:
Storage (per year):
Bandwidth:

Problem 4: Design WhatsApp

Given assumptions:
  • 2B users total
  • 500M DAU
  • Average: 50 messages/day/user
  • Average message: 200 bytes
  • 5% with images (100 KB avg)
Message QPS:
Storage (5 years):
Concurrent Connections:

Interview Tips

The goal of estimation in an interview is not to arrive at a number — it is to arrive at an architectural insight. Every estimate should end with a sentence that connects the number to a design decision.
The Power Move: After every estimate, state the architectural implication. “At 25K QPS with a 50:1 read/write ratio, our write path is only 500 QPS — easily handled by a single PostgreSQL primary — but our read path at 25K QPS needs aggressive caching. I would target a 95%+ cache hit rate, which means our database only sees ~1,250 QPS. That is well within what a single primary with two read replicas can handle, so we do not need sharding.” This single paragraph demonstrates more architectural judgment than most candidates show in an entire interview.

Techniques That Impress

Common Mistakes to Avoid

📊 Quick Reference Tables

Scale Levels

Common Patterns by QPS

Storage Growth Rates

Estimation Interview Questions

These are the exact style of back-of-envelope estimation questions you will face in system design interviews. Each includes a full worked solution showing the thought process, the math, and the insights the interviewer expects you to extract from the numbers.

Question 1: “Estimate the storage requirements for Netflix’s entire video library.”

Difficulty: Foundational What the interviewer is really testing: Can you break a large unknown into smaller estimable pieces? Do you understand video encoding and multi-resolution storage?
“Let me break this down step by step.Step 1 — How many titles does Netflix have? Netflix has approximately 15,000 titles (movies and TV series). A TV series averages 3 seasons at 10 episodes each = 30 episodes. Let’s say 5,000 movies and 10,000 series. Total episodes/movies:
  • 5,000 movies
  • 10,000 series x 30 episodes = 300,000 episodes
  • Total content items: ~305,000
Step 2 — Average duration:
  • Movies: ~100 minutes
  • Episodes: ~40 minutes
  • Total hours: (5,000 x 100 + 300,000 x 40) / 60 = (500K + 12M) / 60 = ~208,000 hours
Step 3 — Storage per hour at each resolution: Netflix encodes each title at multiple quality levels:
  • 240p: ~0.3 GB/hour
  • 480p: ~0.7 GB/hour
  • 720p: ~1.5 GB/hour
  • 1080p: ~3 GB/hour
  • 4K HDR: ~7 GB/hour
Average across 5 quality levels: ~2.5 GB/hour But they also encode multiple audio tracks (5-10 languages) at ~100 MB/hour each: ~0.7 GB/hour for 7 languages.Total per hour of content: ~(0.3 + 0.7 + 1.5 + 3 + 7 + 0.7) = ~13 GB/hourStep 4 — Total storage: 208,000 hours x 13 GB/hour = ~2.7 PBRound to ~3 PB for the core library. With 3x replication across regions: ~9 PB. Plus working copies in the transcoding pipeline: ~10-15 PB total.Key insight to state: The video storage is large but bounded. Netflix’s real infrastructure cost is in bandwidth (serving 200M+ subscribers), not storage. At current S3 pricing (~0.023/GB/month),3PBcostsabout0.023/GB/month), 3 PB costs about 70K/month in storage alone — trivial compared to their $1B+ annual bandwidth spend.”
Red flag answer: “Netflix has a lot of videos, so probably like 100 PB.” Throwing out a number without showing how you derived it gives the interviewer nothing to evaluate. The point is the decomposition, not the final number. Follow-ups:
  1. “Netflix adds 1,000 new titles per month. How does this change their annual storage growth rate?”
  2. “Netflix pre-positions popular content on ISP servers (Open Connect). How much storage does a single ISP appliance need?”
  3. “If Netflix wanted to support 8K streaming, how would that change your storage estimate?”
  4. “What is the cost difference between storing this in S3 vs in a CDN edge cache?”

Question 2: “How many servers does Google need to handle Search?”

Difficulty: Intermediate What the interviewer is really testing: Can you chain multiple estimates together (users -> queries -> QPS -> servers)? Do you account for the difference between peak and average?
“Let me work backwards from the query volume.Step 1 — Query volume: Google processes approximately 8.5 billion searches per day (public data).Step 2 — QPS: QPS = 8.5B / 86,400 ≈ 8.5B / 100K = 85,000 QPS average Peak QPS (3x): ~255,000 QPSStep 3 — How much work is one search query? This is the critical estimation. A single Google search query:
  • Hits the index (likely distributed across 1,000+ shards)
  • Each shard scans its portion of the index and returns top results
  • Results are merged, ranked, and personalized
  • Total compute: roughly equivalent to scanning 1 million documents in ~200ms
A single server can handle a complex search query in about 200ms, meaning one server handles ~5 queries/second (factoring in CPU and I/O).Step 4 — Servers for query serving: 255,000 peak QPS / 5 QPS per server = 51,000 query-serving servers.Step 5 — But that is just the serving layer.
  • Index storage: Google’s index is estimated at 100+ PB. At 10 TB per server (SSD), that is 10,000 servers just for index storage.
  • Crawling and indexing: Continuous pipeline processing billions of pages.
  • Ads serving: A separate system with its own servers.
  • Caching layer: Frequently searched queries (head queries like ‘weather’, ‘facebook’) are cached. This reduces the 255K QPS hitting the full search pipeline to perhaps 100K QPS.
With caching: 100K / 5 = 20,000 query-serving servers.Step 6 — Total estimate:
  • Query serving: ~20,000-50,000 servers
  • Index storage: ~10,000-20,000 servers
  • Crawling/indexing pipeline: ~5,000-10,000 servers
  • Ads, caching, infrastructure: ~10,000-20,000 servers
  • Total for Search alone: ~50,000-100,000 servers
With 2x for redundancy: ~100,000-200,000 servers.Sanity check: Google reportedly operates millions of servers globally across all services. Search being 100-200K servers (5-10% of total fleet) feels reasonable given that Search is one of many products.”
Red flag answer: “Google is huge so they probably need millions of servers for search.” This is vague. The interviewer wants to see you decompose the problem into QPS, per-query cost, and server capacity. Follow-ups:
  1. “What is the cost of running these servers? Estimate the annual infrastructure cost for Google Search.”
  2. “Google serves results in 200ms. Break down the latency budget — where does each millisecond go?”
  3. “How would the server count change if Google moved to a fully GPU-based search ranking pipeline?”
  4. “How does Google handle a 10x traffic spike during a major world event?”

Question 3: “Estimate the bandwidth required for Spotify’s streaming service.”

Difficulty: Intermediate What the interviewer is really testing: Do you understand audio streaming bitrates? Can you distinguish between concurrent users and DAU? Do you think about CDN and edge caching?
“Let me build this up from user behavior.Step 1 — User base and concurrency:
  • Spotify has ~600M total users, ~220M paying subscribers
  • DAU: roughly 200M users
  • Average listening time: 30 minutes/day across all users (some listen for hours, many just a few minutes)
  • Concurrent listeners at peak: ~10% of DAU = 20M concurrent streams
  • Average concurrency: ~5% of DAU = 10M concurrent streams
Step 2 — Bitrate per stream:
  • Free tier (normal quality): 160 kbps
  • Premium (high quality): 320 kbps
  • Let’s assume weighted average of 200 kbps (mix of free and premium, mobile and desktop)
Step 3 — Total bandwidth: Peak: 20M concurrent x 200 kbps = 4,000 Gbps = 4 Tbps Average: 10M concurrent x 200 kbps = 2 TbpsStep 4 — CDN impact: Most popular songs (top 10,000 tracks) account for perhaps 30-40% of all streams. These are cached at the edge, served from CDN PoPs near the user. The CDN absorbs the majority of bandwidth, so origin bandwidth is much lower — maybe 10% of total.
  • Origin bandwidth: ~200-400 Gbps
  • CDN edge bandwidth: ~3.6 Tbps
Step 5 — Storage perspective (bonus): Spotify’s catalog: ~100M tracks x average 4 minutes x 3 quality levels Storage per track per quality: 4 min x 60 sec x 200 kbps / 8 = ~6 MB per quality level 100M x 6 MB x 3 = ~1.8 PB for the entire music catalog With podcasts (5M+ episodes at ~50 MB each): +250 TB Total catalog: ~2 PBKey insight: 4 Tbps is significant but manageable with modern CDN infrastructure. For reference, Netflix peaks at 10+ Tbps. Audio is much less bandwidth-intensive than video — the same CDN infrastructure that serves one Netflix stream could serve 25+ Spotify streams.”
Red flag answer: “600 million users times 320 kbps equals 192 Tbps.” This catastrophic error confuses total users with concurrent users. The interviewer will immediately note that you do not understand the difference between registered users, DAU, and concurrent connections. Follow-ups:
  1. “Spotify is launching in 20 new countries in Africa and Southeast Asia. How does this change your bandwidth estimate and infrastructure requirements?”
  2. “What happens to your bandwidth when Spotify releases a new Taylor Swift album and 50M users try to stream it simultaneously?”
  3. “Compare the infrastructure cost of streaming audio vs streaming video. Why is Spotify profitable while most video streaming services are not?”
  4. “How would Spotify reduce bandwidth costs by 30% without degrading user experience?”

Question 4: “How much storage does Uber need for trip data over 5 years?”

Difficulty: Foundational / Intermediate What the interviewer is really testing: Can you estimate data per event, multiply by event frequency, and account for different data types (structured metadata vs GPS traces)?
“Uber’s trip data has multiple components, each with different storage characteristics.Step 1 — Trip volume:
  • Uber reports ~25 million trips per day globally
  • Per year: 25M x 365 = ~9 billion trips/year
  • Over 5 years: ~45 billion trips
Step 2 — Storage per trip (structured data): A trip record includes:
  • Trip ID, rider ID, driver ID: 3 x 8 bytes = 24 bytes
  • Pickup/dropoff coordinates: 4 x 8 bytes = 32 bytes
  • Start/end timestamps: 2 x 8 bytes = 16 bytes
  • Distance, duration, fare, surge multiplier: 4 x 8 bytes = 32 bytes
  • Payment details, promo codes, ratings: ~100 bytes
  • Status history, cancellation info: ~50 bytes
  • Total per trip record: ~250 bytes, round to 500 bytes with indexes and overhead
Structured trip data: 45B trips x 500 bytes = 22.5 TB With indexes and 3x replication: ~70 TBStep 3 — GPS trace data (the big one): During an active trip, the driver’s phone sends location updates every 4 seconds.
  • Average trip duration: 15 minutes = 900 seconds
  • GPS points per trip: 900 / 4 = 225 points
  • Each point: latitude (8 bytes) + longitude (8 bytes) + timestamp (8 bytes) + speed (4 bytes) + heading (4 bytes) = 32 bytes
  • GPS data per trip: 225 x 32 bytes = 7,200 bytes ≈ 7 KB
GPS data: 45B trips x 7 KB = 315 TB With replication (3x): ~950 TB ≈ 1 PBStep 4 — Driver location data (even bigger): Even when NOT on a trip, available drivers send location updates every 4 seconds.
  • 5M active drivers, ~10 hours/day active
  • Updates: 5M x (10 x 3600 / 4) = 45 billion location pings/day
  • Each ping: 32 bytes
  • Daily: 45B x 32 bytes = 1.44 TB/day
  • Yearly: 1.44 TB x 365 = 526 TB
  • 5 years with replication: ~8 PB
Step 5 — Total:
  • Trip records: ~70 TB
  • GPS traces: ~1 PB
  • Driver location pings: ~8 PB
  • Total: ~9-10 PB over 5 years
Key insight: The structured trip data is tiny. The GPS and location tracking data dominates by 100x. This is why Uber invested heavily in their own time-series database (M3) and geospatial infrastructure — generic databases cannot handle 45 billion location pings per day efficiently.”
Red flag answer: “25 million trips per day, each trip is about 1 KB, so 25 GB per day, 45 TB over 5 years.” This dramatically underestimates because it ignores GPS traces and driver location tracking, which are 100x larger than the trip metadata. Follow-ups:
  1. “Uber wants to retain GPS traces for only 90 days to save storage but keep trip records forever. How does this change the architecture?”
  2. “How would you estimate the QPS for Uber’s location service, given 5M drivers sending updates every 4 seconds?”
  3. “Uber Eats adds 10M food delivery orders per day. Each order has additional data (restaurant, menu items, delivery instructions). How does this change your estimate?”
  4. “At 10 PB, what is Uber’s approximate annual storage cost? Compare cloud vs on-premise.”

Question 5: “Estimate the number of chat messages WhatsApp processes per second.”

Difficulty: Foundational What the interviewer is really testing: Basic QPS estimation. Can you make reasonable assumptions about user behavior and convert daily volume to per-second rate?
“Let me start from users and work toward QPS.Step 1 — User base:
  • WhatsApp has ~2.5 billion monthly active users
  • DAU: roughly 70% of MAU = 1.75 billion DAU (WhatsApp has unusually high daily engagement)
Step 2 — Messages per user per day:
  • This varies wildly by market. In India and Brazil, power users send 100+ messages/day. In the US, maybe 20-30.
  • Weighted global average: ~50 messages/day per active user
  • But messages are bidirectional — for every message sent, someone receives it. The system processes both.
  • Messages processed (sent): 1.75B x 50 = 87.5 billion messages/day
Step 3 — Convert to QPS: QPS = 87.5B / 86,400 ≈ 87.5B / 100K = 875,000 messages/sec Peak (3x): ~2.6 million messages/secStep 4 — Breakdown by type:
  • Text messages (~80%): 700K/sec
  • Image/media messages (~15%): 131K/sec
  • Voice messages (~5%): 44K/sec
Step 5 — Related metrics worth computing:
  • Connections: 1.75B DAU with ~30% concurrent at peak = 525M WebSocket connections
  • Bandwidth for text: 875K/sec x 200 bytes = 175 MB/sec = 1.4 Gbps (trivial)
  • Bandwidth for media: 131K/sec x 200 KB average = 26 GB/sec = 208 Gbps (significant!)
  • Status updates (online/offline, typing, read receipts): probably 5-10x the message rate = 5-9 million events/sec
Key insight: The raw message QPS (875K/sec) is high but manageable. The real engineering challenge is the 525M concurrent WebSocket connections. This is why WhatsApp was famous for handling 2 million connections per server using Erlang/BEAM — they needed roughly 250-500 servers just for the connection layer, which is remarkably lean for 2 billion users.”
Red flag answer: “2.5 billion users times 50 messages is 125 billion, divided by 86,400 is about 1.4 million QPS.” This uses total users instead of DAU, and does not distinguish between peak and average. Also missing the insight about connections being the real bottleneck, not message processing. Follow-ups:
  1. “WhatsApp must deliver messages in order within a conversation. How does this constraint affect their server architecture at 875K messages/sec?”
  2. “During New Year’s Eve, messaging volume spikes 10-20x. How does WhatsApp prepare for this?”
  3. “WhatsApp stores messages on the server only until delivered. Estimate the peak undelivered message storage when 30% of users are offline.”
  4. “Compare WhatsApp’s infrastructure requirements to Slack’s. Why can WhatsApp serve 10x more users with a fraction of the engineering team?”

Question 6: “Estimate the QPS and storage for a system like Instagram.”

Difficulty: Intermediate What the interviewer is really testing: Can you separately estimate reads and writes, account for media storage vs metadata, and identify the read:write ratio as a key architectural signal?
“Instagram is interesting because it has asymmetric read/write patterns and media-heavy storage. Let me break it down.Step 1 — User activity:
  • 2B monthly active users, ~500M DAU
  • Content creation: ~5% of DAU posts daily = 25M posts/day
  • Feed consumption: average user views feed 7 times/day, 20 posts per view = 140 posts viewed/day
  • Likes: average user likes 10 posts/day
  • Comments: average user comments on 1 post/day
Step 2 — Write QPS:
  • Photo/video uploads: 25M / 86,400 ≈ 290/sec (peak: 870/sec)
  • Likes: 500M x 10 / 86,400 ≈ 58,000/sec (peak: 174K/sec)
  • Comments: 500M x 1 / 86,400 ≈ 5,800/sec (peak: 17K/sec)
  • Total write QPS: ~64,000/sec average, ~192K/sec peak
Step 3 — Read QPS:
  • Feed views: 500M x 7 / 86,400 ≈ 40,000 feed requests/sec
  • But each feed request loads 20 posts, each with images: 40K x 20 = 800K image loads/sec
  • Profile views, search, explore page: ~20,000 requests/sec
  • Total read QPS: ~860,000/sec peak
  • Read:Write ratio: roughly 13:1 — read-heavy, which means caching is critical
Step 4 — Storage:Photo storage (the dominant cost):
  • 25M photos/day, average 2 MB after compression
  • Instagram stores 4-5 resolution variants per photo: thumbnail (10 KB), small (100 KB), medium (500 KB), large (1 MB), original (2 MB) = ~3.6 MB total per photo
  • Daily: 25M x 3.6 MB = 90 TB/day
  • Yearly: 90 TB x 365 = 32.8 PB/year
  • 5 years with 3x replication: ~500 PB
Video (Reels):
  • Perhaps 10% of posts are Reels, average 30 seconds at ~5 MB after compression, with 4 quality variants: ~15 MB total
  • 2.5M Reels/day x 15 MB = 37.5 TB/day
  • Yearly: 13.7 PB
Metadata (tiny by comparison):
  • Post metadata: 25M x 1 KB = 25 GB/day
  • User profiles, follows, likes: ~50 GB/day
  • Yearly: ~27 TB
Total annual storage: ~50 PB/year (media dominates by 1000x over metadata)Key insight: Instagram’s costs are overwhelmingly in two areas: (1) blob storage for photos/videos (~50 PB/year) and (2) CDN bandwidth to serve 800K+ image loads per second globally. The application servers and databases are relatively cheap. This is why Meta invested in building their own CDN and custom storage systems rather than using cloud providers — at this scale, cloud storage pricing would be hundreds of millions per year.”
Red flag answer: “500 million users, 1 photo each, 2 MB each, that is 1 PB per day.” This overestimates by 10x because it confuses DAU with daily posters. Only ~5% of users create content on any given day. Also misses the multi-resolution storage multiplication. Follow-ups:
  1. “Instagram Explore page needs to serve personalized content to 500M users. Estimate the recommendation system’s QPS.”
  2. “What percentage of Instagram’s total stored photos are never viewed after the first week? How would you use this to optimize storage costs?”
  3. “Instagram Stories disappear after 24 hours. Estimate the daily storage churn from Stories creation and deletion.”
  4. “Compare Instagram’s infrastructure profile to YouTube’s. Which is more compute-intensive vs storage-intensive, and why?”

Question 7: “How many machines would you need to serve a real-time global leaderboard for a game with 100 million players?”

Difficulty: Senior What the interviewer is really testing: Can you reason about data structure operations per second, memory sizing, and the trade-off between exactness and scalability? Do you distinguish between writes (score updates) and reads (rank lookups)?
“Let me model both the data size and the operation rate.Step 1 — Data size:
  • 100M players, each with: player_id (8 bytes) + score (8 bytes) + metadata (username, avatar_url) = ~50 bytes
  • Total data: 100M x 50 bytes = 5 GB
  • This fits comfortably in a single Redis instance’s memory (modern servers have 128-512 GB RAM)
Step 2 — Operation rate:Score updates (writes):
  • Not all 100M players are active simultaneously. Maybe 10M concurrent.
  • Active players score every ~30 seconds (game action completes)
  • Write QPS: 10M / 30 = 333,000 score updates/sec
  • Peak (tournament event, 2x): 666,000 writes/sec
Rank lookups (reads):
  • Every player checks their rank after each game action, plus periodic checks
  • Read QPS: ~500,000 rank lookups/sec peak
Step 3 — Can a single Redis instance handle this?
  • Redis sorted set operations (ZADD for write, ZREVRANK for rank): each is O(log N) = O(log 100M) = ~27 operations
  • A single Redis instance can handle ~100K-200K sorted set operations/sec
  • We need ~1.2M ops/sec (666K writes + 500K reads)
  • Single instance: No. Need to partition.
Step 4 — Partitioning strategy:Option A — Score-range partitioning:
  • Divide the score range into buckets. Partition 1 handles scores 0-999, Partition 2 handles 1000-1999, etc.
  • Problem: hot partitions. If most players have scores in the 500-2000 range, those partitions are overloaded.
  • Advantage: global rank can be computed by summing player counts in higher partitions + rank within partition.
Option B — Hash partitioning with aggregation:
  • Hash player_id to one of N partitions (N = 10).
  • Each partition maintains its own sorted set of ~10M players.
  • Write: O(1) routing, O(log 10M) per partition = fast.
  • Global rank: Cannot be determined from a single partition. Need a separate aggregation service.
  • Aggregation: maintain a coarse-grained global histogram (score buckets of size 100). Each partition publishes its bucket counts. Global rank = sum of global counts above your score + rank within your bucket.
My recommendation: Option B with aggregation.
  • 10 Redis instances, each handling ~120K ops/sec (well within limits)
  • 1 aggregation service updating the global histogram every 5-10 seconds
  • Approximate global rank (accurate to within ~100 positions for mid-ranked players, exact for top 1000)
  • Top-1000 leaderboard maintained separately with exact rankings
Step 5 — Final machine count:
  • 10 Redis instances (32 GB RAM each, 8 cores): 10 servers
  • 10 Redis replicas for read scaling and failover: 10 servers
  • 2-3 aggregation service instances: 3 servers
  • 2-3 API gateway servers: 3 servers
  • Total: ~26 servers
Add monitoring, load balancers, and headroom: ~30-40 servers.Key insight: 100M players sounds massive, but the actual data is only 5 GB. This is a throughput problem (1.2M ops/sec), not a storage problem. The architecture is shaped by operation rate, not data volume.”
Red flag answer: “100 million players, so we need 100 million rows in a database with an index. Maybe 50 database servers.” This thinks about the problem as a storage/query problem when it is actually a throughput/latency problem. A SQL database cannot serve 500K rank lookups per second with ORDER BY queries. Follow-ups:
  1. “The game runs a 24-hour tournament with 50M concurrent players and score updates every 5 seconds. Re-estimate the infrastructure.”
  2. “How do you display ‘You are ranked #4,847,329 out of 100M players’ when your rank is only approximate? What is the user experience trade-off?”
  3. “The game wants friend-group leaderboards — your rank among your 200 friends. How do you efficiently compute this without N separate sorted sets?”
  4. “During a server failure, some score updates are lost. How do you reconcile scores after recovery?”

Question 8: “Estimate the cost of sending 1 billion push notifications per day.”

Difficulty: Senior What the interviewer is really testing: Do you understand the full cost stack (infrastructure + third-party services)? Can you identify that the dominant cost is not always what you expect? Do you think about failure rates and retries as cost multipliers?
“This is a cost estimation, not just a capacity estimation. I need to think about every layer.Step 1 — Volume breakdown: 1 billion push notifications/day:
  • QPS average: 1B / 86,400 ≈ 11,600/sec
  • Peak QPS (3x): ~35,000/sec
  • Platform split (approximate): 70% Android (FCM), 25% iOS (APNs), 5% Web Push
  • FCM: 700M/day, APNs: 250M/day, Web: 50M/day
Step 2 — Third-party gateway costs (often the biggest cost):
  • FCM (Firebase Cloud Messaging): Free for unlimited push notifications. Google does not charge per message.
  • APNs (Apple Push Notification service): Free. Apple does not charge per message.
  • Web Push: Free (standards-based).
Surprise: the third-party delivery cost is $0. The cost is entirely in YOUR infrastructure to generate, queue, and send these notifications.Step 3 — Infrastructure costs:Compute (generating and dispatching notifications):
  • At 35K/sec peak, assuming each notification takes 5ms of server time (template rendering, user preference check, dedup, API call to FCM/APNs):
  • Concurrent work: 35,000 x 0.005 = 175 concurrent tasks
  • With 8-core servers handling 50 concurrent tasks each: 4 servers
  • With 3x headroom for retries and bursts: 12 servers
  • Cost: 12 x c5.2xlarge at ~250/month=∗∗250/month = **3,000/month**
Message queue (Kafka for buffering):
  • 1B messages/day at 500 bytes each = 500 GB/day throughput
  • 3-day retention: 1.5 TB storage
  • 6-node Kafka cluster with replication
  • Cost: 6 x i3.xlarge at ~300/month=∗∗300/month = **1,800/month**
User preference store (Redis/DynamoDB):
  • Need to check each user’s notification preferences, device tokens, opt-out status
  • 500M users, ~200 bytes per user = 100 GB
  • Redis cluster or DynamoDB
  • Cost: DynamoDB at 35K reads/sec = ~5,000/month(on−demandpricing),orRediscluster:∗∗5,000/month (on-demand pricing), or Redis cluster: **2,000-5,000/month**
Network egress:
  • 1B notifications x 500 bytes = 500 GB/day egress to FCM/APNs
  • AWS egress: 0.09/GB=0.09/GB = 45/day = $1,350/month
Step 4 — Hidden costs most people miss:Retries: FCM and APNs have ~2-5% failure rates (device offline, token expired, rate limited). 1B x 3% retry rate x 2 retries average = 60M extra messages. Adds ~6% to infrastructure costs.Token management: Device tokens change when users reinstall apps. You need a system to handle invalid tokens returned by FCM/APNs, update your database, and stop sending to dead tokens. Sending to invalid tokens wastes compute and can get you rate-limited.Analytics/tracking: Did the notification get delivered? Opened? Each event doubles your event processing pipeline.Step 5 — Total monthly cost estimate:
  • Compute: $3,000
  • Kafka: $1,800
  • User preference store: $3,500
  • Network egress: $1,350
  • Analytics pipeline: $2,000
  • Monitoring and infrastructure overhead: $1,000
  • Total: ~12,000−15,000/monthor 12,000-15,000/month or ~150,000-180,000/year
Key insight: At $0.000015 per notification, the cost is dominated by the user preference lookup and the message queue infrastructure, NOT by the actual sending. This is counterintuitive — the most expensive part of sending a notification is deciding who to send it to and whether they want it, not the sending itself. Companies that skip personalization and blast notifications to everyone pay less per notification but lose users to notification fatigue, which is far more expensive in the long run.”
Red flag answer: “Push notifications are free because APNS and FCM don’t charge.” While technically the delivery is free, this ignores ALL infrastructure costs. It is like saying “making phone calls is free because air transmits sound waves.” The candidate who says this has never operated a notification system. Follow-ups:
  1. “If you switch from FCM free tier to a paid notification service like OneSignal or Airship, how does the cost change?”
  2. “Your notification system has a 5% delivery failure rate. The product team says it should be under 1%. What do you do, and what does it cost?”
  3. “Compare the cost of push notifications vs SMS notifications vs email. At what volume does each channel make economic sense?”
  4. “A bug sends 10x the intended notifications over 2 hours. What is the cost of this incident beyond the extra infrastructure — in terms of user opt-outs and app uninstalls?”

Estimation Template

Use this template in interviews: