> ## Documentation Index
> Fetch the complete documentation index at: https://resources.devweekends.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Estimation Mastery

> Master back-of-envelope calculations for system design interviews

<Warning>
  **Interview Critical**: Every system design interview starts with capacity estimation. Master these calculations to ace the first 5 minutes and build credibility.
</Warning>

## The Art of Estimation

Good estimates are not about being exact -- they are about being *useful*. An estimate of "somewhere between 1,000 and 100,000 QPS" is nearly useless because those two numbers imply completely different architectures. But "roughly 10,000 QPS with peaks around 30,000" is actionable because it tells you: a single database can handle this, but you need a cache layer and probably 20-30 application servers.

The art is in:

1. **Showing structured thinking** -- break complex estimates into simple multiplications, then combine them. Interviewers care more about your process than your final number.
2. **Understanding orders of magnitude** -- the difference between 1 GB and 1 TB is the difference between "fits on a laptop" and "needs a cluster." Getting within 2-3x of the right answer is success; being off by 10x means your architecture is wrong.
3. **Identifying bottlenecks early** -- the estimate often reveals the hardest part of the design. If your storage estimate says 500 PB, storage is your primary constraint, not compute.
4. **Making reasonable assumptions** -- and stating them clearly. "I will assume 200 million DAU with 5 requests per user per day" gives the interviewer a chance to course-correct early.

## 🧠 Numbers You Must Memorize

### Time Constants

```
┌─────────────────────────────────────────────────────────────────┐
│                    TIME REFERENCE CARD                          │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  Seconds in a day:      86,400      ≈ 100,000 (10^5)           │
│  Seconds in a month:    2.6 million ≈ 2.5 × 10^6               │
│  Seconds in a year:     31.5 million ≈ 30 × 10^6               │
│                                                                 │
│  Quick Conversion:                                              │
│  • 1 day = 86,400s ≈ 10^5 s                                    │
│  • 1 week = 604,800s ≈ 6 × 10^5 s                              │
│  • 1 month = 2.6Ms ≈ 2.5 × 10^6 s                              │
│  • 1 year = 31.5Ms ≈ 3 × 10^7 s                                │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘
```

### Latency Numbers (Jeff Dean's Famous List)

```
┌─────────────────────────────────────────────────────────────────┐
│                    LATENCY NUMBERS (2024)                       │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  L1 cache reference                     0.5 ns                  │
│  L2 cache reference                     7 ns                    │
│  Main memory (RAM) reference            100 ns                  │
│  Send 1KB over 1 Gbps network          10 μs                   │
│  SSD random read                        100 μs (0.1 ms)        │
│  Read 1 MB sequentially from memory     250 μs                 │
│  Round trip same datacenter             500 μs (0.5 ms)        │
│  Read 1 MB sequentially from SSD        1 ms                   │
│  HDD seek                               10 ms                  │
│  Read 1 MB sequentially from HDD        20 ms                  │
│  Send packet CA → Netherlands → CA      150 ms                 │
│                                                                 │
│  Rule of Thumb:                                                 │
│  • Memory: 100 ns = 10^-7 s                                    │
│  • SSD: 100 μs = 10^-4 s                                       │
│  • Network (same DC): 1 ms = 10^-3 s                           │
│  • Network (cross-continent): 100 ms = 10^-1 s                 │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘
```

### Data Size Reference

```
┌─────────────────────────────────────────────────────────────────┐
│                    DATA SIZE REFERENCE                          │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  1 char (ASCII)                    1 byte                      │
│  1 char (UTF-8, avg)               2-3 bytes                   │
│  1 Integer (32-bit)                4 bytes                     │
│  1 Long/Timestamp (64-bit)         8 bytes                     │
│  1 UUID                            16 bytes                    │
│                                                                 │
│  Tweet (280 chars + metadata)      ~500 bytes - 1 KB           │
│  Average web page                  2-5 MB                      │
│  Average photo (compressed)        200 KB - 2 MB               │
│  Average short video (1 min)       10-50 MB                    │
│  HD Video (1 hour)                 1-4 GB                      │
│                                                                 │
│  Memory Tiers:                                                  │
│  • 1 KB = 10^3 bytes                                           │
│  • 1 MB = 10^6 bytes                                           │
│  • 1 GB = 10^9 bytes                                           │
│  • 1 TB = 10^12 bytes                                          │
│  • 1 PB = 10^15 bytes                                          │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘
```

### Capacity Numbers

```
┌─────────────────────────────────────────────────────────────────┐
│                    SERVER CAPACITY REFERENCE                    │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  Single Web Server:                                             │
│  • Simple API: 1,000 - 10,000 QPS                              │
│  • Complex API: 100 - 500 QPS                                  │
│  • Concurrent connections: 10K - 100K                          │
│                                                                 │
│  Database Server:                                               │
│  • MySQL/Postgres simple queries: 1,000 - 5,000 QPS            │
│  • MySQL/Postgres complex joins: 100 - 500 QPS                 │
│  • Redis in-memory: 100,000+ QPS                               │
│  • Cassandra write-heavy: 10,000+ QPS                          │
│                                                                 │
│  Message Queue:                                                 │
│  • Kafka per partition: 10,000+ msg/sec                        │
│  • RabbitMQ per queue: 1,000 - 10,000 msg/sec                  │
│                                                                 │
│  Network:                                                       │
│  • 1 Gbps = 125 MB/s                                           │
│  • 10 Gbps = 1.25 GB/s                                         │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘
```

## Estimation Formulas

The formulas below are your toolkit for converting business requirements ("100 million users, 10 actions per day") into infrastructure numbers ("10K QPS, 3TB storage, 20 servers"). The key insight: you almost never need all four estimates (QPS, storage, bandwidth, memory). Identify the *binding constraint* early -- the one number that dictates your architecture -- and spend your interview time on that. For a read-heavy social feed, QPS is the binding constraint. For a video platform, bandwidth and storage are. For a messaging app, concurrent connections are.

### Query Per Second (QPS)

```python theme={null}
# Basic QPS Formula
QPS = (Daily_Active_Users × Actions_Per_User_Per_Day) / Seconds_Per_Day

# With Peak Factor (typically 2x-3x average)
Peak_QPS = QPS × Peak_Factor

# Example: Instagram-like app
# 500M DAU, 20 feed views/day, Peak 3x
QPS = (500M × 20) / 86,400
    = 10B / 86,400
    ≈ 10B / 100K
    = 100,000 QPS

Peak_QPS = 100,000 × 3 = 300,000 QPS
```

### Storage Estimation

```python theme={null}
# Storage Formula
Daily_Storage = Users × Actions_Per_Day × Size_Per_Action
Yearly_Storage = Daily_Storage × 365
Total_Storage = Yearly_Storage × Years × Replication_Factor

# Example: Twitter-like app
# 200M DAU, 2 tweets/day, 500 bytes/tweet, 3x replication, 5 years
Daily = 200M × 2 × 500 bytes = 200 GB
Yearly = 200 GB × 365 = 73 TB
Total = 73 TB × 5 × 3 = 1.1 PB
```

### Bandwidth Estimation

```python theme={null}
# Bandwidth Formula
Bandwidth = QPS × Average_Response_Size

# Example: Video streaming
# 100K concurrent users, 5 Mbps per stream
Total_Bandwidth = 100K × 5 Mbps = 500 Gbps

# For a service:
# 10K QPS, 100 KB average response
Bandwidth = 10K × 100 KB = 1 GB/s = 8 Gbps
```

## 🎯 Step-by-Step Estimation Framework

### Step 1: Identify Key Metrics

Always estimate reads and writes separately -- they have different scaling characteristics and often different bottlenecks. A read-heavy system (100:1 read/write) calls for aggressive caching and read replicas. A write-heavy system (e.g., logging, analytics) calls for append-optimized storage (LSM trees) and message queues for buffering.

```
For any system, estimate these:
1. QPS (read and write separately -- this drives your caching and replication strategy)
2. Storage requirements (this determines single-node vs sharded, and your cost model)
3. Bandwidth requirements (this reveals if you need a CDN or if network is a bottleneck)
4. Memory requirements (for caching -- use the 80/20 rule: cache 20% of data to serve 80% of reads)
```

### Step 2: Make Assumptions Clear

```
"Let me make some assumptions..."

Users:
• Total users: X
• Daily Active Users (DAU): Y% of X
• Concurrent users: Z% of DAU

Actions:
• Actions per user per day: N
• Read:Write ratio: R:W

Data:
• Average data size: S bytes
• Retention period: T years
```

### Step 3: Round Smartly

```
Good rounding (powers of 10):
• 86,400 → 100,000 (10^5)
• 31,536,000 → 30,000,000 (3 × 10^7)
• 2.6M → 2.5M or 3M

Always round to make mental math easier!
```

## 🔥 Practice Problems

### Problem 1: Design a URL Shortener

<Accordion title="Work through the estimation">
  **Given assumptions:**

  * 100M new URLs per month
  * 10:1 read-to-write ratio
  * 5 years of data retention
  * Average URL: 500 bytes

  **Write QPS:**

  ```
  100M URLs / month
  = 100M / (30 days × 24 hours × 3600 seconds)
  = 100M / 2.6M seconds
  ≈ 100M / 3M
  ≈ 33 QPS

  Peak (3x): ~100 QPS
  ```

  **Read QPS:**

  ```
  10:1 ratio = 10 × 33 = 330 QPS
  Peak: ~1000 QPS
  ```

  **Storage (5 years):**

  ```
  URLs = 100M × 12 months × 5 years = 6 Billion URLs
  Storage = 6B × 500 bytes = 3 TB

  With 3x replication: 9 TB
  ```

  **Short URL Length:**

  ```
  62 characters (a-z, A-Z, 0-9)
  62^6 = 56 billion combinations (enough for 6B URLs)
  62^7 = 3.5 trillion (very safe)

  → Use 7 characters for safety
  ```
</Accordion>

### Problem 2: Design Twitter

<Accordion title="Work through the estimation">
  **Given assumptions:**

  * 500M total users
  * 200M DAU
  * Average user follows 200 people
  * 10% of users tweet daily
  * Average tweet: 500 bytes
  * Read:Write ratio: 50:1

  **Tweet Write QPS:**

  ```
  Tweets/day = 200M × 10% × 2 tweets = 40M tweets
  Write QPS = 40M / 86,400 ≈ 40M / 100K = 400 QPS
  Peak: 1,200 QPS
  ```

  **Timeline Read QPS:**

  ```
  Timeline views = 200M DAU × 5 views = 1B views/day
  Read QPS = 1B / 86,400 ≈ 1B / 100K = 10,000 QPS
  Peak: 30,000 QPS
  ```

  **Storage (5 years):**

  ```
  Daily tweets: 40M × 500 bytes = 20 GB
  Yearly: 20 GB × 365 = 7.3 TB
  5 years: 36.5 TB

  Media (10% with images):
  4M images × 1 MB = 4 TB/day = 7.3 PB over 5 years
  ```

  **Timeline Fanout:**

  ```
  Celebrity with 50M followers:
  1 tweet → 50M timeline updates
  This is the "celebrity problem" → hybrid approach needed
  ```
</Accordion>

### Problem 3: Design YouTube

<Accordion title="Work through the estimation">
  **Given assumptions:**

  * 2B monthly active users
  * 500M DAU watching videos
  * Average video watch: 5 min, 5 videos/day
  * 500K creators uploading daily
  * Average video: 10 min, 100 MB (compressed)

  **Video Watch QPS:**

  ```
  Views = 500M × 5 videos = 2.5B views/day
  View QPS = 2.5B / 86,400 ≈ 2.5B / 100K = 25,000 QPS
  Peak: 75,000 QPS
  ```

  **Video Upload QPS:**

  ```
  Uploads = 500K / 86,400 ≈ 6 uploads/second
  Peak: 20 uploads/second
  ```

  **Storage (per year):**

  ```
  Videos/year = 500K × 365 = 182.5M videos
  Storage = 182.5M × 100 MB = 18.25 PB/year

  Multiple resolutions (4x): ~75 PB/year
  ```

  **Bandwidth:**

  ```
  Concurrent streams: 50M users
  Average bitrate: 5 Mbps
  Total: 50M × 5 Mbps = 250 Pbps

  Need massive CDN infrastructure!
  ```
</Accordion>

### Problem 4: Design WhatsApp

<Accordion title="Work through the estimation">
  **Given assumptions:**

  * 2B users total
  * 500M DAU
  * Average: 50 messages/day/user
  * Average message: 200 bytes
  * 5% with images (100 KB avg)

  **Message QPS:**

  ```
  Messages/day = 500M × 50 = 25B messages
  QPS = 25B / 86,400 ≈ 25B / 100K = 250,000 QPS
  Peak: 750,000 QPS

  This is why WhatsApp uses Erlang for concurrency!
  ```

  **Storage (5 years):**

  ```
  Text: 25B × 200 bytes = 5 TB/day
  Images: 25B × 5% × 100 KB = 125 TB/day
  Total: 130 TB/day = 47 PB/year = 235 PB over 5 years
  ```

  **Concurrent Connections:**

  ```
  500M DAU, 20% online at any time = 100M concurrent
  WebSocket connections are long-lived!
  ```
</Accordion>

## Interview Tips

The goal of estimation in an interview is not to arrive at a number -- it is to arrive at an *architectural insight*. Every estimate should end with a sentence that connects the number to a design decision.

<Tip>
  **The Power Move**: After every estimate, state the architectural implication. "At 25K QPS with a 50:1 read/write ratio, our write path is only 500 QPS -- easily handled by a single PostgreSQL primary -- but our read path at 25K QPS needs aggressive caching. I would target a 95%+ cache hit rate, which means our database only sees \~1,250 QPS. That is well within what a single primary with two read replicas can handle, so we do not need sharding." This single paragraph demonstrates more architectural judgment than most candidates show in an entire interview.
</Tip>

### Techniques That Impress

```
1. State assumptions CLEARLY and invite correction
   "I will assume 100 million DAU -- does that match the
   scale you have in mind, or should I adjust?"

2. Round aggressively and explain why
   "86,400 is close enough to 100,000. The goal is order
   of magnitude, and this makes the mental math clean."

3. Work through calculations visibly
   Write them down. Interviewers give credit for showing
   your work even if you make an arithmetic error.

4. Sanity check against known benchmarks
   "Our estimate says 10K QPS -- that is in the same
   ballpark as mid-sized Twitter, which feels right
   for 100M DAU with 10 actions each."

5. Always end with the architectural implication
   "At this scale, the bottleneck is read QPS, not
   storage. That tells me caching is the highest-leverage
   investment."
```

### Common Mistakes to Avoid

```
1. False precision
   Bad: "That is exactly 11,574.074 QPS"
   Good: "Roughly 12K QPS, call it 15K with headroom"

2. Skipping the estimation entirely
   Jumping to architecture without sizing shows you
   do not understand the scale you are designing for.

3. Forgetting peak traffic
   Average QPS is 2.5-3x lower than peak. Design for
   peak, or your system fails during the Super Bowl.

4. Forgetting replication
   Storage usually needs 3x for fault tolerance, plus
   additional space for indexes (add 30-50% for SQL).

5. Confusing MB and Mb (megabytes vs megabits)
   1 MB = 8 Mb. Network bandwidth is measured in bits;
   storage is measured in bytes. Getting this wrong makes
   your bandwidth estimate 8x off.
```

## 📊 Quick Reference Tables

### Scale Levels

| Scale | DAU | QPS | Storage/Year | Complexity |
| - | - | - | - | - |
| Small | 100K | 10-100 | 100 GB | Single server |
| Medium | 10M | 1K-10K | 10 TB | Multi-server |
| Large | 100M | 10K-100K | 100 TB | Distributed |
| Massive | 1B+ | 100K+ | 1 PB+ | Global infrastructure |

### Common Patterns by QPS

| QPS | Architecture | Examples |
| - | - | - |
| \< 100 | Single server | Small SaaS apps |
| 100-1K | Load balancer + servers | Medium businesses |
| 1K-10K | + Caching layer | Social apps |
| 10K-100K | + Sharding + CDN | Large platforms |
| 100K+ | Global distribution | FAANG scale |

### Storage Growth Rates

| Data Type | Per Item | Daily (10M users) | Yearly |
| - | - | - | - |
| Text message | 200 B | 20 GB | 7 TB |
| Tweet | 500 B | 50 GB | 18 TB |
| Photo | 500 KB | 500 TB | 182 PB |
| Video (1 min) | 50 MB | 5 PB | 1.8 EB |

## Estimation Interview Questions

These are the exact style of back-of-envelope estimation questions you will face in system design interviews. Each includes a full worked solution showing the thought process, the math, and the insights the interviewer expects you to extract from the numbers.

***

### Question 1: "Estimate the storage requirements for Netflix's entire video library."

**Difficulty:** Foundational

**What the interviewer is really testing:** Can you break a large unknown into smaller estimable pieces? Do you understand video encoding and multi-resolution storage?

<Accordion title="Strong Answer (Worked Solution)">
  "Let me break this down step by step.

  **Step 1 -- How many titles does Netflix have?**
  Netflix has approximately 15,000 titles (movies and TV series). A TV series averages 3 seasons at 10 episodes each = 30 episodes. Let's say 5,000 movies and 10,000 series. Total episodes/movies:

  * 5,000 movies
  * 10,000 series x 30 episodes = 300,000 episodes
  * Total content items: \~305,000

  **Step 2 -- Average duration:**

  * Movies: \~100 minutes
  * Episodes: \~40 minutes
  * Total hours: (5,000 x 100 + 300,000 x 40) / 60 = (500K + 12M) / 60 = \~208,000 hours

  **Step 3 -- Storage per hour at each resolution:**
  Netflix encodes each title at multiple quality levels:

  * 240p: \~0.3 GB/hour
  * 480p: \~0.7 GB/hour
  * 720p: \~1.5 GB/hour
  * 1080p: \~3 GB/hour
  * 4K HDR: \~7 GB/hour

  Average across 5 quality levels: \~2.5 GB/hour
  But they also encode multiple audio tracks (5-10 languages) at \~100 MB/hour each: \~0.7 GB/hour for 7 languages.

  Total per hour of content: \~(0.3 + 0.7 + 1.5 + 3 + 7 + 0.7) = \~13 GB/hour

  **Step 4 -- Total storage:**
  208,000 hours x 13 GB/hour = \~2.7 PB

  Round to \~3 PB for the core library. With 3x replication across regions: \~9 PB. Plus working copies in the transcoding pipeline: \~10-15 PB total.

  **Key insight to state:** The video storage is large but bounded. Netflix's real infrastructure cost is in bandwidth (serving 200M+ subscribers), not storage. At current S3 pricing (\~$0.023/GB/month), 3 PB costs about $70K/month in storage alone -- trivial compared to their \$1B+ annual bandwidth spend."
</Accordion>

**Red flag answer:** "Netflix has a lot of videos, so probably like 100 PB." Throwing out a number without showing how you derived it gives the interviewer nothing to evaluate. The point is the decomposition, not the final number.

**Follow-ups:**

1. "Netflix adds 1,000 new titles per month. How does this change their annual storage growth rate?"
2. "Netflix pre-positions popular content on ISP servers (Open Connect). How much storage does a single ISP appliance need?"
3. "If Netflix wanted to support 8K streaming, how would that change your storage estimate?"
4. "What is the cost difference between storing this in S3 vs in a CDN edge cache?"

***

### Question 2: "How many servers does Google need to handle Search?"

**Difficulty:** Intermediate

**What the interviewer is really testing:** Can you chain multiple estimates together (users -> queries -> QPS -> servers)? Do you account for the difference between peak and average?

<Accordion title="Strong Answer (Worked Solution)">
  "Let me work backwards from the query volume.

  **Step 1 -- Query volume:**
  Google processes approximately 8.5 billion searches per day (public data).

  **Step 2 -- QPS:**
  QPS = 8.5B / 86,400 ≈ 8.5B / 100K = 85,000 QPS average
  Peak QPS (3x): \~255,000 QPS

  **Step 3 -- How much work is one search query?**
  This is the critical estimation. A single Google search query:

  * Hits the index (likely distributed across 1,000+ shards)
  * Each shard scans its portion of the index and returns top results
  * Results are merged, ranked, and personalized
  * Total compute: roughly equivalent to scanning 1 million documents in \~200ms

  A single server can handle a complex search query in about 200ms, meaning one server handles \~5 queries/second (factoring in CPU and I/O).

  **Step 4 -- Servers for query serving:**
  255,000 peak QPS / 5 QPS per server = 51,000 query-serving servers.

  **Step 5 -- But that is just the serving layer.**

  * Index storage: Google's index is estimated at 100+ PB. At 10 TB per server (SSD), that is 10,000 servers just for index storage.
  * Crawling and indexing: Continuous pipeline processing billions of pages.
  * Ads serving: A separate system with its own servers.
  * Caching layer: Frequently searched queries (head queries like 'weather', 'facebook') are cached. This reduces the 255K QPS hitting the full search pipeline to perhaps 100K QPS.

  With caching: 100K / 5 = 20,000 query-serving servers.

  **Step 6 -- Total estimate:**

  * Query serving: \~20,000-50,000 servers
  * Index storage: \~10,000-20,000 servers
  * Crawling/indexing pipeline: \~5,000-10,000 servers
  * Ads, caching, infrastructure: \~10,000-20,000 servers
  * **Total for Search alone: \~50,000-100,000 servers**

  With 2x for redundancy: \~100,000-200,000 servers.

  **Sanity check:** Google reportedly operates millions of servers globally across all services. Search being 100-200K servers (5-10% of total fleet) feels reasonable given that Search is one of many products."
</Accordion>

**Red flag answer:** "Google is huge so they probably need millions of servers for search." This is vague. The interviewer wants to see you decompose the problem into QPS, per-query cost, and server capacity.

**Follow-ups:**

1. "What is the cost of running these servers? Estimate the annual infrastructure cost for Google Search."
2. "Google serves results in 200ms. Break down the latency budget -- where does each millisecond go?"
3. "How would the server count change if Google moved to a fully GPU-based search ranking pipeline?"
4. "How does Google handle a 10x traffic spike during a major world event?"

***

### Question 3: "Estimate the bandwidth required for Spotify's streaming service."

**Difficulty:** Intermediate

**What the interviewer is really testing:** Do you understand audio streaming bitrates? Can you distinguish between concurrent users and DAU? Do you think about CDN and edge caching?

<Accordion title="Strong Answer (Worked Solution)">
  "Let me build this up from user behavior.

  **Step 1 -- User base and concurrency:**

  * Spotify has \~600M total users, \~220M paying subscribers
  * DAU: roughly 200M users
  * Average listening time: 30 minutes/day across all users (some listen for hours, many just a few minutes)
  * Concurrent listeners at peak: \~10% of DAU = 20M concurrent streams
  * Average concurrency: \~5% of DAU = 10M concurrent streams

  **Step 2 -- Bitrate per stream:**

  * Free tier (normal quality): 160 kbps
  * Premium (high quality): 320 kbps
  * Let's assume weighted average of 200 kbps (mix of free and premium, mobile and desktop)

  **Step 3 -- Total bandwidth:**
  Peak: 20M concurrent x 200 kbps = 4,000 Gbps = 4 Tbps
  Average: 10M concurrent x 200 kbps = 2 Tbps

  **Step 4 -- CDN impact:**
  Most popular songs (top 10,000 tracks) account for perhaps 30-40% of all streams. These are cached at the edge, served from CDN PoPs near the user. The CDN absorbs the majority of bandwidth, so origin bandwidth is much lower -- maybe 10% of total.

  * Origin bandwidth: \~200-400 Gbps
  * CDN edge bandwidth: \~3.6 Tbps

  **Step 5 -- Storage perspective (bonus):**
  Spotify's catalog: \~100M tracks x average 4 minutes x 3 quality levels
  Storage per track per quality: 4 min x 60 sec x 200 kbps / 8 = \~6 MB per quality level
  100M x 6 MB x 3 = \~1.8 PB for the entire music catalog
  With podcasts (5M+ episodes at \~50 MB each): +250 TB
  Total catalog: \~2 PB

  **Key insight:** 4 Tbps is significant but manageable with modern CDN infrastructure. For reference, Netflix peaks at 10+ Tbps. Audio is much less bandwidth-intensive than video -- the same CDN infrastructure that serves one Netflix stream could serve 25+ Spotify streams."
</Accordion>

**Red flag answer:** "600 million users times 320 kbps equals 192 Tbps." This catastrophic error confuses total users with concurrent users. The interviewer will immediately note that you do not understand the difference between registered users, DAU, and concurrent connections.

**Follow-ups:**

1. "Spotify is launching in 20 new countries in Africa and Southeast Asia. How does this change your bandwidth estimate and infrastructure requirements?"
2. "What happens to your bandwidth when Spotify releases a new Taylor Swift album and 50M users try to stream it simultaneously?"
3. "Compare the infrastructure cost of streaming audio vs streaming video. Why is Spotify profitable while most video streaming services are not?"
4. "How would Spotify reduce bandwidth costs by 30% without degrading user experience?"

***

### Question 4: "How much storage does Uber need for trip data over 5 years?"

**Difficulty:** Foundational / Intermediate

**What the interviewer is really testing:** Can you estimate data per event, multiply by event frequency, and account for different data types (structured metadata vs GPS traces)?

<Accordion title="Strong Answer (Worked Solution)">
  "Uber's trip data has multiple components, each with different storage characteristics.

  **Step 1 -- Trip volume:**

  * Uber reports \~25 million trips per day globally
  * Per year: 25M x 365 = \~9 billion trips/year
  * Over 5 years: \~45 billion trips

  **Step 2 -- Storage per trip (structured data):**
  A trip record includes:

  * Trip ID, rider ID, driver ID: 3 x 8 bytes = 24 bytes
  * Pickup/dropoff coordinates: 4 x 8 bytes = 32 bytes
  * Start/end timestamps: 2 x 8 bytes = 16 bytes
  * Distance, duration, fare, surge multiplier: 4 x 8 bytes = 32 bytes
  * Payment details, promo codes, ratings: \~100 bytes
  * Status history, cancellation info: \~50 bytes
  * Total per trip record: \~250 bytes, round to 500 bytes with indexes and overhead

  Structured trip data: 45B trips x 500 bytes = 22.5 TB
  With indexes and 3x replication: \~70 TB

  **Step 3 -- GPS trace data (the big one):**
  During an active trip, the driver's phone sends location updates every 4 seconds.

  * Average trip duration: 15 minutes = 900 seconds
  * GPS points per trip: 900 / 4 = 225 points
  * Each point: latitude (8 bytes) + longitude (8 bytes) + timestamp (8 bytes) + speed (4 bytes) + heading (4 bytes) = 32 bytes
  * GPS data per trip: 225 x 32 bytes = 7,200 bytes ≈ 7 KB

  GPS data: 45B trips x 7 KB = 315 TB
  With replication (3x): \~950 TB ≈ 1 PB

  **Step 4 -- Driver location data (even bigger):**
  Even when NOT on a trip, available drivers send location updates every 4 seconds.

  * 5M active drivers, \~10 hours/day active
  * Updates: 5M x (10 x 3600 / 4) = 45 billion location pings/day
  * Each ping: 32 bytes
  * Daily: 45B x 32 bytes = 1.44 TB/day
  * Yearly: 1.44 TB x 365 = 526 TB
  * 5 years with replication: \~8 PB

  **Step 5 -- Total:**

  * Trip records: \~70 TB
  * GPS traces: \~1 PB
  * Driver location pings: \~8 PB
  * **Total: \~9-10 PB over 5 years**

  **Key insight:** The structured trip data is tiny. The GPS and location tracking data dominates by 100x. This is why Uber invested heavily in their own time-series database (M3) and geospatial infrastructure -- generic databases cannot handle 45 billion location pings per day efficiently."
</Accordion>

**Red flag answer:** "25 million trips per day, each trip is about 1 KB, so 25 GB per day, 45 TB over 5 years." This dramatically underestimates because it ignores GPS traces and driver location tracking, which are 100x larger than the trip metadata.

**Follow-ups:**

1. "Uber wants to retain GPS traces for only 90 days to save storage but keep trip records forever. How does this change the architecture?"
2. "How would you estimate the QPS for Uber's location service, given 5M drivers sending updates every 4 seconds?"
3. "Uber Eats adds 10M food delivery orders per day. Each order has additional data (restaurant, menu items, delivery instructions). How does this change your estimate?"
4. "At 10 PB, what is Uber's approximate annual storage cost? Compare cloud vs on-premise."

***

### Question 5: "Estimate the number of chat messages WhatsApp processes per second."

**Difficulty:** Foundational

**What the interviewer is really testing:** Basic QPS estimation. Can you make reasonable assumptions about user behavior and convert daily volume to per-second rate?

<Accordion title="Strong Answer (Worked Solution)">
  "Let me start from users and work toward QPS.

  **Step 1 -- User base:**

  * WhatsApp has \~2.5 billion monthly active users
  * DAU: roughly 70% of MAU = 1.75 billion DAU (WhatsApp has unusually high daily engagement)

  **Step 2 -- Messages per user per day:**

  * This varies wildly by market. In India and Brazil, power users send 100+ messages/day. In the US, maybe 20-30.
  * Weighted global average: \~50 messages/day per active user
  * But messages are bidirectional -- for every message sent, someone receives it. The system processes both.
  * Messages processed (sent): 1.75B x 50 = 87.5 billion messages/day

  **Step 3 -- Convert to QPS:**
  QPS = 87.5B / 86,400 ≈ 87.5B / 100K = 875,000 messages/sec
  Peak (3x): \~2.6 million messages/sec

  **Step 4 -- Breakdown by type:**

  * Text messages (\~80%): 700K/sec
  * Image/media messages (\~15%): 131K/sec
  * Voice messages (\~5%): 44K/sec

  **Step 5 -- Related metrics worth computing:**

  * Connections: 1.75B DAU with \~30% concurrent at peak = 525M WebSocket connections
  * Bandwidth for text: 875K/sec x 200 bytes = 175 MB/sec = 1.4 Gbps (trivial)
  * Bandwidth for media: 131K/sec x 200 KB average = 26 GB/sec = 208 Gbps (significant!)
  * Status updates (online/offline, typing, read receipts): probably 5-10x the message rate = 5-9 million events/sec

  **Key insight:** The raw message QPS (875K/sec) is high but manageable. The real engineering challenge is the 525M concurrent WebSocket connections. This is why WhatsApp was famous for handling 2 million connections per server using Erlang/BEAM -- they needed roughly 250-500 servers just for the connection layer, which is remarkably lean for 2 billion users."
</Accordion>

**Red flag answer:** "2.5 billion users times 50 messages is 125 billion, divided by 86,400 is about 1.4 million QPS." This uses total users instead of DAU, and does not distinguish between peak and average. Also missing the insight about connections being the real bottleneck, not message processing.

**Follow-ups:**

1. "WhatsApp must deliver messages in order within a conversation. How does this constraint affect their server architecture at 875K messages/sec?"
2. "During New Year's Eve, messaging volume spikes 10-20x. How does WhatsApp prepare for this?"
3. "WhatsApp stores messages on the server only until delivered. Estimate the peak undelivered message storage when 30% of users are offline."
4. "Compare WhatsApp's infrastructure requirements to Slack's. Why can WhatsApp serve 10x more users with a fraction of the engineering team?"

***

### Question 6: "Estimate the QPS and storage for a system like Instagram."

**Difficulty:** Intermediate

**What the interviewer is really testing:** Can you separately estimate reads and writes, account for media storage vs metadata, and identify the read:write ratio as a key architectural signal?

<Accordion title="Strong Answer (Worked Solution)">
  "Instagram is interesting because it has asymmetric read/write patterns and media-heavy storage. Let me break it down.

  **Step 1 -- User activity:**

  * 2B monthly active users, \~500M DAU
  * Content creation: \~5% of DAU posts daily = 25M posts/day
  * Feed consumption: average user views feed 7 times/day, 20 posts per view = 140 posts viewed/day
  * Likes: average user likes 10 posts/day
  * Comments: average user comments on 1 post/day

  **Step 2 -- Write QPS:**

  * Photo/video uploads: 25M / 86,400 ≈ 290/sec (peak: 870/sec)
  * Likes: 500M x 10 / 86,400 ≈ 58,000/sec (peak: 174K/sec)
  * Comments: 500M x 1 / 86,400 ≈ 5,800/sec (peak: 17K/sec)
  * Total write QPS: \~64,000/sec average, \~192K/sec peak

  **Step 3 -- Read QPS:**

  * Feed views: 500M x 7 / 86,400 ≈ 40,000 feed requests/sec
  * But each feed request loads 20 posts, each with images: 40K x 20 = 800K image loads/sec
  * Profile views, search, explore page: \~20,000 requests/sec
  * Total read QPS: \~860,000/sec peak
  * **Read:Write ratio: roughly 13:1** -- read-heavy, which means caching is critical

  **Step 4 -- Storage:**

  **Photo storage (the dominant cost):**

  * 25M photos/day, average 2 MB after compression
  * Instagram stores 4-5 resolution variants per photo: thumbnail (10 KB), small (100 KB), medium (500 KB), large (1 MB), original (2 MB) = \~3.6 MB total per photo
  * Daily: 25M x 3.6 MB = 90 TB/day
  * Yearly: 90 TB x 365 = 32.8 PB/year
  * 5 years with 3x replication: \~500 PB

  **Video (Reels):**

  * Perhaps 10% of posts are Reels, average 30 seconds at \~5 MB after compression, with 4 quality variants: \~15 MB total
  * 2.5M Reels/day x 15 MB = 37.5 TB/day
  * Yearly: 13.7 PB

  **Metadata (tiny by comparison):**

  * Post metadata: 25M x 1 KB = 25 GB/day
  * User profiles, follows, likes: \~50 GB/day
  * Yearly: \~27 TB

  **Total annual storage: \~50 PB/year** (media dominates by 1000x over metadata)

  **Key insight:** Instagram's costs are overwhelmingly in two areas: (1) blob storage for photos/videos (\~50 PB/year) and (2) CDN bandwidth to serve 800K+ image loads per second globally. The application servers and databases are relatively cheap. This is why Meta invested in building their own CDN and custom storage systems rather than using cloud providers -- at this scale, cloud storage pricing would be hundreds of millions per year."
</Accordion>

**Red flag answer:** "500 million users, 1 photo each, 2 MB each, that is 1 PB per day." This overestimates by 10x because it confuses DAU with daily posters. Only \~5% of users create content on any given day. Also misses the multi-resolution storage multiplication.

**Follow-ups:**

1. "Instagram Explore page needs to serve personalized content to 500M users. Estimate the recommendation system's QPS."
2. "What percentage of Instagram's total stored photos are never viewed after the first week? How would you use this to optimize storage costs?"
3. "Instagram Stories disappear after 24 hours. Estimate the daily storage churn from Stories creation and deletion."
4. "Compare Instagram's infrastructure profile to YouTube's. Which is more compute-intensive vs storage-intensive, and why?"

***

### Question 7: "How many machines would you need to serve a real-time global leaderboard for a game with 100 million players?"

**Difficulty:** Senior

**What the interviewer is really testing:** Can you reason about data structure operations per second, memory sizing, and the trade-off between exactness and scalability? Do you distinguish between writes (score updates) and reads (rank lookups)?

<Accordion title="Strong Answer (Worked Solution)">
  "Let me model both the data size and the operation rate.

  **Step 1 -- Data size:**

  * 100M players, each with: player\_id (8 bytes) + score (8 bytes) + metadata (username, avatar\_url) = \~50 bytes
  * Total data: 100M x 50 bytes = 5 GB
  * This fits comfortably in a single Redis instance's memory (modern servers have 128-512 GB RAM)

  **Step 2 -- Operation rate:**

  Score updates (writes):

  * Not all 100M players are active simultaneously. Maybe 10M concurrent.
  * Active players score every \~30 seconds (game action completes)
  * Write QPS: 10M / 30 = 333,000 score updates/sec
  * Peak (tournament event, 2x): 666,000 writes/sec

  Rank lookups (reads):

  * Every player checks their rank after each game action, plus periodic checks
  * Read QPS: \~500,000 rank lookups/sec peak

  **Step 3 -- Can a single Redis instance handle this?**

  * Redis sorted set operations (ZADD for write, ZREVRANK for rank): each is O(log N) = O(log 100M) = \~27 operations
  * A single Redis instance can handle \~100K-200K sorted set operations/sec
  * We need \~1.2M ops/sec (666K writes + 500K reads)
  * Single instance: **No.** Need to partition.

  **Step 4 -- Partitioning strategy:**

  **Option A -- Score-range partitioning:**

  * Divide the score range into buckets. Partition 1 handles scores 0-999, Partition 2 handles 1000-1999, etc.
  * Problem: hot partitions. If most players have scores in the 500-2000 range, those partitions are overloaded.
  * Advantage: global rank can be computed by summing player counts in higher partitions + rank within partition.

  **Option B -- Hash partitioning with aggregation:**

  * Hash player\_id to one of N partitions (N = 10).
  * Each partition maintains its own sorted set of \~10M players.
  * Write: O(1) routing, O(log 10M) per partition = fast.
  * Global rank: Cannot be determined from a single partition. Need a separate aggregation service.
  * Aggregation: maintain a coarse-grained global histogram (score buckets of size 100). Each partition publishes its bucket counts. Global rank = sum of global counts above your score + rank within your bucket.

  **My recommendation: Option B with aggregation.**

  * 10 Redis instances, each handling \~120K ops/sec (well within limits)
  * 1 aggregation service updating the global histogram every 5-10 seconds
  * Approximate global rank (accurate to within \~100 positions for mid-ranked players, exact for top 1000)
  * Top-1000 leaderboard maintained separately with exact rankings

  **Step 5 -- Final machine count:**

  * 10 Redis instances (32 GB RAM each, 8 cores): 10 servers
  * 10 Redis replicas for read scaling and failover: 10 servers
  * 2-3 aggregation service instances: 3 servers
  * 2-3 API gateway servers: 3 servers
  * **Total: \~26 servers**

  Add monitoring, load balancers, and headroom: **\~30-40 servers.**

  **Key insight:** 100M players sounds massive, but the actual data is only 5 GB. This is a throughput problem (1.2M ops/sec), not a storage problem. The architecture is shaped by operation rate, not data volume."
</Accordion>

**Red flag answer:** "100 million players, so we need 100 million rows in a database with an index. Maybe 50 database servers." This thinks about the problem as a storage/query problem when it is actually a throughput/latency problem. A SQL database cannot serve 500K rank lookups per second with `ORDER BY` queries.

**Follow-ups:**

1. "The game runs a 24-hour tournament with 50M concurrent players and score updates every 5 seconds. Re-estimate the infrastructure."
2. "How do you display 'You are ranked #4,847,329 out of 100M players' when your rank is only approximate? What is the user experience trade-off?"
3. "The game wants friend-group leaderboards -- your rank among your 200 friends. How do you efficiently compute this without N separate sorted sets?"
4. "During a server failure, some score updates are lost. How do you reconcile scores after recovery?"

***

### Question 8: "Estimate the cost of sending 1 billion push notifications per day."

**Difficulty:** Senior

**What the interviewer is really testing:** Do you understand the full cost stack (infrastructure + third-party services)? Can you identify that the dominant cost is not always what you expect? Do you think about failure rates and retries as cost multipliers?

<Accordion title="Strong Answer (Worked Solution)">
  "This is a cost estimation, not just a capacity estimation. I need to think about every layer.

  **Step 1 -- Volume breakdown:**
  1 billion push notifications/day:

  * QPS average: 1B / 86,400 ≈ 11,600/sec
  * Peak QPS (3x): \~35,000/sec
  * Platform split (approximate): 70% Android (FCM), 25% iOS (APNs), 5% Web Push
  * FCM: 700M/day, APNs: 250M/day, Web: 50M/day

  **Step 2 -- Third-party gateway costs (often the biggest cost):**

  * FCM (Firebase Cloud Messaging): Free for unlimited push notifications. Google does not charge per message.
  * APNs (Apple Push Notification service): Free. Apple does not charge per message.
  * Web Push: Free (standards-based).

  **Surprise: the third-party delivery cost is \$0.** The cost is entirely in YOUR infrastructure to generate, queue, and send these notifications.

  **Step 3 -- Infrastructure costs:**

  **Compute (generating and dispatching notifications):**

  * At 35K/sec peak, assuming each notification takes 5ms of server time (template rendering, user preference check, dedup, API call to FCM/APNs):
  * Concurrent work: 35,000 x 0.005 = 175 concurrent tasks
  * With 8-core servers handling 50 concurrent tasks each: 4 servers
  * With 3x headroom for retries and bursts: 12 servers
  * Cost: 12 x `c5.2xlarge` at \~$250/month = **$3,000/month\*\*

  **Message queue (Kafka for buffering):**

  * 1B messages/day at 500 bytes each = 500 GB/day throughput
  * 3-day retention: 1.5 TB storage
  * 6-node Kafka cluster with replication
  * Cost: 6 x `i3.xlarge` at \~$300/month = **$1,800/month\*\*

  **User preference store (Redis/DynamoDB):**

  * Need to check each user's notification preferences, device tokens, opt-out status
  * 500M users, \~200 bytes per user = 100 GB
  * Redis cluster or DynamoDB
  * Cost: DynamoDB at 35K reads/sec = \~$5,000/month (on-demand pricing), or Redis cluster: **$2,000-5,000/month\*\*

  **Network egress:**

  * 1B notifications x 500 bytes = 500 GB/day egress to FCM/APNs
  * AWS egress: $0.09/GB = $45/day = **\$1,350/month**

  **Step 4 -- Hidden costs most people miss:**

  **Retries:** FCM and APNs have \~2-5% failure rates (device offline, token expired, rate limited). 1B x 3% retry rate x 2 retries average = 60M extra messages. Adds \~6% to infrastructure costs.

  **Token management:** Device tokens change when users reinstall apps. You need a system to handle invalid tokens returned by FCM/APNs, update your database, and stop sending to dead tokens. Sending to invalid tokens wastes compute and can get you rate-limited.

  **Analytics/tracking:** Did the notification get delivered? Opened? Each event doubles your event processing pipeline.

  **Step 5 -- Total monthly cost estimate:**

  * Compute: \$3,000
  * Kafka: \$1,800
  * User preference store: \$3,500
  * Network egress: \$1,350
  * Analytics pipeline: \$2,000
  * Monitoring and infrastructure overhead: \$1,000
  * **Total: \~$12,000-15,000/month or ~$150,000-180,000/year**

  **Key insight:** At \$0.000015 per notification, the cost is dominated by the user preference lookup and the message queue infrastructure, NOT by the actual sending. This is counterintuitive -- the most expensive part of sending a notification is deciding who to send it to and whether they want it, not the sending itself. Companies that skip personalization and blast notifications to everyone pay less per notification but lose users to notification fatigue, which is far more expensive in the long run."
</Accordion>

**Red flag answer:** "Push notifications are free because APNS and FCM don't charge." While technically the delivery is free, this ignores ALL infrastructure costs. It is like saying "making phone calls is free because air transmits sound waves." The candidate who says this has never operated a notification system.

**Follow-ups:**

1. "If you switch from FCM free tier to a paid notification service like OneSignal or Airship, how does the cost change?"
2. "Your notification system has a 5% delivery failure rate. The product team says it should be under 1%. What do you do, and what does it cost?"
3. "Compare the cost of push notifications vs SMS notifications vs email. At what volume does each channel make economic sense?"
4. "A bug sends 10x the intended notifications over 2 hours. What is the cost of this incident beyond the extra infrastructure -- in terms of user opt-outs and app uninstalls?"

***

## Estimation Template

Use this template in interviews:

```
┌─────────────────────────────────────────────────────────────────┐
│                    ESTIMATION TEMPLATE                          │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  1. USERS                                                       │
│     Total: _______    DAU: _______    Concurrent: _______      │
│                                                                 │
│  2. TRAFFIC                                                     │
│     Actions/day/user: _______                                  │
│     Total/day: _______                                         │
│     QPS: _______ / 86,400 = _______                           │
│     Peak QPS (×3): _______                                     │
│                                                                 │
│  3. STORAGE                                                     │
│     Size per item: _______                                     │
│     Items/day: _______                                         │
│     Daily storage: _______                                     │
│     Yearly storage: _______                                    │
│     5-year storage: _______                                    │
│     With replication (×3): _______                             │
│                                                                 │
│  4. BANDWIDTH                                                   │
│     Response size: _______                                     │
│     Bandwidth = QPS × Size = _______                          │
│                                                                 │
│  5. MEMORY (Cache)                                             │
│     Hot data %: _______                                        │
│     Cache size: _______                                        │
│                                                                 │
│  6. KEY INSIGHTS                                                │
│     Read:Write ratio: _______                                  │
│     Primary bottleneck: _______                                │
│     Scaling strategy: _______                                  │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.