Foundations & First PrinciplesBeginner16 min read4 questions

Back-of-the-Envelope Estimation for System Design

The arithmetic that turns a vague brief into a concrete architecture. Learn to size QPS, storage and bandwidth in under two minutes - and to know instantly whether a design is plausible.

Covers: QPS maths, storage sizing, bandwidth, latency numbers every engineer should know, powers of two, peak factors

Estimation is what separates a design from a drawing. The number 40,000 writes per second tells you a single Postgres primary will not do. The number 40 writes per second tells you it will do comfortably for years. Same diagram, opposite conclusions - and you cannot reach either without arithmetic.

Filter
0/4 mastered
BeginnerestimationqpscapacityAsked at Meta, Google

30-second answer

Take DAU, multiply by actions per user per day to get daily requests, divide by 86,400 seconds to get average QPS, then multiply by a peak factor of 2 to 3. For 100 million DAU doing 20 actions a day: 2 billion requests a day, ÷ 86,400 ≈ 23,000 average QPS, so roughly 50,000–70,000 QPS at peak. Then split that by read/write ratio, because 50,000 reads and 500 writes are completely different engineering problems.

IntermediateestimationstoragebandwidthAsked at Amazon, Dropbox

30-second answer

Storage = objects per day × average size × retention, then multiply by a replication factor (usually 3) and add roughly 30% for indexes, metadata and headroom. Bandwidth = QPS × average payload size, computed separately for ingress and egress - egress is usually far larger for media products and is the line item that actually costs money. Always express the result per day and per year, because the yearly figure is what decides your storage tier.

IntermediatelatencyestimationperformanceAsked at Google, Jane Street

30-second answer

The ladder spans nine orders of magnitude: L1 cache around 1 ns, main memory around 100 ns, an SSD read around 100 µs, a same-datacentre round trip around 500 µs, a disk seek around 10 ms, and a cross-continent round trip around 150 ms. The practical consequence is that memory is roughly a thousand times faster than SSD and a hundred thousand times faster than a cross-ocean hop - so the design question is never 'is this fast?' but 'how many times do I cross which boundary?'.

IntermediatecapacityestimationscalingAsked at Amazon, Netflix

30-second answer

Divide target QPS by realistic per-server throughput, then divide again by your target utilisation, then add redundancy. If a server handles 2,000 QPS and you run at 50% utilisation for headroom, 50,000 QPS needs 50 servers, plus enough extra to survive losing an availability zone - so around 75 across three zones. The number matters less than the reasoning: state per-server capacity, utilisation target and failure budget explicitly.

Showing 4 of 4 questions for back-of-the-envelope-estimation.

Check your understanding

4 questions · no sign-up, nothing stored

0/4 answered
Question 1

1.A service has 50 million DAU, each making 4 requests per day. Roughly what is the peak QPS at a 2× peak factor?

Question 2

2.Roughly how much slower is an SSD read than a main-memory read?

Question 3

3.At 20,000 QPS with an average latency of 50 ms, how many requests are in flight at any moment?

Question 4

4.Why do capacity plans target 60–70% utilisation instead of 95%?

Hands-on challenge

Build it - this is what you talk about in a deep-dive round.

Size a video-sharing platform end to end

Produce a complete capacity plan for a video platform with 20 million DAU. Show every assumption. The goal is a defensible order of magnitude reached in under ten minutes.

Requirements

  • Estimate peak upload QPS and peak view QPS, stating your per-user activity assumptions.
  • Compute storage per day and per year for original files plus three transcoded renditions, with 3× replication.
  • Compute origin egress with and without a CDN at a 95% offload rate.
  • Use Little's Law to size the concurrent request count for the view path.
  • State the single number that most constrains the design, and what you would change because of it.

Stretch goals

  • Add a 30-day hot / 1-year warm / archive-after tiering policy and recompute the storage cost profile.
  • Estimate how many transcoding workers you need if one worker handles 1× real-time and average video length is 8 minutes.