Reliability, Security & ObservabilityAdvanced15 min read1 questions

Designing for Failure: Bulkheads, Degradation and Chaos

Everything fails eventually. The discipline is deciding in advance how it fails - which features degrade, which stay up, and how much data you are willing to lose.

Covers: Failure modes, bulkhead isolation, graceful degradation, blast radius, chaos engineering, RTO and RPO, disaster recovery

Junior designs assume the happy path and add error handling afterwards. Senior designs start from the failure list and work backwards. In an interview, raising failure modes before you are asked is one of the clearest levelling signals available.

Filter
0/1 mastered
AdvancedbulkheaddegradationresilienceAsked at Amazon, Netflix

30-second answer

Isolate resources with bulkheads so one slow dependency cannot consume the whole thread or connection pool, and design explicit graceful degradation so the feature that depends on it disappears rather than the page. The classic failure is a shared thread pool: one dependency slows from 50 ms to 5 s, every thread ends up blocked waiting on it, and requests that never touch that dependency start timing out too. Separate pools per dependency, combined with circuit breakers and a defined degraded mode per feature, contain the blast radius.

Showing 1 of 1 questions for designing-for-failure.

Check your understanding

3 questions · no sign-up, nothing stored

0/3 answered
Question 1

1.One optional dependency slows to 5 s and your entire service starts timing out. What was missing?

Question 2

2.Which feature should fail closed rather than degrade?

Question 3

3.What do RTO and RPO measure?