Reliability — MTBF, MTTR, and error budgets
How to talk about reliability without hand-waving.
Reliability is about failure, availability is about uptime. They're related but distinct. Reliability is measured by two numbers: how often things break (MTBF) and how fast you recover (MTTR). The ratio decides your availability.
MTBF
Mean Time Between Failures. Higher = more reliable components. Improved by: good architecture, defensive coding, chaos testing, feature flags.
MTTR
Mean Time To Recovery. Lower = faster recovery. Improved by: automation, runbooks, monitoring, blameless postmortems.
availability = MTBF / (MTBF + MTTR)
3 nines: MTBF 30 days, MTTR 30s = 99.998%. Or MTBF 5 days, MTTR 30s = 99.99%. Or MTBF 90 days, MTTR 5 min = 99.996%. Multiple paths to the same number.
The error budget concept (Google SRE)
Error budget = 100% - SLO. If SLO is 99.9%, budget is 0.1%. Over 30 days, that's 43 minutes of allowed downtime.
- Budget UNSPENT → ship features aggressively, take risks.
- Budget SPENT → freeze features, focus on reliability improvements.
- Recomputed monthly. Team-level accountability.
- Aligns product velocity with reliability — both PMs and SREs are on the same team.
Blameless postmortems
Every incident triggers a postmortem. Facilitator is not the on-call. The doc has: timeline (with time-zones), impact (users, revenue), root cause, contributing factors, action items with owners. NEVER contains “X should have known.”
Google's SRE book is the reference. Read it if you want to work in reliability engineering.
Practice what you just read
Every foundation concept has a companion quiz to close the loop.