Curriculum

Reliability & Operations

Failure modes and blast radius, SLOs and error budgets, the three signals of observability, the retry patterns that cause outages rather than preventing them, and shipping a schema change without downtime.
Free
5 modules · 18h · 25 lessons · Advanced
Failure Modes & Blast Radius
4 hours
5 lessons
Advanced
Failure Modes & Blast Radius
Free
06.1
SLOs & Error Budgets
3 hours
5 lessons
Advanced
SLOs & Error Budgets
Free
06.2
Observability
4 hours
5 lessons
Advanced
Observability
Free
06.3
Timeouts, Retries & Circuit Breakers
3 hours
5 lessons
Advanced
Timeouts, Retries & Circuit Breakers
Free
06.4
Deployment, Migration & Rollback
4 hours
5 lessons
Advanced
Deployment, Migration & Rollback
Free
06.5
The drills in this track
06.1 · Blast Radius

One dependency goes down. Trace what it takes with it before the pager does.

06.2 · Budget

Spend a quarter of error budget across features, incidents and migrations without running out in month two.

06.3 · Signal

Latency has tripled and nothing is down. Find the cause with the telemetry you already have.

06.4 · Retry Storm

A retry policy is amplifying a small failure into an outage. Fix it without making the service less reliable.

06.5 · Expand/Contract

Rename a column on a 900 million row table with zero downtime and a working rollback at every step.

Every track is free. Start anywhere.