Reliability & Operations





One dependency goes down. Trace what it takes with it before the pager does.
Spend a quarter of error budget across features, incidents and migrations without running out in month two.
Latency has tripled and nothing is down. Find the cause with the telemetry you already have.
A retry policy is amplifying a small failure into an outage. Fix it without making the service less reliable.
Rename a column on a 900 million row table with zero downtime and a working rollback at every step.
The other tracks
The machine and the network under every diagram · 5 modules · 23h
Where the state lives, and what it costs to keep it · 6 modules · 30h
Making the second request cheap · 4 modules · 16h
Decoupling two services without losing a message · 5 modules · 21h
From one box to a fleet, without a rewrite · 5 modules · 17h
The six systems you will actually be asked to draw · 6 modules · 22h
Running the forty-five minutes on purpose · 5 modules · 12h