SherlockLiu Logo SherlockLiu
The SRE Book, Ten Years On
Engineering · Books

The SRE Book, Ten Years On

A 12-part guide to site reliability engineering, ten years after Google's SRE book set the discipline down in writing — what still holds, what's changed, and what the practice looks like in 2026, enriched with modern incidents and tools.

Series progress
12 / 12 parts
Start Reading

All Parts

1

The SRE Book, Ten Years On — Part 1: Hope Is Not a Strategy

Aug 26, 2026 17 min

Google's SRE book invented the error budget and blameless postmortem — most quote it unread. Part 1: what SRE is, and why the sysadmin model was doomed.

2

The SRE Book, Ten Years On — Part 2: 100% Is the Wrong Target

Aug 27, 2026 16 min

SRE's key idea: 100% availability is the wrong target, and the error budget turns that into an incentive system. Part 2, with the burn graph explained.

3

The SRE Book, Ten Years On — Part 3: The War on Toil

Aug 28, 2026 22 min

Five principles for actually getting toil down, each with a real story — the automation ladder, the competence/latency/relevance tradeoff, and Diskerase.

4

The SRE Book, Ten Years On — Part 4: Four Signals and a Pager

Aug 29, 2026 15 min

Latency, traffic, errors, saturation — and the rule that matters most: every page must be actionable. Part 4, with Borgmon's lineage to Prometheus.

5

The SRE Book, Ten Years On — Part 5: On-Call and the 25% Rule

Aug 30, 2026 14 min

The pager arithmetic behind on-call: 25% cap, two incidents per shift, and why operational underload is as dangerous as overload. Part 5, ten years on.

6

The SRE Book, Ten Years On — Part 6: How to Fix a Broken System

Aug 31, 2026 16 min

The hypothetico-deductive loop, 'triage first, root-cause later', and why negative results are magic. Part 6 — plus Meta's 2021 outage as the modern Diskerase.

7

The SRE Book, Ten Years On — Part 7: Run the Incident, Don't Let It Run You

Sep 01, 2026 14 min

ICS roles, the incident-state doc, the blameless postmortem, and the data layer that turns outages into improvement. Part 7, plus the incident.io ecosystem.

8

The SRE Book, Ten Years On — Part 8: Why Systems Fall Over

Sep 02, 2026 13 min

How load balancing works, why 'queries per second' is a terrible metric, and the cascade loop that kills distributed systems. Part 8, with Fastly's 2021 outage.

9

The SRE Book, Ten Years On — Part 9: Agreeing to Agree

Sep 03, 2026 23 min

See leader election, think distributed consensus — anything less is a ticking bomb. Part 9: quorums, split brains, CAP, and the double-launching cron.

10

The SRE Book, Ten Years On — Part 10: Never Trust a Backup You Haven't Restored

Sep 04, 2026 18 min

Pipelines are distributed systems with a moiré problem; backups are a tax and restores are the product. Part 10 — with GitLab 2017 and Atlassian 2022.

11

The SRE Book, Ten Years On — Part 11: Ship It Without Sinking It

Sep 05, 2026 18 min

Hermetic builds, canarying, the launch checklist's one brutal rule — and CrowdStrike 2024 as the lesson in skipping validation. Part 11.

12

The SRE Book, Ten Years On — Part 12: The Cockpit

Sep 06, 2026 17 min

The series finale: how SRE's management practices invented platform engineering a decade early. Training, interrupts, culture — and what the data says now.