Observability Engineering
A 10-part, plain-English field guide to Observability Engineering (2nd Edition) by Charity Majors, Liz Fong-Jones, George Miranda, and Austin Parker — rewritten chapter by chapter, with original diagrams and no jargon left unexplained.
All Parts
Observability Engineering — Part 1: What Observability Actually Means
Observability has a precise, decades-old technical definition that most 'observability platforms' quietly fail to meet — here's what it actually requires.
Observability Engineering — Part 2: The Art of Instrumentation
Why the shape of the data you collect today quietly decides which questions you're allowed to ask tomorrow — and how OpenTelemetry lets you avoid that trap.
Observability Engineering — Part 3: From Data to Insights
Collecting telemetry is only step one — here's how to actually investigate it, build it into your daily workflow, use AI agents wisely, and alert on what users really feel.
Observability Engineering — Part 4: SLO-Based Alerts and Storage That Scales
How error-budget math actually decides when to page you, and why observability data needs a database that plays by different rules than metrics ever did.
Observability Engineering — Part 5: Sampling, Pipelines, and a Shared Language
How engineering teams keep observability affordable at scale — through smarter sampling, telemetry pipelines, and a shared vocabulary that keeps humans and AI agents honest.
Observability Engineering — Part 6: Observability Across CI/CD, Mobile, and Performance
Your build pipeline, your mobile app, and your cloud bill are all systems worth debugging the same rigorous way you'd debug a production outage.
Observability Engineering — Part 7: LLMs, AI Agents, and the Fin Story
How production telemetry and evaluations form a learning flywheel for LLM applications — illustrated by Fin's real turnaround in speed, cost, and reliability.
Observability Engineering — Part 8: Why Learning Speed Is Your Biggest Bottleneck
Why the real constraint on engineering teams in the AI era isn't how fast they can write code, but how fast they can understand what that code just did.
Observability Engineering — Part 9: Making the Business Case and Driving Change
How to fund observability like a strategic bet instead of a cost center, tell whether the money is actually working, and push real change past organizational resistance.
Observability Engineering — Part 10: Build vs. Buy, Vendor Partnerships, and What's Next
The book's closing playbook: a clear-eyed framework for build vs. buy vs. open source, how to run a vendor trial that actually proves something, and where observability goes next.