SherlockLiu Logo SherlockLiu
Observability Engineering
Engineering · Books

Observability Engineering

A 10-part, plain-English field guide to Observability Engineering (2nd Edition) by Charity Majors, Liz Fong-Jones, George Miranda, and Austin Parker — rewritten chapter by chapter, with original diagrams and no jargon left unexplained.

Series progress
10 / 10 parts
Start Reading

All Parts

1

Observability Engineering — Part 1: What Observability Actually Means

Jul 17, 2026 9 min

Observability has a precise, decades-old technical definition that most 'observability platforms' quietly fail to meet — here's what it actually requires.

2

Observability Engineering — Part 2: The Art of Instrumentation

Jul 18, 2026 14 min

Why the shape of the data you collect today quietly decides which questions you're allowed to ask tomorrow — and how OpenTelemetry lets you avoid that trap.

3

Observability Engineering — Part 3: From Data to Insights

Jul 19, 2026 9 min

Collecting telemetry is only step one — here's how to actually investigate it, build it into your daily workflow, use AI agents wisely, and alert on what users really feel.

4

Observability Engineering — Part 4: SLO-Based Alerts and Storage That Scales

Jul 20, 2026 10 min

How error-budget math actually decides when to page you, and why observability data needs a database that plays by different rules than metrics ever did.

5

Observability Engineering — Part 5: Sampling, Pipelines, and a Shared Language

Jul 21, 2026 13 min

How engineering teams keep observability affordable at scale — through smarter sampling, telemetry pipelines, and a shared vocabulary that keeps humans and AI agents honest.

6

Observability Engineering — Part 6: Observability Across CI/CD, Mobile, and Performance

Jul 22, 2026 13 min

Your build pipeline, your mobile app, and your cloud bill are all systems worth debugging the same rigorous way you'd debug a production outage.

7

Observability Engineering — Part 7: LLMs, AI Agents, and the Fin Story

Jul 23, 2026 9 min

How production telemetry and evaluations form a learning flywheel for LLM applications — illustrated by Fin's real turnaround in speed, cost, and reliability.

8

Observability Engineering — Part 8: Why Learning Speed Is Your Biggest Bottleneck

Jul 24, 2026 9 min

Why the real constraint on engineering teams in the AI era isn't how fast they can write code, but how fast they can understand what that code just did.

9

Observability Engineering — Part 9: Making the Business Case and Driving Change

Jul 25, 2026 10 min

How to fund observability like a strategic bet instead of a cost center, tell whether the money is actually working, and push real change past organizational resistance.

10

Observability Engineering — Part 10: Build vs. Buy, Vendor Partnerships, and What's Next

Jul 26, 2026 10 min

The book's closing playbook: a clear-eyed framework for build vs. buy vs. open source, how to run a vendor trial that actually proves something, and where observability goes next.