Read in order

Guides & series

Connected, multi-part deep dives. Each guide takes one hard problem and works it end to end.

12-part series

Distributed Systems Patterns That Hold Up in Production

The recurring patterns behind reliable distributed systems — event replay, idempotency, sharding, timeouts, and backpressure — explained from production experience.

Start reading

9-part series

Kubernetes Operations for Production Platforms

Running Kubernetes for real workloads — multi-tenancy and namespace strategy, health checks that tell the truth, and the operational decisions that keep platforms up.

Start reading

7-part series

Observability for Distributed Systems

Seeing inside a distributed system — distributed tracing that finds the slow hop, and dashboards that speed up incident response instead of decorating a review.

Start reading

6-part series

Microservice Service Design: Boundaries That Hold

Where a service should begin and end — drawing ownership boundaries and bounded contexts that let teams ship independently instead of building a distributed monolith.

Start reading

6-part series

Polyglot Microservices: Choosing the Right Language

Go, Rust, Python, and gRPC across a polyglot fleet — when each language earns its place in production, and where the boundaries between them break.

Start reading

4-part series

Staff Engineer Craft: Design, Influence, and Learning

The non-code skills that define senior and staff engineers — system design interviews, RFCs that survive review, and blameless postmortems that actually fix systems.

Start reading

1-part series

Build Systems and Developer Infrastructure at Scale

The build and CI decisions that decide how fast a large engineering org ships — monorepo build graphs, caching, hermeticity, and when the rigor pays off.

Start reading