Read in order
Guides & series
Connected, multi-part deep dives. Each guide takes one hard problem and works it end to end.
12-part series
Distributed Systems Patterns That Hold Up in Production
The recurring patterns behind reliable distributed systems — event replay, idempotency, sharding, timeouts, and backpressure — explained from production experience.
Start reading9-part series
Kubernetes Operations for Production Platforms
Running Kubernetes for real workloads — multi-tenancy and namespace strategy, health checks that tell the truth, and the operational decisions that keep platforms up.
Start reading7-part series
Observability for Distributed Systems
Seeing inside a distributed system — distributed tracing that finds the slow hop, and dashboards that speed up incident response instead of decorating a review.
Start reading6-part series
Microservice Service Design: Boundaries That Hold
Where a service should begin and end — drawing ownership boundaries and bounded contexts that let teams ship independently instead of building a distributed monolith.
Start reading6-part series
Polyglot Microservices: Choosing the Right Language
Go, Rust, Python, and gRPC across a polyglot fleet — when each language earns its place in production, and where the boundaries between them break.
Start reading4-part series
Staff Engineer Craft: Design, Influence, and Learning
The non-code skills that define senior and staff engineers — system design interviews, RFCs that survive review, and blameless postmortems that actually fix systems.
Start reading1-part series
Build Systems and Developer Infrastructure at Scale
The build and CI decisions that decide how fast a large engineering org ships — monorepo build graphs, caching, hermeticity, and when the rigor pays off.
Start reading