Blog

Writing

Notes on reliability work: what broke, what was measured, and what changed. Each post is written so that I can still follow the reasoning a year later.

Nothing published yet. The first posts are being written.

In progress

  • One unbounded label was enough to run Prometheus out of memory
  • Why Kubernetes rollouts behind an ALB drop requests, and what fixed it
  • Three ways to share an H100 on GKE: MIG, MPS and time-slicing