Blog
Writing
Notes on reliability work: what broke, what was measured, and what changed. Each post is written so that I can still follow the reasoning a year later.
Nothing published yet. The first posts are being written.
In progress
- One unbounded label was enough to run Prometheus out of memory
- Why Kubernetes rollouts behind an ALB drop requests, and what fixed it
- Three ways to share an H100 on GKE: MIG, MPS and time-slicing