Site Reliability Engineer
Will Pan
I run Kubernetes platforms on AWS and GCP, keep infrastructure in code, and follow incidents through to root cause. Measure first. Then change one thing.
- 2020 – 2022Hexiang Digital TechnologyOperations Engineer
- 2022 – 2023Zuopin TechnologyOperations Engineer
- 2023 – 2024Palup.aiSite Reliability Engineer
- 2024 – 2025DBS BankSite Reliability Engineer
- 2025 – presentTRON DAOSenior Site Reliability Engineer
Latest writing
All posts →Nothing published yet. The first posts are being written.
Focus areas
Work history →Reliability and incident response
Rollouts that do not drop requests, incidents followed through to root cause, and every mitigation committed to Git.
Kubernetes platforms
EKS and GKE: version upgrades, autoscaling, disruption budgets, network policy and admission guardrails.
Infrastructure as code and delivery
Terraform and OpenTofu with a plan on every merge request, drift detection, and GitOps delivery through Argo CD.
Observability
Metrics, logs and traces that connect to each other, plus monitoring for the monitoring stack itself.
Tech stack
Kubernetes
- EKS
- GKE
- Helm
- Argo CD
- Cluster Autoscaler
- External Secrets
IaC and automation
- Terraform
- OpenTofu
- Ansible
- Python
- Shell
- Go
Observability
- Prometheus
- Grafana
- Loki
- Tempo
- OpenTelemetry
Cloud
- AWS
- GCP
- Cloudflare
- PostgreSQL
- Redis