Site Reliability Engineer

Will Pan

I run Kubernetes platforms on AWS and GCP, keep infrastructure in code, and follow incidents through to root cause. Measure first. Then change one thing.

  1. 2020 – 2022Hexiang Digital TechnologyOperations Engineer
  2. 2022 – 2023Zuopin TechnologyOperations Engineer
  3. 2023 – 2024Palup.aiSite Reliability Engineer
  4. 2024 – 2025DBS BankSite Reliability Engineer
  5. 2025 – presentTRON DAOSenior Site Reliability Engineer

Latest writing

All posts →

Nothing published yet. The first posts are being written.

    Focus areas

    Work history →

    Reliability and incident response

    Rollouts that do not drop requests, incidents followed through to root cause, and every mitigation committed to Git.

    Kubernetes platforms

    EKS and GKE: version upgrades, autoscaling, disruption budgets, network policy and admission guardrails.

    Infrastructure as code and delivery

    Terraform and OpenTofu with a plan on every merge request, drift detection, and GitOps delivery through Argo CD.

    Observability

    Metrics, logs and traces that connect to each other, plus monitoring for the monitoring stack itself.

    Tech stack

    Kubernetes

    • EKS
    • GKE
    • Helm
    • Argo CD
    • Cluster Autoscaler
    • External Secrets

    IaC and automation

    • Terraform
    • OpenTofu
    • Ansible
    • Python
    • Shell
    • Go

    Observability

    • Prometheus
    • Grafana
    • Loki
    • Tempo
    • OpenTelemetry

    Cloud

    • AWS
    • GCP
    • Cloudflare
    • PostgreSQL
    • Redis