Work

Work history

Operations and reliability work since 2020, most recent first.

  1. 2025 – present

    TRON DAO

    Senior Site Reliability Engineer

    SRE for two product lines: a live-streaming platform and an AI API and chat platform on EKS.

    • Made API rollouts safe behind an AWS load balancer with probes, readiness gates, preStop hooks and graceful shutdown; 5XX errors went from 2,000–3,000 per release to about ten in a typical release.
    • Brought 50+ existing AWS components under OpenTofu by import with no production impact, then added a plan on every merge request, daily drift detection and a security-scan gate.
    • Root-caused a Prometheus out-of-memory failure to a single unbounded label (about 2 million series) and replaced it with a bounded label and recording rules.
    • Traced a production Node.js memory leak from an HPA alert to unbounded retries after upstream HTTP 429 responses; p99 latency, which had risen to 4.4 s, was back to 0.65 s after mitigation.
    • Ran the database side of converting two billing tables (1.46 TB, 690 million rows) to time partitioning under production traffic.
    • Moved a staging database off a shared production PostgreSQL instance that could not be restarted, using trigger-based change data capture: no write freeze, and every table verified by checksum.
    • Introduced Cluster Autoscaler and 18 production PodDisruptionBudgets and consolidated node groups into three pools (workload, CI and platform), which released idle nodes and cut node churn.
    • Built a parallel EKS 1.34 cluster in Terraform as the upgrade path for staging, and performed in-place upgrades of Argo CD, Loki and kube-state-metrics with live validation.
    • Operate 20+ platform Helm charts through Argo CD (kube-prometheus-stack, Loki, Tempo, OpenTelemetry Collector, cluster-autoscaler), with CI render checks on every merge request.
    • Built a meta-monitor in Python on AWS Lambda, independent of Prometheus and Grafana, to watch the monitoring stack itself, and deployed a Tempo trace backend so engineers can jump from a Loki log line to its trace.
    • Proposed and delivered a multi-domain rollout for a live-streaming site after ISP-level DNS interference cut users off from the primary domain: CloudFront distributions, certificates and DNS in Terraform, staging first, then production.
    • Owned the infrastructure side of the ordered shutdown of a live-streaming platform (EKS, RDS, ElastiCache, MongoDB, OpenSearch, MSK): staging before production, with a snapshot or cold backup before every deletion.
  2. 2024 – 2025

    DBS Bank

    Site Reliability Engineer

    • Stabilized shell-based deployment workflows in a regulated banking environment and added version traceability.
    • Implemented compliant Nginx gateway configuration for third-party API integration.
  3. 2023 – 2024

    Palup.ai

    Site Reliability Engineer

    Infrastructure for an AI product on GKE, working directly with the ML teams.

    • Built GKE clusters and GPU node pools in Terraform for ML teams, including NVIDIA H100 pools on Spot VMs (8 × H100 80 GB per node).
    • Worked with ML teams on GPU sharing (MIG, then time-slicing) and on model storage through the Cloud Storage FUSE CSI driver.
    • Rolled out platform components to new clusters through Argo CD ApplicationSets and built the shared CUDA base image for GPU services.
    • Replaced the production n2-standard-32 node pool with n2-standard-8 and n1-standard-4 pools through a two-step replacement.
  4. 2022 – 2023

    Zuopin Technology

    Operations Engineer

    • Maintained GitLab and Jenkins pipelines and Shell automation, and deployed services from binaries.
    • Built Prometheus and ELK monitoring and log collection.
  5. 2020 – 2022

    Hexiang Digital Technology

    Operations Engineer

    • Operated 200+ Linux VMs across AWS, GCP, Alibaba Cloud, Tencent Cloud and other providers: Nginx, certificates and host monitoring.