Work
Work history
Operations and reliability work since 2020, most recent first.
- 2025 – present
TRON DAO
Senior Site Reliability Engineer
SRE for two product lines: a live-streaming platform and an AI API and chat platform on EKS.
- Made API rollouts safe behind an AWS load balancer with probes, readiness gates, preStop hooks and graceful shutdown; 5XX errors went from 2,000–3,000 per release to about ten in a typical release.
- Brought 50+ existing AWS components under OpenTofu by import with no production impact, then added a plan on every merge request, daily drift detection and a security-scan gate.
- Root-caused a Prometheus out-of-memory failure to a single unbounded label (about 2 million series) and replaced it with a bounded label and recording rules.
- Traced a production Node.js memory leak from an HPA alert to unbounded retries after upstream HTTP 429 responses; p99 latency, which had risen to 4.4 s, was back to 0.65 s after mitigation.
- Ran the database side of converting two billing tables (1.46 TB, 690 million rows) to time partitioning under production traffic.
- Moved a staging database off a shared production PostgreSQL instance that could not be restarted, using trigger-based change data capture: no write freeze, and every table verified by checksum.
- Introduced Cluster Autoscaler and 18 production PodDisruptionBudgets and consolidated node groups into three pools (workload, CI and platform), which released idle nodes and cut node churn.
- Built a parallel EKS 1.34 cluster in Terraform as the upgrade path for staging, and performed in-place upgrades of Argo CD, Loki and kube-state-metrics with live validation.
- Operate 20+ platform Helm charts through Argo CD (kube-prometheus-stack, Loki, Tempo, OpenTelemetry Collector, cluster-autoscaler), with CI render checks on every merge request.
- Built a meta-monitor in Python on AWS Lambda, independent of Prometheus and Grafana, to watch the monitoring stack itself, and deployed a Tempo trace backend so engineers can jump from a Loki log line to its trace.
- Proposed and delivered a multi-domain rollout for a live-streaming site after ISP-level DNS interference cut users off from the primary domain: CloudFront distributions, certificates and DNS in Terraform, staging first, then production.
- Owned the infrastructure side of the ordered shutdown of a live-streaming platform (EKS, RDS, ElastiCache, MongoDB, OpenSearch, MSK): staging before production, with a snapshot or cold backup before every deletion.
- 2024 – 2025
DBS Bank
Site Reliability Engineer
- Stabilized shell-based deployment workflows in a regulated banking environment and added version traceability.
- Implemented compliant Nginx gateway configuration for third-party API integration.
- 2023 – 2024
Palup.ai
Site Reliability Engineer
Infrastructure for an AI product on GKE, working directly with the ML teams.
- Built GKE clusters and GPU node pools in Terraform for ML teams, including NVIDIA H100 pools on Spot VMs (8 × H100 80 GB per node).
- Worked with ML teams on GPU sharing (MIG, then time-slicing) and on model storage through the Cloud Storage FUSE CSI driver.
- Rolled out platform components to new clusters through Argo CD ApplicationSets and built the shared CUDA base image for GPU services.
- Replaced the production n2-standard-32 node pool with n2-standard-8 and n1-standard-4 pools through a two-step replacement.
- 2022 – 2023
Zuopin Technology
Operations Engineer
- Maintained GitLab and Jenkins pipelines and Shell automation, and deployed services from binaries.
- Built Prometheus and ELK monitoring and log collection.
- 2020 – 2022
Hexiang Digital Technology
Operations Engineer
- Operated 200+ Linux VMs across AWS, GCP, Alibaba Cloud, Tencent Cloud and other providers: Nginx, certificates and host monitoring.