🔭 Staff AI Infrastructure Engineer | LLM Inference, GPU Platforms & Distributed Data Infrastructure
I build and scale production systems for LLM inference, GPU orchestration, MLOps, and distributed data platforms on Kubernetes (GCP & AWS).
Recently, I've been shipping Kubernetes-native LLM inference (vLLM, SGLang, llm-d), GPU sharing and governance (HAMi, KAI Scheduler), GPU-to-token observability, and streaming backbones (Kafka, Flink, Spark, Iceberg) serving 10M+ users and 700M+ daily events.
300+ merged PRs across 25+ organizations · AI Inference, Agentic Frameworks, AI Data Curation & Distributed Data Infrastructure
- Kubernetes LLM Inference: Running vLLM, SGLang, and llm-d on GKE and EKS with the Gateway API Inference Extension, prefix-cache routing, and disaggregated prefill/decode.
- GPU Scheduling & Sharing (HAMi & KAI Scheduler): Multi-tenant GPU sharing with hard VRAM limits, dynamic resource allocation (DRA), and fractional GPU slicing.
- Data Platforms & Streaming: Operating Kafka, Flink, Spark, and Iceberg pipelines (up to 700M events/day), with Airflow orchestration and ClickHouse for low-latency queries.
- Agent Memory & Runtime Systems (OpenViking & Honcho): Enforcing token-budget limits and hybrid retrieval (pgvector, BM25) to keep agent contexts bounded and retrieval latency low.
- Reliable LLM Inference Stack (3+ Years Production Experience): Owned AI infrastructure for Goodnotes 6 AI features (AI typing and handwriting features) serving 21M+ monthly users. Shipped one of the earliest production LLM stacks in 2023 with vLLM 0.2 (PagedAttention) on EKS GPU clusters, including an inference gateway, caching layer (KV + Redis), scheduler, routing, AI gateway (Envoy + Gloo Gateway), and AI safety, reducing latency and cutting costs by 97% versus other SaaS options (Read -> AWS case study).
- GPU-to-Token Observability: Tracking inference end to end across hardware metrics (DCGM), KV cache pressure, token latency (TTFT, ITL), and per-tenant cost attribution using FOCUS 1.0.
- AI Infrastructure & LLM Inference:
vLLM · SGLang · llm-d · Ray Serve · KServe · HAMi / HAMi-core · NVIDIA KAI Scheduler · Kubernetes Gateway API Inference Extension · PagedAttention · Prefix-Cache Routing · Disaggregated Prefill/Decode · RoCEv2 / InfiniBand RDMA - GPU Platforms & MLOps:
Kubernetes (EKS / GKE) · Karpenter · Kueue · Karmada · KubeRay · JobSet · LeaderWorkerSet (LWS) · Dynamic Resource Allocation (DRA v1.31) · NVIDIA DCGM · PyTorch · NCCL · MLflow · SageMaker · Terraform · Helm - Distributed Data & Data Engineering:
Apache Kafka · Apache Flink · Apache Spark · Apache Iceberg · Apache Airflow (14K+ DAG runs/day) · NVIDIA NeMo Curator · NeMo DataDesigner · Trino/Presto · ClickHouse · dbt · Great Expectations · Feast · Databricks · Snowflake · VectorChord / pgvector · Redis - Observability & Economics:
GPU-to-Token (8-Layer Stack) · OpenTelemetry (OTel) · Prometheus & Grafana · FOCUS 1.0 / OpenCost · Arize Phoenix - Languages:
Python · Go · Rust · SQL · Scala




