Skip to content
View chethanuk's full-sized avatar
🌴
On Vacation — Back soon (No AI plumbing ATM)
🌴
On Vacation — Back soon (No AI plumbing ATM)

Organizations

@trinodb

Block or report chethanuk

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
chethanuk/README.md

Typing SVG

🔭 Staff AI Infrastructure Engineer | LLM Inference, GPU Platforms & Distributed Data Infrastructure

I build and scale production systems for LLM inference, GPU orchestration, MLOps, and distributed data platforms on Kubernetes (GCP & AWS).

Recently, I've been shipping Kubernetes-native LLM inference (vLLM, SGLang, llm-d), GPU sharing and governance (HAMi, KAI Scheduler), GPU-to-token observability, and streaming backbones (Kafka, Flink, Spark, Iceberg) serving 10M+ users and 700M+ daily events.

LinkedIn Email GitHub

github contribution grid snake animation

Open Source Contributions

300+ merged PRs across 25+ organizations · AI Inference, Agentic Frameworks, AI Data Curation & Distributed Data Infrastructure
  • Inference Routing & GPU Scheduling
    llm-d llm-d router vLLM AIBrix vLLM Semantic Router Gateway API Inference Extension Mooncake prime-rl

  • AI / ML & Agent Runtimes
    OpenViking.AI DB Alibaba OpenCodeReview Google ADK-Go Google ADK-Python Hindsight Honcho MCPGateway AutoFlow.AI Swarms OpenWebUIMCPO

  • AI Data Curation & Synthetic Data
    NeMo Curator NeMo DataDesigner

  • Big Data and Data Frameworks
    Apache Airflow Apache Pinot Apache Beam Flink K8s Operator dbt

  • Data Infrastructure
    Kubeflow Spark Operator Trino KubeFlow ZenML KONG DAPR SDKMAN

  • Cloud
    Google Tunix Data on EKS python-deequ Google Cloud Dataproc

What I've Been Doing Recently ⚙️

  • Kubernetes LLM Inference: Running vLLM, SGLang, and llm-d on GKE and EKS with the Gateway API Inference Extension, prefix-cache routing, and disaggregated prefill/decode.
  • GPU Scheduling & Sharing (HAMi & KAI Scheduler): Multi-tenant GPU sharing with hard VRAM limits, dynamic resource allocation (DRA), and fractional GPU slicing.
  • Data Platforms & Streaming: Operating Kafka, Flink, Spark, and Iceberg pipelines (up to 700M events/day), with Airflow orchestration and ClickHouse for low-latency queries.
  • Agent Memory & Runtime Systems (OpenViking & Honcho): Enforcing token-budget limits and hybrid retrieval (pgvector, BM25) to keep agent contexts bounded and retrieval latency low.
  • Reliable LLM Inference Stack (3+ Years Production Experience): Owned AI infrastructure for Goodnotes 6 AI features (AI typing and handwriting features) serving 21M+ monthly users. Shipped one of the earliest production LLM stacks in 2023 with vLLM 0.2 (PagedAttention) on EKS GPU clusters, including an inference gateway, caching layer (KV + Redis), scheduler, routing, AI gateway (Envoy + Gloo Gateway), and AI safety, reducing latency and cutting costs by 97% versus other SaaS options (Read -> AWS case study).
  • GPU-to-Token Observability: Tracking inference end to end across hardware metrics (DCGM), KV cache pressure, token latency (TTFT, ITL), and per-tenant cost attribution using FOCUS 1.0.

GitHub Streak

Technical Skills & Platforms

Pinned Loading

  1. Computer-Vision---Facial-Keypoint-Detection Computer-Vision---Facial-Keypoint-Detection Public

    Computer Vision - Facial Keypoint Detection

    HTML

  2. Lane-Finding-using-Computer-Vision Lane-Finding-using-Computer-Vision Public

    Lane Finding using Computer Vision

    HTML 1