AI Infrastructure · GPU Programming · LLM Inference
I'm an undergraduate focused on the systems behind efficient AI: from GPU kernels and hardware backends to inference engines and serving infrastructure.
I care about understanding how these systems work end to end, then turning that understanding into implementations that are measurable, reproducible, and useful.
CUDA Triton TileLang C++ Python
- GPU kernels: operator implementation, correctness, profiling, and optimization
- LLM inference: KV cache, scheduling, continuous batching, and memory management
- Inference infrastructure: hardware backends, framework integration, and serving systems
- Exploring reproducible autotuning workflows for GPU kernels
- Building inference components from kernels upward to understand the full stack
Performance engineering, systems programming, developer tooling, and the boundary between machine learning frameworks and accelerator hardware.

