聚焦 AI 基础设施、CUDA Kernel 与高性能系统工程
🔬 Focus: AI Infrastructure · CUDA Kernels · LLM Inference · HPC Systems
🌱 Currently: Building high-throughput inference pipelines and GPU-first systems
🤝 Open to: AI infrastructure, performance engineering, research collaboration, and open-source collaboration
I build AI infrastructure and GPU-first high-performance systems with C++/CUDA, Python, and Go. 主要聚焦 AI 基础设施、GPU 算子优化与高性能系统工程实践。
- 🔥 GPU Kernel Engineering — CUDA/Triton kernels for FlashAttention, GEMM, quantization, and memory-aware operator design
GPU 算子工程 — FlashAttention、GEMM、量化与内存感知算子设计 - 🧠 AI Inference Systems — lightweight LLM runtimes, KV Cache, W8A16/FP8 quantization, and inference path optimization
AI 推理系统 — 轻量 LLM 运行时、KV Cache、量化方案与推理路径优化 - ⚡ High-Performance Computing — simulation, rendering, and image-processing pipelines tuned for throughput and scalability
高性能计算 — 面向吞吐与可扩展性的仿真、渲染与图像处理流水线 - 🌐 Real-time Systems — RTC signaling, streaming applications, and digital human platforms with system-level integration
实时系统 — RTC 信令、流媒体应用与数字人平台的系统级集成
Currently / 当前关注: inference acceleration, kernel fusion, and end-to-end GPU system design.
推理加速、算子融合与端到端 GPU 系统设计。
Featured Projects / 核心项目 — Start here for the quickest overview of my work in bioinformatics, HPC, AI inference, and developer tooling.
如果你想快速判断我的技术重心与代表作,建议先看下面 4 个项目。
Best entry points for collaboration, hiring conversations, and technical review.
|
High-performance FASTQ compression with 3.97x ratio and O(1) random access. C++23, ABC+SCM algorithms. |
High-performance FASTQ QC toolkit (stat/filter/trim); zero-copy I/O, TBB pipeline, C++23. |
|
End-to-end Metagenomic Intelligence and Comprehensive Omics Suite (Mammoth Cup 2024) |
Systematic knowledge base for bioinformatics (Chinese community) |
Original, sole-authored engineering portfolio — from CUDA kernels to a working inference engine, with tests, benchmarks & honest measurement records.
从 CUDA 算子到可运行推理引擎的个人原创作品集(含测试、基准与诚实记录的测量口径)。入口:open-infra-ai/aicl-lab
|
CUDA operator engineering path: SGEMM ladder to reusable inference components |
FlashAttention fwd/bwd from scratch in CUDA C++ (FP16/BF16 WMMA) |
Triton kernel library (RMSNorm+RoPE / SwiGLU / FlashAttn / SGEMM) + torch.library registration |
|
CUDA-native C++ inference engine: GGUF loading, W8A16 quantization, paged-KV strategy; 170 tests |
PagedAttention-style paged KV + continuous batching control plane (Rust), e2e-verified vs llama.cpp |
Portfolio landing page: evidence pack, benchmarks methodology & interview storytelling |
|
12-week AI Infra transition plan (2026-08-24 ~ 2026-11-15): skill matrix, weekly plans, interview matrix & application pipeline. |
Standalone navigation center: full repo inventory, fork/upstream audit, migration log & deep-dive guides. |
|
Systematic knowledge base for bioinformatics (Chinese community) |
End-to-end Metagenomic Intelligence and Comprehensive Omics Suite |
|
High-performance FASTQ compression with 3.97x ratio and O(1) random access. C++23, ABC+SCM. |
High-performance FASTQ QC toolkit (stat/filter/trim); zero-copy I/O, TBB pipeline, C++23. |
|
Curated bioinformatics algorithms knowledge base with complexity analysis, CLI tools, and bilingual docs. |
|
CUDA image processing experiments: parallel filters and pipelines on GPU. |
High-performance C++ optimization guide with lock-free data structures, SIMD, and memory optimization. |
|
Header-only C++23 bit manipulation library with SIMD acceleration (SSE2/AVX2/AVX-512/NEON). |
Classic lossless compression algorithms in C++17, Go, and Rust with cross-language binary verification. |
|
My solutions to TensorTonic problems (tensor/GPU computing drills). |
Compression Knowledge Base: Algorithm Theory, Performance Benchmarks & C++ Examples |
|
Cursor AI 编程规则精选集 | 132+ 规则,覆盖前端/后端/AI/DevOps 等 32 个领域 |
Local-first bookmark manager with structured storage and import/export. |
|
Notes syncing utility built around Brave browser data. |
Offline-first bookmark cleaner: rules-first, ML-assisted, LLM-optional |
|
Multi-Model Real-Time Visual Recognition System with REST API and WebSocket Streaming |
Privacy-first diagram editor with local WASM rendering, Kroki full mode, sharing, and export. |
|
Browser-native 3D digital human engine with voice, vision & dialogue. Zero-config, offline-ready. |
Lightweight WebRTC Demo: Go Signaling Server + Vanilla JavaScript Client, OpenSpec-Driven |
|
Browser-based memory training PWA with FSRS-4.5 spaced repetition, N-back training, and adaptive difficulty |
|
Computer Science related background. / 计算机科学相关背景 |
Engineering across medical imaging, RTC systems, and genomic-scale data workflows. / 覆盖医疗影像、实时音视频系统与基因数据工程。 |
聚焦与核心项目强相关的技术:AI Infrastructure · CUDA Kernel Engineering · LLM Inference · HPC Systems
| Category | Technologies |
|---|---|
| AI Infrastructure | |
| CUDA Kernel Engineering | |
| LLM Inference Optimization | |
| HPC Performance Engineering |
Reach out if you're building AI infrastructure, inference acceleration, GPU systems, or performance-critical tooling.
欢迎联系我交流 AI 基础设施、推理加速、GPU 系统,以及对性能敏感的工程项目。
Open to technical collaboration, engineering roles, research discussions, and thoughtful open-source work.




