Performance portable parallel programming in Python backed by Kokkos
-
Updated
Oct 9, 2026 - Python
Performance portable parallel programming in Python backed by Kokkos
CG-Kit: Code Generation tool-Kit for algorithmic and performance portability
Measuring what a GPU DSL forecloses, at a fixed algorithm: CUDA vs Triton vs NVIDIA Warp vs Gluon on NFA/DFA simulation. Artifact for HPEC 2026.
Does an attention backend chosen on one GPU stay the right choice on another? A measurement study over six NVIDIA GPUs, forward prefill and KV-cache decode.
To associate your repository with the performance-portability topic, visit your repo's landing page and select "manage topics."