Skip to content
View ColdSlither's full-sized avatar

Block or report ColdSlither

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Stars

GPU ish

15 repositories

GPU Memory Reservation Library

C++ 55 26 Updated Jul 17, 2026

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Rust 183 26 Updated Jul 22, 2026

Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.

Python 1,475 152 Updated Mar 19, 2026

A Fusion Code Generator for NVIDIA GPUs (commonly known as "nvFuser")

C++ 396 82 Updated May 31, 2026

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance…

Python 3,436 777 Updated Jul 21, 2026

NVIDIA Inference Xfer Library (NIXL)

C++ 1,143 373 Updated Jul 22, 2026

Library for reducing tail latency in RAM reads

C++ 2,774 158 Updated Apr 11, 2026

A minimal GPU design in Verilog to learn how GPUs work from the ground up

SystemVerilog 12,759 1,222 Updated Aug 18, 2024

Official Kaggle CLI

Python 7,466 1,393 Updated Jul 21, 2026

A runtime substrate that turns an agent's execution into a reversible, Git-like trace, so meta-agents can observe, fork, replay, and revert any run. Couples agent and environments in a copy-on-writ…

Python 1,522 119 Updated Jul 21, 2026

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

C 17,840 1,706 Updated Jul 22, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 18,912 1,476 Updated Jul 20, 2026
JavaScript 19 13 Updated Jul 20, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,286 509 Updated Jul 22, 2026

Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.

HTML 20,470 1,363 Updated Jul 22, 2026