H100 LLM KV eval: cuMem 2MiB leaf pack (42% VRAM @70%) + vLLM 0.21 TurboQuant 4bit KV (-58% bytes, ~0.86x decode). Eval-only repro.
cuda inference memory-allocator gpu-memory nvidia-gpu kv-cache long-context vllm llm-inference h100 qwen paged-attention turboquant composable-kv cumem moon-xq
-
Updated
May 27, 2026 - C