Skip to content
#

prefix-caching

Here are 30 public repositories matching this topic...

Serving Qwen3.8-27B-FP8 on a single DGX Spark (GB10): 7.88 to 58.5 tok/s single-stream from decode strategy alone, weights untouched. Speculative decoding and prefix caching benchmarked, plus DFlash 2 — the only Qwen3.8-27B build that can serve it under vLLM.

  • Updated Aug 19, 2026
  • Python

Context engineering toolkit for LLMs — pack, cache, debug, red-team, and orchestrate context windows. Council of Experts, adversarial testing, immune system, context compiler, drift detection, multi-agent entanglement. TypeScript + Python.

  • Updated Aug 17, 2026
  • Python

Improve this page

Add a description, image, and links to the prefix-caching topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the prefix-caching topic, visit your repo's landing page and select "manage topics."

Learn more