Skip to content
#

glm-5-3-flash

Here are 33 public repositories matching this topic...

Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.

  • Updated Sep 15, 2026
  • Python

NVFP4 BIZ: nvidia/GLM-5.3-Flash-NVFP4 on two DGX Spark-class GB10 systems with a pinned vLLM (TP=2, Marlin W4A16, FA2 prefill, MTP, prefix caching, image input, bit-reproducible completions); accepted for routine use, one active sequence. Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.

  • Updated Sep 25, 2026
  • Python

🐑 Fleece Radar — 薅羊毛雷达 - Automated radar & passive intelligence pipeline for 200+ Chinese AI gateways and free API relays. Zero-auth model probing, free tier telemetry, and instant OmniRoute upstream exports.

  • Updated Sep 19, 2026
  • Python

Single-file GPU/RAM sizing calculator for LLM agent sessions on vLLM: KV-cache offload to RAM, prefix caching, MLA/DSA and hybrid models, TP chosen per model × GPU pair (H100–B300). Runs in the browser, no dependencies.

  • Updated Sep 24, 2026
  • HTML
GLM-5.3-Flash-Free-Z-AI

GLM 5.3 Flash Free Z.AI - free glm 5.3 flash download on z.ai. glm 5.3 vs 5.3 flash, huggingface, gguf, ollama glm 5.3, openrouter, glm 5.3 api, opencode. Windows Mac Linux zip. Official free download. Download:🡇

  • Updated Sep 16, 2026
  • C++

Add this topic to your repo

To associate your repository with the glm-5-3-flash topic, visit your repo's landing page and select "manage topics."

Learn more