Skip to content
#

bf16

Here are 25 public repositories matching this topic...

Systematic 24-hour benchmark study of Qwen3.6-27B inference on dual NVIDIA RTX PRO 6000 Blackwell SM120 (TP=2). 8 experiments comparing repne/vllm fork vs upstream vLLM across FP8/BF16/NVFP4/Q8_0 quants and MTP/DFlash speculative decoding. Peak: 2,083 tok/s at c=32. Quality: KLD vs BF16 = 0.0018 (noise floor).

  • Updated Jun 3, 2026
  • Python

Python implementations for multi-precision quantization in computer vision and sensor fusion workloads, targeting the XR-NPE Mixed-Precision SIMD Neural Processing Engine. The code includes visual inertial odometry (VIO), object classification, and eye gaze extraction code in FP4, FP8, Posit4, Posit8, and BF16 formats.

  • Updated Aug 17, 2025
  • Jupyter Notebook

DeepSeek-OCR-experimental is an advanced, multi-purpose visual document intelligence and object localization sandbox. Powered by the unredacted prithivMLmods/DeepSeek-OCR-Latest-BF16.I64-v2.0 architecture, this suite is designed to deliver highly accurate, structure-aware image text extractions.

  • Updated Jun 27, 2026
  • Python

Distributed GPT-2 fine-tuning with PyTorch FSDP and BF16 mixed precision, INT8 post-training quantisation, a custom Triton quantisation kernel achieving 1.4x throughput over unfused PyTorch, and a full FP32 vs BF16 vs INT8 benchmark suite. 13 tests passing.

  • Updated May 23, 2026
  • Python

Add this topic to your repo

To associate your repository with the bf16 topic, visit your repo's landing page and select "manage topics."

Learn more