#
l40s
Here are 3 public repositories matching this topic...
Fixes NCCL hangs on NVIDIA L40S GPUs by resolving IOMMU-induced PCIe P2P communication issues. Includes reproducible tests, architecture explanation, and production-ready solution.
machine-learning performance deep-learning hpc gpu linux-kernel cuda nvidia pcie distributed-training nccl iommu ai-infrastructure l40s
-
Updated
Apr 30, 2026 - Shell
Reproducible LLM inference benchmark scaffold for NVIDIA L40S and OpenAI-compatible servers.
-
Updated
Jun 10, 2026 - Python
Improve this page
Add a description, image, and links to the l40s topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the l40s topic, visit your repo's landing page and select "manage topics."