Early-warning bottleneck profiler for GPU training nodes: GPU + RDMA fabric telemetry with job-level classification
-
Updated
Jun 13, 2026 - Python
Early-warning bottleneck profiler for GPU training nodes: GPU + RDMA fabric telemetry with job-level classification
Packet-level simulator (ns-3 + RoCEv2/DCQCN/PFC/ECN) comparing fat-tree vs rail-optimized topologies for AI training fabrics. Validated against NCCL on real GPUs.
Ultra-Ethernet compliance test architecture eliminating expensive HBM memory buffers. Design utilizes FPGAs, PRBS payloads, state hash tables, and virtual RDMA.
A simple experiment applying compressible flow principles to soften distributed gradient communication stalls.
A distributed hardware-software co-design fabric for MoE models (DeepSeek-V3, Mixtral) that eradicates NCCL All-to-All communication stalls via distributed RoCEv2 RDMA virtual address MUX and JAX/XLA SPMD sharding.
RoCEv2 control plane simulator — DCQCN congestion control on fat-tree GPU clusters with NCCL Ring AllReduce workloads
Lossless-Ethernet benchmark harness for RoCEv2 fabrics: sweeps queue thresholds, ECN marking, and oversubscription to locate throughput collapse and PFC pressure.
Add a description, image, and links to the rocev2 topic page so that developers can more easily learn about it.
To associate your repository with the rocev2 topic, visit your repo's landing page and select "manage topics."