NVIDIA Rack-Scale AI Momentum: d-Matrix Adopts NVLink Fusion as Vera Rubin NVL72 Debuts in MLPerf Inference v6.1
d-Matrix will integrate its Raptor inference XPUs with NVIDIA NVLink Fusion, MGX racks and Spectrum-X networking for rack-scale AI factory deployment, while NVIDIA's Vera Rubin NVL72 made its first MLPerf Inference v6.1 submission with up to 3.7x higher…
Two NVIDIA announcements in September 2026 detail progress in rack-scale AI hardware. On 2026-09-10, inference chipmaker d-Matrix announced adoption of NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's scale-up and scale-out networking, MGX rack architecture, and broader AI factory platform. NVIDIA claims sixth-generation NVLink delivers 3x lower XPU-to-XPU latency than off-the-shelf Ethernet and 3 TB/s per-XPU all-to-all bandwidth. d-Matrix plans to integrate Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-X Ethernet, with its racks able to work alongside Vera Rubin NVL72 GPU systems; other NVLink Fusion partners include AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung and Cadence. On 2026-09-16, NVIDIA reported that Vera Rubin NVL72 debuted in MLPerf Inference v6.1 preview submissions, achieving up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1. A separate 288-GPU GB300 NVL72 submission spanning four racks reached 99% scaling efficiency on the DeepSeek-R1 offline benchmark. Software optimizations leveraging TensorRT-LLM, vLLM, Dynamo, disaggregated serving and NVFP4 precision delivered up to 1.6x gains over v6.0 submissions, and Nebius also submitted Vera Rubin NVL72 preview results.
- 2026-09-10: d-Matrix announced adoption of NVIDIA NVLink Fusion to connect its next-generation Raptor inference XPUs via NVLink scale-up and Spectrum-X scale-out in MGX racks.
- NVIDIA claims sixth-generation NVLink provides 3x lower XPU-to-XPU latency than off-the-shelf Ethernet and 3 TB/s per-XPU all-to-all bandwidth.
- d-Matrix plans include Vera CPUs, ConnectX-9 SuperNICs and BlueField-4 DPUs, with racks able to work alongside Vera Rubin NVL72 GPU systems.
- Other NVLink Fusion partners include AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung and Cadence.
- 2026-09-16: Vera Rubin NVL72 debuted in MLPerf Inference v6.1 preview with up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1.
- A 288-GPU GB300 NVL72 submission across four racks achieved 99% scaling efficiency on the DeepSeek-R1 offline benchmark.
- Software optimizations using TensorRT-LLM, vLLM, Dynamo, disaggregated serving and NVFP4 precision yielded up to 1.6x performance gains over v6.0 submissions.
- Nebius also submitted Vera Rubin NVL72 preview results in MLPerf Inference v6.1.
Coverage timelineoldest first · each row is one article
- · 6d agod-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
NVIDIA Blog· 40
d-Matrix will integrate its Raptor inference XPUs with NVIDIA NVLink Fusion, MGX racks and Spectrum-X networking for rack-scale AI factory deployment.
- · 6h agoNVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
NVIDIA Blog· 42
NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x higher throughput than GB300 NVL72 and 99% scaling efficiency at 288 GPUs.