ZeroHour
Story · 1 source · 2 articlesfirst updated ()

NVIDIA Rack-Scale AI Momentum: d-Matrix Adopts NVLink Fusion as Vera Rubin NVL72 Debuts in MLPerf Inference v6.1

infoAI industryimportance 42
What's new: First merged summary for this dashboard. New developments: d-Matrix joined the NVLink Fusion partner roster (alongside AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung and Cadence), committing to integrate Raptor XPUs with MGX racks, Spectrum-X networking and Vera Rubin NVL72 systems; and Vera Rubin NVL72 made its first MLPerf Inference submission (v6.1 preview), posting up to 3.7x throughput…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

d-Matrix will integrate its Raptor inference XPUs with NVIDIA NVLink Fusion, MGX racks and Spectrum-X networking for rack-scale AI factory deployment, while NVIDIA's Vera Rubin NVL72 made its first MLPerf Inference v6.1 submission with up to 3.7x higher…

Two NVIDIA announcements in September 2026 detail progress in rack-scale AI hardware. On 2026-09-10, inference chipmaker d-Matrix announced adoption of NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's scale-up and scale-out networking, MGX rack architecture, and broader AI factory platform. NVIDIA claims sixth-generation NVLink delivers 3x lower XPU-to-XPU latency than off-the-shelf Ethernet and 3 TB/s per-XPU all-to-all bandwidth. d-Matrix plans to integrate Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-X Ethernet, with its racks able to work alongside Vera Rubin NVL72 GPU systems; other NVLink Fusion partners include AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung and Cadence. On 2026-09-16, NVIDIA reported that Vera Rubin NVL72 debuted in MLPerf Inference v6.1 preview submissions, achieving up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1. A separate 288-GPU GB300 NVL72 submission spanning four racks reached 99% scaling efficiency on the DeepSeek-R1 offline benchmark. Software optimizations leveraging TensorRT-LLM, vLLM, Dynamo, disaggregated serving and NVFP4 precision delivered up to 1.6x gains over v6.0 submissions, and Nebius also submitted Vera Rubin NVL72 preview results.

  • 2026-09-10: d-Matrix announced adoption of NVIDIA NVLink Fusion to connect its next-generation Raptor inference XPUs via NVLink scale-up and Spectrum-X scale-out in MGX racks.
  • NVIDIA claims sixth-generation NVLink provides 3x lower XPU-to-XPU latency than off-the-shelf Ethernet and 3 TB/s per-XPU all-to-all bandwidth.
  • d-Matrix plans include Vera CPUs, ConnectX-9 SuperNICs and BlueField-4 DPUs, with racks able to work alongside Vera Rubin NVL72 GPU systems.
  • Other NVLink Fusion partners include AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung and Cadence.
  • 2026-09-16: Vera Rubin NVL72 debuted in MLPerf Inference v6.1 preview with up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1.
  • A 288-GPU GB300 NVL72 submission across four racks achieved 99% scaling efficiency on the DeepSeek-R1 offline benchmark.
  • Software optimizations using TensorRT-LLM, vLLM, Dynamo, disaggregated serving and NVFP4 precision yielded up to 1.6x performance gains over v6.0 submissions.
  • Nebius also submitted Vera Rubin NVL72 preview results in MLPerf Inference v6.1.

Coverage timeline

  1. · 6d ago
    NVIDIA Blog· 40
    d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

    d-Matrix will integrate its Raptor inference XPUs with NVIDIA NVLink Fusion, MGX racks and Spectrum-X networking for rack-scale AI factory deployment.

  2. · 6h ago
    NVIDIA Blog· 42
    NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

    NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x higher throughput than GB300 NVL72 and 99% scaling efficiency at 288 GPUs.