CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production; NVIDIA Pitches AI Factory ROI With 30x Throughput-per-Megawatt Claim
CoreWeave announced production availability of NVIDIA Vera Rubin NVL72 with Spectrum-X 102.4T Ethernet, with first customer Cognition reporting up to 4.8x Devin token throughput versus GB200 NVL72; NVIDIA separately claims, citing SemiAnalysis AgentX, over…
CoreWeave announced production availability of NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet, with Cognition as the first customer running production workloads on the platform. Cognition reported up to a 4.8x increase in total token throughput for Devin SWE-2 inference versus a GB200 NVL72 baseline. CoreWeave also plans to offer the NVIDIA Vera CPU, where tests showed more than 3x faster agent sandbox startups and a 1.7x Terminal-Bench gain, and introduced CoreWeave Forge, which combines Weights & Biases, OpenPipe, and marimo for continuous model and agent improvement, including generally available ARIA and Sandboxes plus the new Agent Lens. In a follow-up post, NVIDIA argued that AI factory operators, who spend about $60 million per megawatt, should judge returns by earning capacity, useful hardware life, and token demand. Citing SemiAnalysis AgentX, NVIDIA claims Vera Rubin NVL72 delivers over 30 times the throughput per megawatt of GB300 NVL72 and up to 45 times lower cost per million tokens on DeepSeek V4 Pro; note that the 4.8x and 30x figures use different baselines (GB200 NVL72 vs GB300 NVL72) and different metrics (total token throughput vs throughput per megawatt), so they are not directly comparable. NVIDIA also pointed to hardware longevity, noting A100 GPUs shipped in 2020 remain in service with CoreWeave bookings extended through 2029, and that Microsoft's V100 fleet ran 8.4 years versus a six-year book life, with CUDA enabling one platform to span GPU generations and many workload types.
- CoreWeave announced production availability of NVIDIA Vera Rubin NVL72 with Spectrum-X 102.4T Ethernet on 2026-09-30.
- Cognition is the first customer running production workloads on the platform.
- Cognition reported up to 4.8x total token throughput for Devin SWE-2 inference versus a GB200 NVL72 baseline.
- On the planned NVIDIA Vera CPU, tests showed more than 3x faster agent sandbox startups and a 1.7x Terminal-Bench gain.
- CoreWeave Forge combines Weights & Biases, OpenPipe, and marimo for continuous model and agent improvement; ARIA and Sandboxes are generally available and Agent Lens is new.
- NVIDIA says AI factory operators spend about $60 million per megawatt and should judge returns by earning capacity, useful hardware life, and token demand.
- Citing SemiAnalysis AgentX, NVIDIA claims Vera Rubin NVL72 delivers over 30x the throughput per megawatt of GB300 NVL72.
- NVIDIA claims up to 45x lower cost per million tokens on DeepSeek V4 Pro.
Coverage timelineoldest first · each row is one article
- · 1d agoFrom Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI
NVIDIA Blog· 68
CoreWeave now offers NVIDIA Vera Rubin NVL72, and Cognition reports up to 4.8x higher Devin token throughput.
- · 8h agoProductive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment
NVIDIA Blog· 36
NVIDIA argues AI factories maximize ROI through higher tokens per megawatt, longer GPU life, and workload fungibility.