From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
NVIDIA detailed DSX power-management results: Lambda gained 24% token throughput at fixed power, and an AI factory auto-shed 1MW via Emerald AI's grid program.
NVIDIA says Lambda's first validation of DSX MaxLPS on HGX B200 servers ran 19 nodes within a 16-node power budget, lifting cluster token throughput 24% (roughly 4M to 5M tokens/second) and improving performance per watt by 23%. NVIDIA projects DSX MaxLPS can enable up to 40% more GPU capacity for Vera Rubin NVL72 factories within the same megawatt budget. Emerald AI's Conductor platform, running at NVIDIA's Eos factory with Silicon Valley Power, responded to over 200 utility demand signals, automatically dropping power from 4MW to 3MW without interrupting priority workloads. The first dedicated DSX Flex commercial deployment is planned at a 96-megawatt Manassas, Virginia facility.
- Lambda: 24% more token throughput and 23% better performance per watt
- DSX MaxLPS enables up to 40% more GPUs per megawatt on Vera Rubin
- Emerald AI Conductor handled 200+ utility demand signals automatically
- Santa Clara factory shed 1MW while high-priority AI jobs kept running
Full article1,213 words · extracted from blogs.nvidia.com · click to collapse
On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption.
Varun Sivaram was watching on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers at the data center and people from the utility itself. Nobody touched anything.
Emerald AI’s Conductor platform — a grid-orchestration platform from NVIDIA partner Emerald AI, and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver — receives signals about grid conditions and adjusts the data center’s flexible computing workloads. Work that can wait is slowed or rescheduled, while higher-priority services continue operating.
The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads— exactly what Silicon Valley Power needed,
When the reduction showed on screen, everyone cheered.
“We were watching with bated breath,” Sivaram said. “It was our first time deploying across thousands of NVIDIA GPUs.” His head of product, Mansi Shah, was emotional. “This feels kind of like a SpaceX rocket launch,” she said.
Silicon Valley Power has since sent more than 200 demand signals to that AI factory. It worked every single time.

This is grid flexibility in production. And it points at something much bigger than one facility in Santa Clara: a path to unlocking the power America’s AI factories need, without waiting a decade to build new transmission lines.
At the AI Infra Summit on Tuesday, Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote.
Results from cloud provider Lambda’s first validation in a deployment environment, released the same day, put numbers to it: a fixed power budget can support 24% more token throughput when managed intelligently.
“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” said Dave Ward, president of cloud services at Lambda. “NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.”
That August evening, when SVP called, Conductor executed against a predefined workload hierarchy: lowest-priority jobs yielded, high-priority inference kept running, and power fell from four megawatts to three. Automated. No operator required.
In the AI factory economy, power is the constraint. Work per gigawatt is the metric. Data center operators are meticulous about efficiency — every watt put to work is a watt delivering productive compute, and the industry has driven remarkable gains at every layer of the stack, from facility design to rack-level power conversion.
DSX extends that discipline into the AI workload itself. Smarter rack provisioning puts power where workloads actually need it. Operational intelligence — tighter scheduling, faster restarts, leaner checkpointing — keeps GPUs running rather than waiting. The goal is the same one operators have always pursued: more work from the power you have.
“A one-gigawatt factory will never become a two-gigawatt factory,” NVIDIA founder and CEO Jensen Huang has said.
The answer engineers reach when systems hit physical limits is always the same: stop optimizing the parts and start designing the whole.
Introduced at GTC Taipei in May, NVIDIA DSX is that answer for the AI factory — and the early deployments are already proving it out.
The full platform spans networking, cooling, water efficiency and facility design; the sections below focus on some of the results so far in power management and grid participation.
More Compute, Same Budget: DSX MaxLPS
Lambda’s results, released at the AI Infra Summit, are the first validation of DSX MaxLPS on NVIDIA HGX B200 GPU Servers.
DSX MaxLPS monitors GPU and rack-level power consumption and reallocates headroom across nodes based on workload type, recovering capacity that static provisioning would leave stranded. Training and inference draw power differently; MaxLPS optimizes allocation in AI factories running both.
Lambda, a GPU cloud provider serving more than 10,000 customers from AI-native startups to hyperscalers, ran the software on a five-rack, 19-node cluster.
What they found: by running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24% more cluster-wide token throughput — from roughly 4 million tokens per second to 5 million. Performance per watt improved by 23%.
Based on NVIDIA’s projections, DSX MaxLPS can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment environments.
Automated Demand Response, Proven in Production
The Santa Clara story isn’t a DSX Flex installation — it’s something earlier and more important: proof that the concept works at commercial scale.<
NVIDIA’s Eos AI factory is running Emerald AI Conductor as a participant in Silicon Valley Power‘s Flexible Load Interconnect Program, the first commercial grid utility program designed to treat AI factories as dispatchable resources.
When Silicon Valley Power sends a signal, Conductor responds in under a minute. The factory that’s willing to flex gets to run bigger.
That’s the pattern DSX Flex is built to generalize — with Emerald AI Conductor integrating into DSX Flex as the platform matures. The first dedicated DSX Flex commercial deployment will be the Manassas, Virginia, facility: a 96-megawatt Vera Rubin AI factory at NVIDIA’s AI Factory Research Center, building on five prior demonstrations across two continents.
The Next Power Architecture Layer: 800V DC Power Architecture
The gains inside today’s AI factory are real and deployable now. The next layer is how power is delivered to denser accelerated computing racks.
As AI factories scale, traditional lower-voltage power paths add conversion complexity and distribution constraints.
NVIDIA’s 800 VDC architecture is designed to reduce conversion complexity, improve power delivery efficiency and support denser accelerated computing racks.
NVIDIA DSX is incorporating 800V DC into its reference designs.
The Whole Factory, Not the Parts
No single component can optimize an AI factory on its own. A faster GPU still waits on the network. Power can be stranded by bad provisioning. Cooling overhead still diverts electricity from GPUs; GB200 NVL72 racks running direct liquid cooling carry ~120 kW of heat that has to go somewhere before that power reaches compute.
The only reliable path to more tokens per megawatt is to optimize the whole factory — DSX Sim before the first rack goes in, DSX OS and DSX Exchange once it’s running, DSX Reference Designs so builders start from a validated architecture rather than from scratch. (See sidebar for the full DSX suite at a glance.)
The Gigawatt Infrastructure Standard
It all comes down to one question: how much useful work does the factory produce per megawatt consumed?
NVIDIA DSX gives infrastructure builders the reference designs, simulation tools, operational software, and power-management technology to compete on that metric, on current hardware and into the next generation.
When the grid needed relief, the factory gave it without dropping a job, without asking for more power.
With NVIDIA DSX, that’s the new baseline for what an AI factory is supposed to do.
Text extracted automatically; images, tables and formatting may be missing. Original: https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/