ZeroHour

Search: “ai factory”

23 stories

AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

At AI Infra Summit, NVIDIA showcased Vera Rubin and DSX gains up to 1.4x tokens per megawatt, plus Annapurna, d-Matrix, and Pinterest partnerships.

Ian Buck's AI Infra Summit keynote before 8,000+ attendees emphasized validated agentic tokens per megawatt as the emerging AI infrastructure metric. Announcements include Amazon's Annapurna Labs collaborating on NVHBM custom high-bandwidth memory, d-Matrix integrating NVLink Fusion with Raptor XPUs, and Pinterest using Blackwell plus Dynamo inference software for conversational visual discovery. Lambda reported 23% better performance per watt with DSX MaxLPS on Blackwell servers, running 19 nodes on a 16-node power budget. NVIDIA says DSX MaxLPS combined with Groq 3 LPX on Vera Rubin NVL72 targets up to 35X token throughput per megawatt versus GB200 NVL72 for 2-trillion-plus-parameter models.

NVIDIA Blog · 6h agoAI industry

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

NVIDIA detailed DSX power-management results: Lambda gained 24% token throughput at fixed power, and an AI factory auto-shed 1MW via Emerald AI's grid program.

NVIDIA says Lambda's first validation of DSX MaxLPS on HGX B200 servers ran 19 nodes within a 16-node power budget, lifting cluster token throughput 24% (roughly 4M to 5M tokens/second) and improving performance per watt by 23%. NVIDIA projects DSX MaxLPS can enable up to 40% more GPU capacity for Vera Rubin NVL72 factories within the same megawatt budget. Emerald AI's Conductor platform, running at NVIDIA's Eos factory with Silicon Valley Power, responded to over 200 utility demand signals, automatically dropping power from 4MW to 3MW without interrupting priority workloads. The first dedicated DSX Flex commercial deployment is planned at a 96-megawatt Manassas, Virginia facility.

NVIDIA Blog · 6h agoAI industry

d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

d-Matrix will integrate its Raptor inference XPUs with NVIDIA NVLink Fusion, MGX racks and Spectrum-X networking for rack-scale AI factory deployment.

Inference chipmaker d-Matrix announced adoption of NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's scale-up and scale-out networking, MGX rack architecture, and broader AI factory platform. NVIDIA claims 3x lower XPU-to-XPU latency than off-the-shelf Ethernet and 3 TB/s per-XPU all-to-all bandwidth via sixth-generation NVLink. d-Matrix plans to integrate Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-X Ethernet, with racks able to work alongside Vera Rubin NVL72 GPU systems. Other NVLink Fusion partners include AWS, Arm, Intel, Fujitsu, Marvell, MediaTek, Samsung and Cadence.

NVIDIA Blog · 5d agoAI industry

How XPUs Meet a World-Class AI Factory

NVIDIA argues AI factories with custom XPUs and NVLink Fusion connectivity must optimize tokens-per-second, tokens-per-watt, cost and uptime.

NVIDIA published a blog explaining that AI factories running continuously are economically defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. It argues hyperscalers and AI-native companies building custom XPUs need infrastructure designed as a complete factory rather than collections of individual accelerators. The piece promotes NVIDIA's NVLink Fusion and full-stack XPU connectivity as the foundation for such world-class AI factory builds.

NVIDIA Blog · 22d agoAI industry

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

NVIDIA partners with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to mobilize over $500 billion of third-party capital for AI infrastructure financing.

NVIDIA announced partnerships with major financial firms including Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR. The independent financing platforms are designed to mobilize more than $500 billion of third-party capital to support AI infrastructure buildout. NVIDIA frames the move as positioning AI factory compute as an investable asset class.

NVIDIA Blog · Aug 12, 2026AI industry

Securing the Infrastructure of Intelligence

NVIDIA positions AI factories combining chips, networking, power and data as the defining infrastructure of the AI economy needing full-stack security.

NVIDIA's blog argues that AI factories are the defining infrastructure of the AI era, transforming energy and data into intelligence that powers businesses and countries. It frames compute as revenue and lists the full stack of critical resources required: advanced chips, packaging, memory, networking, land and power. The piece is a corporate positioning article about securing this infrastructure, with no specific incident or product announcement detailed in the excerpt.

NVIDIA Blog · 29d agoAI industry

Why Scaling AI Compute Performance Requires a New Power Architecture

NVIDIA argues AI factories need 800 VDC power distribution as dense GPU racks outgrow traditional AC-based delivery.

NVIDIA's blog contends each generation of accelerated computing demands higher rack density and more efficient, scalable power distribution. It frames the bottleneck as how power moves from the grid to the GPU rather than raw wattage, and describes limitations of traditional AC power delivery. NVIDIA advocates a new 800 VDC power architecture for AI factories.

NVIDIA Blog · Aug 11, 2026AI industry

PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors

Top AI open source projects like Vercel, Astro, Flue, and tldraw are restricting external PRs and using agent-based software factories to triage, fix, and review contributions.

Several prominent AI-native open source projects are closing or limiting external pull requests, largely because submissions are often AI-generated. Vercel built a multi-agent software factory for its AI SDK (over 20 million weekly npm downloads) that now authors 25-35% of merged PRs and closes 70-80% of issues. Astro adopted similar auto-triage automation, Fred Schott created the Flue framework with automatic PR-to-issue conversion, and tldraw automatically closes external PRs.

Latent Space · 14d agoAI industry

Powering AI is an architecture problem

Sponsored analysis argues AI data centers need medium-voltage, inline power architecture after Virginia grid faults knocked over 3GW of load offline.

A sponsored MIT Technology Review piece recounts a July 22, 2026 transmission fault in Ashburn, Virginia that shed more than 3 GW of data center load, and a 2024 incident where one failed surge arrester dropped about 60 facilities and 1,500 MW. It argues legacy UPS-based power stacks fail at AI scale because campuses can swing 70% of load in milliseconds and trip offline during grid disturbances. The proposed fix moves protection to medium voltage (13.8 kV and above) in inline enclosures near substations, improving density, permitting timelines, and backup power economics. A full-scale system tested at the DOE National Laboratory of the Rockies cleared ERCOT large-load ride-through requirements.

MIT Technology Review · AI · 5d agoAI industry1

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

NVIDIA puts Groq 3 LPX into full production and extends Vera Rubin NVL72 rack-scale systems for fast token generation in agentic AI inference.

NVIDIA announced that Groq 3 LPX is in full production as part of an extension of the Vera Rubin NVL72 rack-scale platform aimed at agentic AI inference. The announcement frames the next era of inference as full-stack AI factory co-design across chips, networking, and systems rather than a single component breakthrough. The focus is improving token generation speed for agent workloads.

NVIDIA Blog · 22d agoAI industry

Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

NVIDIA details a three-computer robotaxi platform that Uber, Lyft, May Mobility, Mercedes-Benz and others are adopting to scale autonomous fleets.

NVIDIA says every major commercial robotaxi program runs on its stack, spanning training (DGX with Alpamayo VLA models), simulation and validation (Omniverse, Cosmos, AlpaSim on RTX PRO), and in-vehicle compute (DRIVE Hyperion 10 with dual DRIVE AGX Thor chips). Adding meta-action and chain-of-thought reasoning data to a VLA model reduced minimum average displacement error by 43%, from 2.08 to 1.18. Uber plans NVIDIA DRIVE Hyperion-based fleets across 28 cities by 2028, partnering with Autobrains, Avride, Lucid, May Mobility, Mercedes-Benz, Momenta, Nissan, Nuro, Pony.ai, Stellantis, Waabi, Wayve, WeRide and Zoox. DRIVE Hyperion 10 combines 14 cameras, nine radars, three lidars and 12 ultrasonics with redundant compute and NVIDIA Halos safety validation.

NVIDIA Blog · 5d agoAI industry

[AINews] Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud

NVIDIA struck a $12B deal with AI coding startup Poolside, licensing its Model Factory and hiring 109 of its technical employees.

NVIDIA spent roughly $12B in an unusual reverse-execuhire of Poolside, licensing the company's Model Factory while hiring 109 of its ~115 technical staff; founders retain a $1B stake and employees receive about $6B. Poolside had raced to raise $2B to fund a 40,000 GB300 cluster after missing a six-week funding window, and founders argue frontier-scale training now requires an order of magnitude more compute plus contracted data center space. An infrastructure arm spun out in January 2026 is scaling toward 7GW as a neocloud. The newsletter also recaps OpenAI and Anthropic agent-platform releases.

Latent Space · 25d agoAI industry

There Gonna Be a Shortage of Everything

An essay argues AI capex ($1.0–1.3T in 2026) now rivals global human-benefit spending and predicts shortages spreading from DRAM/HBM to cars and appliances.

The essay claims AI demand has caused DRAM and NAND shortages, doubled US gas turbine prices, and pushed memory makers to shift DRAM capacity to HBM. It cites the Hugging Face hack, where models self-organized and persisted information, and ARC-AGI 3 being solved in six months by OpenAI's general-purpose LLM ('Astra') as evidence of rapid capability gains. The author predicts mass humanoid robot production within five years will inflate prices of cars and appliances, and notes AI capital investment ($1.0–1.3T in 2026) now matches the ~$1.3T spent globally on human-benefit research and aid.

NVIDIA to Acquire Hugging Face

NVIDIA agreed to acquire Hugging Face for $12.93 billion while pledging to keep the platform open, multi-cloud and vendor-neutral.

NVIDIA announced an agreement to acquire Hugging Face for $12,930,300,000. Hugging Face hosts more than 3 million models, 500,000 datasets and 1 million applications used by over 18 million developers and 200,000 companies. NVIDIA says the platform will remain open, with no requirement to use NVIDIA compute, and will continue supporting multi-cloud and multi-accelerator development and deployment. NVIDIA is already Hugging Face's largest contributor of open models and datasets, with more than 500 models and 250 open datasets released.

NVIDIA Blog · 12d agoAI industry

Agility Robotics says its new Digit 5 robot can work next to people without safety fences

Agility Robotics unveiled Digit 5, a humanoid warehouse robot that can work safely alongside people without fences, with deliveries from early 2027.

Agility Robotics announced Digit 5, a humanoid robot for warehouses and factories that uses AI and sensors to detect people and stop or step aside without safety fences. It lifts up to 22.7 kg (40% more than Digit 4), charges in 9 minutes for 90 minutes of runtime, and is the first partner for Nvidia's Halos robotics safety platform. Agility cites over $300 million in orders, with first deliveries in early 2027 and production capacity of up to 10,000 units per year in Salem, Oregon.

The Decoder · 10h agoAI industry

OpenAI Announced $1B in Defensive Tools for Water Utilities

OpenAI pledges $1 billion in subsidized Daybreak cyber models and training for water utilities, grid operators, and other critical-infrastructure defenders.

OpenAI announced Daybreak for Frontline Defenders on September 3, 2026, committing $1 billion in product credits and subsidized access to its Daybreak cyber models, training, and technical support for under-resourced defenders. Priority access goes to water and wastewater utilities, electric grid operators, state and local governments, community banks, nonprofits, and open-source maintainers; around 2,000 organizations already use Daybreak, which includes Daybreak Blue and Daybreak Red tiers. The program includes an MS-ISAC pilot, the Daybreak Defense Network with 35+ partner products (including HackerOne), and publication of OpenAI's Defense Factory automated vulnerability discovery architecture; it launched the same day OpenAI shipped a model it internally classifies as Critical for cyber capability.

Security Affairs · 10d agoAI industry

OpenAI Pledges $1bn to Bring its AI Cybersecurity Tools to Essential Services

OpenAI pledged $1bn to subsidize Daybreak cybersecurity model access for water, power, banking, government and nonprofit defenders, starting in the US with an MS-ISAC pilot.

OpenAI announced a $1 billion pledge to subsidize access to its Daybreak cyber models for essential services including water, electricity, local governments, nonprofits and banking, starting in the US and expanding to partner countries. The Daybreak for Frontline Defenders initiative includes a pilot with the Multi-State Information Sharing and Analysis Center (MS-ISAC) pairing model access with guided training for public sector and water system defenders. OpenAI unveiled Daybreak in May 2026, deploying frontier LLMs and its Codex coding assistant for defender tasks, and split it into Daybreak Red and Daybreak Blue tiers in August. The pledge follows an August 27 open letter from more than 100 tech and cybersecurity companies warning of a narrowing window before AI-enabled attacks escalate.

Infosecurity Magazine · 11d agoAI industry1

Maven Robotics wants to steal your robot deployment deal

Warehouse robotics startup Maven Robotics emerges from stealth with $100 million to build 250 third-generation palletizing robots.

Maven Robotics, founded in 2024 by former Apple special projects engineer Hamza Derbas and his brother Khalid, emerged from stealth after raising $100 million from RoboStrategy, LocalGlobe, Vine Ventures, and XTX Markets Ventures. Its wheeled dual-arm robots, moving 10 mph and lifting up to 30 kg, perform mixed palletizing in distribution centers, with up to eight units reportedly running 16 hours a day at 99%+ uptime. The company plans to build 250 third-generation robots, start design on a fourth-generation platform, and expand toward material handling and fabrication, positioning itself against rivals like Agility, which is going public via a $2.4 billion SPAC deal.

TechCrunch · AI · 5d agoAI industry