High-severity Nvidia bug could crash GPU monitoring on exposed servers
Nvidia patched CVE-2026-47483, a high-severity flaw that can crash exposed DCGM Exporter GPU monitoring.
Nvidia fixed CVE-2026-47483, a CVSS 8.2 flaw in DCGM Exporter that lets unauthenticated attackers crash GPU monitoring by exhausting memory with concurrent requests, potentially disrupting AI workload visibility. Lava's scans between March and May found about 2,100 internet-exposed DCGM hosts, covering roughly 12,000 GPU UUIDs at about 300 organizations, including Blackwell Ultra B300, H200, H100, and RTX 5090 and 4090 systems. About 25 percent of those hosts also exposed Go's /debug/pprof profiler. Researchers separately found 12,096 public Prometheus Node Exporter instances and said Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean worked with customers after notification. Nvidia's fix is in version 4.8.2 or later.
- CVE-2026-47483 scores CVSS 8.2 and is fixed in DCGM Exporter 4.8.2.
- Concurrent unauthenticated requests can exhaust memory and crash the exporter.
- About 2,100 exposed hosts revealed 12,000 GPU UUIDs without authentication.
- Roughly 25 percent also exposed Go's unauthenticated debug profiler.
- Separately, 12,096 public Node Exporter hosts leaked OS, firmware, and hardware details.
Vulnerabilities mentionedAll →
- CVE-2026-474838.2<1%NVIDIA DCGM Exporter for all platforms contains a vulnerability in the /debug/pprof endpoints, where an attacker could cause uncontrolled resource consumption…published
| CVE | Vulnerability | CVSS | EPSS | Flags | Affected | Exposure | Published |
|---|---|---|---|---|---|---|---|
| CVE-2026-47483 | NVIDIA DCGM Exporter for all platforms contains a vulnerability in the /debug/pprof endpoints, where an attacker could cause uncontrolled resource consumption… NVIDIA DCGM Exporter for all platforms contains a vulnerability in the /debug/pprof endpoints, where an attacker could cause uncontrolled resource consumption by submitting concurrent unauthenticated profiling requests. A successful exploit of this vulnerability might lead to denial of service and information disclosure. |
Full article615 words · extracted from theregister.com · click to collapse
security
The GPU giant released a fix for the flaw, tracked as CVE-2026-47483
Researchers found thousands of GPU servers exposing Nvidia's DCGM Exporter to the internet, with hundreds potentially vulnerable to a high-severity flaw that could let unauthenticated attackers crash the GPU monitoring service and disrupt AI workloads.
DCGM Exporters read telemetry from the GPUs on a host, including its hardware, utilization, memory usage, power consumption, and error events. Each GPU has its own unique ID, or UUID, and all of these metrics are exposed in plaintext over HTTP.
This exposure provides would-be attackers with detailed information useful for reconnaissance, including mapping GPU infrastructure, identifying potentially vulnerable systems, and monitoring workload activity.
REG AD
Michael Katchinskiy, a researcher at datacenter security startup Lava, found and reported the bug in the GPU health and performance monitoring service. In September, the GPU giant Nvidia released a fix for the flaw, tracked as CVE-2026-47483, and gave it an 8.2 CVSS high-severity rating.
REG AD
“Once we realized how much these endpoints revealed, the next question was: How many of them are exposed to the internet?” Katchinskiy said in a Thursday blog.
So the researchers started scanning the internet for exposed DCGM Exporters. And the scale of exposure proved “especially significant,” he wrote.
Over the course of four scans between March and May, the threat hunters found about 2,100 GPU servers exposing DCGM Exporter metrics to the open internet. These included 12,000 GPU UUIDs. None of these required authentication.
The hosts belonged to about 300 organizations, according to Katchinskiy, and nearly half - 5,274 of the exposed GPUs, or 44 percent of the total - were located in the US.
These GPUs represented about $100 million in hardware, and included Nvidia Blackwell Ultra B300 GPUs, H200s, and H100s - used to run large-scale AI workloads - plus consumer RTX 5090 and 4090 systems.
While investigating the exposed systems, the Lava team found that about 25 percent of the exposed DCGM hosts also revealed data from Go’s /debug/pprof/ built-in profiling tool. The profiler collects and exposes runtime performance data such as CPU and memory usage for running Go applications. This includes CPU usage, memory allocations, goroutine states, and blocking events.
“With enough concurrent unauthenticated requests, the exporter could run out of memory and crash, cutting off visibility into GPU health and activity,” Katchinskiy wrote.
The CPU and memory pressure could also affect AI training or inference workloads.
REG AD
Nvidia fixed the issue in version 4.8.2, and operators should upgrade to that version or later.
In addition to GPU telemetry, Lava looked into Prometheus Node Exporter, which monitors server hardware and operating systems, and also exposes metrics over HTTP. The team found 12,096 public Node Exporter hosts exposing data on server models, operating systems, firmware versions, hostnames, storage paths and networking hardware commonly used in GPU clusters.
This information reveals how environments are built and configured, which could also be used by attackers for reconnaissance, matching the system to known vulnerabilities.
The publicly exposed monitoring services affected customer infrastructure across neocloud and GPU cloud providers including Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean. Lava reported all of this to the affected providers, and Katchinskiy says that these providers worked with customers to address the exposures.
For operators: the security shop says Nvidia DCGM Exporter, Node Exporter and Prometheus services should not be directly reachable from the public internet and recommends restricting them to authorized monitoring infrastructure.
“The findings highlight a growing security gap in AI infrastructure: companies are spending millions on GPUs while leaving critical systems exposed,” Katchinskiy wrote. “Those exposures can reveal how AI environments are built and, in some cases, allow attackers to disrupt them.” ®