The infrastructure behind an AI model can reveal a surprising amount before anyone breaks into it. A public monitoring endpoint may disclose the GPUs in a server, their utilization and the software around them. A flaw in that same monitoring service can turn visibility into an availability risk.
New research from Lava, released October 8, describes both problems. The security company identified roughly 2,100 publicly accessible NVIDIA DCGM Exporter hosts reporting more than 12,000 unique GPUs without authentication. During its investigation, Lava also discovered a high-severity vulnerability that could let an unauthenticated attacker exhaust resources and crash GPU monitoring.
NVIDIA has assigned the issue CVE-2026-47483, rated it 8.2, High, and issued an update. The findings put a less glamorous part of AI infrastructure in the spotlight: the services used to observe expensive compute need protection of their own.
What the researchers found—and what the numbers mean
Lava’s original research by Michael Katchinskiy describes four scans conducted between March and May 2026. The totals therefore represent observations across that research period, rather than a live count of systems still exposed today.
The hosts returned GPU telemetry without authentication. Lava observed data-center accelerators, including H100s, H200s and Blackwell Ultra B300s, as well as RTX 4090 and 5090 systems. The company estimated that the observed GPUs represented more than $100 million in hardware, based on approximate market values. That figure describes hardware value, not losses from an attack.
About a quarter of the exposed DCGM hosts also made internal Go profiling endpoints accessible. This subset is important: an exposed metrics endpoint and a reachable vulnerable profiling interface are related but distinct findings. It would be misleading to describe all 12,000-plus GPUs as confirmed victims of this vulnerability.
Lava says it reproduced resource exhaustion in a controlled environment, rather than attacking the public deployments. The research demonstrates a potential attack path; it does not establish that the observed organizations suffered exploitation or that their model data was stolen.
Why GPU monitoring reveals more than a status light
DCGM stands for Data Center GPU Manager. NVIDIA’s DCGM Exporter documentation explains that the exporter collects selected GPU telemetry fields and serves them in a format Prometheus can consume. Its metrics endpoint is typically used by monitoring systems to track the condition and activity of GPU nodes.
Temperature, utilization, memory usage, power consumption and error events are useful to operators because they describe how compute is behaving. When the same information is accessible to strangers, it becomes an inventory and reconnaissance source.
The exposed responses can reveal hardware models and operational details. Repeated readings can provide clues about busy periods and recurring activity. Those clues are not proof that a particular model is being trained or served, but they can help an outsider narrow down what an environment contains and when it is active.
That distinction is worth preserving. Reading GPU telemetry is not the same as reading a model’s weights, training data or prompts. Yet infrastructure information can still be valuable: an attacker who learns which components and versions are present has a more specific starting point than someone facing an opaque server.
The vulnerability targets the monitoring service
NVIDIA’s security bulletin locates the flaw in DCGM Exporter’s /debug/pprof endpoints. Concurrent unauthenticated profiling requests can cause uncontrolled resource consumption, with potential denial of service and information disclosure. The advisory credits Lava’s Michael Katchinskiy for reporting it.
Profiling is a legitimate diagnostic capability. It helps developers investigate CPU and memory behavior inside an application. The security problem arises when a potentially expensive internal function becomes reachable by an untrusted caller without suitable controls.
According to Lava, researchers initially suspected an operator configuration error, then reproduced the behavior with NVIDIA’s official container. They demonstrated that resource exhaustion could crash the exporter, removing visibility into GPU health. CPU and memory pressure could also affect training or inference workloads sharing the server.
Crashing an exporter does not necessarily stop the GPU workload itself. The immediate effect is loss of monitoring; interference with neighboring workloads depends on resource isolation and the deployment. This is a software-service vulnerability around GPU infrastructure, rather than evidence of a flaw in the GPU silicon.
The distinction matters operationally. If monitoring disappears during a workload slowdown, responders need to investigate whether the observation system is itself failing. Treating every missing metric as an instrumentation inconvenience could delay recognition of a resource-consumption incident.
The exposure extends beyond the GPU layer
Lava’s announcement also describes 12,096 publicly accessible Node Exporter hosts. Node Exporter reports server and operating-system information rather than serving the same role as DCGM Exporter. The exposed data included hardware and software details that could help outsiders understand the systems surrounding GPU workloads.
Those counts should remain separate. The Node Exporter observations are a broader infrastructure exposure finding, not another count of hosts confirmed vulnerable to CVE-2026-47483. Combining the figures would obscure which service and risk each number represents.
The wider implication is that AI security needs to include the monitoring and management layer. Model access controls do not automatically protect a metrics service deployed beside the model. An organization can secure its inference API while leaving another service on the same infrastructure open to the internet.
Patching and restricting access address different problems
The security update is already available. NVIDIA’s bulletin identifies DCGM Exporter 4.8.2 as an updated version and also lists DCGM 4.5.3. Operators should consult the current advisory and supported release pairing for their deployment rather than treating those two component version numbers as interchangeable.
Upgrading addresses the disclosed flaw. It does not, by itself, establish that the metrics endpoint is appropriately restricted. A patched exporter can still disclose telemetry if it remains publicly reachable without access controls.
The Prometheus security model explicitly cautions against exposing component HTTP endpoints to public networks without appropriate measures. Its guidance covers metrics, APIs and Go profiling interfaces, and recognizes the possibility of overloading these services.
For teams reviewing their AI infrastructure, that suggests a practical sequence:
- Inventory deployed monitoring services. Establish which exporters, Prometheus servers and diagnostic interfaces are running, who owns them and how they are reachable.
- Apply the vendor’s security updates. Check the actual deployed software or container version, not just a configuration file that has not yet been rolled out.
- Limit monitoring access. Use private networking and appropriate firewall, security-group and access controls so telemetry is available to the monitoring infrastructure that needs it.
- Review profiling requirements. Lava recommends leaving
--enable-pprofdisabled unless profiling is explicitly needed; in current versions, it is opt-in. - Verify visibility after remediation. Confirm that authorized collection still works and that unexpected exporter failures are noticed.
These steps address separate questions: whether the software contains the flaw, whether an untrusted party can reach it, and whether a monitoring failure will be detected. Solving one does not settle the others.
AI infrastructure needs an explicit security owner
GPU capacity often spans provider-operated infrastructure and customer-deployed services. A useful security review identifies who maintains each component, who controls network exposure and who responds when a public endpoint is reported. Without those assignments, a monitoring service can sit between two teams that each expect the other to secure it.
The central lesson of Lava’s research is practical: protecting AI compute includes protecting the systems that measure and manage it. The new findings document significant historical exposure, while NVIDIA’s advisory provides a remediation path for the disclosed vulnerability. For operators, the priority is to verify their present deployment, apply the fix and keep internal observation services within their intended trust boundary.




