GPU Monitoring

Application

Datadog's Infrastructure-family product for shared GPU fleets across cloud, on-prem and neocloud providers: it links device health, cost and performance to the workloads and teams using them, alerts on unmet GPU requests, thermal throttling and ECC/XID errors, forecasts GPU demand and recommends optimisations such as reclaiming GPUs held by zombie processes. The page markets alerting, forecasting and recommendations but does not name a model or AI mechanism behind them.

Find alternatives to GPU Monitoring

TypeApplication
RoleBundled component
DeploymentSaaS
Intended forDeveloperSelf-reported · 8 Sep 2026
SecurityNot recorded
StatusActive
Sold withinDatadog Infrastructure — Included with the parentOfficial documentation · 5 Sep 2026
PricingNot recorded
Sources—

No charging model published. The product page has no price line (CTAs are 'Start free trial' for 'the entire Datadog product suite', 'Request a demo', 'View documentation'); the pricing page's Infrastructure sidebar lists Infrastructure, Storage Management, Edge Device Monitoring, Kubernetes Autoscaling, Serverless, Network, Disaster Recovery and Cloud Cost Management but no GPU Monitoring entry, and ?product=gpu-monitoring redirects to the Infrastructure section; docs (https://docs.datadoghq.com/gpu_monitoring/) and the launch press release state no price.

What it does

datadoghq.com

Sources

Description

Official documentation · 5 Sep 2026

Fetched. Title 'GPU Monitoring for AI Workloads | Datadog'; 'Datadog GPU Monitoring delivers end-to-end visibility across shared GPU fleets by linking device health, cost, and performance directly to the workloads and teams using them ... whether deployed in cloud, on-prem, or neocloud environments'; 'With proactive alerting and actionable recommendations, GPU Monitoring helps teams optimize efficiency, resolve stalled or failed AI workloads, prevent hardware issues, and reduce wasted spend.'; 'Forecast GPU demand'; 'guided optimization actions like reclaiming GPUs tied up by zombie processes'; 'alerts on workloads or clusters with unmet GPU requests'; 'Detect and remediate ECC/XID errors proactively with built-in alerts and prescriptive next steps'; 'just by toggling a single configuration flag in the Datadog Agent'. Docs H1 'GPU Monitoring': 'provides a centralized view into your GPU fleet's health, cost, and performance.' No sentence on either page names an AI/ML mechanism.

Sold within

Official documentation · 5 Sep 2026

The site's Product menu files 'GPU Monitoring' under the 'Infrastructure' group beside Infrastructure Monitoring, Metrics, Network Monitoring, Container Monitoring, Kubernetes Autoscaling, Serverless, Cloud Cost Management, Cloudcraft and Storage Management; the page's 'Works great with' box names Agent Observability and Infrastructure Monitoring, and data is collected 'by toggling a single configuration flag in the Datadog Agent'. No page states a plan that includes it.

Also from Datadog

Open Datadog in the directory

Something wrong here? Send a correction — quote this product id: datadog-gpu-monitoring.