GPU Monitoring
Datadog's Infrastructure-family product for shared GPU fleets across cloud, on-prem and neocloud providers: it links device health, cost and performance to the workloads and teams using them, alerts on unmet GPU requests, thermal throttling and ECC/XID errors, forecasts GPU demand and recommends optimisations such as reclaiming GPUs held by zombie processes. The page markets alerting, forecasting and recommendations but does not name a model or AI mechanism behind them.
Find alternatives to GPU Monitoring
No charging model published. The product page has no price line (CTAs are 'Start free trial' for 'the entire Datadog product suite', 'Request a demo', 'View documentation'); the pricing page's Infrastructure sidebar lists Infrastructure, Storage Management, Edge Device Monitoring, Kubernetes Autoscaling, Serverless, Network, Disaster Recovery and Cloud Cost Management but no GPU Monitoring entry, and ?product=gpu-monitoring redirects to the Infrastructure section; docs (https://docs.datadoghq.com/gpu_monitoring/) and the launch press release state no price.
What it does
- Evaluation and observability
- Data analysis
Sources
- Description
-
Official documentation · 5 Sep 2026
Fetched. Title 'GPU Monitoring for AI Workloads | Datadog'; 'Datadog GPU Monitoring delivers end-to-end visibility across shared GPU fleets by linking device health, cost, and performance directly to the workloads and teams using them ... whether deployed in cloud, on-prem, or neocloud environments'; 'With proactive alerting and actionable recommendations, GPU Monitoring helps teams optimize efficiency, resolve stalled or failed AI workloads, prevent hardware issues, and reduce wasted spend.'; 'Forecast GPU demand'; 'guided optimization actions like reclaiming GPUs tied up by zombie processes'; 'alerts on workloads or clusters with unmet GPU requests'; 'Detect and remediate ECC/XID errors proactively with built-in alerts and prescriptive next steps'; 'just by toggling a single configuration flag in the Datadog Agent'. Docs H1 'GPU Monitoring': 'provides a centralized view into your GPU fleet's health, cost, and performance.' No sentence on either page names an AI/ML mechanism.
- Sold within
-
Official documentation · 5 Sep 2026
The site's Product menu files 'GPU Monitoring' under the 'Infrastructure' group beside Infrastructure Monitoring, Metrics, Network Monitoring, Container Monitoring, Kubernetes Autoscaling, Serverless, Cloud Cost Management, Cloudcraft and Storage Management; the page's 'Works great with' box names Agent Observability and Infrastructure Monitoring, and data is collected 'by toggling a single configuration flag in the Datadog Agent'. No page states a plan that includes it.
Also from Datadog
-
Agent Observability Developer tool
Traces, evaluates and monitors LLM and AI-agent applications in production, with offline experimentation on datasets built from real traces.
-
Bits Chat Application
Conversational AI interface for querying Datadog metrics, logs, traces and monitors in natural language and generating dashboards and notebooks, accessible from Datadog, Slack or mobile.
-
Bits Investigation AI agent
AI SRE agent that autonomously investigates every alert the moment it fires, explores multiple root-cause hypotheses in parallel and reports findings into Slack, Jira, ServiceNow or GitHub.
-
Bits Code AI agent
Coding agent that triages production errors, regressions and vulnerabilities from Datadog telemetry, generates fixes with unit tests grounded in logs, traces and runtime variables, and opens pull requests for review.
-
Bits Agent Builder Agent platform
No-code builder for custom AI agents that investigate, decide and act inside Datadog to automate incident response, observability, security and operational workflows.
-
Bits Security Analyst AI agent
Always-on AI SOC analyst that autonomously triages and investigates security alerts and delivers written investigation results to Datadog, Slack or Jira within minutes.
-
Datadog MCP Server Developer tool
Model Context Protocol server that gives AI coding agents such as Claude Code, Cursor and Codex secure real-time access to Datadog logs, metrics and traces under existing RBAC controls.
-
AI Impact Application
Detects AI-assisted pull requests from coding assistants such as Claude Code, Cursor and GitHub Copilot and compares them on adoption, PR throughput, cycle time, change failure rate and daily cost per active user.
-
Watchdog Application
Datadog's built-in AI engine: it continuously analyses metrics, traces and logs across the platform to raise anomaly alerts without configuration, detect faulty deployments by comparing code versions, run automated root-cause analysis on critical failures and surface tag-based insights and impact analysis. Available inside Infrastructure Monitoring, APM, Log Management and RUM rather than sold on its own.
Something wrong here? Send a correction — quote this product id: datadog-gpu-monitoring.