Skip to content

GPU Metrics

This page lists some commonly used GPU metrics.

Cluster Level

Metric Name Description
Number of GPUs Total number of GPUs in the cluster
Average GPU Utilization Average compute utilization of all GPUs in the cluster
Average GPU Memory Utilization Average memory utilization of all GPUs in the cluster
GPU Power Power consumption of all GPUs in the cluster
GPU Temperature Temperature of all GPUs in the cluster
GPU Utilization Details 24-hour usage details of all GPUs in the cluster (includes max, avg, current)
GPU Memory Usage Details 24-hour memory usage details of all GPUs in the cluster (includes min, max, avg, current)
GPU Memory Bandwidth Utilization For example, an Nvidia V100 GPU has a maximum memory bandwidth of 900 GB/sec. If the current memory bandwidth is 450 GB/sec, the utilization is 50%

Node Level

Metric Name Description
GPU Mode Usage mode of GPUs on the node, including full-card mode, MIG mode, vGPU mode
Number of Physical GPUs Total number of physical GPUs on the node
Number of Virtual GPUs Number of vGPU devices created on the node
Number of MIG Instances Number of MIG instances created on the node
GPU Memory Allocation Rate Memory allocation rate of all GPUs on the node
Average GPU Utilization Average compute utilization of all GPUs on the node
Average GPU Memory Utilization Average memory utilization of all GPUs on the node
GPU Driver Version Driver version information of GPUs on the node
GPU Utilization Details 24-hour usage details of each GPU on the node (includes max, avg, current)
GPU Memory Usage Details 24-hour memory usage details of each GPU on the node (includes min, max, avg, current)

Troubleshoot GPU-related issues based on XID status

The XID message is an error report that the NVIDIA driver prints to the operating system kernel log or event log. XID messages are used to identify GPU error events, providing information such as the error type, error location, and error code in the GPU hardware, NVIDIA software, or application. If the XID exception on the GPU node in the check items is empty, it indicates that there is no XID message. If there is one, you can use the table below to troubleshoot and resolve the problem yourself, or view more XID messages.

XID Message Description
13 Graphics Engine Exception. Usually caused by an array out-of-bounds access or an instruction error, and in rare cases by a hardware issue.
31 GPU memory page fault. Usually caused by an illegal address access by the application, and in very rare cases by a driver or hardware issue.
32 Invalid or corrupted push buffer stream. This event is reported by the DMA controller on the PCIe bus that manages the communication between the NVIDIA driver and the GPU. It is usually caused by a PCI quality issue rather than by your program.
38 Driver firmware error. Usually a driver firmware error rather than a hardware issue.
43 GPU stopped processing. Usually caused by an error in your own application rather than a hardware issue.
45 Preemptive cleanup, due to previous errors -- Most likely to see when running multiple cuda applications and hitting a DBE. Usually caused by a GPU application exit due to your manual exit or another failure (hardware, resource limits, etc.). XID 45 only provides a result, and the specific cause usually requires further log analysis.
48 Double Bit ECC Error (DBE). When an uncorrectable error occurs on the GPU, this event is reported, and the error is also fed back to your application. Resetting the GPU or restarting the node is usually required to clear this error.
61 Internal micro-controller breakpoint/warning. The GPU internal engine has stopped working, and your business has been affected.
62 Internal micro-controller halt. Similar to the trigger scenario of XID 61.
63 ECC page retirement or row remapping recording event. When an application encounters a GPU memory hardware error, the NVIDIA self-correction mechanism retires or remaps the faulty memory region. The retirement and remapping information must be recorded to infoROM to take effect permanently. Volt architecture: the ECC page retirement event is successfully recorded to infoROM. Ampere architecture: the row remapping event is successfully recorded to infoROM.
64 ECC page retirement or row remapper recording failure. Similar to the trigger scenario of XID 63. However, XID 63 indicates that the retirement and remapping information was successfully recorded to infoROM, while XID 64 indicates that the recording operation failed.
68 NVDEC0 Exception. Usually a hardware or driver issue.
74 NVLINK Error. An XID generated by an NVLink hardware error, indicating that the GPU already has a serious hardware failure and needs to be taken offline for repair.
79 GPU has fallen off the bus. The GPU hardware has detected a card drop, and the GPU cannot be detected on the bus, indicating that the GPU already has a serious hardware failure and needs to be taken offline for repair.
92 High single-bit ECC error rate. Hardware or driver failure.
94 Contained ECC error. When an application encounters an uncorrectable GPU memory ECC error, the NVIDIA error containment mechanism attempts to contain the error to the application with the hardware failure to prevent the error from affecting other applications running on the GPU node. When the containment mechanism successfully contains the error, this event is generated, and only the application with the uncorrectable ECC error is affected.
95 Uncontained ECC error. Similar to the trigger scenario of XID 94. However, XID 94 indicates successful containment, while XID 95 indicates failed containment, meaning that all applications running on this GPU have been affected.

Pod Level

Category Metric Name Description
Application Overview GPU - Compute & Memory Pod GPU Utilization Compute utilization of the GPUs used by the current Pod
Pod GPU Memory Utilization Memory utilization of the GPUs used by the current Pod
Pod GPU Memory Usage Memory usage of the GPUs used by the current Pod
Memory Allocation Memory allocation of the GPUs used by the current Pod
Pod GPU Memory Copy Ratio Memory copy ratio of the GPUs used by the current Pod
GPU - Engine Overview GPU Graphics Engine Activity Percentage Percentage of time the Graphics or Compute engine is active during a monitoring cycle
GPU Memory Bandwidth Utilization Memory bandwidth utilization (Memory BW Utilization) indicates the fraction of cycles during which data is sent to or received from the device memory. This value represents the average over the interval, not an instantaneous value. A higher value indicates higher utilization of device memory.
A value of 1 (100%) indicates that a DRAM instruction is executed every cycle during the interval (in practice, a peak of about 0.8 (80%) is the maximum achievable).
A value of 0.2 (20%) indicates that 20% of the cycles during the interval are spent reading from or writing to device memory.
Tensor Core Utilization Percentage of time the Tensor Core pipeline is active during a monitoring cycle
FP16 Engine Utilization Percentage of time the FP16 pipeline is active during a monitoring cycle
FP32 Engine Utilization Percentage of time the FP32 pipeline is active during a monitoring cycle
FP64 Engine Utilization Percentage of time the FP64 pipeline is active during a monitoring cycle
GPU Decode Utilization Decode engine utilization of the GPU
GPU Encode Utilization Encode engine utilization of the GPU
GPU - Temperature & Power GPU Temperature Temperature of all GPUs in the cluster
GPU Power Power consumption of all GPUs in the cluster
GPU Total Power Consumption Total power consumption of the GPUs
GPU - Clock GPU Memory Clock Memory clock frequency
GPU Application SM Clock Application SM clock frequency
GPU Application Memory Clock Application memory clock frequency
GPU Video Engine Clock Video engine clock frequency
GPU Throttle Reasons Reasons for GPU throttling
GPU - Other Details Graphics Engine Activity The fraction of time any part of the graphics or compute engine is active. The graphics engine is active if a graphics/compute context is bound and the graphics/compute pipeline is busy. This value represents the average over the interval, not an instantaneous value.
SM Activity The fraction of time at least one warp is active on the multiprocessor, averaged over all multiprocessors. Note that "active" does not necessarily mean a warp is actively computing. For example, a warp waiting on a memory request is considered active. This value represents the average over the interval, not an instantaneous value. A value of 0.8 or greater is necessary but not sufficient for effective GPU use. A value less than 0.5 may indicate inefficient GPU use. To give a simplified view of the GPU architecture, if the GPU has N SMs, a kernel using N blocks and running for the entire interval corresponds to activity 1 (100%). A kernel using N/5 blocks and running for the entire interval corresponds to activity 0.2 (20%). A kernel using N blocks and running for one-fifth of the interval will also have activity 0.2 (20%) if the SMs are idle. This value is independent of the number of threads per block (see DCGM_FI_PROF_SM_OCCUPANCY).
SM Occupancy The ratio of warps resident on the multiprocessor relative to the maximum number of concurrent warps supported on the multiprocessor. This value represents the average over the interval, not an instantaneous value. Higher occupancy does not necessarily mean higher GPU utilization. For GPU memory bandwidth-bound workloads (see DCGM_FI_PROF_DRAM_ACTIVE), higher occupancy indicates higher GPU utilization. However, if the workload is compute-bound (that is, not limited by GPU memory bandwidth or latency), higher occupancy is not necessarily correlated with higher GPU utilization. Compute occupancy is not simple; it depends on factors such as GPU attributes, the number of threads per block, registers per thread, and shared memory per block. Use the CUDA Occupancy Calculator to explore various occupancy scenarios.
Tensor Activity The fraction of cycles during which the tensor (HMMA / IMMA) pipeline is active. This value represents the average over the interval, not an instantaneous value. A higher value indicates higher tensor core utilization. Activity 1 (100%) corresponds to issuing a tensor instruction every other cycle during the entire interval. Activity 0.2 (20%) may indicate that 20% of the SMs are utilized at 100% for the entire interval, 100% of the SMs are utilized at 20% for the entire interval, 100% of the SMs are utilized at 100% for 20% of the interval, or any combination in between (see DCGM_FI_PROF_SM_ACTIVE to help disambiguate these possibilities).
FP64 Engine Activity The fraction of cycles during which the FP64 (double precision) pipeline is active. This value represents the average over the interval, not an instantaneous value. A higher value indicates higher FP64 core utilization. Activity 1 (100%) corresponds to executing one FP64 instruction per SM every four cycles on Volta during the entire interval. Activity 0.2 (20%) may indicate that 20% of the SMs are utilized at 100% for the entire interval, 100% of the SMs are utilized at 20% for the entire interval, 100% of the SMs are utilized at 100% for 20% of the interval, or any combination in between (see DCGM_FI_PROF_SM_ACTIVE to help disambiguate these possibilities).
FP32 Engine Activity The fraction of cycles during which the FMA (FP32 (single precision) and integer) pipeline is active. This value represents the average over the interval, not an instantaneous value. A higher value indicates higher FP32 core utilization. Activity 1 (100%) corresponds to executing one FP32 instruction every other cycle during the entire interval. Activity 0.2 (20%) may indicate that 20% of the SMs are utilized at 100% for the entire interval, 100% of the SMs are utilized at 20% for the entire interval, 100% of the SMs are utilized at 100% for 20% of the interval, or any combination in between (see DCGM_FI_PROF_SM_ACTIVE to help disambiguate these possibilities).
FP16 Engine Activity The fraction of cycles during which the FP16 (half precision) pipeline is active. This value represents the average over the interval, not an instantaneous value. A higher value indicates higher FP16 core utilization. Activity 1 (100%) corresponds to executing one FP16 instruction every other cycle during the entire interval. Activity 0.2 (20%) may indicate that 20% of the SMs are utilized at 100% for the entire interval, 100% of the SMs are utilized at 20% for the entire interval, 100% of the SMs are utilized at 100% for 20% of the interval, or any combination in between (see DCGM_FI_PROF_SM_ACTIVE to help disambiguate these possibilities).
Memory Bandwidth Utilization The fraction of cycles during which data is sent to or received from the device memory. This value represents the average over the interval, not an instantaneous value. A higher value indicates higher device memory utilization. Activity rate 1 (100%) corresponds to executing one DRAM instruction every cycle during the entire interval (in practice, a peak of about 0.8 (80%) is the maximum achievable). Activity rate 0.2 (20%) means that 20% of the cycles during the interval are reading from or writing to device memory.
NVLink Bandwidth The rate of data transmitted/received over NVLink (excluding protocol headers), in bytes per second. This value represents the average over a period, not an instantaneous value. The rate is an average over a period. For example, if 1 GB of data is transferred in 1 second, the rate is 1 GB/s regardless of whether the data is transferred at a constant rate or in bursts. Theoretically, the maximum NVLink Gen2 bandwidth per link per direction is 25 GB/s.
PCIe Bandwidth The rate of data transmitted/received over the PCIe bus, including protocol headers and data payload, in bytes per second. This value represents the average over a period, not an instantaneous value. The rate is an average over a period. For example, if 1 GB of data is transferred in 1 second, the rate is 1 GB/s regardless of whether the data is transferred at a constant rate or in bursts. Theoretically, the maximum PCIe Gen3 bandwidth is 985 MB/s per lane.
PCIe Transfer Rate Data transfer rate of the GPU through the PCIe bus
PCIe Receive Rate Data receive rate of the GPU through the PCIe bus

Comments