RDMA Latency Monitoring Guide¶
Background¶
What Is RDMA Latency Monitoring¶
RDMA latency monitoring is a network performance monitoring feature provided by Unifabric. It continuously monitors the latency performance of the RDMA network in a Kubernetes cluster. By automatically running RDMA latency tests between cluster nodes, it helps operators:
- Detect network performance issues early: Continuously monitor latency changes and quickly identify network performance degradation.
- Validate network topology: Verify that the network topology is configured correctly through latency tests.
- Locate faulty nodes: Pinpoint the nodes and NICs that show abnormal latency.
- Optimize network configuration: Provide data support for network optimization.
How It Works¶
The unifabric-agent on each node contains a standalone rdma-latency container, which is responsible for:
- Starting the RDMA test server: Starting the
ib_write_latandib_read_latserver processes on each RDMA NIC. - Generating test tasks: Automatically generating latency test tasks between nodes based on the configured policy (
none/topo). - Running latency tests: Acting as a client to connect to the server on other nodes and running RDMA Write/Read latency tests.
- Exporting monitoring metrics: Exposing test results in Prometheus metric format for collection by the monitoring system.
Key Features¶
- Automated testing: Automatically generates test tasks based on ScaleoutLeafGroup topology information.
- Multiple test policies: Supports the
none(simple pairing) andtopo(topology-based) policies. - Full coverage: Covers multiple network paths such as intra-group same-rail, intra-group cross-rail, and cross-group same-rail.
- Bidirectional testing: Tests both RDMA Write and RDMA Read latency.
- Prometheus integration: Test results are exposed as Prometheus metrics for easy monitoring and alerting.
- Continuous monitoring: Runs tests periodically to detect performance degradation in time.
Test Plan¶
Test Policy Overview¶
RDMA latency monitoring supports two test policies:
| Policy | Applicable scenario | Prerequisites | Test complexity |
|---|---|---|---|
| none | Small clusters, simple network topology | None | Low |
| topo | Large clusters, fine-grained monitoring | Requires the RDMA neighbor discovery feature (ScaleoutLeafGroup resource) | High |
None Policy (Simple Pairing)¶
Test Logic¶
- Sort all nodes in lexicographical order.
- Pair nodes for testing (0-1, 2-3, 4-5...).
- If the total number of nodes is odd:
- 3 nodes: node 0 tests nodes 1 and 2.
- More than 3 nodes: the first node additionally tests the last node.
Client and Server Roles¶
Under the none policy, within each node pair:
- Server: The node that comes later in lexicographical order (for example, node-2, node-4).
- Client: The node that comes earlier in lexicographical order (for example, node-1, node-3).
Example (4 Nodes)¶
Client node-1 -> Server node-2 (tests all RDMA NIC pairs on the same subnet)
Client node-3 -> Server node-4 (tests all RDMA NIC pairs on the same subnet)
Topo Policy (Topology-Based)¶
Prerequisites¶
The ScaleoutLeafGroup resource must exist in the cluster. This resource is created automatically through LLDP neighbor discovery and defines the grouping and topology relationship of nodes.
Test Scenarios¶
The topo policy contains three test scenarios covering different network paths:
Scenario 1: Intra-Group Same-Rail Test¶
Test goal: Verify the latency between nodes on the same subnet (rail) under the same Leaf switch.
Rail definition: NICs with the same IP subnet belong to the same rail, or NICs with the same name belong to the same rail.
- For example: the 172.17.3.0/24 subnet is rail 1, and the 172.17.4.0/24 subnet is rail 2.
Test scope: All nodes in the group are paired with each other to test NICs on the same rail.
Client and server:
- Each node acts as both client and server.
- Node A acts as the client to test node B.
- Node B acts as the client to test node A.
Network path: Traffic passes through the same Leaf switch.
Example:
ScaleoutLeafGroup: group-1 (connected to Leaf-1)
Nodes: node-1, node-2, node-3, node-4
Test pairs (rail 1: 172.17.3.0/24):
Client node-1 (172.17.3.141) -> Server node-2 (172.17.3.142)
Client node-3 (172.17.3.143) -> Server node-4 (172.17.3.144)
Metric label: latency_type="sameleaf_samerail"
Scenario 2: Intra-Group Cross-Rail Test¶
Test goal: Verify the latency between nodes on different subnets (rails) under the same Leaf switch.
Test scope: For each group, select the first and the last node in lexicographical order for cross-rail testing.
Client and server:
- Client: The first node (for example, node-1).
- Server: The last node (for example, node-4).
Network path: Traffic may pass through the Spine switch (depending on the Leaf switch configuration).
Example:
ScaleoutLeafGroup: group-1
Nodes: node-1, node-2, node-3, node-4
Test pairs (cross-rail):
Client node-1 (rail 1: 172.17.1.141) -> Server node-4 (rail 2: 172.17.2.144)
Client node-1 (rail 1: 172.17.1.141) -> Server node-4 (rail 3: 172.17.3.144)
Client node-1 (rail 1: 172.17.1.141) -> Server node-4 (rail 4: 172.17.4.144)
Client node-1 (rail 2: 172.17.2.142) -> Server node-4 (rail 1: 172.17.1.144)
Client node-1 (rail 2: 172.17.2.142) -> Server node-4 (rail 3: 172.17.3.144)
Client node-1 (rail 2: 172.17.2.142) -> Server node-4 (rail 4: 172.17.4.144)
...
Metric label: latency_type="sameleaf_crossrail"
Scenario 3: Cross-Group Same-Rail Test¶
Test goal: Verify the latency between nodes on the same subnet (rail) across different Leaf switches.
Prerequisites: At least 2 ScaleoutLeafGroups.
Test scope: Select the first node in lexicographical order from each group, and test it against the first node of the other groups.
Client and server:
- The first node in lexicographical order of each group is the client.
- The first node of other groups is the server.
- Nodes in different groups test each other.
Network path: Traffic passes through the Spine switch.
Example:
ScaleoutLeafGroup: group-1 (connected to Leaf-1)
Nodes: node-1, node-2, node-3, node-4
Test node: node-1
ScaleoutLeafGroup: group-2 (connected to Leaf-2)
Nodes: node-5, node-6, node-7, node-8
Test node: node-5
Test pairs:
Rail 1: 172.17.1.0/24:
Client node-1 (172.17.1.141) -> Server node-5 (172.17.1.145)
Rail 2: 172.17.2.0/24:
Client node-1 (172.17.2.142) -> Server node-5 (172.17.2.145)
Rail 3: 172.17.3.0/24:
Client node-1 (172.17.3.143) -> Server node-5 (172.17.3.145)
Rail 4: 172.17.4.0/24:
Client node-1 (172.17.4.144) -> Server node-5 (172.17.4.145)
Metric label: latency_type="crossgroup_samerail"
Monitoring Metric Format¶
All test results are exposed as Prometheus metrics:
Metric Name¶
Metric Type¶
Gauge - the average latency (in microseconds) of the most recent test.
Metric Labels¶
| Label | Description | Example value |
|---|---|---|
source_node | Source node (client) name | node-1 |
source_rdma_device | Source node RDMA device name | mlx5_0 |
source_ifname | Source node NIC interface name | ens841np0 |
source_ip | Source node IP address | 172.17.3.141 |
target_node | Target node (server) name | node-2 |
target_rdma_device | Target node RDMA device name | mlx5_0 |
target_ifname | Target node NIC interface name | ens841np0 |
target_ip | Target node IP address | 172.17.3.142 |
source_group | ScaleoutLeafGroup of the source node | group-1 |
target_group | ScaleoutLeafGroup of the target node | group-1 |
latency_type | Latency test type | sameleaf_samerail / sameleaf_crossrail / crossgroup_samerail |
test_strategy | Test policy | pairwise / topo |
state | Test state | success / failed |
error_msg | Error message (only on failure) | connection timeout |
Metric Examples¶
Successful same-rail test:
unifaric_rdma_latency_us{
source_node="node-1",
source_rdma_device="mlx5_0",
source_ifname="ens841np0",
source_ip="172.17.3.141",
target_node="node-2",
target_rdma_device="mlx5_0",
target_ifname="ens841np0",
target_ip="172.17.3.142",
source_group="group-1",
target_group="group-1",
latency_type="sameleaf_samerail",
test_strategy="topo",
state="success",
error_msg=""
} 2.45
Failed cross-rail test:
unifaric_rdma_latency_us{
source_node="node-3",
source_rdma_device="mlx5_0",
source_ifname="ens841np0",
source_ip="172.17.3.143",
target_node="node-4",
target_rdma_device="mlx5_1",
target_ifname="ens842np0",
target_ip="172.17.4.144",
source_group="group-1",
target_group="group-1",
latency_type="sameleaf_crossrail",
test_strategy="topo",
state="failed",
error_msg="connection timeout"
} 0
Configuration and Installation¶
Configuration Examples¶
Small-Scale Cluster (none Policy)¶
features:
rdmaLatencyDetection:
interval: 3m
workers: 3
gpuNetwork:
enabled: true
strategy: pairwise
testServerPorts: "20000-20017"
Large-Scale Cluster (topo Policy)¶
features:
rdmaNeighbor:
enabled: true
enableScaleOutLeafGroup: true
gpuRdmaNicFilter: "interface=ens*"
rdmaLatencyDetection:
interval: 5m
workers: 10
gpuNetwork:
enabled: true
strategy: topo
testServerPorts: "20000-20017"
Helm configuration parameters:
| Parameter | Type | Default value | Description |
|---|---|---|---|
interval | string | "5m" | Test interval, supports the s/m/h units |
workers | int | 5 | Number of workers that run tests concurrently |
gpuNetwork.enabled | bool | true | Whether to enable GPU RDMA network latency detection |
gpuNetwork.strategy | string | "topo" | GPU network test policy: pairwise (pairwise pairing) or topo (topology-based) |
storageNetwork.enabled | bool | true | Whether to enable storage RDMA network latency detection |
testServerPorts | string | "20000-20010" | Port range of the RDMA test server |
metrics.port | int | 5027 | Port that exposes Prometheus metrics |
Note
- If
strategyis set totopo, make surefeatures.rdmaNeighbor.enableScaleOutLeafGroupis set totrue. - To start the RDMA latency test server properly, make sure
testServerPortsreserves a sufficient port range and avoids port conflicts. Reserve at least two ports for each RDMA NIC: one forib_write_latand one forib_read_lat. For example, if there are 9 RDMA NICs, at least 18 ports are required.
Installation Steps¶
Install or Upgrade the Helm Chart¶
# Or upgrade an existing installation
helm upgrade unifabric ./chart \
--namespace unifabric-system \
-f values.yaml
By default, RDMA latency detection is based on the
topopolicy. Make surefeatures.rdmaNeighbor.enableScaleOutLeafGroupis set totrue.
Verification¶
Container Architecture¶
The unifabric-agent Pod contains 3 containers:
unifabric-agent Pod
├── lldpd (container) # LLDP neighbor discovery service
├── agent (container) # Main business logic, manages FabricNode and other resources
└── rdma-latency (container) # RDMA latency detection service (standalone container)
Responsibilities of the rdma-latency container:
- Starts the RDMA test servers (
ib_write_lat,ib_read_lat). - Runs latency test tasks.
- Exposes Prometheus metrics (port 5027).
Check Pod Status¶
# Check Pod status
kubectl get pods -n unifabric-system -l app=unifabric-agent
# Expected output:
NAME READY STATUS RESTARTS AGE
unifabric-agent-xxxxx 3/3 Running 0 5m
Note: The READY column should show 3/3, meaning all 3 containers are ready.
Check Whether the ScaleoutLeafGroup CRD Is Working¶
A ScaleoutLeafGroup object represents a group of nodes that share the same GPU network topology.
~# kubectl get scaleoutleafgroups.unifabric.io
NAME healthyNodes totalNodes healthy AGE
409bf491f7136a8a 4 4 true 6h1m
Check whether healthy of the scaleoutleafgroups is true. This indicates that Unifabric has correctly collected the GPU network topology information of all nodes.
View Metrics¶
Method 1: View Through Port-Forward¶
# Forward the metrics port
kubectl port-forward -n unifabric-system <pod-name> 5027:5027
# View metrics in another terminal
curl http://localhost:5027/metrics | grep unifaric_rdma_latency_us
...
unifabric_rdma_latency_us{error_msg="",latency_type="sameleaf_samerail",source_group="409bf491f7136a8a",source_ifname="ens834np0",source_ip="172.17.1.141",source_node="sh-cube-master-1",source_rdma_device="mlx5_4",state="success",target_group="409bf491f7136a8a",target_ifname="ens834np0",target_ip="172.17.1.142",target_node="sh-cube-master-2",target_rdma_device="mlx5_4",test_strategy="topo"} 2.994644
unifabric_rdma_latency_us{error_msg="",latency_type="sameleaf_samerail",source_group="409bf491f7136a8a",source_ifname="ens834np0",source_ip="172.17.1.142",source_node="sh-cube-master-2",source_rdma_device="mlx5_4",state="success",target_group="409bf491f7136a8a",target_ifname="ens834np0",target_ip="172.17.1.141",target_node="sh-cube-master-1",target_rdma_device="mlx5_4",test_strategy="topo"} 5.515484
...
The output above shows:
unifaric_rdma_latency_us: the RDMA latency metric is 2.994644 us.error_msg: error message (only on failure).latency_type: test type.source_group: source node group.source_ifname: source NIC interface name.source_ip: source node IP.source_node: source node name.source_rdma_device: source RDMA device name.state: test state.target_group: target node group.target_ifname: target NIC interface name.target_ip: target node IP.target_node: target node name.target_rdma_device: target RDMA device name.test_strategy: test policy.latency_type: test type.
Method 2: View Through Prometheus¶
Run the following queries in the Prometheus UI:
# View all metrics
unifaric_rdma_latency_us
# View successful tests
unifaric_rdma_latency_us{state="success"}
# View failed tests
unifaric_rdma_latency_us{state="failed"}
Verify Same-Rail Metrics¶
Same-rail test metric characteristics:
latency_type="sameleaf_samerail".source_ipandtarget_ipare on the same subnet.- The latency value is usually low (< 5 microseconds).
Example output:
unifaric_rdma_latency_us{
source_node="node-1",
source_ip="172.17.3.141",
target_node="node-2",
target_ip="172.17.3.142",
latency_type="sameleaf_samerail",
state="success"
} 2.3
Verify Cross-Rail Metrics¶
Cross-rail test metric characteristics:
latency_type="sameleaf_crossrail".source_ipandtarget_ipare on different subnets.- The latency value is usually higher (5-15 microseconds).
# Query cross-rail tests
unifaric_rdma_latency_us{latency_type="sameleaf_crossrail", state="success"}
Example output:
unifaric_rdma_latency_us{
source_node="node-3",
source_ip="172.17.3.143",
target_node="node-4",
target_ip="172.17.4.144",
latency_type="sameleaf_crossrail",
state="success"
} 8.7
Troubleshooting¶
PromQL Query Examples¶
-
View all successful latency tests
-
View the latency of a specific node
-
View intra-group same-rail latency
-
View cross-rail latency (may be higher)
-
View failed tests
-
Calculate the average latency
-
Find paths with abnormally high latency (>10 microseconds)
The Container Fails to Start¶
Symptoms¶
kubectl get pods -n unifabric-system -l app=unifabric-agent
# Output shows: READY 2/3 or CrashLoopBackOff
Troubleshooting Steps¶
Step 1: Check the container status
# View Pod details
kubectl describe pod -n unifabric-system <pod-name>
# View the logs of the rdma-latency container
kubectl logs -n unifabric-system <pod-name> -c rdmalatency --tail=100
Step 2: Common errors and solutions
Error 1: Port conflict¶
Error message:
Cause: The configured port range conflicts with other services.
Solution:
# Modify values.yaml and adjust the port range
features:
rdmaLatencyDetection:
testServerPorts: "21000-22000" # Use a different port range
Error 2: The RDMA device does not exist¶
Error message:
Cause: The node has no RDMA NIC or the NIC is not properly recognized.
Troubleshooting:
# Check the RDMA devices on the node
kubectl exec -n unifabric-system <pod-name> -c rdmalatency -- ibv_devices
# Check the FabricNode resource
kubectl get fabricnode <node-name> -o yaml | grep -A 20 gpuNics
Error 3: Insufficient permissions¶
Error message:
Cause: The Pod does not have sufficient permissions to access the RDMA devices.
Solution: Make sure the DaemonSet is configured with hostNetwork: true and an appropriate securityContext.
Adjust the Log Level¶
If you need more detailed logs for troubleshooting:
# Update the configuration
helm upgrade unifabric ./chart -n unifabric-system -f values.yaml
# View detailed logs
kubectl logs -n unifabric-system <pod-name> -c rdma-latency -f
No Metrics Are Generated¶
Possible causes:
- RDMA latency detection is not enabled.
- No RDMA NIC is found.
- The
topopolicy is used but no ScaleoutLeafGroup exists.
Troubleshooting steps:
# 1. Check the configuration
kubectl get cm -n unifabric-system unifabric -o yaml
# 2. Check whether FabricNode contains RDMA NIC information
kubectl get fabricnode <node-name> -o yaml
# 3. View the logs and check whether the Metrics Server started properly
kubectl logs -n unifabric-system daemonset/unifabric-agent -c rdmalatency
Metrics Show Failed Tests (state=failed)¶
Symptoms¶
Troubleshooting Steps¶
Step 1: View the error message
Common error messages:
Error 1: connection timeout¶
Cause: The server on the target node is not started, or the network is unreachable.
Troubleshooting:
# 1. Check the status of the rdma-latency container on the target node
kubectl get pods -n unifabric-system -l app=unifabric-agent -o wide | grep <target-node>
# 2. Check the server process on the target node
kubectl exec -n unifabric-system <target-pod-name> -c rdma-latency -- \
ps aux | grep ib_write_lat
# 3. Check port listening
kubectl exec -n unifabric-system <target-pod-name> -c rdma-latency -- \
netstat -tlnp | grep 20000
# 4. Manually test the RDMA connection
kubectl exec -n unifabric-system <source-pod-name> -c rdma-latency -- \
ib_write_lat -d mlx5_0 -p 20000 <target-ip>
Error 2: no route to host¶
Cause: The network is unreachable or blocked by a firewall.
Troubleshooting:
# 1. Test network connectivity
kubectl exec -n unifabric-system <source-pod-name> -c rdma-latency -- \
ping -c 3 <target-ip>
# 2. Check firewall rules
kubectl exec -n unifabric-system <source-pod-name> -c rdma-latency -- \
iptables -L -n | grep 20000
# 3. Check the route
kubectl exec -n unifabric-system <source-pod-name> -c rdma-latency -- \
ip route get <target-ip>