EtcdMemberCommunicationSlow
EtcdMemberCommunicationSlow
Description
This alert fires when the 99th percentile round-trip time for peer-to-peer communication between etcd members exceeds the configured threshold (default: 0.15s).
etcd members constantly exchange Raft messages (heartbeats, log replication, vote requests) to keep the cluster consistent and to maintain a stable leader. Slow inter-member communication delays log replication and heartbeats, which can trigger unnecessary leader elections (EtcdHighNumberOfLeaderChanges), increase write latency, and — if sustained — risk quorum instability.
Possible Causes:
- Network latency or congestion between nodes hosting etcd members
- Packet loss or an unstable network link between etcd peers
- CPU starvation on one or more etcd nodes, delaying message processing
- Etcd members located across availability zones or regions with inherently higher latency than the configured timeouts assume
- Network policies, firewalls, or security groups throttling traffic on the etcd peer port (2380)
- Overloaded network interface due to competing traffic on the same node
- DNS resolution delays affecting peer connection setup
- One specific member with a hardware or network fault (identifiable via the
Tolabel)
Severity estimation
Medium to High severity — communication degradation precedes harder failures.
- Medium if round-trip time is elevated for a single peer pair but the cluster remains stable with a consistent leader
- High if multiple peer pairs are affected, or if it is accompanied by increased leader elections
- Critical if paired with
EtcdHighNumberOfLeaderChanges,EtcdNoLeader, orEtcdInsufficientMembers— communication issues have escalated to cluster instability
Troubleshooting steps
-
Identify which member pair is affected
- Command / Action:
- Query the metric behind the alert, broken down by source instance and destination
-
histogram_quantile(0.99, sum by (instance, To, le) (rate(etcd_network_peer_round_trip_time_seconds_bucket{job=~".etcd."}[5m])))
- Expected result:
- Value below
0.15s; alert labelsinstanceandToidentify the source and destination members
- Value below
- additional info:
- If only one
Totarget is slow across multiple sources, the problem lies with that specific member or its network path
- If only one
- Command / Action:
-
Check network latency directly between the affected peers
- Command / Action:
- From the source etcd node, ping/traceroute the destination peer IP
-
ping <etcd-peer-ip>
-
mtr -rw <etcd-peer-ip>
- Expected result:
- Round-trip latency under ~5ms for same-datacenter deployments; no packet loss
- additional info:
- Higher baseline latency is expected for multi-AZ/multi-region clusters; compare against the cluster’s known topology and configured
--heartbeat-interval/--election-timeout
- Higher baseline latency is expected for multi-AZ/multi-region clusters; compare against the cluster’s known topology and configured
- Command / Action:
-
Check for packet loss or link instability
- Command / Action:
- Run a sustained connectivity test between the affected nodes
-
mtr –report –report-cycles 100 <etcd-peer-ip>
- Expected result:
- 0% packet loss across the path
- additional info:
- Any non-zero loss on the direct path is a strong signal of a network fault; escalate to network/infra team if loss is confirmed
- Command / Action:
-
Check CPU and load on both etcd nodes
- Command / Action:
- Verify neither node is CPU-starved, which can delay message send/receive processing
-
kubectl top node <node-name>
-
kubectl describe node <node-name>
- Expected result:
- CPU usage within normal bounds; no
MemoryPressureorDiskPressureconditions
- CPU usage within normal bounds; no
- additional info:
- A CPU-starved etcd process can appear as slow peer communication even when the network itself is healthy
- Command / Action:
-
Check for network policies or firewall rules affecting peer port 2380
- Command / Action:
- Verify no network policy, security group, or firewall rule is throttling or rate-limiting peer traffic
-
kubectl get networkpolicy -n kube-system
- Check cloud provider security group rules for the etcd peer port (default 2380)
- Expected result:
- No restrictive policies affecting inter-member traffic
- additional info:
- Recently applied network policies or security group changes are a common cause of sudden communication degradation
- Command / Action:
-
Check for competing network traffic on the same node/interface
- Command / Action:
- Check network interface utilization on the etcd nodes
-
iftop -i <interface>
-
kubectl top pod -n kube-system -l component=etcd
- Expected result:
- Network interface is not saturated; no noisy-neighbor workloads competing for bandwidth
- additional info:
- If etcd nodes are shared with other high-bandwidth workloads, consider dedicating nodes to etcd or applying QoS/traffic shaping
- Command / Action:
-
Check etcd logs for related warnings
- Command / Action:
- Review etcd logs for slow peer or heartbeat-related messages
-
kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘peer|heartbeat|slow|took too long’
- Expected result:
- No recurring warnings about slow peer connections
- additional info:
- Messages like
failed to send out heartbeat on timereinforce that peer communication delays are impacting cluster stability
- Messages like
- Command / Action:
-
Verify round-trip time has returned to normal
- Command / Action:
- Re-run the Prometheus query used in step 1
-
histogram_quantile(0.99, sum by (instance, To, le) (rate(etcd_network_peer_round_trip_time_seconds_bucket{job=~".etcd."}[5m])))
- Expected result:
- Value has dropped back below the
0.15sthreshold for all peer pairs
- Value has dropped back below the
- additional info:
- Also confirm no new leader elections occurred during the incident:
increase(etcd_server_leader_changes_seen_total[1h])
- Also confirm no new leader elections occurred during the incident:
- Command / Action:
Additional resources
- etcd tuning — heartbeat interval and election timeout
- etcd hardware and network recommendations
- Kubernetes etcd cluster administration
- Related alert: EtcdHighNumberOfLeaderChanges
- Related alert: EtcdNoLeader
- Related alert: EtcdInsufficientMembers