Alert Runbooks

EtcdMemberCommunicationSlow

EtcdMemberCommunicationSlow

Description

This alert fires when the 99th percentile round-trip time for peer-to-peer communication between etcd members exceeds the configured threshold (default: 0.15s).

etcd members constantly exchange Raft messages (heartbeats, log replication, vote requests) to keep the cluster consistent and to maintain a stable leader. Slow inter-member communication delays log replication and heartbeats, which can trigger unnecessary leader elections (EtcdHighNumberOfLeaderChanges), increase write latency, and — if sustained — risk quorum instability.


Possible Causes:


Severity estimation

Medium to High severity — communication degradation precedes harder failures.


Troubleshooting steps

  1. Identify which member pair is affected

    • Command / Action:
      • Query the metric behind the alert, broken down by source instance and destination
      • histogram_quantile(0.99, sum by (instance, To, le) (rate(etcd_network_peer_round_trip_time_seconds_bucket{job=~".etcd."}[5m])))

    • Expected result:
      • Value below 0.15s; alert labels instance and To identify the source and destination members
    • additional info:
      • If only one To target is slow across multiple sources, the problem lies with that specific member or its network path

  1. Check network latency directly between the affected peers

    • Command / Action:
      • From the source etcd node, ping/traceroute the destination peer IP
      • ping <etcd-peer-ip>

      • mtr -rw <etcd-peer-ip>

    • Expected result:
      • Round-trip latency under ~5ms for same-datacenter deployments; no packet loss
    • additional info:
      • Higher baseline latency is expected for multi-AZ/multi-region clusters; compare against the cluster’s known topology and configured --heartbeat-interval/--election-timeout

  1. Check for packet loss or link instability

    • Command / Action:
      • Run a sustained connectivity test between the affected nodes
      • mtr –report –report-cycles 100 <etcd-peer-ip>

    • Expected result:
      • 0% packet loss across the path
    • additional info:
      • Any non-zero loss on the direct path is a strong signal of a network fault; escalate to network/infra team if loss is confirmed

  1. Check CPU and load on both etcd nodes

    • Command / Action:
      • Verify neither node is CPU-starved, which can delay message send/receive processing
      • kubectl top node <node-name>

      • kubectl describe node <node-name>

    • Expected result:
      • CPU usage within normal bounds; no MemoryPressure or DiskPressure conditions
    • additional info:
      • A CPU-starved etcd process can appear as slow peer communication even when the network itself is healthy

  1. Check for network policies or firewall rules affecting peer port 2380

    • Command / Action:
      • Verify no network policy, security group, or firewall rule is throttling or rate-limiting peer traffic
      • kubectl get networkpolicy -n kube-system

      • Check cloud provider security group rules for the etcd peer port (default 2380)
    • Expected result:
      • No restrictive policies affecting inter-member traffic
    • additional info:
      • Recently applied network policies or security group changes are a common cause of sudden communication degradation

  1. Check for competing network traffic on the same node/interface

    • Command / Action:
      • Check network interface utilization on the etcd nodes
      • iftop -i <interface>

      • kubectl top pod -n kube-system -l component=etcd

    • Expected result:
      • Network interface is not saturated; no noisy-neighbor workloads competing for bandwidth
    • additional info:
      • If etcd nodes are shared with other high-bandwidth workloads, consider dedicating nodes to etcd or applying QoS/traffic shaping

  1. Check etcd logs for related warnings

    • Command / Action:
      • Review etcd logs for slow peer or heartbeat-related messages
      • kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘peer|heartbeat|slow|took too long’

    • Expected result:
      • No recurring warnings about slow peer connections
    • additional info:
      • Messages like failed to send out heartbeat on time reinforce that peer communication delays are impacting cluster stability

  1. Verify round-trip time has returned to normal

    • Command / Action:
      • Re-run the Prometheus query used in step 1
      • histogram_quantile(0.99, sum by (instance, To, le) (rate(etcd_network_peer_round_trip_time_seconds_bucket{job=~".etcd."}[5m])))

    • Expected result:
      • Value has dropped back below the 0.15s threshold for all peer pairs
    • additional info:
      • Also confirm no new leader elections occurred during the incident: increase(etcd_server_leader_changes_seen_total[1h])

Additional resources