Alert Runbooks

etcdHighNumberOfFailedProposals

etcdHighNumberOfFailedProposals

Description

This alert fires when an etcd member’s rate of failed Raft proposals over the last 15 minutes exceeds the configured threshold (default: 5).

A “proposal” in etcd is a Raft consensus operation — typically triggered by a write request — that must be agreed upon by a quorum of members before being committed. Failed proposals mean writes are not being successfully committed to the cluster, which directly affects the Kubernetes API server’s ability to persist cluster state (pods, deployments, configmaps, secrets, etc.).


Possible Causes:


Severity estimation

High to Critical severity — failed proposals mean writes are not durable.

Impact assessment:


Troubleshooting steps

  1. Check the current failed proposal rate

    • Command / Action:
      • Query the metric behind the alert
      • rate(etcd_server_proposals_failed_total{job=~".etcd."}[15m])

    • Expected result:
      • Value at or near 0; alert fires when sustained above 5
    • additional info:
      • Break down by (instance) to identify whether one member or the whole cluster is affected

  1. Check etcd cluster health and current leader

    • Command / Action:
      • Verify a stable leader exists and all members are healthy
      • kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint health –cluster –write-out=table

      • kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint status –cluster –write-out=table

    • Expected result:
      • All members healthy: true; exactly one member shows IS LEADER: true
    • additional info:
      • No leader or a flapping leader means proposals will fail until stability returns — refer to EtcdNoLeader and EtcdHighNumberOfLeaderChanges

  1. Check recent leader election activity

    • Command / Action:
      • Correlate proposal failures with leader change events
      • increase(etcd_server_leader_changes_seen_total[1h])

    • Expected result:
      • Value close to 0; a non-zero value correlating with the failed proposal spike points to election-related failures
    • additional info:
      • Proposals submitted during an election window are expected to fail and retry; this is normal if brief, but a problem if elections are frequent

  1. Check member communication latency

    • Command / Action:
      • Verify peer round-trip time is within acceptable bounds
      • histogram_quantile(0.99, rate(etcd_network_peer_round_trip_time_seconds_bucket{job=~".etcd."}[5m]))

    • Expected result:
      • Below 0.15s; higher values indicate network issues preventing quorum agreement
    • additional info:
      • Refer to EtcdMemberCommunicationSlow for detailed network troubleshooting steps

  1. Check disk I/O and WAL fsync latency

    • Command / Action:
      • etcd proposals require a durable WAL write to commit; slow disks cause failures under load
      • histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket{job=~".etcd."}[5m]))

      • iostat -x 2 5

    • Expected result:
      • Fsync duration below 0.5s; disk await in single-digit milliseconds
    • additional info:
      • Refer to etcdHighFsyncDurations for detailed disk troubleshooting steps

  1. Check etcd member and quorum status

    • Command / Action:
      • Confirm enough members are up to maintain quorum
      • kubectl get pods -n kube-system -l component=etcd -o wide

      • kubectl exec -n kube-system <etcd-pod> – etcdctl member list –write-out=table

    • Expected result:
      • (N/2)+1 members healthy and reachable for an N-member cluster
    • additional info:
      • If quorum is lost, all proposals will fail until enough members recover — refer to EtcdInsufficientMembers

  1. Check etcd logs for proposal-related errors

    • Command / Action:
      • Look for specific failure reasons in etcd logs
      • kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘proposal|failed|timeout|quorum’

    • Expected result:
      • No recurring proposal failure messages
    • additional info:
      • Look for proposal dropped or raft: proposal is queued for a while patterns and correlate with timestamps of resource pressure or network issues

  1. Check storage quota and database size

    • Command / Action:
      • Verify the database is not near its quota, which causes writes (and their proposals) to be rejected
      • kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint status –write-out=table

      • kubectl exec -n kube-system <etcd-pod> – etcdctl alarm list

    • Expected result:
      • DB SIZE well below the configured --quota-backend-bytes; no active alarms
    • additional info:
      • A NOSPACE alarm puts etcd in a read-only state, causing all write proposals to fail — refer to EtcdDatabaseQuotaLowSpace

  1. Check CPU and memory pressure on etcd members

    • Command / Action:
      • Verify no member is resource-starved
      • kubectl top pod -n kube-system -l component=etcd

    • Expected result:
      • CPU and memory usage within normal bounds; no OOMKilled events
    • additional info:
      • etcd should generally not have restrictive CPU limits applied, as throttling directly increases proposal processing time

  1. Verify the failed proposal rate has returned to normal

    • Command / Action:
      • Re-run the Prometheus query used in step 1
      • rate(etcd_server_proposals_failed_total{job=~".etcd."}[15m])

    • Expected result:
      • Value has dropped back to at or near 0
    • additional info:
      • Also verify the Kubernetes API server is responding normally: kubectl get nodes and kubectl get pods -A return quickly without timeout errors

Additional resources