etcdHighNumberOfFailedProposals
etcdHighNumberOfFailedProposals
Description
This alert fires when an etcd member’s rate of failed Raft proposals over the last 15 minutes exceeds the configured threshold (default: 5).
A “proposal” in etcd is a Raft consensus operation — typically triggered by a write request — that must be agreed upon by a quorum of members before being committed. Failed proposals mean writes are not being successfully committed to the cluster, which directly affects the Kubernetes API server’s ability to persist cluster state (pods, deployments, configmaps, secrets, etc.).
Possible Causes:
- Leader election in progress or recently completed — proposals fail while no stable leader exists
- Network partition or high latency between members preventing quorum agreement (see
EtcdMemberCommunicationSlow) - Disk I/O saturation causing slow WAL writes and proposal timeouts
- etcd member overloaded — too many concurrent proposals exceeding server capacity
- Quorum loss — insufficient healthy members to commit proposals (see
EtcdInsufficientMembers) - Clock skew between members interfering with Raft timing
- A member falling significantly behind on log replication (slow follower)
- Storage quota exhausted, causing writes to be rejected at the proposal stage
Severity estimation
High to Critical severity — failed proposals mean writes are not durable.
- High if failures are intermittent and correlate with a recent, resolved leader election
- Critical if failures are sustained and the Kubernetes API server is experiencing write errors or timeouts
- Critical if paired with
EtcdNoLeader,EtcdInsufficientMembers, orEtcdMemberCommunicationSlow— indicates active quorum or availability issues
Impact assessment:
- Kubernetes API server write operations (create/update/delete) fail or time out
- Controllers cannot update object status, causing reconciliation loops to stall
- New pod scheduling, scaling, and rollout operations may fail
- Sustained failures risk cluster-wide control plane unavailability
Troubleshooting steps
-
Check the current failed proposal rate
- Command / Action:
- Query the metric behind the alert
-
rate(etcd_server_proposals_failed_total{job=~".etcd."}[15m])
- Expected result:
- Value at or near 0; alert fires when sustained above
5
- Value at or near 0; alert fires when sustained above
- additional info:
- Break down
by (instance)to identify whether one member or the whole cluster is affected
- Break down
- Command / Action:
-
Check etcd cluster health and current leader
- Command / Action:
- Verify a stable leader exists and all members are healthy
-
kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint health –cluster –write-out=table
-
kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint status –cluster –write-out=table
- Expected result:
- All members
healthy: true; exactly one member showsIS LEADER: true
- All members
- additional info:
- No leader or a flapping leader means proposals will fail until stability returns — refer to
EtcdNoLeaderandEtcdHighNumberOfLeaderChanges
- No leader or a flapping leader means proposals will fail until stability returns — refer to
- Command / Action:
-
Check recent leader election activity
- Command / Action:
- Correlate proposal failures with leader change events
-
increase(etcd_server_leader_changes_seen_total[1h])
- Expected result:
- Value close to 0; a non-zero value correlating with the failed proposal spike points to election-related failures
- additional info:
- Proposals submitted during an election window are expected to fail and retry; this is normal if brief, but a problem if elections are frequent
- Command / Action:
-
Check member communication latency
- Command / Action:
- Verify peer round-trip time is within acceptable bounds
-
histogram_quantile(0.99, rate(etcd_network_peer_round_trip_time_seconds_bucket{job=~".etcd."}[5m]))
- Expected result:
- Below
0.15s; higher values indicate network issues preventing quorum agreement
- Below
- additional info:
- Refer to
EtcdMemberCommunicationSlowfor detailed network troubleshooting steps
- Refer to
- Command / Action:
-
Check disk I/O and WAL fsync latency
- Command / Action:
- etcd proposals require a durable WAL write to commit; slow disks cause failures under load
-
histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket{job=~".etcd."}[5m]))
-
iostat -x 2 5
- Expected result:
- Fsync duration below
0.5s; diskawaitin single-digit milliseconds
- Fsync duration below
- additional info:
- Refer to
etcdHighFsyncDurationsfor detailed disk troubleshooting steps
- Refer to
- Command / Action:
-
Check etcd member and quorum status
- Command / Action:
- Confirm enough members are up to maintain quorum
-
kubectl get pods -n kube-system -l component=etcd -o wide
-
kubectl exec -n kube-system <etcd-pod> – etcdctl member list –write-out=table
- Expected result:
(N/2)+1members healthy and reachable for an N-member cluster
- additional info:
- If quorum is lost, all proposals will fail until enough members recover — refer to
EtcdInsufficientMembers
- If quorum is lost, all proposals will fail until enough members recover — refer to
- Command / Action:
-
Check etcd logs for proposal-related errors
- Command / Action:
- Look for specific failure reasons in etcd logs
-
kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘proposal|failed|timeout|quorum’
- Expected result:
- No recurring proposal failure messages
- additional info:
- Look for
proposal droppedorraft: proposal is queued for a whilepatterns and correlate with timestamps of resource pressure or network issues
- Look for
- Command / Action:
-
Check storage quota and database size
- Command / Action:
- Verify the database is not near its quota, which causes writes (and their proposals) to be rejected
-
kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint status –write-out=table
-
kubectl exec -n kube-system <etcd-pod> – etcdctl alarm list
- Expected result:
DB SIZEwell below the configured--quota-backend-bytes; no active alarms
- additional info:
- A
NOSPACEalarm puts etcd in a read-only state, causing all write proposals to fail — refer toEtcdDatabaseQuotaLowSpace
- A
- Command / Action:
-
Check CPU and memory pressure on etcd members
- Command / Action:
- Verify no member is resource-starved
-
kubectl top pod -n kube-system -l component=etcd
- Expected result:
- CPU and memory usage within normal bounds; no OOMKilled events
- additional info:
- etcd should generally not have restrictive CPU limits applied, as throttling directly increases proposal processing time
- Command / Action:
-
Verify the failed proposal rate has returned to normal
- Command / Action:
- Re-run the Prometheus query used in step 1
-
rate(etcd_server_proposals_failed_total{job=~".etcd."}[15m])
- Expected result:
- Value has dropped back to at or near 0
- additional info:
- Also verify the Kubernetes API server is responding normally:
kubectl get nodesandkubectl get pods -Areturn quickly without timeout errors
- Also verify the Kubernetes API server is responding normally:
- Command / Action:
Additional resources
- etcd Raft consensus overview
- etcd tuning — heartbeat interval and election timeout
- etcd maintenance — compaction and defragmentation
- Related alert: EtcdNoLeader
- Related alert: EtcdInsufficientMembers
- Related alert: EtcdMemberCommunicationSlow
- Related alert: etcdHighFsyncDurations
- Related alert: EtcdHighNumberOfLeaderChanges