etcdHighFsyncDurations
etcdHighFsyncDurations
Description
This alert fires when the 99th percentile duration of etcd’s write-ahead log (WAL) fsync operations exceeds the configured threshold (default: 0.5s).
Every write to etcd must be durably persisted to the WAL via an fsync call before it is acknowledged. Because this happens on the critical path of every write, etcd is extremely sensitive to disk latency. High fsync duration directly increases write latency for the entire cluster, and if sustained, can cause heartbeat timeouts, unnecessary leader elections, and failed proposals.
Possible Causes:
- Slow underlying disk — HDDs or network-attached storage with high write latency
- Disk I/O saturation from noisy-neighbor workloads sharing the same disk as etcd’s data directory
- Overloaded storage backend (e.g. cloud provider EBS volume hitting IOPS/throughput limits)
- etcd data directory placed on the wrong storage tier (e.g. non-SSD, or shared network storage instead of local SSD/NVMe)
- Filesystem or kernel-level write barriers/journaling adding overhead
- Virtual machine live migration or hypervisor-level I/O stalls
- Insufficient disk write cache or write-back caching disabled
- RAID controller or storage array degraded state (e.g. a failed disk in a RAID array)
Severity estimation
High severity — this is one of the most direct indicators of etcd disk health.
- High if fsync duration is elevated but the cluster maintains a stable leader and writes are still succeeding (with higher latency)
- Critical if fsync duration is severely elevated and correlates with failed proposals, leader elections, or API server timeouts
- Critical if the underlying disk is failing or degraded — this risks data loss and cluster-wide unavailability
Impact assessment:
- Increased latency for all etcd writes, and therefore all Kubernetes API write operations
- Risk of missed heartbeats leading to unnecessary leader elections (see
EtcdHighNumberOfLeaderChanges) - Risk of failed proposals under sustained load (see
etcdHighNumberOfFailedProposals) - In severe cases, cascading control plane slowness affecting the entire cluster
Troubleshooting steps
-
Check the current WAL fsync duration
- Command / Action:
- Query the metric behind the alert
-
histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket{job=~".etcd."}[5m]))
- Expected result:
- Value below
0.5s, ideally in single-digit milliseconds
- Value below
- additional info:
- Break down
by (instance)to identify whether one member or the whole cluster is affected
- Break down
- Command / Action:
-
Check disk I/O latency on the etcd node
- Command / Action:
- Measure real-time disk latency and utilization
-
iostat -x 2 5
- Expected result:
awaitunder ~1ms for the disk backing etcd’s data directory;%utilnot consistently near 100%
- additional info:
- The disk device to check is the one backing etcd’s
--data-dir(default/var/lib/etcd); confirm withdf -h /var/lib/etcdandlsblk
- The disk device to check is the one backing etcd’s
- Command / Action:
-
Identify the etcd data directory’s underlying storage type
- Command / Action:
- Confirm etcd is running on local SSD/NVMe, not HDD or network-attached storage
-
lsblk -d -o NAME,ROTA,TYPE,SIZE
-
mount | grep etcd
- Expected result:
ROTA=0(non-rotational, i.e. SSD/NVMe) for the etcd data disk
- additional info:
ROTA=1indicates a spinning HDD, which cannot reliably meet etcd’s latency requirements; network-attached storage (NFS, some cloud block storage under load) is also a common culprit
- Command / Action:
-
Check for noisy-neighbor workloads sharing the same disk
- Command / Action:
- Identify other processes or pods generating disk I/O on the same node/volume
-
iotop -oa
-
kubectl get pods -A –field-selector spec.nodeName=<node-name> -o wide
- Expected result:
- etcd’s disk is not shared with high-I/O workloads
- additional info:
- If the node is shared, consider dedicating nodes to etcd via taints (
node-role.kubernetes.io/etcd=:NoSchedule) or moving etcd to a dedicated disk
- If the node is shared, consider dedicating nodes to etcd via taints (
- Command / Action:
-
Check cloud provider disk throughput/IOPS limits
- Command / Action:
- If running on a cloud provider, check whether the disk is hitting provisioned IOPS or throughput limits
- Review cloud provider monitoring for the volume backing etcd’s data directory
- Expected result:
- Disk usage well within provisioned IOPS/throughput limits
- additional info:
- Under-provisioned volumes (e.g. gp2/gp3 with low baseline IOPS) commonly cause fsync spikes under load; consider upgrading to a higher-performance volume type
- Command / Action:
-
Run a direct disk latency benchmark
- Command / Action:
- Use
fioto directly benchmark write latency on the etcd data directory (on a non-production node/volume, or during a maintenance window) -
fio –rw=write –ioengine=sync –fdatasync=1 –directory=/var/lib/etcd –size=22m –bs=2300 –name=etcd-fsync-test
- Use
- Expected result:
- 99th percentile fdatasync duration under ~10ms, matching etcd’s hardware recommendations
- additional info:
- This is etcd’s own recommended benchmark methodology; results below this threshold indicate the disk itself is not the bottleneck, pointing to noisy neighbors or contention instead
- Command / Action:
-
Check for hardware/RAID degradation
- Command / Action:
- Check for disk or RAID array health issues
-
smartctl -a /dev/<disk>
-
cat /proc/mdstat (if using software RAID)
- Expected result:
- No SMART errors, reallocated sectors, or degraded RAID state
- additional info:
- A failing disk can cause intermittent latency spikes before fully failing; treat any SMART warnings as urgent
- Command / Action:
-
Check etcd logs for slow fsync warnings
- Command / Action:
- Review etcd logs for explicit slow-disk warnings
-
kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘slow fdatasync|wal: sync duration|took too long’
- Expected result:
- No recurring slow fsync warnings
- additional info:
- etcd logs a warning whenever a WAL fsync takes longer than expected; frequency and magnitude of these warnings help gauge severity
- Command / Action:
-
Mitigate: move etcd data directory to faster storage
- Command / Action:
- If the disk is confirmed to be the bottleneck, migrate the etcd member to a node/volume with local SSD/NVMe storage
- Follow etcd member replacement procedure: remove the affected member, provision new storage, rejoin as a new member
- Expected result:
- fsync duration drops significantly after migration to faster storage
- additional info:
- This is a disruptive operation — perform one member at a time to maintain quorum throughout, and take a backup/snapshot before starting
- Command / Action:
-
Verify fsync duration has returned to normal
- Command / Action:
- Re-run the Prometheus query used in step 1
-
histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket{job=~".etcd."}[5m]))
- Expected result:
- Value has dropped back below the
0.5sthreshold
- Value has dropped back below the
- additional info:
- Also confirm commit durations and leader election rate have stabilized, since sustained fsync latency often cascades into both
- Command / Action:
Additional resources
- etcd disk performance benchmarking with fio
- etcd hardware recommendations
- etcd tuning — heartbeat interval and election timeout
- Related alert: etcdHighCommitDurations
- Related alert: etcdHighNumberOfFailedProposals
- Related alert: EtcdHighNumberOfLeaderChanges