Alert Runbooks

etcdHighFsyncDurations

etcdHighFsyncDurations

Description

This alert fires when the 99th percentile duration of etcd’s write-ahead log (WAL) fsync operations exceeds the configured threshold (default: 0.5s).

Every write to etcd must be durably persisted to the WAL via an fsync call before it is acknowledged. Because this happens on the critical path of every write, etcd is extremely sensitive to disk latency. High fsync duration directly increases write latency for the entire cluster, and if sustained, can cause heartbeat timeouts, unnecessary leader elections, and failed proposals.


Possible Causes:


Severity estimation

High severity — this is one of the most direct indicators of etcd disk health.

Impact assessment:


Troubleshooting steps

  1. Check the current WAL fsync duration

    • Command / Action:
      • Query the metric behind the alert
      • histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket{job=~".etcd."}[5m]))

    • Expected result:
      • Value below 0.5s, ideally in single-digit milliseconds
    • additional info:
      • Break down by (instance) to identify whether one member or the whole cluster is affected

  1. Check disk I/O latency on the etcd node

    • Command / Action:
      • Measure real-time disk latency and utilization
      • iostat -x 2 5

    • Expected result:
      • await under ~1ms for the disk backing etcd’s data directory; %util not consistently near 100%
    • additional info:
      • The disk device to check is the one backing etcd’s --data-dir (default /var/lib/etcd); confirm with df -h /var/lib/etcd and lsblk

  1. Identify the etcd data directory’s underlying storage type

    • Command / Action:
      • Confirm etcd is running on local SSD/NVMe, not HDD or network-attached storage
      • lsblk -d -o NAME,ROTA,TYPE,SIZE

      • mount | grep etcd

    • Expected result:
      • ROTA=0 (non-rotational, i.e. SSD/NVMe) for the etcd data disk
    • additional info:
      • ROTA=1 indicates a spinning HDD, which cannot reliably meet etcd’s latency requirements; network-attached storage (NFS, some cloud block storage under load) is also a common culprit

  1. Check for noisy-neighbor workloads sharing the same disk

    • Command / Action:
      • Identify other processes or pods generating disk I/O on the same node/volume
      • iotop -oa

      • kubectl get pods -A –field-selector spec.nodeName=<node-name> -o wide

    • Expected result:
      • etcd’s disk is not shared with high-I/O workloads
    • additional info:
      • If the node is shared, consider dedicating nodes to etcd via taints (node-role.kubernetes.io/etcd=:NoSchedule) or moving etcd to a dedicated disk

  1. Check cloud provider disk throughput/IOPS limits

    • Command / Action:
      • If running on a cloud provider, check whether the disk is hitting provisioned IOPS or throughput limits
      • Review cloud provider monitoring for the volume backing etcd’s data directory
    • Expected result:
      • Disk usage well within provisioned IOPS/throughput limits
    • additional info:
      • Under-provisioned volumes (e.g. gp2/gp3 with low baseline IOPS) commonly cause fsync spikes under load; consider upgrading to a higher-performance volume type

  1. Run a direct disk latency benchmark

    • Command / Action:
      • Use fio to directly benchmark write latency on the etcd data directory (on a non-production node/volume, or during a maintenance window)
      • fio –rw=write –ioengine=sync –fdatasync=1 –directory=/var/lib/etcd –size=22m –bs=2300 –name=etcd-fsync-test

    • Expected result:
      • 99th percentile fdatasync duration under ~10ms, matching etcd’s hardware recommendations
    • additional info:
      • This is etcd’s own recommended benchmark methodology; results below this threshold indicate the disk itself is not the bottleneck, pointing to noisy neighbors or contention instead

  1. Check for hardware/RAID degradation

    • Command / Action:
      • Check for disk or RAID array health issues
      • smartctl -a /dev/<disk>

      • cat /proc/mdstat (if using software RAID)

    • Expected result:
      • No SMART errors, reallocated sectors, or degraded RAID state
    • additional info:
      • A failing disk can cause intermittent latency spikes before fully failing; treat any SMART warnings as urgent

  1. Check etcd logs for slow fsync warnings

    • Command / Action:
      • Review etcd logs for explicit slow-disk warnings
      • kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘slow fdatasync|wal: sync duration|took too long’

    • Expected result:
      • No recurring slow fsync warnings
    • additional info:
      • etcd logs a warning whenever a WAL fsync takes longer than expected; frequency and magnitude of these warnings help gauge severity

  1. Mitigate: move etcd data directory to faster storage

    • Command / Action:
      • If the disk is confirmed to be the bottleneck, migrate the etcd member to a node/volume with local SSD/NVMe storage
      • Follow etcd member replacement procedure: remove the affected member, provision new storage, rejoin as a new member
    • Expected result:
      • fsync duration drops significantly after migration to faster storage
    • additional info:
      • This is a disruptive operation — perform one member at a time to maintain quorum throughout, and take a backup/snapshot before starting

  1. Verify fsync duration has returned to normal

    • Command / Action:
      • Re-run the Prometheus query used in step 1
      • histogram_quantile(0.99, rate(etcd_disk_wal_fsync_duration_seconds_bucket{job=~".etcd."}[5m]))

    • Expected result:
      • Value has dropped back below the 0.5s threshold
    • additional info:
      • Also confirm commit durations and leader election rate have stabilized, since sustained fsync latency often cascades into both

Additional resources