Alert Runbooks

etcdHighCommitDurations

etcdHighCommitDurations

Description

This alert fires when the 99th percentile duration of etcd’s backend (bolt/bbolt) commit operations exceeds the configured threshold (default: 0.25s).

Backend commit duration measures how long it takes etcd’s storage engine to commit a batch of writes to its on-disk key-value store (boltdb), separately from the WAL fsync step. High commit durations indicate the backend database layer — not just the WAL — is struggling to keep up with the write load, which is often driven by database size, fragmentation, or the same underlying disk latency issues that affect WAL fsync.


Possible Causes:


Severity estimation

High severity — commit duration is a core indicator of backend storage health.

Impact assessment:


Troubleshooting steps

  1. Check the current backend commit duration

    • Command / Action:
      • Query the metric behind the alert
      • histogram_quantile(0.99, rate(etcd_disk_backend_commit_duration_seconds_bucket{job=~".etcd."}[5m]))

    • Expected result:
      • Value below 0.25s, ideally in single-digit milliseconds
    • additional info:
      • Break down by (instance) to identify whether one member or the whole cluster is affected

  1. Check disk I/O latency on the etcd node

    • Command / Action:
      • Backend commits are also disk-bound, so check the same disk metrics as for fsync latency
      • iostat -x 2 5

    • Expected result:
      • await under ~1ms for the disk backing etcd’s data directory
    • additional info:
      • If both commit and fsync durations are high together, the root cause is very likely disk latency — refer to etcdHighFsyncDurations for detailed disk troubleshooting

  1. Check etcd database size and fragmentation ratio

    • Command / Action:
      • Inspect current DB size and how much of it is fragmented (free but unreclaimed) space
      • kubectl exec -n kube-system <etcd-pod> – etcdctl endpoint status –cluster –write-out=table

      • histogram_quantile(0.99, rate(etcd_disk_backend_commit_duration_seconds_bucket[5m])) combined with etcd_mvcc_db_total_size_in_bytes / etcd_mvcc_db_total_size_in_use_in_bytes

    • Expected result:
      • DB size well below quota; fragmentation ratio (total size vs. in-use size) close to 1
    • additional info:
      • A high ratio (e.g. total size 2x the in-use size) indicates the database needs defragmentation — refer to EtcdDatabaseHighFragmentationRatio

  1. Check write volume and churn

    • Command / Action:
      • Identify which resources are generating high write load
      • sum(rate(etcd_mvcc_put_total[5m]))

      • kubectl get events -A | wc -l

    • Expected result:
      • Write rate consistent with normal cluster activity; no runaway controller/operator generating excessive updates
    • additional info:
      • A misbehaving controller doing tight reconcile loops with frequent status updates is a common source of excessive write churn; check kubectl get --raw /metrics on the API server for per-resource request rates if available

  1. Check for large objects increasing commit overhead

    • Command / Action:
      • Look for unusually large ConfigMaps, Secrets, or CRDs
      • kubectl get configmaps -A -o json | jq -r ‘.items[] | select((.data|tostring|length) > 100000) | “(.metadata.namespace)/(.metadata.name)”’

    • Expected result:
      • No excessively large objects being frequently updated
    • additional info:
      • etcd has a default 1.5MB request size limit; objects approaching this size and updated frequently significantly increase commit cost

  1. Manually defragment the etcd member if fragmentation is confirmed

    • Command / Action:
      • Run defragmentation on one member at a time (never all members simultaneously)
      • kubectl exec -n kube-system <etcd-pod> – etcdctl defrag –endpoints=https://127.0.0.1:2379

      • Verify the member is healthy before moving to the next
    • Expected result:
      • DB size shrinks to closer match the in-use size; commit duration improves
    • additional info:
      • Defragmentation blocks the member being defragmented and can briefly increase latency during the operation — perform during low-traffic periods, one member at a time

  1. Check CPU and memory pressure on the etcd node

    • Command / Action:
      • Verify the node is not resource-starved, which slows down the backend’s memory-mapped I/O
      • kubectl top pod -n kube-system -l component=etcd

      • kubectl describe node <node-name>

    • Expected result:
      • CPU and memory usage within normal bounds; no OOMKilled events
    • additional info:
      • etcd’s backend relies on the OS page cache; insufficient memory forces more disk reads during commit, increasing latency

  1. Check etcd logs for commit-related warnings

    • Command / Action:
      • Review etcd logs for slow commit or backend warnings
      • kubectl logs -n kube-system <etcd-pod> –tail=200 | grep -iE ‘commit|backend|slow|took too long’

    • Expected result:
      • No recurring slow commit warnings
    • additional info:
      • etcd logs a warning when a backend commit takes longer than expected; correlate timing with disk and fragmentation findings above

  1. Verify commit duration has returned to normal

    • Command / Action:
      • Re-run the Prometheus query used in step 1
      • histogram_quantile(0.99, rate(etcd_disk_backend_commit_duration_seconds_bucket{job=~".etcd."}[5m]))

    • Expected result:
      • Value has dropped back below the 0.25s threshold
    • additional info:
      • Also confirm fsync durations and failed proposal rate have stabilized, since these three metrics commonly move together

Additional resources