KubeletPlegDurationHigh
KubeletPlegDurationHigh
Description
This alert fires when the kubelet’s Pod Lifecycle Event Generator (PLEG) has a 99th percentile relist duration at or above the configured threshold (default: 10s).
The PLEG is the kubelet subsystem responsible for periodically listing containers via the container runtime, detecting state changes (start, stop, crash), and generating pod lifecycle events. It is on the critical path for every pod status update. When PLEG relisting takes too long, the kubelet falls behind on detecting container state changes, pod status updates are delayed, and the kubelet may eventually report the node as NotReady due to PLEG health checks failing.
Possible Causes:
- Container runtime (containerd/CRI-O/Docker) is slow to respond to
ListContainers/ListPodSandboxcalls - Too many pods/containers scheduled on the node, increasing the work done per relist cycle
- Disk I/O saturation on the node, slowing down runtime and cgroup filesystem reads
- CPU starvation on the node, delaying the kubelet’s relist loop
- Container runtime experiencing internal issues (e.g. containerd shim leaks, hung processes)
- Zombie or defunct processes accumulating on the node
- A large number of terminated/evicted pods not yet garbage collected, still visible to the runtime
- Kernel or cgroup driver issues causing slow cgroup stat reads
Severity estimation
Medium to High severity — the node is degraded but usually still functional.
- Medium if relist duration is elevated but the node remains
Readyand pod status updates are only mildly delayed - High if relist duration keeps climbing and the node is at risk of, or has started, flapping to
NotReady - Critical if paired with
KubeNodeNotReadyorKubeNodeReadinessFlapping— PLEG is unhealthy enough to affect node status reporting
Impact assessment:
- Delayed detection of container crashes and restarts
- Stale pod status reported to the API server
- New pod scheduling on the node may be delayed
- Sustained high PLEG duration can trip the kubelet’s internal PLEG health check, marking the node
NotReady
Troubleshooting steps
-
Check the current PLEG relist duration for the affected node
- Command / Action:
- Query the metric behind the alert to see the current value and trend
-
histogram_quantile(0.99, sum by (instance, le) (rate(kubelet_pleg_relist_duration_seconds_bucket{instance="<node-instance>"}[5m])))
- Expected result:
- Value well below the
10sthreshold; healthy nodes are typically under 1s
- Value well below the
- additional info:
- The alert label
nodeidentifies the affected node; the metric labelinstanceis the kubelet’s scrape target
- The alert label
- Command / Action:
-
Check node resource pressure
- Command / Action:
- Verify the node is not under CPU, memory, or disk pressure
-
kubectl describe node <node-name>
-
kubectl top node <node-name>
- Expected result:
- No
MemoryPressure,DiskPressure, orPIDPressureconditions; CPU is not pegged at 100%
- No
- additional info:
- PLEG relist duration correlates strongly with node CPU saturation and disk latency
- Command / Action:
-
Check number of pods and containers on the node
- Command / Action:
- Count pods scheduled on the affected node
-
kubectl get pods -A –field-selector spec.nodeName=<node-name> -o wide | wc -l
- Expected result:
- Pod count within the node’s expected capacity (default max is 110 pods per node)
- additional info:
- Each relist cycle scans every container on the node; nodes running near their pod limit are more prone to PLEG slowdowns
- Command / Action:
-
Check container runtime health and responsiveness
- Command / Action:
- SSH to the node and check the container runtime service
-
systemctl status containerd
-
crictl ps -a | wc -l
-
time crictl ps -a
- Expected result:
- Runtime service is active and
crictl psreturns quickly (well under a second)
- Runtime service is active and
- additional info:
- A slow
crictl psmirrors what the kubelet experiences during relist; this directly points to the runtime as the bottleneck
- A slow
- Command / Action:
-
Check container runtime logs for errors or slow operations
- Command / Action:
- Review runtime logs for timeouts or slow shim/API calls
-
journalctl -u containerd -n 200 | grep -iE ’timeout|slow|error|deadline'
- Expected result:
- No recurring timeout or deadline-exceeded errors
- additional info:
- Look for containerd-shim processes that are hung or leaked;
ps aux | grep containerd-shimcan reveal an abnormally high count
- Look for containerd-shim processes that are hung or leaked;
- Command / Action:
-
Check disk I/O latency on the node
- Command / Action:
- Verify the disk backing the container runtime and kubelet state is not saturated
-
iostat -x 2 5
- Expected result:
awaitand%utilwithin normal bounds; no sustained high I/O wait
- additional info:
- Slow disks delay cgroup and container state file reads, directly increasing relist duration
- Command / Action:
-
Check kubelet logs for PLEG-related warnings
- Command / Action:
- Review kubelet logs on the affected node
-
journalctl -u kubelet -n 200 | grep -i pleg
- Expected result:
- No
PLEG is not healthymessages
- No
- additional info:
PLEG is not healthy: pleg was last seen activemessages indicate the relist loop has stalled beyond the kubelet’s internal threshold (default 3 minutes), which will mark the nodeNotReady
- Command / Action:
-
Check for zombie or leaked processes on the node
- Command / Action:
- Look for accumulated defunct processes that can slow down process listing
-
ps aux | grep -c defunct
- Expected result:
- Zero or near-zero zombie processes
- additional info:
- A high number of zombie processes is often a symptom of a container runtime or init process bug
- Command / Action:
-
Reduce pod density or cordon the node if overloaded
- Command / Action:
- If the node is legitimately over capacity, cordon it and rebalance workloads
-
kubectl cordon <node-name>
-
kubectl drain <node-name> –ignore-daemonsets –delete-emptydir-data
- Expected result:
- Pods reschedule to other nodes; PLEG duration on the affected node decreases as load drops
- additional info:
- Only drain if there is sufficient cluster capacity to absorb the workload; consider this a mitigation, not a root-cause fix
- Command / Action:
-
Restart the container runtime or kubelet if a leak/hang is confirmed
- Command / Action:
- Restart the runtime first, then the kubelet if needed
-
systemctl restart containerd
-
systemctl restart kubelet
- Expected result:
- PLEG relist duration returns to normal after restart
- additional info:
- Restarting containerd briefly disrupts container management on the node; running pods are not restarted but new operations pause during the restart
- Command / Action:
-
Verify PLEG duration has returned to normal
- Command / Action:
- Re-run the Prometheus query used in step 1
-
histogram_quantile(0.99, sum by (instance, le) (rate(kubelet_pleg_relist_duration_seconds_bucket{instance="<node-instance>"}[5m])))
- Expected result:
- Value has dropped back below the
10sthreshold and node isReady
- Value has dropped back below the
- additional info:
- Continue monitoring for at least one full alert grace period (
5m) to confirm the alert clears
- Continue monitoring for at least one full alert grace period (
- Command / Action:
Additional resources
- Kubelet source of truth — Pod Lifecycle Event Generator
- Debugging kubelet PLEG issues
- containerd troubleshooting
- Related alert: KubeletPodStartUpLatencyHigh
- Related alert: KubeNodeNotReady
- Related alert: KubeNodeReadinessFlapping