Alert Runbooks

KubeletPodStartUpLatencyHigh

KubeletPodStartUpLatencyHigh

Description

This alert fires when the kubelet’s 99th percentile pod startup latency is above the configured threshold (default: 60s) on a node.

Pod startup latency measures the time the kubelet’s pod worker spends bringing a pod up — pulling images, creating the sandbox, starting containers, and running readiness-related setup. Consistently high startup latency means users and controllers (Deployments, Jobs, HPA scale-ups) experience slow pod creation, which delays rollouts, scaling, and recovery from failures.


Possible Causes:


Severity estimation

Medium to High severity — depends on impact to rollouts and autoscaling.

Impact assessment:


Troubleshooting steps

  1. Check current pod startup latency for the affected node

    • Command / Action:
      • Query the metric behind the alert
      • histogram_quantile(0.99, sum by (instance, le) (rate(kubelet_pod_worker_duration_seconds_bucket{instance="<node-instance>"}[5m])))

    • Expected result:
      • Value well below the 60s threshold
    • additional info:
      • The alert label node identifies the affected node

  1. Identify which pods are slow to start

    • Command / Action:
      • List recently created pods on the node and their age/status
      • kubectl get pods -A –field-selector spec.nodeName=<node-name> -o wide –sort-by=.metadata.creationTimestamp

      • kubectl describe pod -n <namespace> <pod-name>

    • Expected result:
      • Pod events show a clear timeline; look for long gaps between Scheduled, Pulling, Pulled, and Started events
    • additional info:
      • The event timeline is the fastest way to isolate which startup phase is slow (image pull vs. sandbox creation vs. container start)

  1. Check for slow image pulls

    • Command / Action:
      • Look for Pulling image and Successfully pulled image events and the time between them
      • kubectl get events -n <namespace> –field-selector involvedObject.name=<pod-name>

      • crictl images

    • Expected result:
      • Image pulls complete within a few seconds for cached images, or a reasonable time proportional to image size for cold pulls
    • additional info:
      • Large images, slow/rate-limited registries, or missing local caching are common root causes; consider image pull policy and registry mirrors

  1. Check node resource pressure

    • Command / Action:
      • Verify the node is not under CPU, memory, or disk pressure
      • kubectl describe node <node-name>

      • kubectl top node <node-name>

    • Expected result:
      • No MemoryPressure, DiskPressure, or PIDPressure; CPU not saturated
    • additional info:
      • Resource-starved nodes slow down every stage of pod startup, not just one

  1. Check container runtime performance

    • Command / Action:
      • Test how quickly the runtime responds to basic operations
      • time crictl ps -a

      • journalctl -u containerd -n 200 | grep -iE ’timeout|slow|error'

    • Expected result:
      • Fast response time; no recurring runtime errors
    • additional info:
      • Also check KubeletPlegDurationHigh — a slow PLEG relist and slow pod startup often share the same runtime-level root cause

  1. Check CNI plugin / network setup latency

    • Command / Action:
      • Review CNI plugin logs for delays in IP allocation or network setup
      • journalctl -u kubelet -n 200 | grep -iE ‘cni|network’

      • Check CNI daemon pod logs, e.g.: >kubectl logs -n kube-system <cni-pod>
    • Expected result:
      • No repeated delays or errors in IP allocation, route programming, or iptables/nftables updates
    • additional info:
      • IP address exhaustion in the node’s pod CIDR range or slow IPAM plugins are common culprits

  1. Check volume mount times

    • Command / Action:
      • Look for slow AttachVolume/MountVolume events
      • kubectl describe pod -n <namespace> <pod-name> | grep -A5 Volumes

      • kubectl get events -n <namespace> –field-selector involvedObject.name=<pod-name> | grep -i volume

    • Expected result:
      • Volumes attach and mount within a few seconds
    • additional info:
      • Network-attached storage (EBS, network file shares) and CSI driver issues are common sources of slow mounts; check the CSI driver pod logs if present

  1. Check for concurrent pod startup contention

    • Command / Action:
      • Count how many pods are starting simultaneously on the node
      • kubectl get pods -A –field-selector spec.nodeName=<node-name> -o json | jq ‘[.items[] | select(.status.phase==“Pending” or (.status.containerStatuses[]?.state.waiting))] | length’

    • Expected result:
      • A small, manageable number of concurrent pod starts
    • additional info:
      • Large scale-up events or node replacements can cause many pods to start at once, saturating image pull bandwidth, CPU, and disk I/O; consider --serialize-image-pulls (default true) and --registry-burst/--registry-qps kubelet flags

  1. Check for slow init containers or admission webhooks

    • Command / Action:
      • Review pod spec and events for long-running init containers or webhook delays
      • kubectl describe pod -n <namespace> <pod-name> | grep -A10 “Init Containers”

    • Expected result:
      • Init containers complete quickly; no webhook timeout warnings
    • additional info:
      • Slow validating/mutating webhooks add latency before the pod is even scheduled to the kubelet; check kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations

  1. Mitigate: reduce concurrent workload or scale out

    • Command / Action:
      • If the node is contending under load, cordon and rebalance, or add capacity
      • kubectl cordon <node-name>

      • Consider adding nodes or a node pool with pre-pulled images
    • Expected result:
      • Pod startup latency decreases as contention drops
    • additional info:
      • For predictable slow-image issues, consider pre-pulling critical images via a DaemonSet or imagePullPolicy: IfNotPresent with a warm image cache

  1. Verify pod startup latency has returned to normal

    • Command / Action:
      • Re-run the Prometheus query used in step 1
      • histogram_quantile(0.99, sum by (instance, le) (rate(kubelet_pod_worker_duration_seconds_bucket{instance="<node-instance>"}[5m])))

    • Expected result:
      • Value has dropped back below the 60s threshold
    • additional info:
      • Continue monitoring for at least one full alert grace period (15m) to confirm the alert clears

Additional resources