KubeletPodStartUpLatencyHigh
KubeletPodStartUpLatencyHigh
Description
This alert fires when the kubelet’s 99th percentile pod startup latency is above the configured threshold (default: 60s) on a node.
Pod startup latency measures the time the kubelet’s pod worker spends bringing a pod up — pulling images, creating the sandbox, starting containers, and running readiness-related setup. Consistently high startup latency means users and controllers (Deployments, Jobs, HPA scale-ups) experience slow pod creation, which delays rollouts, scaling, and recovery from failures.
Possible Causes:
- Slow image pulls — large images, slow registry, or no local image cache
- Container runtime overloaded or slow to create sandboxes/containers
- Node under CPU, memory, or disk pressure, slowing every startup step
- Slow CNI plugin — network setup (IP allocation, route/iptables programming) is a common bottleneck
- Volume mount delays — slow CSI drivers, network-attached storage, or large volume initialization
- Image pull backoff due to registry authentication or rate-limiting issues
- Too many pods starting concurrently on the same node, contending for resources
- kubelet itself under resource pressure or busy with PLEG relist (see
KubeletPlegDurationHigh) - Admission webhooks or init containers with long-running logic
Severity estimation
Medium to High severity — depends on impact to rollouts and autoscaling.
- Medium if latency is elevated on a subset of pods/nodes but deployments and jobs still complete within acceptable time
- High if latency is delaying critical rollouts, autoscaling response, or crash-loop recovery
- Critical if paired with widespread image pull failures or node resource exhaustion affecting many nodes
Impact assessment:
- Slower Deployment/StatefulSet rollouts
- Delayed response to HPA scale-up events, risking under-provisioning during load spikes
- Slower recovery when crash-looping pods are rescheduled
- Increased time-to-ready for newly scheduled workloads
Troubleshooting steps
-
Check current pod startup latency for the affected node
- Command / Action:
- Query the metric behind the alert
-
histogram_quantile(0.99, sum by (instance, le) (rate(kubelet_pod_worker_duration_seconds_bucket{instance="<node-instance>"}[5m])))
- Expected result:
- Value well below the
60sthreshold
- Value well below the
- additional info:
- The alert label
nodeidentifies the affected node
- The alert label
- Command / Action:
-
Identify which pods are slow to start
- Command / Action:
- List recently created pods on the node and their age/status
-
kubectl get pods -A –field-selector spec.nodeName=<node-name> -o wide –sort-by=.metadata.creationTimestamp
-
kubectl describe pod -n <namespace> <pod-name>
- Expected result:
- Pod events show a clear timeline; look for long gaps between
Scheduled,Pulling,Pulled, andStartedevents
- Pod events show a clear timeline; look for long gaps between
- additional info:
- The event timeline is the fastest way to isolate which startup phase is slow (image pull vs. sandbox creation vs. container start)
- Command / Action:
-
Check for slow image pulls
- Command / Action:
- Look for
Pulling imageandSuccessfully pulled imageevents and the time between them -
kubectl get events -n <namespace> –field-selector involvedObject.name=<pod-name>
-
crictl images
- Look for
- Expected result:
- Image pulls complete within a few seconds for cached images, or a reasonable time proportional to image size for cold pulls
- additional info:
- Large images, slow/rate-limited registries, or missing local caching are common root causes; consider image pull policy and registry mirrors
- Command / Action:
-
Check node resource pressure
- Command / Action:
- Verify the node is not under CPU, memory, or disk pressure
-
kubectl describe node <node-name>
-
kubectl top node <node-name>
- Expected result:
- No
MemoryPressure,DiskPressure, orPIDPressure; CPU not saturated
- No
- additional info:
- Resource-starved nodes slow down every stage of pod startup, not just one
- Command / Action:
-
Check container runtime performance
- Command / Action:
- Test how quickly the runtime responds to basic operations
-
time crictl ps -a
-
journalctl -u containerd -n 200 | grep -iE ’timeout|slow|error'
- Expected result:
- Fast response time; no recurring runtime errors
- additional info:
- Also check
KubeletPlegDurationHigh— a slow PLEG relist and slow pod startup often share the same runtime-level root cause
- Also check
- Command / Action:
-
Check CNI plugin / network setup latency
- Command / Action:
- Review CNI plugin logs for delays in IP allocation or network setup
-
journalctl -u kubelet -n 200 | grep -iE ‘cni|network’
- Check CNI daemon pod logs, e.g.: >kubectl logs -n kube-system <cni-pod>
- Expected result:
- No repeated delays or errors in IP allocation, route programming, or iptables/nftables updates
- additional info:
- IP address exhaustion in the node’s pod CIDR range or slow IPAM plugins are common culprits
- Command / Action:
-
Check volume mount times
- Command / Action:
- Look for slow
AttachVolume/MountVolumeevents -
kubectl describe pod -n <namespace> <pod-name> | grep -A5 Volumes
-
kubectl get events -n <namespace> –field-selector involvedObject.name=<pod-name> | grep -i volume
- Look for slow
- Expected result:
- Volumes attach and mount within a few seconds
- additional info:
- Network-attached storage (EBS, network file shares) and CSI driver issues are common sources of slow mounts; check the CSI driver pod logs if present
- Command / Action:
-
Check for concurrent pod startup contention
- Command / Action:
- Count how many pods are starting simultaneously on the node
-
kubectl get pods -A –field-selector spec.nodeName=<node-name> -o json | jq ‘[.items[] | select(.status.phase==“Pending” or (.status.containerStatuses[]?.state.waiting))] | length’
- Expected result:
- A small, manageable number of concurrent pod starts
- additional info:
- Large scale-up events or node replacements can cause many pods to start at once, saturating image pull bandwidth, CPU, and disk I/O; consider
--serialize-image-pulls(default true) and--registry-burst/--registry-qpskubelet flags
- Large scale-up events or node replacements can cause many pods to start at once, saturating image pull bandwidth, CPU, and disk I/O; consider
- Command / Action:
-
Check for slow init containers or admission webhooks
- Command / Action:
- Review pod spec and events for long-running init containers or webhook delays
-
kubectl describe pod -n <namespace> <pod-name> | grep -A10 “Init Containers”
- Expected result:
- Init containers complete quickly; no webhook timeout warnings
- additional info:
- Slow validating/mutating webhooks add latency before the pod is even scheduled to the kubelet; check
kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations
- Slow validating/mutating webhooks add latency before the pod is even scheduled to the kubelet; check
- Command / Action:
-
Mitigate: reduce concurrent workload or scale out
- Command / Action:
- If the node is contending under load, cordon and rebalance, or add capacity
-
kubectl cordon <node-name>
- Consider adding nodes or a node pool with pre-pulled images
- Expected result:
- Pod startup latency decreases as contention drops
- additional info:
- For predictable slow-image issues, consider pre-pulling critical images via a DaemonSet or
imagePullPolicy: IfNotPresentwith a warm image cache
- For predictable slow-image issues, consider pre-pulling critical images via a DaemonSet or
- Command / Action:
-
Verify pod startup latency has returned to normal
- Command / Action:
- Re-run the Prometheus query used in step 1
-
histogram_quantile(0.99, sum by (instance, le) (rate(kubelet_pod_worker_duration_seconds_bucket{instance="<node-instance>"}[5m])))
- Expected result:
- Value has dropped back below the
60sthreshold
- Value has dropped back below the
- additional info:
- Continue monitoring for at least one full alert grace period (
15m) to confirm the alert clears
- Continue monitoring for at least one full alert grace period (
- Command / Action:
Additional resources
- Kubelet architecture and pod lifecycle
- Optimizing image pulls
- CNI plugin troubleshooting
- Related alert: KubeletPlegDurationHigh
- Related alert: KubeNodeNotReady
- Related alert: KubePodNotReady