CertmanagerCertificateExpiration
CertmanagerCertificateExpiration
Description
A Certmanager issued certificate is expiring soon or has expired.
These are serving and peer certificates — Ingress TLS, admission webhooks, internal mTLS, dashboards, streaming endpoints.
Possible Causes:
- CertificateRequest stuck unapproved, or Order / Challenge stuck
- Renewal-failure alert fired earlier and was acknowledged but not resolved
- cert-manager controller not running, crash-looping, or wedged
- ACME HTTP-01 failing: /.well-known/acme-challenge/ not routed, wrong ingress class, challenge blocked by auth or WAF, public DNS not pointing at the cluster
- ACME DNS-01 failing: provider credentials expired, insufficient permissions, propagation timeout, broken CNAME delegation ACME rate limit reached
- Certificate renewal process not configured or not working properly
- Clock skew issues causing premature expiration detection
- Let’s Encrypt allows 5 duplicate certificates per week per identifier set and 5 failed validations per hostname per hour; usually self-inflicted by repeated retries
- RBAC denials preventing cert-manager from writing the Secret or creating challenge resources
Severity estimation
High to Critical severity, depending on which certificate is expiring.
- High if internal-only or non-production endpoint is affected
- High if the certificate is expiring within 30 days and affects non-critical components
- Critical if the certificate is expiring within 7 days or has already expired
Impact assessment:
- Expired serving certificates cause TLS handshake failures for every client of the endpoint. Browsers, players, SDKs and API consumers get hard errors, not warnings
- An expired webhook certificate is the dangerous case: the API server cannot reach the webhook, admission fails, and creates/updates of affected resources are blocked cluster-wide. This can block the remediation itself
- Expired mTLS certificates break service-to-service traffic as connections churn, often silently at first
- Public endpoints are customer-visible immediately on expiry
Troubleshooting steps
-
Identify which certificates are expiring
- Command / Action:
- Check alert labels to identify the affected certificate
-
kubectl get certificate -n
- Review certificate expiration dates across the cluster
-
kubectl get certificate -A
- Expected result:
- List of certificates with READY status
- Identification of certificates expiring soon or expired
- Command / Action:
-
Check certificate resource
- Command / Action:
-
kubectl describe certificate -n
-
- Expected result:
- Certificate expiration date and time displayed
- Eventlog at the bottom which names the failure (
Failed to create Order, waiting for CertificateRequest to be signed, presenting challenge)
- Command / Action:
-
Walk the issuance chain
- Command / Action:
- check cert request, order and challenge
-
kubectl get certificaterequest,order,challenge -n
-
kubectl describe challenge -n
- Expected result:
- the Challenge object’s
Reasonfield states exactly why validation failed: (HTTP 404 on the challenge path, DNS record not found, propagation timeout, rate limit response from the ACME server)
- the Challenge object’s
- Command / Action:
-
Check the issuer
- Command / Action:
- check if the wanted dns zone is configured in the issuer
-
kubectl describe clusterissuer <issuer_name>
-
kubectl describe issuer -n <issuer_name>
- Expected result:
- Issuer status:
Ready=True
- Issuer status:
- additional info:
- the DNS zone for the certificate should be listed in the
Solverssection of the issuer config
- the DNS zone for the certificate should be listed in the
- Command / Action:
-
Check cert-manager
- Command / Action:
- list cert-manager pods
-
kubectl -n cert-manager get pods
- check certmanager logs
-
kubectl -n cert-manager logs deploy/cert-manager –since=1h | grep -Ei ’error|failed|rate limit'
- Expected result:
- Cert-manager, cert-manager-webhook and cert-manager-cainjector all running
- Command / Action:
-
Verify the Secret against the metric
- Command / Action:
- The metric reflects cert-manager’s view. Confirm against the real material
- get secret name
-
kubectl get certificate -n -o jsonpath=’{.spec.secretName}{"\n"}'
-
kubectl get secret -n -o jsonpath=’{.data.tls.crt}’ | base64 -d | openssl x509 -noout -dates -subject -issuer
- Expected result:
notAftermatches the metric. A mismatch means the Secret was replaced out-of-band, or the consumer mounts a different Secret than the one cert-manager manages
- Command / Action:
-
Fix root cause, then renew once
- Command / Action:
- delete the certificate request to trigger a new one
- Expected result:
- new certificate request will be created and certificate gets updated
- additional info:
Do not loop an this- Every attempt counts against ACME rate limits, and repeated forced renewals against a broken issuer will lock the domain out for a week
- Command / Action:
Additional resources
- cert-manager documentation
- Related alert: CertmanagerCertificateRenewalFailed