Alert Runbooks

CertmanagerCertificateExpiration

CertmanagerCertificateExpiration

Description

A Certmanager issued certificate is expiring soon or has expired.

These are serving and peer certificates — Ingress TLS, admission webhooks, internal mTLS, dashboards, streaming endpoints.


Possible Causes:


Severity estimation

High to Critical severity, depending on which certificate is expiring.

Impact assessment:


Troubleshooting steps

  1. Identify which certificates are expiring

    • Command / Action:
      • Check alert labels to identify the affected certificate
      • kubectl get certificate -n

      • Review certificate expiration dates across the cluster
      • kubectl get certificate -A

    • Expected result:
      • List of certificates with READY status
      • Identification of certificates expiring soon or expired

  1. Check certificate resource

    • Command / Action:
      • kubectl describe certificate -n

    • Expected result:
      • Certificate expiration date and time displayed
      • Eventlog at the bottom which names the failure (Failed to create Order, waiting for CertificateRequest to be signed, presenting challenge)

  1. Walk the issuance chain

    • Command / Action:
      • check cert request, order and challenge
      • kubectl get certificaterequest,order,challenge -n

      • kubectl describe challenge -n

    • Expected result:
      • the Challenge object’s Reason field states exactly why validation failed: (HTTP 404 on the challenge path, DNS record not found, propagation timeout, rate limit response from the ACME server)

  1. Check the issuer

    • Command / Action:
      • check if the wanted dns zone is configured in the issuer
      • kubectl describe clusterissuer <issuer_name>

      • kubectl describe issuer -n <issuer_name>

    • Expected result:
      • Issuer status: Ready=True
    • additional info:
      • the DNS zone for the certificate should be listed in the Solvers section of the issuer config

  1. Check cert-manager

    • Command / Action:
      • list cert-manager pods
      • kubectl -n cert-manager get pods

      • check certmanager logs
      • kubectl -n cert-manager logs deploy/cert-manager –since=1h | grep -Ei ’error|failed|rate limit'

    • Expected result:
      • Cert-manager, cert-manager-webhook and cert-manager-cainjector all running

  1. Verify the Secret against the metric

    • Command / Action:
      • The metric reflects cert-manager’s view. Confirm against the real material
      • get secret name
      • kubectl get certificate -n -o jsonpath=’{.spec.secretName}{"\n"}'

      • kubectl get secret -n -o jsonpath=’{.data.tls.crt}’ | base64 -d | openssl x509 -noout -dates -subject -issuer

    • Expected result:
      • notAfter matches the metric. A mismatch means the Secret was replaced out-of-band, or the consumer mounts a different Secret than the one cert-manager manages

  1. Fix root cause, then renew once

    • Command / Action:
      • delete the certificate request to trigger a new one
    • Expected result:
      • new certificate request will be created and certificate gets updated
    • additional info:
      • Do not loop an this
      • Every attempt counts against ACME rate limits, and repeated forced renewals against a broken issuer will lock the domain out for a week

Additional resources