Cert-Manager Deep-Dive Skill
Deep analysis of cert-manager certificates, issuers, and TLS lifecycle management.
MANDATORY: Discovery-First Pattern
Always check cert-manager installation and issuers before inspecting certificates.
Phase 1: Discovery
#!/bin/bash
echo "=== Cert-Manager Version ==="
kubectl get deployment cert-manager -n cert-manager -o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null
echo ""
echo ""
echo "=== Cert-Manager Pods ==="
kubectl get pods -n cert-manager -o custom-columns='NAME:.metadata.name,STATUS:.status.phase,RESTARTS:.status.containerStatuses[0].restartCount,AGE:.metadata.creationTimestamp' 2>/dev/null
echo ""
echo "=== ClusterIssuers ==="
kubectl get clusterissuers -o custom-columns='NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,AGE:.metadata.creationTimestamp' 2>/dev/null
echo ""
echo "=== Issuers (all namespaces) ==="
kubectl get issuers --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status' 2>/dev/null | head -20
echo ""
echo "=== Certificates (all namespaces) ==="
kubectl get certificates --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,EXPIRY:.status.notAfter,RENEWAL:.status.renewalTime' 2>/dev/null | head -20
Phase 2: Analysis
#!/bin/bash
echo "=== Not Ready Certificates ==="
kubectl get certificates --all-namespaces -o json 2>/dev/null | jq -r '
.items[] |
select(.status.conditions[]? | select(.type == "Ready" and .status != "True")) |
"\(.metadata.namespace)/\(.metadata.name)\t\(.status.conditions[] | select(.type == "Ready") | .reason): \(.message // "no message")"
' | column -t | head -15
echo ""
echo "=== Expiring Soon (within 30 days) ==="
THRESHOLD=$(date -u -d '+30 days' +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -v+30d +%Y-%m-%dT%H:%M:%SZ)
kubectl get certificates --all-namespaces -o json 2>/dev/null | jq -r --arg threshold "$THRESHOLD" '
.items[] |
select(.status.notAfter // "" < $threshold and .status.notAfter // "" != "") |
"\(.metadata.namespace)/\(.metadata.name)\tExpires:\(.status.notAfter)"
' | head -15
echo ""
echo "=== Certificate Requests ==="
kubectl get certificaterequests --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,APPROVED:.status.conditions[?(@.type=="Approved")].status,READY:.status.conditions[?(@.type=="Ready")].status' 2>/dev/null | head -15
echo ""
echo "=== Pending Challenges ==="
kubectl get challenges --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,STATE:.status.state,DOMAIN:.spec.dnsName,TYPE:.spec.type' 2>/dev/null | head -15
echo ""
echo "=== Pending Orders ==="
kubectl get orders --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,STATE:.status.state' 2>/dev/null | head -10
echo ""
echo "=== Cert-Manager Logs (errors) ==="
kubectl logs deployment/cert-manager -n cert-manager --tail=20 2>/dev/null | grep -i "error\|fail\|warn" | head -10
Output Format
- Target ≤50 lines per output
- Use
-o custom-columns for targeted field extraction
- Show certificate expiry dates in ISO format
- Aggregate certificate counts by issuer when many exist
- Never dump full certificate specs -- show status and expiry only
Anti-Hallucination Rules
- NEVER assume resource names — always discover via CLI/API in Phase 1 before referencing in Phase 2.
- NEVER fabricate metric names or dimensions — verify against the service documentation or
--help output.
- NEVER mix CLI commands between service versions — confirm which version/API you are targeting.
- ALWAYS use the discovery → verify → analyze chain — every resource referenced must have been discovered first.
- ALWAYS handle empty results gracefully — an empty response is valid data, not an error to retry.
Counter-Rationalizations
| Shortcut |
Counter |
Why |
| "I'll skip discovery and check known resources" |
Always run Phase 1 discovery first |
Resource names change, new resources appear — assumed names cause errors |
| "The user only asked for a quick check" |
Follow the full discovery → analysis flow |
Quick checks miss critical issues; structured analysis catches silent failures |
| "Default configuration is probably fine" |
Audit configuration explicitly |
Defaults often leave logging, security, and optimization features disabled |
| "Metrics aren't needed for this" |
Always check relevant metrics when available |
API/CLI responses show current state; metrics reveal trends and intermittent issues |
| "I don't have access to that" |
Try the command and report the actual error |
Assumed permission failures prevent useful investigation; actual errors are informative |
Common Pitfalls
- Issuer vs ClusterIssuer: Issuers are namespaced; ClusterIssuers are cluster-wide -- certificates reference one or the other
- ACME challenges: HTTP-01 requires ingress access; DNS-01 requires DNS provider credentials -- check challenge type
- Challenge stuck: Pending challenges often indicate DNS propagation or ingress routing issues
- Renewal timing: Cert-manager renews at 2/3 of certificate lifetime by default -- check
renewBefore annotation
- Rate limits: Let's Encrypt has rate limits (50 certs/domain/week) -- check for rate limit errors in logs
- Secret not created: Certificate Ready=True but no Secret means the Secret was manually deleted -- check events
- Webhook failures: cert-manager webhook must be healthy for certificate creation -- check webhook pod
- Order states: valid, pending, failed, errored -- failed orders need manual investigation
1---2name: managing-k8s-cert-manager-deep3description: Use when working with K8S Cert Manager Deep — cert-manager deep-dive management for Kubernetes TLS certificate lifecycle. Covers certificate inventory, issuer health, certificate requests, challenges, orders, ACME configuration, and renewal status. Use when debugging certificate issuance failures, auditing TLS configurations, reviewing issuer setups, or monitoring certificate expiration.4---56# Cert-Manager Deep-Dive Skill78Deep analysis of cert-manager certificates, issuers, and TLS lifecycle management.910## MANDATORY: Discovery-First Pattern1112**Always check cert-manager installation and issuers before inspecting certificates.**1314### Phase 1: Discovery1516```bash17#!/bin/bash1819echo "=== Cert-Manager Version ==="20kubectl get deployment cert-manager -n cert-manager -o jsonpath='{.spec.template.spec.containers[0].image}' 2>/dev/null21echo ""2223echo ""24echo "=== Cert-Manager Pods ==="25kubectl get pods -n cert-manager -o custom-columns='NAME:.metadata.name,STATUS:.status.phase,RESTARTS:.status.containerStatuses[0].restartCount,AGE:.metadata.creationTimestamp' 2>/dev/null2627echo ""28echo "=== ClusterIssuers ==="29kubectl get clusterissuers -o custom-columns='NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,AGE:.metadata.creationTimestamp' 2>/dev/null3031echo ""32echo "=== Issuers (all namespaces) ==="33kubectl get issuers --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status' 2>/dev/null | head -203435echo ""36echo "=== Certificates (all namespaces) ==="37kubectl get certificates --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,EXPIRY:.status.notAfter,RENEWAL:.status.renewalTime' 2>/dev/null | head -2038```3940### Phase 2: Analysis4142```bash43#!/bin/bash4445echo "=== Not Ready Certificates ==="46kubectl get certificates --all-namespaces -o json 2>/dev/null | jq -r '47 .items[] |48 select(.status.conditions[]? | select(.type == "Ready" and .status != "True")) |49 "\(.metadata.namespace)/\(.metadata.name)\t\(.status.conditions[] | select(.type == "Ready") | .reason): \(.message // "no message")"50' | column -t | head -155152echo ""53echo "=== Expiring Soon (within 30 days) ==="54THRESHOLD=$(date -u -d '+30 days' +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -v+30d +%Y-%m-%dT%H:%M:%SZ)55kubectl get certificates --all-namespaces -o json 2>/dev/null | jq -r --arg threshold "$THRESHOLD" '56 .items[] |57 select(.status.notAfter // "" < $threshold and .status.notAfter // "" != "") |58 "\(.metadata.namespace)/\(.metadata.name)\tExpires:\(.status.notAfter)"59' | head -156061echo ""62echo "=== Certificate Requests ==="63kubectl get certificaterequests --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,APPROVED:.status.conditions[?(@.type=="Approved")].status,READY:.status.conditions[?(@.type=="Ready")].status' 2>/dev/null | head -156465echo ""66echo "=== Pending Challenges ==="67kubectl get challenges --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,STATE:.status.state,DOMAIN:.spec.dnsName,TYPE:.spec.type' 2>/dev/null | head -156869echo ""70echo "=== Pending Orders ==="71kubectl get orders --all-namespaces -o custom-columns='NAMESPACE:.metadata.namespace,NAME:.metadata.name,STATE:.status.state' 2>/dev/null | head -107273echo ""74echo "=== Cert-Manager Logs (errors) ==="75kubectl logs deployment/cert-manager -n cert-manager --tail=20 2>/dev/null | grep -i "error\|fail\|warn" | head -1076```7778## Output Format7980- Target ≤50 lines per output81- Use `-o custom-columns` for targeted field extraction82- Show certificate expiry dates in ISO format83- Aggregate certificate counts by issuer when many exist84- Never dump full certificate specs -- show status and expiry only8586## Anti-Hallucination Rules87881. **NEVER assume resource names** — always discover via CLI/API in Phase 1 before referencing in Phase 2.892. **NEVER fabricate metric names or dimensions** — verify against the service documentation or `--help` output.903. **NEVER mix CLI commands between service versions** — confirm which version/API you are targeting.914. **ALWAYS use the discovery → verify → analyze chain** — every resource referenced must have been discovered first.925. **ALWAYS handle empty results gracefully** — an empty response is valid data, not an error to retry.9394## Counter-Rationalizations9596| Shortcut | Counter | Why |97|----------|---------|-----|98| "I'll skip discovery and check known resources" | Always run Phase 1 discovery first | Resource names change, new resources appear — assumed names cause errors |99| "The user only asked for a quick check" | Follow the full discovery → analysis flow | Quick checks miss critical issues; structured analysis catches silent failures |100| "Default configuration is probably fine" | Audit configuration explicitly | Defaults often leave logging, security, and optimization features disabled |101| "Metrics aren't needed for this" | Always check relevant metrics when available | API/CLI responses show current state; metrics reveal trends and intermittent issues |102| "I don't have access to that" | Try the command and report the actual error | Assumed permission failures prevent useful investigation; actual errors are informative |103104## Common Pitfalls105106- **Issuer vs ClusterIssuer**: Issuers are namespaced; ClusterIssuers are cluster-wide -- certificates reference one or the other107- **ACME challenges**: HTTP-01 requires ingress access; DNS-01 requires DNS provider credentials -- check challenge type108- **Challenge stuck**: Pending challenges often indicate DNS propagation or ingress routing issues109- **Renewal timing**: Cert-manager renews at 2/3 of certificate lifetime by default -- check `renewBefore` annotation110- **Rate limits**: Let's Encrypt has rate limits (50 certs/domain/week) -- check for rate limit errors in logs111- **Secret not created**: Certificate Ready=True but no Secret means the Secret was manually deleted -- check events112- **Webhook failures**: cert-manager webhook must be healthy for certificate creation -- check webhook pod113- **Order states**: valid, pending, failed, errored -- failed orders need manual investigation