VibeInfra Incident Solver Skill
Use this skill when diagnosing, mitigating, and verifying hands-on production incidents in VibeInfra's flight simulators.
The Canonical 7-Step Production Incident Lifecycle
Every intermediate, advanced, and expert production incident follows this strict lifecycle:
Traffic Shedding & Triage:
- Enable gateway rate-limiting or circuit breaker to prevent client retry storms.
- Inspect edge proxy metrics (
P99 latency,5xx error rate).
Blast Radius Isolation:
- Cordon degraded Kubernetes nodes / isolate poisoned backend containers.
- Prevent cascading failures to upstream dependencies.
Data Tier & Pool Unclog:
- Check PostgreSQL
pg_stat_activity, active locks, and unindexed sequential scans. - Terminate hanging idle-in-transaction connections and increase connection pool headroom.
- Check PostgreSQL
Root Cause Patch:
- Apply targeted configuration fix (Nginx
proxy_pass,keepalive_timeout, KubernetesreadinessProbe, AWS IAM ARN policy, or sysctlnf_conntrack_max).
- Apply targeted configuration fix (Nginx
Queue Drain & State Reconciliation:
- Purge Dead Letter Queue (DLQ) if corrupted, flush poisoned cache keys in Redis.
Controlled Progressive Scale-Up:
- Gradually increase replica count (1 -> 3 -> 5) to prevent thundering herd on cold databases.
Sustained Traffic Gate & AAR:
- Maintain $\ge 30\text{s}$ continuous loadgen verification with $0%$ error rate.
- Review diagnostic findings in After-Action Review.