CoreWeave Incident Runbook
Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Overview
Respond to GPU workload incidents by stabilizing customer impact, preserving
redacted evidence, and restoring a known-good state. The incident commander owns
communications and escalation; responders use only the access required for triage.
Prerequisites
- An incident ID, named commander, affected namespace/service, and on-call route.
- Authorized read-only cluster access plus a documented production rollback revision.
- A redaction policy for logs, prompts, model artifacts, and credentials.
Instructions
- Declare severity and scope, then capture pod, event, node, and service status.
- Stabilize impact with the documented scale, failover, or rollback action before root-cause work.
- Collect only redacted, bounded diagnostics and escalate hardware or capacity issues through CoreWeave support.
- Verify recovery against the service SLO, update stakeholders, and create follow-up work for root cause and prevention.
Triage Steps
# 1. Check pod status
kubectl get pods -l app=inference -o wide
# 2. Check recent events
kubectl get events --sort-by=.lastTimestamp | tail -20
# 3. Check node status
kubectl get nodes -l gpu.nvidia.com/class -o wide
# 4. Check GPU health
kubectl exec -it $(kubectl get pod -l app=inference -o name | head -1) -- nvidia-smi
Common Incidents
Inference Service Down
- Check pod status and events
- If OOMKilled: reduce batch size or upgrade GPU
- If ImagePullBackOff: check registry credentials
- If Pending: check GPU quota and availability
GPU Node Failure
- Pods will be rescheduled automatically
- If no capacity: scale down non-critical workloads
- Contact CoreWeave support for extended outages
Model Loading Failure
- Check HuggingFace token secret exists
- Verify model name spelling
- Check PVC has sufficient storage
- Review container logs for download errors
Rollback
kubectl rollout undo deployment/inference
Output
- A time-stamped incident record with scope, owner, stabilization action, and redacted evidence.
- A verified recovery or an explicit escalation with a safe customer-impact mitigation.
Error Handling
| Incident complication |
Required response |
| Rollback fails |
Stop repeated deploy attempts, escalate to the platform owner, and preserve events. |
| Diagnostics include a secret |
Restrict distribution, rotate the secret, and recollect redacted evidence. |
| No GPU capacity is available |
Prioritize critical services under approved policy; do not remove tenant quotas. |
| Recovery cannot meet SLO |
Declare the continuing impact and use the approved fallback or communication path. |
Examples
For a failing production rollout, stabilize first and capture only the relevant evidence:
kubectl -n inference-prod rollout undo deployment/inference
kubectl -n inference-prod rollout status deployment/inference --timeout=10m
kubectl -n inference-prod get events --sort-by=.lastTimestamp | tail -30
Record the revision, outcome, and redacted events in the incident. Do not retry a broken image or expose a troubleshooting endpoint publicly.
Resources
Next Steps
For data handling, see coreweave-data-handling.
1---2name: coreweave-incident-runbook3description: Incident response runbook for CoreWeave GPU workload failures. Use when inference services are down, GPUs are unavailable, or responding to production incidents on CoreWeave. Trigger with phrases like "coreweave incident", "coreweave outage", "coreweave runbook", "coreweave service down".4license: MIT5---6# CoreWeave Incident Runbook
7
8> **Community-contributed.** Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
9
10## Overview
11
12Respond to GPU workload incidents by stabilizing customer impact, preserving
13redacted evidence, and restoring a known-good state. The incident commander owns
14communications and escalation; responders use only the access required for triage.
15
16## Prerequisites
17
18- An incident ID, named commander, affected namespace/service, and on-call route.
19- Authorized read-only cluster access plus a documented production rollback revision.
20- A redaction policy for logs, prompts, model artifacts, and credentials.
21
22## Instructions
23
241. Declare severity and scope, then capture pod, event, node, and service status.
252. Stabilize impact with the documented scale, failover, or rollback action before root-cause work.
263. Collect only redacted, bounded diagnostics and escalate hardware or capacity issues through CoreWeave support.
274. Verify recovery against the service SLO, update stakeholders, and create follow-up work for root cause and prevention.
28
29## Triage Steps
30
31```bash
32# 1. Check pod status
33kubectl get pods -l app=inference -o wide
34
35# 2. Check recent events
36kubectl get events --sort-by=.lastTimestamp | tail -20
37
38# 3. Check node status
39kubectl get nodes -l gpu.nvidia.com/class -o wide
40
41# 4. Check GPU health
42kubectl exec -it $(kubectl get pod -l app=inference -o name | head -1) -- nvidia-smi
43```
44
45## Common Incidents
46
47### Inference Service Down
48
491. Check pod status and events
502. If OOMKilled: reduce batch size or upgrade GPU
513. If ImagePullBackOff: check registry credentials
524. If Pending: check GPU quota and availability
53
54### GPU Node Failure
55
561. Pods will be rescheduled automatically
572. If no capacity: scale down non-critical workloads
583. Contact CoreWeave support for extended outages
59
60### Model Loading Failure
61
621. Check HuggingFace token secret exists
632. Verify model name spelling
643. Check PVC has sufficient storage
654. Review container logs for download errors
66
67## Rollback
68
69```bash
70kubectl rollout undo deployment/inference
71```
72
73## Output
74
75- A time-stamped incident record with scope, owner, stabilization action, and redacted evidence.
76- A verified recovery or an explicit escalation with a safe customer-impact mitigation.
77
78## Error Handling
79
80| Incident complication | Required response |
81|---|---|
82| Rollback fails | Stop repeated deploy attempts, escalate to the platform owner, and preserve events. |
83| Diagnostics include a secret | Restrict distribution, rotate the secret, and recollect redacted evidence. |
84| No GPU capacity is available | Prioritize critical services under approved policy; do not remove tenant quotas. |
85| Recovery cannot meet SLO | Declare the continuing impact and use the approved fallback or communication path. |
86
87## Examples
88
89For a failing production rollout, stabilize first and capture only the relevant evidence:
90
91```bash
92kubectl -n inference-prod rollout undo deployment/inference
93kubectl -n inference-prod rollout status deployment/inference --timeout=10m
94kubectl -n inference-prod get events --sort-by=.lastTimestamp | tail -30
95```
96
97Record the revision, outcome, and redacted events in the incident. Do not retry a broken image or expose a troubleshooting endpoint publicly.
98
99## Resources
100
101- CoreWeave Support
102- [CoreWeave Status](https://status.coreweave.com)
103
104## Next Steps
105
106For data handling, see `coreweave-data-handling`.