CoreWeave Upgrade & Migration
Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Overview
CoreWeave is a GPU-specialized cloud provider running Kubernetes-native infrastructure. Migrations involve upgrading between GPU instance types (A100 to H100), updating CUDA driver versions, and handling Kubernetes API version changes across namespaces. Tracking API versions is critical because CoreWeave's instance type labels and resource quotas change between platform releases, and deploying to a deprecated instance class will cause scheduling failures.
Version Detection
import { KubeConfig, CoreV1Api } from "@kubernetes/client-node";
async function detectCoreWeaveVersion(): Promise<void> {
const kc = new KubeConfig();
kc.loadFromDefault();
const k8sApi = kc.makeApiClient(CoreV1Api);
// Check current namespace GPU allocations
const pods = await k8sApi.listNamespacedPod("my-namespace");
for (const pod of pods.body.items) {
const gpuClass = pod.spec?.nodeSelector?.["gpu.nvidia.com/class"];
const cudaVersion = pod.metadata?.labels?.["cuda-version"];
console.log(`Pod ${pod.metadata?.name}: GPU=${gpuClass}, CUDA=${cudaVersion}`);
}
// Detect deprecated instance types
const deprecated = ["A100_PCIE_40GB", "V100_PCIE_16GB", "RTX_A5000"];
const activeGpus = pods.body.items
.map((p) => p.spec?.nodeSelector?.["gpu.nvidia.com/class"])
.filter(Boolean);
const stale = activeGpus.filter((g) => deprecated.includes(g!));
if (stale.length > 0) console.warn(`Deprecated GPU types in use: ${stale.join(", ")}`);
}
Migration Checklist
Schema Migration
// CoreWeave instance type labels changed in 2025 platform update
// Old: gpu.nvidia.com/class: "A100_PCIE_80GB"
// New: gpu.nvidia.com/class: "H100_SXM5_80GB"
interface DeploymentMigration {
oldSelector: Record<string, string>;
newSelector: Record<string, string>;
cudaMinVersion: string;
}
const GPU_MIGRATIONS: DeploymentMigration[] = [
{
oldSelector: { "gpu.nvidia.com/class": "A100_PCIE_80GB" },
newSelector: { "gpu.nvidia.com/class": "H100_SXM5_80GB" },
cudaMinVersion: "12.4",
},
{
oldSelector: { "gpu.nvidia.com/class": "A100_SXM4_80GB" },
newSelector: { "gpu.nvidia.com/class": "H100_SXM5_80GB" },
cudaMinVersion: "12.4",
},
];
function migrateNodeSelector(manifest: any, migration: DeploymentMigration): any {
const selector = manifest.spec?.template?.spec?.nodeSelector;
if (!selector) return manifest;
for (const [key, oldVal] of Object.entries(migration.oldSelector)) {
if (selector[key] === oldVal) {
selector[key] = migration.newSelector[key];
}
}
return manifest;
}
Rollback Strategy
Prerequisites
- A versioned manifest, compatible CUDA/image matrix, and capacity confirmation for the target GPU or cluster change.
- Staging evaluation evidence, named migration owner, and a verified previous deployment revision.
- Backup/retention approval for any PVC, checkpoint, or model-artifact move.
Instructions
- Inventory current selectors, images, API versions, volumes, quotas, and service SLOs.
- Update the manifest in staging, validate server-side, and run compatibility and performance checks.
- Promote through a bounded canary only after the readiness, quality, latency, and integrity gates pass.
- Use the prior immutable revision on any failed gate; do not delete old data or nodes until the observation period closes.
import { AppsV1Api, KubeConfig } from "@kubernetes/client-node";
async function rollbackDeployment(namespace: string, name: string): Promise<void> {
const kc = new KubeConfig();
kc.loadFromDefault();
const appsApi = kc.makeApiClient(AppsV1Api);
// Kubernetes rollout undo — reverts to previous revision
const deployment = await appsApi.readNamespacedDeployment(name, namespace);
const currentRevision = deployment.body.metadata?.annotations?.["deployment.kubernetes.io/revision"];
console.log(`Rolling back ${name} from revision ${currentRevision}`);
// Patch to trigger rollback via revision annotation
await appsApi.patchNamespacedDeployment(name, namespace, {
spec: { template: { metadata: { annotations: { "kubectl.kubernetes.io/restartedAt": new Date().toISOString() } } } },
}, undefined, undefined, undefined, undefined, undefined, { headers: { "Content-Type": "application/strategic-merge-patch+json" } });
console.log(`Rollback initiated for ${name} in ${namespace}`);
}
Error Handling
| Migration Issue |
Symptom |
Fix |
| GPU class not schedulable |
Pod stuck in Pending with Insufficient nvidia.com/gpu |
Verify instance type exists in target region; check quota |
| CUDA version mismatch |
Container crashes with CUDA driver version is insufficient |
Rebuild container with CUDA matching target GPU driver |
| Namespace quota exceeded |
Forbidden: exceeded quota on deployment |
Request quota increase for new instance type via CoreWeave dashboard |
| PVC migration failure |
VolumeAttachment timeout on new node |
Detach old PVC, recreate in target availability zone |
| API version deprecated |
no matches for kind "Deployment" in version "extensions/v1beta1" |
Update manifest to apps/v1 and adjust spec fields |
Output
- A reviewed migration plan with compatibility results, capacity decision, owner, and rollback revision.
- A staged/canary receipt proving readiness, SLO, and artifact-integrity checks.
- A reversible fallback that preserves the old workload and approved data until sign-off.
Examples
Test the target manifest in staging before moving a production selector:
kubectl -n inference-staging apply --dry-run=server -f migration.yaml
kubectl -n inference-staging apply -f migration.yaml
kubectl -n inference-staging rollout status deployment/inference --timeout=15m
If CUDA compatibility, scheduling, or the evaluation gate fails, undo the staging revision and stop promotion. Do not change production affinity or delete the old PVC to make a migration appear complete.
Resources
Next Steps
For CI/CD pipeline integration, see coreweave-ci-integration.
1---2name: coreweave-upgrade-migration3description: Upgrade CoreWeave deployments and migrate between GPU types. Use when migrating from A100 to H100, upgrading CUDA versions, or updating inference server versions. Trigger with phrases like "upgrade coreweave", "coreweave gpu migration", "coreweave cuda upgrade", "migrate coreweave".4license: MIT5---6# CoreWeave Upgrade & Migration
7
8> **Community-contributed.** Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
9
10## Overview
11
12CoreWeave is a GPU-specialized cloud provider running Kubernetes-native infrastructure. Migrations involve upgrading between GPU instance types (A100 to H100), updating CUDA driver versions, and handling Kubernetes API version changes across namespaces. Tracking API versions is critical because CoreWeave's instance type labels and resource quotas change between platform releases, and deploying to a deprecated instance class will cause scheduling failures.
13
14## Version Detection
15
16```typescript
17import { KubeConfig, CoreV1Api } from "@kubernetes/client-node";
18
19async function detectCoreWeaveVersion(): Promise<void> {
20 const kc = new KubeConfig();
21 kc.loadFromDefault();
22 const k8sApi = kc.makeApiClient(CoreV1Api);
23
24 // Check current namespace GPU allocations
25 const pods = await k8sApi.listNamespacedPod("my-namespace");
26 for (const pod of pods.body.items) {
27 const gpuClass = pod.spec?.nodeSelector?.["gpu.nvidia.com/class"];
28 const cudaVersion = pod.metadata?.labels?.["cuda-version"];
29 console.log(`Pod ${pod.metadata?.name}: GPU=${gpuClass}, CUDA=${cudaVersion}`);
30 }
31
32 // Detect deprecated instance types
33 const deprecated = ["A100_PCIE_40GB", "V100_PCIE_16GB", "RTX_A5000"];
34 const activeGpus = pods.body.items
35 .map((p) => p.spec?.nodeSelector?.["gpu.nvidia.com/class"])
36 .filter(Boolean);
37 const stale = activeGpus.filter((g) => deprecated.includes(g!));
38 if (stale.length > 0) console.warn(`Deprecated GPU types in use: ${stale.join(", ")}`);
39}
40```
41
42## Migration Checklist
43
44- [ ] Review CoreWeave release notes for deprecated instance types
45- [ ] Audit all deployments for `gpu.nvidia.com/class` node selectors
46- [ ] Verify CUDA version compatibility with target GPU (see matrix below)
47- [ ] Update container base images to match new CUDA/cuDNN requirements
48- [ ] Test inference latency on new GPU type in staging namespace
49- [ ] Update resource requests (`nvidia.com/gpu`) for new instance memory
50- [ ] Migrate persistent volumes if switching regions or availability zones
51- [ ] Update Kubernetes API version in manifests (e.g., `apps/v1` changes)
52- [ ] Validate HPA scaling behavior on new instance type throughput
53- [ ] Run canary deployment with traffic split before full cutover
54
55## Schema Migration
56
57```typescript
58// CoreWeave instance type labels changed in 2025 platform update
59// Old: gpu.nvidia.com/class: "A100_PCIE_80GB"
60// New: gpu.nvidia.com/class: "H100_SXM5_80GB"
61
62interface DeploymentMigration {
63 oldSelector: Record<string, string>;
64 newSelector: Record<string, string>;
65 cudaMinVersion: string;
66}
67
68const GPU_MIGRATIONS: DeploymentMigration[] = [
69 {
70 oldSelector: { "gpu.nvidia.com/class": "A100_PCIE_80GB" },
71 newSelector: { "gpu.nvidia.com/class": "H100_SXM5_80GB" },
72 cudaMinVersion: "12.4",
73 },
74 {
75 oldSelector: { "gpu.nvidia.com/class": "A100_SXM4_80GB" },
76 newSelector: { "gpu.nvidia.com/class": "H100_SXM5_80GB" },
77 cudaMinVersion: "12.4",
78 },
79];
80
81function migrateNodeSelector(manifest: any, migration: DeploymentMigration): any {
82 const selector = manifest.spec?.template?.spec?.nodeSelector;
83 if (!selector) return manifest;
84 for (const [key, oldVal] of Object.entries(migration.oldSelector)) {
85 if (selector[key] === oldVal) {
86 selector[key] = migration.newSelector[key];
87 }
88 }
89 return manifest;
90}
91```
92
93## Rollback Strategy
94
95## Prerequisites
96
97- A versioned manifest, compatible CUDA/image matrix, and capacity confirmation for the target GPU or cluster change.
98- Staging evaluation evidence, named migration owner, and a verified previous deployment revision.
99- Backup/retention approval for any PVC, checkpoint, or model-artifact move.
100
101## Instructions
102
1031. Inventory current selectors, images, API versions, volumes, quotas, and service SLOs.
1042. Update the manifest in staging, validate server-side, and run compatibility and performance checks.
1053. Promote through a bounded canary only after the readiness, quality, latency, and integrity gates pass.
1064. Use the prior immutable revision on any failed gate; do not delete old data or nodes until the observation period closes.
107
108```typescript
109import { AppsV1Api, KubeConfig } from "@kubernetes/client-node";
110
111async function rollbackDeployment(namespace: string, name: string): Promise<void> {
112 const kc = new KubeConfig();
113 kc.loadFromDefault();
114 const appsApi = kc.makeApiClient(AppsV1Api);
115
116 // Kubernetes rollout undo — reverts to previous revision
117 const deployment = await appsApi.readNamespacedDeployment(name, namespace);
118 const currentRevision = deployment.body.metadata?.annotations?.["deployment.kubernetes.io/revision"];
119 console.log(`Rolling back ${name} from revision ${currentRevision}`);
120
121 // Patch to trigger rollback via revision annotation
122 await appsApi.patchNamespacedDeployment(name, namespace, {
123 spec: { template: { metadata: { annotations: { "kubectl.kubernetes.io/restartedAt": new Date().toISOString() } } } },
124 }, undefined, undefined, undefined, undefined, undefined, { headers: { "Content-Type": "application/strategic-merge-patch+json" } });
125 console.log(`Rollback initiated for ${name} in ${namespace}`);
126}
127```
128
129## Error Handling
130
131| Migration Issue | Symptom | Fix |
132|----------------|---------|-----|
133| GPU class not schedulable | Pod stuck in `Pending` with `Insufficient nvidia.com/gpu` | Verify instance type exists in target region; check quota |
134| CUDA version mismatch | Container crashes with `CUDA driver version is insufficient` | Rebuild container with CUDA matching target GPU driver |
135| Namespace quota exceeded | `Forbidden: exceeded quota` on deployment | Request quota increase for new instance type via CoreWeave dashboard |
136| PVC migration failure | `VolumeAttachment` timeout on new node | Detach old PVC, recreate in target availability zone |
137| API version deprecated | `no matches for kind "Deployment" in version "extensions/v1beta1"` | Update manifest to `apps/v1` and adjust spec fields |
138
139## Output
140
141- A reviewed migration plan with compatibility results, capacity decision, owner, and rollback revision.
142- A staged/canary receipt proving readiness, SLO, and artifact-integrity checks.
143- A reversible fallback that preserves the old workload and approved data until sign-off.
144
145## Examples
146
147Test the target manifest in staging before moving a production selector:
148
149```bash
150kubectl -n inference-staging apply --dry-run=server -f migration.yaml
151kubectl -n inference-staging apply -f migration.yaml
152kubectl -n inference-staging rollout status deployment/inference --timeout=15m
153```
154
155If CUDA compatibility, scheduling, or the evaluation gate fails, undo the staging revision and stop promotion. Do not change production affinity or delete the old PVC to make a migration appear complete.
156
157## Resources
158
159- [CoreWeave GPU Instances](https://docs.coreweave.com/docs/platform/instances/gpu-instances)
160- CoreWeave Kubernetes Docs
161- [CoreWeave Changelog](https://docs.coreweave.com/docs/release-notes)
162
163## Next Steps
164
165For CI/CD pipeline integration, see `coreweave-ci-integration`.