Debugging CockroachDB Operator Migrations
Debugs migration from Helm StatefulSet or public operator v1alpha1 workloads to CockroachDB Operator v1beta1 CrdbNode management. Use this after collecting the general baseline from diagnosing-cockroachdb-helm-deployments.
When to Use This Skill
- A Helm StatefulSet to operator migration is stuck or unclear
- A public operator v1alpha1 to v1beta1 migration has conversion or webhook issues
- Migration labels remain
startorfinalized status.migration.phaseorstatus.migration.messagereports an error- Source StatefulSet ownership, pod ownerReferences, or PVC ownerReferences are unclear
- The operator stops reconciling after migration
Safety Considerations
- Do not delete the source StatefulSet until the migration phase is
finalizedand the customer is ready to complete migration. - Do not modify the source StatefulSet spec during migration.
- Do not remove the migration label
crdb.io/migrateduring an active migration. - Do not manually create
CrdbNoderesources during migration. - Do not delete migrated PVCs. They contain CockroachDB data.
- If the operator stalls after migration, collect the escalation packet before restarting it.
Execution Discipline
- Execute one step at a time and inspect the output before moving on. Migration phase, source workload state, and ownership determine which later checks are safe.
- Do not delete the source StatefulSet, patch labels, patch mode, restart the operator, run Helm upgrades, or change ownerReferences unless the user explicitly approves the action for the target cluster.
- Do not run interactive
kubectl execshells or debug containers unless the user approves them and the customer policy allows the image/source. - In production or when the migration phase is ambiguous, involve TSE or the operator team before changing source or target resources.
Required Inputs
- Migration type: Helm StatefulSet to operator, or public operator v1alpha1 to v1beta1
- Source resource name and namespace
- Target
CrdbClustername and namespace - Operator namespace and version
- Values file or migration command used
- Current migration label and phase
- Whether the source StatefulSet has already been deleted
Step 1: Confirm Migration Documentation and Versions
kubectl -n <operator-namespace> get deploy cockroach-operator -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
kubectl get crd crdbclusters.crdb.cockroachlabs.com -o jsonpath='{.spec.versions[*].name}{"\n"}'
helm -n <cockroachdb-namespace> history <cockroachdb-release> || true
Check the local migration guide and chart changelog for version-specific migration fixes:
Step 2: Inspect Migration State
kubectl -n <cockroachdb-namespace> get crdbcluster <cockroachdb-release> -o json | jq '{
mode: .spec.mode,
migrateLabel: .metadata.labels["crdb.io/migrate"],
migrationStatus: .metadata.labels["crdb.cockroachlabs.com/migration"],
migrationPhase: .status.migration.phase,
migrationMessage: .status.migration.message,
initialized: [.status.conditions[]? | select(.type=="Initialized" or .type=="ClusterInitialized")],
generation: .metadata.generation,
observedGeneration: .status.observedGeneration
}'
kubectl -n <operator-namespace> logs -l app=cockroach-operator --tail=300 | grep -Ei 'migrationctrl|migration|phase|cert' || true
kubectl -n <cockroachdb-namespace> get crdbnodes -o wide
kubectl -n <cockroachdb-namespace> get crdbnodes -o yaml
Step 3: Inspect Source Workload
For Helm StatefulSet migration:
kubectl -n <cockroachdb-namespace> get sts <source-statefulset> -o json | jq '{
replicas: .spec.replicas,
readyReplicas: .status.readyReplicas,
migrateLabel: .metadata.labels["crdb.io/migrate"],
ownerReferences: .metadata.ownerReferences
}'
kubectl -n <cockroachdb-namespace> get sts <source-statefulset> -o yaml
kubectl -n <cockroachdb-namespace> get pods -l app.kubernetes.io/name=cockroachdb -o yaml | grep -E 'name:|ownerReferences:|kind:|uid:|controller:' -A8
kubectl -n <cockroachdb-namespace> get pvc -o yaml | grep -E 'name:|ownerReferences:|kind:|uid:|controller:' -A8
For public operator v1alpha1 migration:
kubectl -n <cockroachdb-namespace> get crdbcluster.crdb.cockroachlabs.com <cluster-name> -o yaml
kubectl -n <cockroachdb-namespace> get crdbcluster.v1beta1.crdb.cockroachlabs.com <cluster-name> -o yaml 2>&1 || true
kubectl -n <operator-namespace> get svc cockroach-webhook-service
kubectl -n <operator-namespace> get endpoints cockroach-webhook-service
kubectl get validatingwebhookconfigurations | grep cockroach
The operator supports both v1alpha1 and v1beta1 through conversion webhooks. If conversion fails, check webhook service health and operator logs.
Step 4: Interpret Migration Phases
| Phase | What Happens | What to Check |
|---|---|---|
| Init | Operator validates prerequisites and prepares migration | Migration label is start; check operator logs for validation errors |
| CertMigration | Operator migrates or creates TLS certificate resources | Check certificate Secrets, CA resources, cert-manager resources, and cert metadata |
| PodMigration | Operator creates CrdbNode resources to take over pods from the StatefulSet |
Check CrdbNode creation and pod ownerReferences |
| Finalization | Operator sets mode=MutableOnly and stops source migration work |
CrdbCluster mode should be MutableOnly; migration label should be finalized |
| Complete | Customer deletes the source StatefulSet and the delete event triggers final reconcile | StatefulSet must be manually deleted only after finalization; migration label should become complete |
Step 5: Debug Common Migration Stalls
Migration Not Progressing
kubectl -n <cockroachdb-namespace> get crdbcluster <cockroachdb-release> -o jsonpath='{.status.migration}{"\n"}'
kubectl -n <operator-namespace> logs -l app=cockroach-operator --tail=300 | grep -Ei 'migration|phase|cert|error' || true
kubectl -n <operator-namespace> get deploy cockroach-operator -o jsonpath='{.spec.template.spec.containers[0].args}{"\n"}'
Check:
- The migration feature is enabled in operator args, if the deployed version uses a feature flag.
- The source StatefulSet still exists when phase is before
finalized. - The source StatefulSet spec has not been changed mid-migration.
- Certificate prerequisites are present and trusted.
CrdbNodeobjects are being created and observed.
CertMigration Issues
Use configuring-cockroachdb-helm-tls. Confirm:
- CA ConfigMap or Secret exists.
- Node, HTTP, and root client certificate resources exist.
- Certificate expiry, issuer, subject, and SANs match generated pod DNS names.
- Cert-manager
Certificateresources are Ready, when cert-manager is used. - Secret keys are present without printing private key values.
PodMigration Issues
kubectl -n <cockroachdb-namespace> get pods -l app.kubernetes.io/name=cockroachdb -o wide
kubectl -n <cockroachdb-namespace> describe pods -l app.kubernetes.io/name=cockroachdb
kubectl -n <cockroachdb-namespace> get crdbnodes -o custom-columns=NAME:.metadata.name,PHASE:.status.phase,NODE_ID:.status.nodeID,OBSERVED:.status.observedGeneration
kubectl -n <cockroachdb-namespace> get pvc -o json | jq '.items[] | {name: .metadata.name, ownerReferences: .metadata.ownerReferences, storageClass: .spec.storageClassName, capacity: .status.capacity}'
Check:
CrdbNoderesources are created for expected pods.- Pods and PVCs have expected ownership after the operator takes over.
- PVCs are not orphaned before deleting source resources.
Finalization or Completion Stuck
Check:
modeisMutableOnly, notDisabled.- migration label is
finalizedbefore deleting the source StatefulSet. - source StatefulSet deletion happened only after finalization.
- migration label eventually becomes
complete. - operator logs show a final reconcile after the source StatefulSet delete event.
Step 6: Post-Migration Operator Stall
If the operator stops reconciling after migration completes:
- Check the initialization condition. If missing on an already-initialized cluster, the operator may try to initialize and block.
- Verify
modeisMutableOnly. - Verify the migration label is
complete, notstartorfinalized. - Use collecting-cockroachdb-operator-escalation-packet to collect pprof goroutine dump, metrics, and logs before restart.
Output Format
Return findings in this order:
- Migration type and current phase
- Source workload state
- Target
CrdbClustermode, migration labels, and migration message CrdbNode, pod, and PVC ownership state- Current blocker and likely cause
- Safe next action
- Whether escalation packet collection is required