CCB Self Recover
Use this skill for runtime recovery after diagnosis. Mutations must go through
CCB control-plane commands. Raw tmux mutation and direct runtime-file writes are
forbidden.
Recovery Gates
Before any mutation:
- Confirm maintenance intent from the user, such as "fix", "recover",
"restart if safe", "apply this config", or "make CCB healthy".
- Read the current mounted daemon graph. The target must be a current
daemon-graph agent for clear or restart-like actions.
- Check busy/pending state:
ccb ps
ccb queue --detail <agent|all>
ccb pend --inbox --detail <agent>
ccb trace <id> when the issue involves active or pending lineage
- Check
ccb fault list. If active fault-injection rules affect the target,
treat them as diagnostic evidence. Clear them with
ccb fault clear <rule_id|all> only when the user intended maintenance and
the rules are known test residue. If a rule is recent, or its task/reason
fields could represent an active drill, ask before clearing unless the user
explicitly asked to clear fault-injection rules.
- Choose the least disruptive supported action.
- Report exact commands, gates, affected agents, blockers, and what remains
unchanged.
If the target is unknown, busy, has queued work, has pending reply delivery, or
has a pending callback continuation, stop and report blockers.
Provider/API Or Startup-Input Recovery
For provider/API failures or changes that affect provider process, model, base
URL, environment, provider profile, command template, role assets, or startup
context, use this exact flow:
- Gather evidence without reading secrets.
- Use built-in
ccb-config to edit .ccb/ccb.config only when the fallback
provider/model/base URL/profile/env-var reference is already configured or
explicitly supplied by the user as a safe reference.
- Run
ccb config validate.
- Run
ccb reload --dry-run.
- If validation and dry-run pass, and the user intended materialization, run
ccb reload.
- Re-check the current daemon graph and affected agent status.
- Decide whether affected running agents still use stale provider process,
environment, model, base URL, role asset, or context state.
- If runtime refresh is still needed, restart only one affected current-graph
agent at a time with
ccb restart <agent>, and only when busy checks pass.
- If
ccb restart <agent> returns blocked or failed, report the blockers.
Do not emulate restart with tmux commands. The remaining user-level options
are to continue with unaffected agents or explicitly stop and restart the
project with ccb kill then ccb; do not run project shutdown
autonomously as a substitute for single-agent restart.
ccb reload is not the recovery finish line. It materializes config into the
daemon graph; running provider processes may still hold old startup inputs.
Supported Actions
ccb clear <agent>: provider-native conversation/context clear. Run the
pre-mutation checks in the Recovery Gates section first; use only when
context clearing is the right fix and pending-work checks pass.
ccb reload --dry-run: no-mutation config reload plan. Always safe in
maintenance workflows.
ccb reload: config materialization after ccb config validate,
ccb reload --dry-run, supported plan, and explicit user materialization
intent.
ccb restart <agent>: one configured pane-backed current-graph agent after
busy checks pass. The command itself must report restart_status, blockers,
restartable agents, busy gate evidence, and old/new runtime evidence.
ccb roles update agentroles.ccb_self or ccb roles sync <path>: role asset
repair when the user is repairing ccb_self itself, there is no active
maintenance operation that depends on the current role assets, and the target
role/source version is clear.
Handoffs
- Use
ccb-self-chain first when trace evidence shows message/reply lineage is
the primary problem. Restart is not the first repair for a broken job chain.
- Use built-in
ccb-config for disk config edits and affected-agent reporting.
After reload, this skill owns guarded runtime refresh decisions.
- Return the original business work to the original target agent unless the
user explicitly retargets it.
Red Lines
- Never restart all agents or unrelated agents.
- Never use
tmux kill-pane, kill-window, kill-server, respawn-pane,
send-keys, manual pane creation, or other raw tmux mutation.
- Never write lifecycle, lease, runtime, mailbox, provider session, or tmux
authority files directly.
- Never read, print, store, search for, scrape, borrow, or use API keys or
credentials.
- Never treat
.ccb/agents/*, disk config, pid files, or tmux panes as live
restart target authority.
1---2name: ccb-self-recover3description: Recover CCB agents, panes, mounts, provider contexts, API/provider failures, config reload aftermath, clear operations, and guarded single-agent restarts. Use when the user asks to fix, recover, restart if safe, clear context, reload, remount, or keep work going after provider/API failure.4---5
6# CCB Self Recover
7
8Use this skill for runtime recovery after diagnosis. Mutations must go through
9CCB control-plane commands. Raw tmux mutation and direct runtime-file writes are
10forbidden.
11
12## Recovery Gates
13
14Before any mutation:
15
161. Confirm maintenance intent from the user, such as "fix", "recover",
17 "restart if safe", "apply this config", or "make CCB healthy".
182. Read the current mounted daemon graph. The target must be a current
19 daemon-graph agent for clear or restart-like actions.
203. Check busy/pending state:
21 - `ccb ps`
22 - `ccb queue --detail <agent|all>`
23 - `ccb pend --inbox --detail <agent>`
24 - `ccb trace <id>` when the issue involves active or pending lineage
254. Check `ccb fault list`. If active fault-injection rules affect the target,
26 treat them as diagnostic evidence. Clear them with
27 `ccb fault clear <rule_id|all>` only when the user intended maintenance and
28 the rules are known test residue. If a rule is recent, or its task/reason
29 fields could represent an active drill, ask before clearing unless the user
30 explicitly asked to clear fault-injection rules.
315. Choose the least disruptive supported action.
326. Report exact commands, gates, affected agents, blockers, and what remains
33 unchanged.
34
35If the target is unknown, busy, has queued work, has pending reply delivery, or
36has a pending callback continuation, stop and report blockers.
37
38## Provider/API Or Startup-Input Recovery
39
40For provider/API failures or changes that affect provider process, model, base
41URL, environment, provider profile, command template, role assets, or startup
42context, use this exact flow:
43
441. Gather evidence without reading secrets.
452. Use built-in `ccb-config` to edit `.ccb/ccb.config` only when the fallback
46 provider/model/base URL/profile/env-var reference is already configured or
47 explicitly supplied by the user as a safe reference.
483. Run `ccb config validate`.
494. Run `ccb reload --dry-run`.
505. If validation and dry-run pass, and the user intended materialization, run
51 `ccb reload`.
526. Re-check the current daemon graph and affected agent status.
537. Decide whether affected running agents still use stale provider process,
54 environment, model, base URL, role asset, or context state.
558. If runtime refresh is still needed, restart only one affected current-graph
56 agent at a time with `ccb restart <agent>`, and only when busy checks pass.
579. If `ccb restart <agent>` returns `blocked` or `failed`, report the blockers.
58 Do not emulate restart with tmux commands. The remaining user-level options
59 are to continue with unaffected agents or explicitly stop and restart the
60 project with `ccb kill` then `ccb`; do not run project shutdown
61 autonomously as a substitute for single-agent restart.
62
63`ccb reload` is not the recovery finish line. It materializes config into the
64daemon graph; running provider processes may still hold old startup inputs.
65
66## Supported Actions
67
68- `ccb clear <agent>`: provider-native conversation/context clear. Run the
69 pre-mutation checks in the Recovery Gates section first; use only when
70 context clearing is the right fix and pending-work checks pass.
71- `ccb reload --dry-run`: no-mutation config reload plan. Always safe in
72 maintenance workflows.
73- `ccb reload`: config materialization after `ccb config validate`,
74 `ccb reload --dry-run`, supported plan, and explicit user materialization
75 intent.
76- `ccb restart <agent>`: one configured pane-backed current-graph agent after
77 busy checks pass. The command itself must report `restart_status`, blockers,
78 restartable agents, busy gate evidence, and old/new runtime evidence.
79- `ccb roles update agentroles.ccb_self` or `ccb roles sync <path>`: role asset
80 repair when the user is repairing `ccb_self` itself, there is no active
81 maintenance operation that depends on the current role assets, and the target
82 role/source version is clear.
83
84## Handoffs
85
86- Use `ccb-self-chain` first when trace evidence shows message/reply lineage is
87 the primary problem. Restart is not the first repair for a broken job chain.
88- Use built-in `ccb-config` for disk config edits and affected-agent reporting.
89 After reload, this skill owns guarded runtime refresh decisions.
90- Return the original business work to the original target agent unless the
91 user explicitly retargets it.
92
93## Red Lines
94
95- Never restart all agents or unrelated agents.
96- Never use `tmux kill-pane`, `kill-window`, `kill-server`, `respawn-pane`,
97 `send-keys`, manual pane creation, or other raw tmux mutation.
98- Never write lifecycle, lease, runtime, mailbox, provider session, or tmux
99 authority files directly.
100- Never read, print, store, search for, scrape, borrow, or use API keys or
101 credentials.
102- Never treat `.ccb/agents/*`, disk config, pid files, or tmux panes as live
103 restart target authority.