Error Recovery Playbook
Standardized procedures for every known failure mode. When something goes wrong, find the matching scenario and follow the playbook. Do not improvise unless no playbook exists — and if that happens, document the new scenario afterward.
Agent Unresponsive
Symptoms: Agent does not respond within expected time window. No output, no error, no acknowledgment.
Playbook:
- Retry (immediate): Resend the same
sessions_send message. Agents may have missed the initial message due to transient issues.
- Wait 3 minutes. Some tasks take longer than expected. Give the agent a reasonable buffer.
- First nudge: Send a status check message: "Status update requested — are you still working on {task}?"
- Wait another expected-duration window.
- Second nudge: Send with urgency: "Second status check — {task} is blocking the pipeline. Please respond with current status or any blockers."
- Check with ClawExpert: Send ClawExpert a diagnostic request: "Is {agent name} reachable? Any infrastructure issues affecting agent availability?" ClawExpert may identify system-level problems.
- Escalate to Nick: If still no response after 2 nudges and ClawExpert check, report to Nick:
- Which agent is unresponsive.
- What task they were assigned.
- How long they've been unresponsive.
- What you've tried.
- Recommended action (wait longer, reassign task, manual intervention).
V0 API Failure
Symptoms: Pixel reports V0 API is down, timing out, or returning errors. Design pipeline cannot proceed.
Playbook:
- Retry (immediate): Ask Pixel to retry the V0 request. Transient failures are common.
- Wait 2 minutes and retry again. V0 outages are often brief.
- Stitch fallback: If V0 remains unavailable, instruct Pixel to switch to Stitch for design generation. Note in the project file that Stitch was used (reduced fidelity).
- Text-based specs fallback: If both V0 and Stitch are unavailable, instruct Pixel to produce detailed text-based design specifications:
- Layout descriptions with exact measurements.
- Color values and typography specs.
- Component hierarchy and behavior descriptions.
- Wireframe-level sketches if possible.
- Note the fallback in the project file and in the handoff to Maks so he knows designs are text-based and should use best judgment for visual details.
- Report to Nick if the fallback significantly impacts quality or timeline.
Forge Blocks 3+ Times (Circuit Breaker)
Symptoms: Maks and Forge have completed 3 fix loops and Forge still has not approved. The circuit breaker has triggered.
Playbook:
Stop the loop immediately. Do not send a 4th fix request to Maks.
Compile a summary:
- Original issues Forge raised in loop 1.
- What Maks fixed in each loop.
- Remaining issues Forge is still flagging.
- Pattern analysis: Are these the same issues recurring, or new issues each time?
Escalate to Nick with 3 options:
Option A: Ship with known issues.
The remaining issues are [{list}]. None are P0 security/data-loss risks. We document them as known issues and address in a follow-up iteration.
Option B: Bring in additional context.
Forge and Maks may be misaligned on [{specific area}]. Nick can provide clarification or additional requirements that resolve the disagreement.
Option C: Redesign the problematic section.
The issues stem from [{root cause — e.g., a specific component, architectural decision}]. Send the affected screen(s) back to Pixel for redesign, then rebuild.
Wait for Nick's decision. Do not proceed until Nick picks an option or provides an alternative direction.
Build Fails to Deploy
Symptoms: Maks reports a deployment failure. Preview or production URL is not accessible.
Playbook:
- Get the error. Ask Maks for the exact error message, stack trace, or deployment log output.
- Categorize the error:
- Code error (build fails, type errors, missing imports): Maks debugs and retries. This is within Maks's domain.
- Infrastructure error (DNS, SSL, Vercel config, environment variables, permissions): Consult ClawExpert.
- Dependency error (package conflicts, version mismatches): Maks resolves, or consult ClawExpert if it's a system-level issue.
- If code error: Tell Maks the specific error and ask for a fix. Allow 2 attempts.
- If infrastructure error: Send ClawExpert the error with context:
- What was being deployed.
- The deployment platform and configuration.
- The exact error message.
- Wait for ClawExpert's diagnostic and fix.
- If still failing after 2 Maks attempts + ClawExpert consultation: Escalate to Nick with the error details and what's been tried.
Design Mismatch After Build
Symptoms: QA or visual inspection reveals that the build doesn't match Pixel's designs for specific screens.
Playbook:
- Identify the specific mismatches. Document which screens don't match and what's different (layout, colors, spacing, missing elements, wrong behavior).
- Assess severity:
- Minor (slight spacing, color shade): Note for Maks with specific corrections. No need to involve Pixel.
- Major (wrong layout, missing sections, broken user flow): Send back to Maks with Pixel's original specs highlighted.
- Send Maks targeted fix requests for each mismatched screen:
- Reference the specific Pixel design (V0 chat ID, preview URL).
- Point out exactly what's different.
- Ask for a match to the design, not a "close enough."
- If Maks cannot match the design due to technical constraints, notify Pixel and ask for an alternative design approach that's feasible to implement.
- Re-QA the fixed screens before proceeding.
Nick Changes Requirements Mid-Build
Symptoms: Nick sends updated requirements while the project is in BUILD, REVIEW, or later phases.
Playbook:
- Acknowledge immediately. Tell Nick you've received the changes and are assessing impact.
- Assess the change type:
- Cosmetic (copy changes, color tweaks, minor layout adjustments): Send directly to Maks as a patch. No need to loop back to earlier phases.
- Structural (new screens, removed screens, different user flows, new features, changed backend requirements): Requires re-evaluation.
- For cosmetic changes:
- Send Maks the specific changes.
- Note the change in the project file.
- Continue the pipeline from current phase.
- For structural changes:
- Pause the current pipeline.
- Update the project file with the new requirements.
- Determine which phase to re-enter:
- New screens → Back to DESIGN for those screens only.
- New features/backend → Back to DESIGN or BUILD depending on scope.
- Changed ICP/positioning → Back to RESEARCH.
- Report to Nick: "Structural change received. This resets us to {phase}. New ETA: {estimate}."
- Never silently absorb structural changes. Always report the impact to Nick.
Gateway Crash / MaksPM Restart
Symptoms: MaksPM loses context — conversation resets, memory is gone, no awareness of active projects.
Playbook:
- Read active-projects/. List all project files and their current states.
- For each active project:
- Read the project file to determine current phase and last activity.
- Check the Phase Log for the most recent transition.
- Identify what was in progress when the crash occurred.
- Resume from the project file. The project file is the source of truth. Pick up from the last completed phase:
- If a brief was sent but no response received → Resend the brief.
- If a response was received but the gate wasn't checked → Run the gate check.
- If a phase was complete but the next wasn't started → Start the next phase.
- Send Nick a recovery report:
- "MaksPM restarted. Recovered {n} active project(s). Current status: {summary per project}. Resuming from last known state."
- Re-establish contact with any agents that had pending tasks via
sessions_send status checks.
Subagent Spawn Failed or Timed Out
- Check if the spawn completed: look at the auto-announce message for errors
- If timeout (no response after runTimeoutSeconds): spawn again with longer timeout
- If agent errored: read the error, fix the task prompt, spawn again
- After 2 failed spawns → escalate to Nick: "[Agent] failed twice on [task]. Error: [details]. Options: (A) retry with different approach (B) skip this phase (C) you intervene"
Subagent Returns Incomplete Work
- Check what's missing against the quality gate
- Spawn the same agent again with: "Previous output was incomplete. Missing: [specific items]. Complete only the missing items."
- If still incomplete after retry → proceed with what you have + note the gap to Nick
Auto-Announce Didn't Arrive
- Check sessions_list for the spawned session — is it still running?
- If still running → wait (don't re-spawn and create duplicates)
- If completed but no announce → read result via sessions_history with the session key
- If session doesn't exist → spawn again
1---2name: error-recovery3description: Error Recovery Playbook4---5# Error Recovery Playbook67Standardized procedures for every known failure mode. When something goes wrong, find the matching scenario and follow the playbook. Do not improvise unless no playbook exists — and if that happens, document the new scenario afterward.89---1011## Agent Unresponsive1213**Symptoms:** Agent does not respond within expected time window. No output, no error, no acknowledgment.1415**Playbook:**16171. **Retry (immediate):** Resend the same `sessions_send` message. Agents may have missed the initial message due to transient issues.182. **Wait 3 minutes.** Some tasks take longer than expected. Give the agent a reasonable buffer.193. **First nudge:** Send a status check message: "Status update requested — are you still working on {task}?"204. **Wait another expected-duration window.**215. **Second nudge:** Send with urgency: "Second status check — {task} is blocking the pipeline. Please respond with current status or any blockers."226. **Check with ClawExpert:** Send ClawExpert a diagnostic request: "Is {agent name} reachable? Any infrastructure issues affecting agent availability?" ClawExpert may identify system-level problems.237. **Escalate to Nick:** If still no response after 2 nudges and ClawExpert check, report to Nick:24 - Which agent is unresponsive.25 - What task they were assigned.26 - How long they've been unresponsive.27 - What you've tried.28 - Recommended action (wait longer, reassign task, manual intervention).2930---3132## V0 API Failure3334**Symptoms:** Pixel reports V0 API is down, timing out, or returning errors. Design pipeline cannot proceed.3536**Playbook:**37381. **Retry (immediate):** Ask Pixel to retry the V0 request. Transient failures are common.392. **Wait 2 minutes and retry again.** V0 outages are often brief.403. **Stitch fallback:** If V0 remains unavailable, instruct Pixel to switch to Stitch for design generation. Note in the project file that Stitch was used (reduced fidelity).414. **Text-based specs fallback:** If both V0 and Stitch are unavailable, instruct Pixel to produce detailed text-based design specifications:42 - Layout descriptions with exact measurements.43 - Color values and typography specs.44 - Component hierarchy and behavior descriptions.45 - Wireframe-level sketches if possible.465. **Note the fallback** in the project file and in the handoff to Maks so he knows designs are text-based and should use best judgment for visual details.476. **Report to Nick** if the fallback significantly impacts quality or timeline.4849---5051## Forge Blocks 3+ Times (Circuit Breaker)5253**Symptoms:** Maks and Forge have completed 3 fix loops and Forge still has not approved. The circuit breaker has triggered.5455**Playbook:**56571. **Stop the loop immediately.** Do not send a 4th fix request to Maks.582. **Compile a summary:**59 - Original issues Forge raised in loop 1.60 - What Maks fixed in each loop.61 - Remaining issues Forge is still flagging.62 - Pattern analysis: Are these the same issues recurring, or new issues each time?633. **Escalate to Nick with 3 options:**6465 > **Option A: Ship with known issues.**66 > The remaining issues are [{list}]. None are P0 security/data-loss risks. We document them as known issues and address in a follow-up iteration.6768 > **Option B: Bring in additional context.**69 > Forge and Maks may be misaligned on [{specific area}]. Nick can provide clarification or additional requirements that resolve the disagreement.7071 > **Option C: Redesign the problematic section.**72 > The issues stem from [{root cause — e.g., a specific component, architectural decision}]. Send the affected screen(s) back to Pixel for redesign, then rebuild.73744. **Wait for Nick's decision.** Do not proceed until Nick picks an option or provides an alternative direction.7576---7778## Build Fails to Deploy7980**Symptoms:** Maks reports a deployment failure. Preview or production URL is not accessible.8182**Playbook:**83841. **Get the error.** Ask Maks for the exact error message, stack trace, or deployment log output.852. **Categorize the error:**86 - **Code error** (build fails, type errors, missing imports): Maks debugs and retries. This is within Maks's domain.87 - **Infrastructure error** (DNS, SSL, Vercel config, environment variables, permissions): Consult ClawExpert.88 - **Dependency error** (package conflicts, version mismatches): Maks resolves, or consult ClawExpert if it's a system-level issue.893. **If code error:** Tell Maks the specific error and ask for a fix. Allow 2 attempts.904. **If infrastructure error:** Send ClawExpert the error with context:91 - What was being deployed.92 - The deployment platform and configuration.93 - The exact error message.94 - Wait for ClawExpert's diagnostic and fix.955. **If still failing after 2 Maks attempts + ClawExpert consultation:** Escalate to Nick with the error details and what's been tried.9697---9899## Design Mismatch After Build100101**Symptoms:** QA or visual inspection reveals that the build doesn't match Pixel's designs for specific screens.102103**Playbook:**1041051. **Identify the specific mismatches.** Document which screens don't match and what's different (layout, colors, spacing, missing elements, wrong behavior).1062. **Assess severity:**107 - **Minor** (slight spacing, color shade): Note for Maks with specific corrections. No need to involve Pixel.108 - **Major** (wrong layout, missing sections, broken user flow): Send back to Maks with Pixel's original specs highlighted.1093. **Send Maks targeted fix requests** for each mismatched screen:110 - Reference the specific Pixel design (V0 chat ID, preview URL).111 - Point out exactly what's different.112 - Ask for a match to the design, not a "close enough."1134. **If Maks cannot match the design** due to technical constraints, notify Pixel and ask for an alternative design approach that's feasible to implement.1145. **Re-QA the fixed screens** before proceeding.115116---117118## Nick Changes Requirements Mid-Build119120**Symptoms:** Nick sends updated requirements while the project is in BUILD, REVIEW, or later phases.121122**Playbook:**1231241. **Acknowledge immediately.** Tell Nick you've received the changes and are assessing impact.1252. **Assess the change type:**126 - **Cosmetic** (copy changes, color tweaks, minor layout adjustments): Send directly to Maks as a patch. No need to loop back to earlier phases.127 - **Structural** (new screens, removed screens, different user flows, new features, changed backend requirements): Requires re-evaluation.1283. **For cosmetic changes:**129 - Send Maks the specific changes.130 - Note the change in the project file.131 - Continue the pipeline from current phase.1324. **For structural changes:**133 - Pause the current pipeline.134 - Update the project file with the new requirements.135 - Determine which phase to re-enter:136 - New screens → Back to DESIGN for those screens only.137 - New features/backend → Back to DESIGN or BUILD depending on scope.138 - Changed ICP/positioning → Back to RESEARCH.139 - Report to Nick: "Structural change received. This resets us to {phase}. New ETA: {estimate}."1405. **Never silently absorb structural changes.** Always report the impact to Nick.141142---143144## Gateway Crash / MaksPM Restart145146**Symptoms:** MaksPM loses context — conversation resets, memory is gone, no awareness of active projects.147148**Playbook:**1491501. **Read active-projects/.** List all project files and their current states.1512. **For each active project:**152 - Read the project file to determine current phase and last activity.153 - Check the Phase Log for the most recent transition.154 - Identify what was in progress when the crash occurred.1553. **Resume from the project file.** The project file is the source of truth. Pick up from the last completed phase:156 - If a brief was sent but no response received → Resend the brief.157 - If a response was received but the gate wasn't checked → Run the gate check.158 - If a phase was complete but the next wasn't started → Start the next phase.1594. **Send Nick a recovery report:**160 - "MaksPM restarted. Recovered {n} active project(s). Current status: {summary per project}. Resuming from last known state."1615. **Re-establish contact with any agents that had pending tasks** via `sessions_send` status checks.162163---164165## Subagent Spawn Failed or Timed Out1661. Check if the spawn completed: look at the auto-announce message for errors1672. If timeout (no response after runTimeoutSeconds): spawn again with longer timeout1683. If agent errored: read the error, fix the task prompt, spawn again1694. After 2 failed spawns → escalate to Nick: "[Agent] failed twice on [task]. Error: [details]. Options: (A) retry with different approach (B) skip this phase (C) you intervene"170171## Subagent Returns Incomplete Work1721. Check what's missing against the quality gate1732. Spawn the same agent again with: "Previous output was incomplete. Missing: [specific items]. Complete only the missing items."1743. If still incomplete after retry → proceed with what you have + note the gap to Nick175176## Auto-Announce Didn't Arrive1771. Check sessions_list for the spawned session — is it still running?1782. If still running → wait (don't re-spawn and create duplicates)1793. If completed but no announce → read result via sessions_history with the session key1804. If session doesn't exist → spawn again