ZERG Architecture
Zero-Effort Rapid Growth - Parallel Claude Code Execution System
ZERG is a distributed software development system that coordinates multiple Claude Code instances to build features in parallel. It combines spec-driven development (GSD methodology), level-based task execution, and git worktrees for isolated execution.
Table of Contents
- Core Principles
- System Layers
- Execution Flow
- Module Reference
- Cross-Cutting Capabilities
- Zergling Execution Model
- Resilience
- State Management
- Claude Code Task Integration
- Context Engineering
- Diagnostics Engine
- Quality Gates
- Pre-commit Hooks
- Security Model
- Configuration
Core Principles
Spec as Memory
Zerglings do not share conversation context. They share:
requirements.md— what to builddesign.md— how to build ittask-graph.json— atomic work units
This makes zerglings stateless. Any zergling can pick up any task. Crash recovery is trivial.
Exclusive File Ownership
Each task declares which files it creates or modifies. The design phase ensures no overlap within a level. This eliminates merge conflicts without runtime locking.
{
"id": "TASK-001",
"files": {
"create": ["src/models/user.py"],
"modify": [],
"read": ["src/config.py"]
}
}
Level-Based Execution
Tasks are organized into dependency levels:
| Level | Name | Description |
|---|---|---|
| 1 | Foundation | Types, schemas, config |
| 2 | Core | Business logic, services |
| 3 | Integration | Wiring, endpoints |
| 4 | Testing | Unit and integration tests |
| 5 | Quality | Docs, cleanup |
All zerglings complete Level N before any proceed to N+1. The orchestrator merges all branches, runs quality gates, then signals zerglings to continue.
Git Worktrees for Isolation
Each zergling operates in its own git worktree with its own branch:
.zerg-worktrees/{feature}/worker-0/ -> branch: zerg/{feature}/worker-0
.zerg-worktrees/{feature}/worker-1/ -> branch: zerg/{feature}/worker-1
Zerglings commit independently. No filesystem conflicts.
System Layers
What Are System Layers?
Think of ZERG like a factory assembly line. Raw materials (your requirements) enter at one end, and finished software comes out the other. But unlike a physical factory, you can't just dump everything in at once—each stage transforms the work before passing it to the next.
ZERG organizes this transformation into four distinct layers, each with a specific job. Understanding these layers helps you know what's happening at any point in the build process and where to look when something goes wrong.
Why Do Layers Exist?
Without clear boundaries, parallel workers would step on each other's toes. Imagine five cooks trying to simultaneously shop for ingredients, prep vegetables, and plate dishes—chaos. Layers enforce order: you finish planning before you start designing, and you finish designing before workers start building.
Each layer also serves as a checkpoint. If requirements are unclear, you find out during planning—not after three workers have already built the wrong thing. This "fail early" philosophy saves hours of wasted work.
Layer Architecture
+---------------------------------------------------------------------+
| Layer 1: Planning |
| requirements.md + INFRASTRUCTURE.md |
+---------------------------------------------------------------------+
|
v
+---------------------------------------------------------------------+
| Layer 2: Design |
| design.md + task-graph.json |
+---------------------------------------------------------------------+
|
v
+---------------------------------------------------------------------+
| Layer 3: Orchestration |
| Zergling lifecycle - Level sync - Branch merging - Monitoring |
+---------------------------------------------------------------------+
| | |
v v v
+-------------+ +-------------+ +-------------+
| Zergling 0 | | Zergling 1 | | Zergling N |
| (worktree) | | (worktree) | | (worktree) |
+-------------+ +-------------+ +-------------+
| | |
+-------------------+-------------------+
|
v
+---------------------------------------------------------------------+
| Layer 4: Quality Gates |
| Lint - Type-check - Test - Merge to main |
+---------------------------------------------------------------------+
How the Layers Connect
Layer 1 (Planning) captures what you want to build. The /zerg:plan command guides you through discovery and produces requirements.md—a document that workers will read to understand their mission.
Layer 2 (Design) breaks requirements into buildable pieces. The /zerg:design command analyzes your requirements and produces both an architecture document (design.md) and a task graph (task-graph.json) that lists every atomic piece of work.
Layer 3 (Orchestration) manages the actual building. It spawns workers (zerglings), assigns them tasks, monitors progress, and coordinates the merge dance at level boundaries. This is where the parallel magic happens.
Layer 4 (Quality Gates) ensures nothing broken reaches your main branch. After each level completes, gates run linting, type checking, and tests. Only code that passes all gates gets merged—protecting your codebase from half-finished work.
Plugin System
What Is the Plugin System?
ZERG's plugin system lets you extend the build process without modifying ZERG's core code. Think of it like adding apps to your phone—the phone works fine out of the box, but plugins let you customize it for your specific needs.
Plugins come in three flavors: quality gates (add new validation checks), lifecycle hooks (react to events like "task completed"), and launchers (change how workers run). Each type has a clear contract defining what it can do and what information it receives.
Why Does the Plugin System Exist?
Every team has unique requirements. Maybe you need to run security scans, post Slack notifications, or deploy workers to Kubernetes instead of local Docker. Rather than bloating ZERG with every possible feature, the plugin system lets you add exactly what you need.
This separation also makes ZERG more maintainable. Core orchestration logic stays simple while specialized behaviors live in plugins that can be developed, tested, and updated independently.
Plugin Architecture
PluginRegistry
+-- hooks: dict[str, list[Callable]] # LifecycleHookPlugin callbacks
+-- gates: dict[str, QualityGatePlugin] # Named quality gate plugins
+-- launchers: dict[str, LauncherPlugin] # Named launcher plugins
QualityGatePlugin (ABC)
+-- name: str
+-- run(ctx: GateContext) -> GateRunResult
LifecycleHookPlugin (ABC)
+-- name: str
+-- on_event(event: LifecycleEvent) -> None
LauncherPlugin (ABC)
+-- name: str
+-- create_launcher(config) -> WorkerLauncher
How Plugins Connect
The PluginRegistry is the central catalog—it knows every plugin that's loaded. When the orchestrator needs to run quality gates, it asks the registry for all registered gate plugins. When a task completes, the registry dispatches the event to all lifecycle hooks.
PluginHookEvent lifecycle (8 events): TASK_STARTED, TASK_COMPLETED, LEVEL_COMPLETE, MERGE_COMPLETE, RUSH_FINISHED, QUALITY_GATE_RUN, WORKER_SPAWNED, WORKER_EXITED
Integration points:
orchestrator.py— emits lifecycle events, runs plugin gates after mergeworker_protocol.py— emitsTASK_STARTED/TASK_COMPLETEDeventsgates.py— delegates to plugin gates registered in the registrylauncher.py— resolves launcher plugins by name viaget_plugin_launcher()
Discovery: Plugins are loaded via importlib.metadata entry points (group: zerg.plugins) or YAML-configured shell command hooks in .zerg/config.yaml.
Configuration models: PluginsConfig -> HookConfig, PluginGateConfig, LauncherPluginConfig (see zerg/plugin_config.py).
Execution Flow
Understanding how ZERG executes a feature build is like understanding how a relay race works. Each phase hands off to the next, and workers run their legs in parallel within each level. Let's trace the entire journey from "I have an idea" to "the feature is merged."
Planning Phase (/zerg:plan)
What Is Planning?
Planning is the conversation phase where ZERG helps you articulate what you want to build. Rather than accepting vague requirements and guessing wrong, ZERG asks probing questions until both you and the system have a shared understanding.
Why Does Planning Exist?
The most expensive bugs are requirements bugs—building the wrong thing entirely. By investing time upfront in clarifying what you want, ZERG avoids the frustration of workers building features you didn't actually need. The output is a requirements.md file that workers will read as their mission brief.
Planning Flow
User Requirements -> [Socratic Discovery] -> requirements.md
|
v
.gsd/specs/{feature}/requirements.md
Design Phase (/zerg:design)
What Is Design?
Design is where ZERG transforms "what to build" into "how to build it." The design phase analyzes your requirements, proposes an architecture, and breaks the work into atomic tasks that can be executed in parallel without stepping on each other.
Why Does Design Exist?
Parallel execution requires careful coordination. If two workers both try to modify the same file, you get merge conflicts. If a worker starts building a feature that depends on code that doesn't exist yet, it fails. Design solves both problems by assigning exclusive file ownership and organizing tasks into dependency levels.
Design Flow
requirements.md -> [Architecture Analysis] -> task-graph.json + design.md
|
+---------------------------+---------------------------+
v v v
Level 1 Tasks Level 2 Tasks Level N Tasks
Rush Phase (/zerg:rush)
What Is Rush?
Rush is the execution phase—when workers actually write code. The orchestrator spawns multiple Claude Code instances (zerglings), each in its own isolated workspace, and coordinates their work through level-based execution.
Why Does Rush Exist This Way?
Traditional development is sequential: one developer finishes, then another starts. ZERG flips this by running workers in parallel wherever possible. But parallelism needs coordination. Rush implements that coordination: spawning workers, assigning tasks, monitoring progress, merging results, and running quality gates.
The level-based approach ensures dependencies are respected. All Level 1 tasks (foundations like types and schemas) complete and merge before any Level 2 tasks (business logic that uses those types) can start.
Rush Flow
[Orchestrator Start]
|
v
[Load task-graph.json] -> [Assign tasks to zerglings]
|
v
[Create git worktrees]
|
v
[Spawn N zergling processes]
|
v
+---------------------------------------------------------------------+
| FOR EACH LEVEL: |
| 1. Zerglings execute tasks in PARALLEL |
| 2. Poll until all level tasks complete |
| 3. MERGE PROTOCOL: |
| - Merge all zergling branches -> staging |
| - Run quality gates |
| - Promote staging -> main |
| 4. Rebase zergling branches |
| 5. Advance to next level |
+---------------------------------------------------------------------+
|
v
[All tasks complete]
How the Phases Connect
The phases form a pipeline: planning feeds design, design feeds rush. But it's not just data flow—it's also quality gates between each transition. You must approve the requirements before design starts. You must approve the task graph before rush starts. This prevents expensive mistakes from propagating through the pipeline.
Zergling Protocol
Each zergling:
- Loads
requirements.md,design.md,task-graph.json - Reads
worker-assignments.jsonfor its tasks - For each level:
- Pick next assigned task at current level
- Read all dependency files
- Implement the task
- Run verification command
- On pass: commit, mark complete
- On fail: retry 3x, then mark blocked
- After level complete: wait for merge signal
- Pull merged changes
- Continue to next level
- At 70% context: commit WIP, exit (orchestrator restarts)
Module Reference
ZERG is composed of 80+ Python modules organized into functional groups.
Core Modules (zerg/)
| Module | Purpose |
|---|---|
orchestrator.py |
Fleet management, level transitions, merge triggers |
levels.py |
Level-based execution control, dependency enforcement |
state.py |
Thread-safe file-based state persistence |
worker_protocol.py |
Zergling-side execution, Claude Code invocation |
launcher.py |
Abstract worker spawning (subprocess/container) |
launcher_configurator.py |
Launcher mode detection and configuration |
worker_manager.py |
Worker lifecycle management, health tracking |
level_coordinator.py |
Cross-level coordination and synchronization |
Task Management
| Module | Purpose |
|---|---|
assign.py |
Task-to-zergling assignment with load balancing |
parser.py |
Parse and validate task graphs |
verify.py |
Execute task verification commands |
task_sync.py |
ClaudeTask model, TaskSyncBridge (JSON state to Claude Tasks) |
task_retry_manager.py |
Retry policy and management for failed tasks |
backlog.py |
Backlog generation and tracking |
dependency_checker.py |
Validate task dependencies before claiming |
graph_validation.py |
Task graph structure validation |
Resilience & Flow Control
| Module | Purpose |
|---|---|
backpressure.py |
Load shedding and flow control under pressure |
circuit_breaker.py |
Circuit breaker pattern for failing operations |
retry_backoff.py |
Exponential backoff for retry strategies |
risk_scoring.py |
Risk assessment for task and merge operations |
whatif.py |
What-if analysis for execution planning |
preflight.py |
Pre-execution validation checks |
heartbeat.py |
Worker heartbeat monitoring for stall detection |
escalation.py |
Worker escalation and human-in-the-loop requests |
progress_reporter.py |
Real-time progress reporting from workers |
Cross-Cutting Capabilities
| Module | Purpose |
|---|---|
capability_resolver.py |
Resolve CLI flags + config into ResolvedCapabilities |
depth_tiers.py |
Analysis depth tiers (quick/standard/think/think-hard/ultrathink) |
modes.py |
Behavioral modes (precision/speed/exploration/refactor/debug) |
loops.py |
Iterative improvement loop controller |
efficiency.py |
Token efficiency zones and compact formatting |
mcp_router.py |
MCP server auto-routing based on task signals |
mcp_telemetry.py |
MCP routing telemetry and analytics |
tdd.py |
TDD enforcement and test-first workflow |
verification_tiers.py |
Verification gate tier configuration |
verification_gates.py |
Per-level verification gate runner |
Git & Merge
| Module | Purpose |
|---|---|
git_ops.py |
Low-level git operations |
worktree.py |
Git worktree management for zergling isolation |
merge.py |
Branch merging after each level |
Quality & Security
| Module | Purpose |
|---|---|
gates.py |
Execute quality gates (lint, typecheck, test) |
security.py |
Security validation, hook patterns |
validation.py |
Task graph and ID validation |
command_executor.py |
Safe command execution with argument parsing |
Configuration & Types
| Module | Purpose |
|---|---|
config.py |
Pydantic configuration management |
constants.py |
Enumerations (TaskStatus, WorkerStatus, GateResult) |
types.py |
TypedDict and dataclass definitions |
schemas/ |
JSON schema definitions |
Plugin System
| Module | Purpose |
|---|---|
plugins.py |
Plugin ABCs (QualityGatePlugin, LifecycleHookPlugin, LauncherPlugin, ContextPlugin), PluginRegistry |
plugin_config.py |
Pydantic models for plugin YAML configuration |
context_plugin.py |
ContextEngineeringPlugin for token-budgeted task context |
command_splitter.py |
Split large command files into .core.md + .details.md |
Logging & Metrics
| Module | Purpose |
|---|---|
log_writer.py |
StructuredLogWriter (per-worker JSONL), TaskArtifactCapture |
log_aggregator.py |
Read-side aggregation, time-sorted queries across workers |
logging.py |
Logging setup, Python logging bridge, LogPhase/LogEvent enums |
metrics.py |
Duration, percentile calculations, metric type definitions |
worker_metrics.py |
Per-task execution metrics (timing, context usage, retries) |
render_utils.py |
Output formatting and display utilities |
status_formatter.py |
Format status output for CLI display |
event_emitter.py |
Event streaming for real-time status updates |
token_tracker.py |
Track token usage across sessions |
token_counter.py |
Estimate token counts for text |
token_aggregator.py |
Aggregate token usage statistics |
Container Management
| Module | Purpose |
|---|---|
containers.py |
ContainerManager, ContainerInfo for Docker lifecycle |
Context & Execution
| Module | Purpose |
|---|---|
context_tracker.py |
Heuristic token counting, checkpoint decisions |
spec_loader.py |
Load and truncate GSD specs (requirements.md, design.md) |
dryrun.py |
Dry-run simulation for /zerg:rush --dry-run |
worker_main.py |
Worker process entry point |
ports.py |
Port allocation for worker processes (range 49152-65535) |
exceptions.py |
Exception hierarchy (ZergError -> Task/Worker/Git/Gate errors) |
state_sync_service.py |
State synchronization across distributed workers |
state_reconciler.py |
Reconcile state conflicts between workers |
adaptive_detail.py |
Adaptive detail levels based on context pressure |
step_generator.py |
Generate execution steps from task specs |
step_executor.py |
Execute generated steps |
claude_tasks_reader.py |
Read Claude Code Task system state |
Project Initialization
| Module | Purpose |
|---|---|
charter.py |
Project charter generation |
inception.py |
Inception mode (empty directory -> project scaffold) |
tech_selector.py |
Technology stack recommendation |
devcontainer_features.py |
Devcontainer feature configuration |
security_rules.py |
Security rules fetching from TikiTribe |
architecture.py |
Project architecture analysis |
architecture_gate.py |
Architecture compliance quality gate |
repo_map.py |
Generate repository structure maps |
repo_map_js.py |
JavaScript-specific repo mapping |
ast_analyzer.py |
AST analysis for code understanding |
ast_cache.py |
Cache AST analysis results |
test_scope.py |
Determine test scope for changes |
formatter_detector.py |
Detect code formatters in project |
Diagnostics (zerg/diagnostics/)
| Module | Purpose |
|---|---|
error_intel.py |
Multi-language error parsing, fingerprinting, chain analysis |
hypothesis_engine.py |
Bayesian hypothesis testing with prior/posterior scoring |
knowledge_base.py |
30+ known failure patterns with calibrated probabilities |
log_correlator.py |
Cross-worker log correlation, temporal clustering |
log_analyzer.py |
Log pattern analysis and trend detection |
code_fixer.py |
Code-aware fix suggestions, import chain analysis |
recovery.py |
Recovery plan generation with risk-rated steps |
env_diagnostics.py |
Environment checks (Python, Docker, resources, config) |
state_introspector.py |
Deep state file analysis and corruption detection |
system_diagnostics.py |
System-level checks (disk, ports, worktrees) |
types.py |
Diagnostic type definitions |
Performance Analysis (zerg/performance/)
| Module | Purpose |
|---|---|
stack_detector.py |
Auto-detect project language and framework |
tool_registry.py |
Registry of available analysis tools |
catalog.py |
Tool catalog with install instructions |
aggregator.py |
Combine results from multiple analysis tools |
formatters.py |
Format analysis output (markdown, SARIF, JSON) |
types.py |
Performance analysis type definitions |
Tool Adapters (zerg/performance/adapters/)
| Adapter | Tool | Analysis Type |
|---|---|---|
radon_adapter.py |
radon | Cyclomatic complexity |
lizard_adapter.py |
lizard | Multi-language complexity |
vulture_adapter.py |
vulture | Dead code detection |
jscpd_adapter.py |
jscpd | Copy-paste detection |
cloc_adapter.py |
cloc | Lines of code counting |
semgrep_adapter.py |
semgrep | Semantic code analysis |
trivy_adapter.py |
trivy | Vulnerability scanning |
deptry_adapter.py |
deptry | Dependency health |
pipdeptree_adapter.py |
pipdeptree | Dependency tree analysis |
hadolint_adapter.py |
hadolint | Dockerfile linting |
dive_adapter.py |
dive | Docker image analysis |
CLI
| Module | Purpose |
|---|---|
cli.py |
CLI entry point (zerg command), install/uninstall subcommands |
__main__.py |
Package entry point for python -m zerg |
CLI Commands (zerg/commands/)
Commands are implemented as Python modules in zerg/commands/ and/or as markdown spec files in zerg/data/commands/. Some commands (marked "spec only") have no Python implementation and are interpreted directly by Claude Code.
| Command | Module | Purpose |
|---|---|---|
/zerg:init |
init.py |
Project initialization (Inception/Discovery modes) |
/zerg:brainstorm |
(spec only) | Feature discovery and GitHub issue creation |
/zerg:plan |
plan.py |
Capture requirements (Socratic discovery) |
/zerg:design |
design.py |
Generate architecture and task graph |
/zerg:rush |
rush.py |
Launch parallel zerglings |
/zerg:status |
status.py |
Progress monitoring dashboard |
/zerg:stop |
stop.py |
Stop zerglings (graceful/force) |
/zerg:retry |
retry.py |
Retry failed tasks |
/zerg:logs |
logs.py |
View and aggregate zergling logs |
/zerg:merge |
merge_cmd.py |
Manual merge control |
/zerg:cleanup |
cleanup.py |
Remove artifacts |
/zerg:debug |
debug.py |
Deep diagnostic investigation |
/zerg:build |
build.py |
Build orchestration with error recovery |
/zerg:test |
test_cmd.py |
Test execution with coverage |
/zerg:analyze |
analyze.py |
Static analysis and metrics |
/zerg:review |
review.py |
Code review (spec compliance + quality) |
/zerg:security |
security_rules_cmd.py |
Vulnerability scanning |
/zerg:refactor |
refactor.py |
Automated code improvement |
/zerg:git |
git_cmd.py |
Intelligent git operations |
/zerg:plugins |
(spec only) | Plugin system management |
/zerg:document |
document.py |
Documentation generation for components |
/zerg:estimate |
(spec only) | Effort estimation with PERT intervals |
/zerg:explain |
(spec only) | Educational code explanations |
/zerg:index |
(spec only) | Project documentation wiki generation |
/zerg:select-tool |
(spec only) | Intelligent tool routing |
/zerg:worker |
worker_protocol.py |
Zergling execution protocol |
/zerg:wiki |
wiki.py |
Generate wiki documentation pages |
install_commands.py |
Install/uninstall slash commands | |
loop_mixin.py |
Shared improvement loop mixin for commands |
Zergling Execution Model
What Is a Zergling?
A zergling is a single Claude Code worker instance—an AI that reads the spec files and writes code to complete its assigned tasks. The name comes from the idea of overwhelming a feature with many small, focused workers rather than one monolithic process.
Each zergling operates independently. It doesn't share memory with other zerglings, doesn't have access to their conversation history, and works in its own isolated directory. This independence is a feature, not a limitation—it makes the system robust against crashes and easy to scale.
Why Independent Execution?
If zerglings shared state, a crash in one could corrupt data for all others. Shared conversation history would mean workers waiting for context windows to sync. Shared file systems would mean merge conflicts constantly.
By keeping zerglings fully isolated, ZERG achieves: crash recovery (just restart the failed worker), horizontal scaling (add more workers without coordination overhead), and deterministic behavior (same inputs always produce same outputs).
Isolation Strategy
+---------------------------------------------------------------------+
| ZERGLING ISOLATION LAYERS |
+---------------------------------------------------------------------+
| 1. Git Worktree: .zerg-worktrees/{feature}-worker-{id}/ |
| - Independent file system |
| - Separate git history |
| - Own branch: zerg/{feature}/worker-{id} |
+---------------------------------------------------------------------+
| 2. Process Isolation |
| - Separate process per zergling |
| - Independent memory space |
| - Communication via state files |
+---------------------------------------------------------------------+
| 3. Spec-Driven Execution |
| - No conversation history sharing |
| - Read specs fresh each time |
| - Stateless, restartable |
+---------------------------------------------------------------------+
How Isolation Layers Work Together
Layer 1 (Git Worktree) gives each worker its own copy of the codebase. Git worktrees are a built-in feature that creates independent working directories sharing the same repository. Worker 0 can edit files in its worktree while Worker 1 edits different files in its worktree—no conflicts possible.
Layer 2 (Process Isolation) means each worker runs as a separate OS process. If Worker 2 crashes, Workers 0, 1, and 3 keep running. Communication happens through files on disk (the state JSON), not shared memory.
Layer 3 (Spec-Driven Execution) is the key to restartability. Workers don't remember previous conversations—they read the spec files fresh every time. If a worker crashes mid-task, just restart it. It will read the specs, see the task is incomplete, and pick up where it left off.
Launcher Abstraction
What Is a Launcher?
A launcher is the mechanism that starts and manages worker processes. ZERG supports multiple execution environments: local processes (subprocess), Docker containers, or plugin-provided environments like Kubernetes. The launcher abstraction hides these differences so the orchestrator doesn't need to know how workers run—just that they do.
Why Abstract the Launcher?
Different environments have different requirements. During development, you might run workers as local processes for fast iteration. In CI, you might run them in containers for reproducibility. In production, you might use Kubernetes for scaling.
Rather than hard-coding one approach, ZERG lets you swap launchers. Your orchestration logic stays the same; only the "how workers run" changes.
Launcher Architecture
WorkerLauncher (ABC)
+-- SubprocessLauncher
| +-- spawn() -> subprocess.Popen
| +-- monitor() -> Check process status
| +-- terminate() -> Kill process
|
+-- ContainerLauncher
| +-- spawn() -> docker run
| +-- monitor() -> Check container status
| +-- terminate() -> Stop/kill container
|
+-- LauncherConfigurator
+-- detect_mode() -> auto-select launcher
+-- configure() -> build launcher with settings
How Launchers Connect
The LauncherConfigurator decides which launcher to use based on your environment and CLI flags. If you have Docker and a devcontainer, it picks ContainerLauncher. Otherwise, it defaults to SubprocessLauncher. You can override this with --mode subprocess or --mode container.
Each launcher implements the same interface: spawn(), monitor(), terminate(). The orchestrator calls these methods without knowing whether it's managing processes or containers underneath.
Execution Modes
| Mode | Launcher Class | How Workers Run |
|---|---|---|
subprocess |
SubprocessLauncher |
Local processes running zerg.worker_main |
container |
ContainerLauncher |
Docker containers with mounted worktrees |
task |
Plugin-provided | Claude Code Task sub-agents (slash command context) |
Auto-detection logic (in launcher_configurator.py):
- If
--modeis explicitly set -> use that mode - If
.devcontainer/devcontainer.jsonexists AND Docker is available ->container - If running inside a Claude Code slash command context ->
task - Otherwise ->
subprocess
Plugin launchers are resolved via get_plugin_launcher(name, registry) which delegates to a LauncherPlugin.create_launcher() call.
Context Management
- Monitor token usage via
ContextTracker - Checkpoint at 70% context threshold (configurable)
- Zergling exits gracefully (code 2)
- Orchestrator restarts zergling from checkpoint
Cross-Cutting Capabilities
What Are Cross-Cutting Capabilities?
Cross-cutting capabilities are behaviors that affect the entire system, not just one module. Think of them as dials and switches that tune how ZERG operates. Want workers to think more deeply? Turn up the analysis depth. Need faster iteration? Enable speed mode. Want test-driven development? Flip the TDD switch.
These capabilities "cut across" all phases and all workers—hence the name. A capability like "compact output" affects planning, design, and rush equally.
Why Do Cross-Cutting Capabilities Exist?
Different situations call for different approaches. A quick prototype needs speed, not exhaustive analysis. A critical production feature needs deep thinking and full verification. Rather than building separate tools for each scenario, ZERG provides one tool with tunable behavior.
This also enables progressive enhancement. Start with defaults, then turn up verification as you approach release. The same command works for both; only the capability settings change.
Capability Architecture
ZERG includes 8 cross-cutting capabilities that influence worker behavior across all phases:
┌─────────────────────────────────────────────────────────────────────┐
│ Cross-Cutting Capability Flow │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ CLI Flags ──┬──► CapabilityResolver ──► ResolvedCapabilities │
│ │ ▲ │ │
│ Config ─────┤ │ ▼ │
│ │ Task Graph Worker Env Vars │
│ Task ───────┴─── Analysis (ZERG_DEPTH, etc.) │
│ │ │
│ ▼ │
│ ContextEngineeringPlugin │
│ │ │
│ ▼ │
│ Task-Scoped Context │
└─────────────────────────────────────────────────────────────────────┘
How Capabilities Flow
CLI flags and configuration merge in the CapabilityResolver, which produces a ResolvedCapabilities object. This gets translated into environment variables that workers inherit. Workers then use these settings to adjust their behavior: how deeply they analyze, which MCP servers they enable, whether they run improvement loops, etc.
The ContextEngineeringPlugin uses these resolved capabilities to build task-scoped context. A task running with --ultrathink gets more context budget than one running with --quick.
Capability Matrix
| Capability | CLI Flag | Config Section | Module |
|---|---|---|---|
| Analysis Depth | --quick/--think/--think-hard/--ultrathink |
— | depth_tiers.py |
| Token Efficiency | --no-compact (ON by default) |
efficiency |
efficiency.py |
| Behavioral Modes | --mode |
behavioral_modes |
modes.py |
| MCP Auto-Routing | --mcp/--no-mcp |
mcp_routing |
mcp_router.py |
| Engineering Rules | — | rules |
(config-driven) |
| Improvement Loops | --no-loop/--iterations (ON by default) |
improvement_loops |
loops.py |
| Verification Gates | — | verification |
verification_tiers.py |
| TDD Enforcement | --tdd |
tdd |
tdd.py |
Analysis Depth Tiers
| Tier | Token Budget | MCP Servers | Use Case |
|---|---|---|---|
quick |
~1,000 | None | Fast, surface-level analysis |
standard |
~2,000 | None | Balanced default |
think |
~4,000 | sequential | Structured multi-step analysis |
think-hard |
~10,000 | sequential, context7 | Deep architectural analysis |
ultrathink |
~32,000 | sequential, context7, playwright, morphllm | Maximum depth |
Behavioral Modes
| Mode | Description | Verification Level |
|---|---|---|
precision |
Careful, thorough execution | Full |
speed |
Fast iteration, minimal overhead | Minimal |
exploration |
Broad discovery and analysis | None |
refactor |
Code transformation focus | Full |
debug |
Diagnostic, verbose logging | Verbose |
Improvement Loops
What Is an Improvement Loop?
An improvement loop is an automated "make it better" cycle. After a worker completes a task, the loop runs quality gates (tests, linting, type checks), measures the quality score, and if the score isn't good enough, asks the worker to improve the code. This repeats until quality converges or the loop detects it's not making progress.
Why Do Improvement Loops Exist?
First drafts are rarely perfect. A human developer naturally iterates: write code, run tests, fix issues, repeat. Improvement loops automate this cycle. Instead of hoping the first attempt is good enough, ZERG keeps refining until quality gates pass.
Loops also detect dead ends. If the quality score stops improving (plateau) or gets worse (regression), the loop stops rather than wasting time on fruitless iteration.
Loop Architecture
The LoopController enables iterative refinement:
Initial Score ──► Run Gates ──► Check Convergence ──► Iterate or Stop
│ │
│ ┌───────┴───────┐
│ ▼ ▼
│ Converged Plateau/Regressed
│ (score ≥ 0.99) (no improvement)
│ │ │
▼ ▼ ▼
Improve Code Complete Complete
Loop status values: RUNNING, CONVERGED, PLATEAU, REGRESSED, MAX_ITERATIONS, ABORTED
How Loops Connect to Workers
Improvement loops wrap worker execution. When loops are enabled (default), each task execution enters the loop. The loop runs quality gates, records the score, and decides: improve or complete? This happens transparently—workers don't need special loop-awareness.
Resilience
What Is Resilience?
Resilience is the system's ability to handle failures gracefully. In a distributed system with multiple workers, things will go wrong: API rate limits hit, workers crash, verification commands time out. Resilient systems anticipate these failures and have strategies to recover automatically.
Why Does Resilience Matter?
Without resilience, a single failure can cascade. One worker fails, retries endlessly, consumes all the API quota, and causes all other workers to fail too. Soon you have ten workers all failing, retrying, and making things worse.
ZERG's resilience mechanisms—circuit breakers, backpressure, and intelligent retry—prevent these cascades. Failures get contained, the system adapts, and progress continues even when some components struggle.
ZERG includes comprehensive resilience mechanisms for fault tolerance:
Circuit Breaker Pattern
What Is a Circuit Breaker?
A circuit breaker stops calling a failing service before it can cause more damage. Just like an electrical circuit breaker trips to prevent fires, a software circuit breaker "trips" when too many failures occur, giving the failing component time to recover.
Why Do Circuit Breakers Exist?
When an external service (like the Claude API) has a temporary outage, repeatedly retrying makes things worse. You waste time, consume quota, and potentially trigger rate limits. Circuit breakers recognize "this is broken" and stop trying for a cooldown period, then test with a single request to see if recovery happened.
Circuit Breaker Flow
Closed ──► Failures exceed threshold ──► Open
▲ │
│ │
└── Cooldown expires, test succeeds ◄───┘
│
Half-Open
In the Closed state, requests flow normally. After enough failures hit the threshold, the breaker Opens—all requests fail fast without even trying. After a cooldown period, it goes Half-Open and allows one test request. If that succeeds, back to Closed. If it fails, back to Open.
Configuration (.zerg/config.yaml):
error_recovery:
circuit_breaker:
enabled: true
failure_threshold: 5
cooldown_seco
…(truncated)