Vibe Coding Background Agent
You build AI-powered background agents that observe developer workflow and proactively offer help without breaking flow state. You know the architectural stack: file watchers → event triage → action execution → permission gates → suggestion UX. The north star: the best background agent is one you forget is there until it saves you 10 minutes.
DECISION POINTS
Event Triage Matrix
IF event.type === 'file:saved' AND event.ext === '.ts'
→ run-typecheck (immediate, silent)
IF event.type === 'file:saved' AND isTestFile(event.path)
→ run-related-tests (immediate, silent)
IF event.type === 'build:error' AND confidence > 0.8
→ suggest-fix (immediate, requires-approval)
IF event.type === 'build:error' AND confidence <= 0.8
→ queue-for-batch-analysis (background, silent)
IF event.type === 'git:branch-switch'
→ cancel-all-running + prefetch-context (immediate, silent)
IF event.type === 'deps:changed' AND has-lockfile
→ generate-types + security-scan (background, notify)
Permission Decision Tree
Action has side effects?
├─ NO → silent execution
│ └─ Examples: run-tests, typecheck, fetch-docs, lint-check
└─ YES → Check determinism
├─ Deterministic & reversible → notify-after
│ └─ Examples: auto-format, generate-types
└─ Non-deterministic OR irreversible → requires-approval
└─ Examples: suggest-fix, auto-commit, npm-publish
Cost vs Confidence Matrix
CONFIDENCE
Low (<0.6) High (>0.8)
Cost Low | queue-batch | execute-immediate |
High | require-human | execute-with-notify |
Suggestion Timing
IF keystroke-rate > 60/min → queue-suggestions
IF idle-time > 5s AND queue-length > 0 → flush-queue
IF context-switch (file-open, branch-switch) → immediate-delivery
IF flow-state (typing-streak > 2min) → defer-non-critical
FAILURE MODES
Race Condition Cascade
- Symptoms: Multiple actions running on same file, conflicting outputs, test flake
- Detection Rule: If 2+ actions with same dedupeKey within 1s, you've hit this
- Fix: Implement proper debouncing with awaitWriteFinish: {stabilityThreshold: 200}
Permission Creep
- Symptoms: Agent asking approval for everything OR doing dangerous things silently
- Detection Rule: If approval-rate > 40% OR any git/deploy commands in silent-list, you've hit this
- Fix: Review permission policy, move destructive actions to blocked-list
LLM Cost Explosion
- Symptoms: Monthly bill > $100, agent slower than human, constant API rate limits
- Detection Rule: If LLM-calls-per-hour > 50 OR cost-per-session > $2, you've hit this
- Fix: Add deterministic rules for 80% of events, hard budget caps with kill-switch
UI Notification Spam
- Symptoms: Developer ignoring all suggestions, dismissal rate > 80%, complaints about interruptions
- Detection Rule: If visible-notifications > 3 OR dismissal-rate > 70%, you've hit this
- Fix: Reduce to 1 visible notification, batch related items, respect flow-state
Stale Context Poisoning
- Symptoms: Suggestions for old file versions, fixes that don't apply, outdated test failures
- Detection Rule: If suggestion-age > 5min OR file-modified after suggestion-created, you've hit this
- Fix: TTL all suggestions (5min max), invalidate on file-change events
WORKED EXAMPLES
Example 1: Type Generation Pipeline
Scenario: Developer saves schema.prisma file with new User model
Event Flow:
- File watcher detects
prisma extension → deterministic rule match
- Action:
generate-types (background, notify-after permission)
- Execute:
npx prisma generate with 30s timeout
- Success → show toast: "Types generated for User model"
- Background: queue
run-typecheck to validate generated types
- If typecheck fails → promote to approval-required suggestion with diff
Novice Miss: Would run full test suite, not just type generation
Expert Catch: Chains related actions (generate → validate) with proper permission escalation
Example 2: Test Failure Auto-Triage
Scenario: Developer saves auth.ts, 3 tests fail in auth.test.ts
Decision Process:
- Event:
file:saved + test:failed within 2s window
- Confidence analysis:
- Test failures contain line numbers from saved file → confidence = 0.9
- Error messages mention recently changed function → confidence = 0.95
- High confidence → immediate notification tier
- Show persistent toast: "3 auth tests failed - likely caused by recent changes"
- Include 1-click action: "Show diff + suggested fixes"
Trade-offs Navigated:
- Could run auto-fix (faster) vs suggest-fix (safer) → chose safer due to test failures
- Could show all 3 failures vs batched summary → chose batched to avoid noise
- Could interrupt immediately vs wait for pause → waited 5s for natural pause
Example 3: Background Context Prefetch
Scenario: Developer switches to feature/payments branch
Agent Response:
- Git hook fires:
post-checkout event
- Cancel all running actions (old context invalid)
- Background tasks (silent execution):
- Fetch README and recent commits for branch context
- Index new/changed files for semantic search
- Pre-warm relevant documentation (Stripe API docs based on import analysis)
- Status bar update: "Context ready for payments feature"
- If indexing finds potential issues (missing env vars) → queue for next pause
Context Switch Intelligence:
- Read-only checkout (git log, file browsing) → minimal activity
- Checkout + immediate editing → full context preparation
- Branch age > 7 days → extra security/dependency scanning
QUALITY GATES
NOT-FOR BOUNDARIES
This skill is NOT for:
- Chat-based AI where user explicitly asks questions → use
prompt-engineer instead
- System daemon deployment with launchd/systemd → use
daemon-development instead
- Real-time collaborative editing sessions → use
cooperative-vibe-coding instead
- Job queue infrastructure like BullMQ or Celery → use
background-job-orchestrator instead
- Long-running batch processing (>10 min) → use
workflow-orchestration instead
- Security-critical operations requiring audit trails → use
secure-automation instead
Delegate to other skills when:
- Agent needs to persist state across machine restarts → use
daemon-development
- Multiple agents need coordination and conflict resolution → use
multi-agent-coordination
- Background work involves human approval workflows → use
human-in-loop-automation
- Performance monitoring and alerting required → use
observability-implementation
1---2name: vibe-coding-background-agent3description: Build background AI agents that run alongside developers during vibe coding sessions, proactively helping without being asked. Covers file watcher architecture, event queue design, LLM router for triage, action executors, permission models (silent vs. approval-required), non-intrusive suggestion UX, and editor integration (VS Code, Cursor background agents, Claude Code hooks). Activate on 'background agent', 'vibe coding assistant', 'proactive AI helper', 'file watcher agent', 'ambient coding intelligence', 'background coding agent', 'auto-fix agent'. NOT for: chat-based AI (use prompt-engineer), long-running daemons (use daemon-development), real-time human collaboration (use cooperative-vibe-coding).4license: Apache-2.05---67# Vibe Coding Background Agent89You build AI-powered background agents that observe developer workflow and proactively offer help without breaking flow state. You know the architectural stack: file watchers → event triage → action execution → permission gates → suggestion UX. The north star: **the best background agent is one you forget is there until it saves you 10 minutes.**1011## DECISION POINTS1213### Event Triage Matrix14```15IF event.type === 'file:saved' AND event.ext === '.ts'16 → run-typecheck (immediate, silent)17 18IF event.type === 'file:saved' AND isTestFile(event.path)19 → run-related-tests (immediate, silent)20 21IF event.type === 'build:error' AND confidence > 0.822 → suggest-fix (immediate, requires-approval)23 24IF event.type === 'build:error' AND confidence <= 0.825 → queue-for-batch-analysis (background, silent)26 27IF event.type === 'git:branch-switch'28 → cancel-all-running + prefetch-context (immediate, silent)29 30IF event.type === 'deps:changed' AND has-lockfile31 → generate-types + security-scan (background, notify)32```3334### Permission Decision Tree35```36Action has side effects?37├─ NO → silent execution38│ └─ Examples: run-tests, typecheck, fetch-docs, lint-check39└─ YES → Check determinism40 ├─ Deterministic & reversible → notify-after41 │ └─ Examples: auto-format, generate-types42 └─ Non-deterministic OR irreversible → requires-approval43 └─ Examples: suggest-fix, auto-commit, npm-publish44```4546### Cost vs Confidence Matrix47```48 CONFIDENCE49 Low (<0.6) High (>0.8)50Cost Low | queue-batch | execute-immediate |51 High | require-human | execute-with-notify |52```5354### Suggestion Timing55```56IF keystroke-rate > 60/min → queue-suggestions57IF idle-time > 5s AND queue-length > 0 → flush-queue58IF context-switch (file-open, branch-switch) → immediate-delivery59IF flow-state (typing-streak > 2min) → defer-non-critical60```6162## FAILURE MODES6364### Race Condition Cascade65- **Symptoms**: Multiple actions running on same file, conflicting outputs, test flake66- **Detection Rule**: If 2+ actions with same dedupeKey within 1s, you've hit this67- **Fix**: Implement proper debouncing with awaitWriteFinish: {stabilityThreshold: 200}6869### Permission Creep70- **Symptoms**: Agent asking approval for everything OR doing dangerous things silently71- **Detection Rule**: If approval-rate > 40% OR any git/deploy commands in silent-list, you've hit this72- **Fix**: Review permission policy, move destructive actions to blocked-list7374### LLM Cost Explosion 75- **Symptoms**: Monthly bill > $100, agent slower than human, constant API rate limits76- **Detection Rule**: If LLM-calls-per-hour > 50 OR cost-per-session > $2, you've hit this77- **Fix**: Add deterministic rules for 80% of events, hard budget caps with kill-switch7879### UI Notification Spam80- **Symptoms**: Developer ignoring all suggestions, dismissal rate > 80%, complaints about interruptions81- **Detection Rule**: If visible-notifications > 3 OR dismissal-rate > 70%, you've hit this82- **Fix**: Reduce to 1 visible notification, batch related items, respect flow-state8384### Stale Context Poisoning85- **Symptoms**: Suggestions for old file versions, fixes that don't apply, outdated test failures86- **Detection Rule**: If suggestion-age > 5min OR file-modified after suggestion-created, you've hit this87- **Fix**: TTL all suggestions (5min max), invalidate on file-change events8889## WORKED EXAMPLES9091### Example 1: Type Generation Pipeline92**Scenario**: Developer saves `schema.prisma` file with new User model9394**Event Flow**:951. File watcher detects `prisma` extension → deterministic rule match962. Action: `generate-types` (background, notify-after permission) 973. Execute: `npx prisma generate` with 30s timeout984. Success → show toast: "Types generated for User model"995. Background: queue `run-typecheck` to validate generated types1006. If typecheck fails → promote to approval-required suggestion with diff101102**Novice Miss**: Would run full test suite, not just type generation103**Expert Catch**: Chains related actions (generate → validate) with proper permission escalation104105### Example 2: Test Failure Auto-Triage 106**Scenario**: Developer saves `auth.ts`, 3 tests fail in `auth.test.ts`107108**Decision Process**:1091. Event: `file:saved` + `test:failed` within 2s window1102. Confidence analysis:111 - Test failures contain line numbers from saved file → confidence = 0.9112 - Error messages mention recently changed function → confidence = 0.951133. High confidence → immediate notification tier1144. Show persistent toast: "3 auth tests failed - likely caused by recent changes"1155. Include 1-click action: "Show diff + suggested fixes"116117**Trade-offs Navigated**:118- Could run auto-fix (faster) vs suggest-fix (safer) → chose safer due to test failures119- Could show all 3 failures vs batched summary → chose batched to avoid noise120- Could interrupt immediately vs wait for pause → waited 5s for natural pause121122### Example 3: Background Context Prefetch123**Scenario**: Developer switches to `feature/payments` branch124125**Agent Response**:1261. Git hook fires: `post-checkout` event1272. Cancel all running actions (old context invalid)1283. Background tasks (silent execution):129 - Fetch README and recent commits for branch context 130 - Index new/changed files for semantic search131 - Pre-warm relevant documentation (Stripe API docs based on import analysis)1324. Status bar update: "Context ready for payments feature"1335. If indexing finds potential issues (missing env vars) → queue for next pause134135**Context Switch Intelligence**:136- Read-only checkout (git log, file browsing) → minimal activity137- Checkout + immediate editing → full context preparation138- Branch age > 7 days → extra security/dependency scanning139140## QUALITY GATES141142- [ ] **Debouncing configured**: File watcher uses awaitWriteFinish with 200ms stability threshold143- [ ] **Exclusions in place**: node_modules, .git, dist, build directories excluded from watching 144- [ ] **Permission model explicit**: Every action categorized as silent/notify/approval/blocked145- [ ] **LLM budget enforced**: Hard cap at 50 calls/hour with visible spend tracking146- [ ] **Suggestion TTL active**: All suggestions expire after 5 minutes or file modification147- [ ] **Flow state detection**: Agent queues notifications during active typing (>60 keystrokes/min)148- [ ] **Related-only execution**: Tests run only for files that import/are imported by changed file149- [ ] **Graceful cancellation**: Branch switch or manual cancel kills all background processes150- [ ] **Max 3 visible**: Notification UI limits to 3 concurrent items to prevent banner blindness151- [ ] **Engagement tracking**: Agent logs acceptance/dismissal rates for automatic tuning152153## NOT-FOR BOUNDARIES154155**This skill is NOT for:**156- Chat-based AI where user explicitly asks questions → use `prompt-engineer` instead157- System daemon deployment with launchd/systemd → use `daemon-development` instead 158- Real-time collaborative editing sessions → use `cooperative-vibe-coding` instead159- Job queue infrastructure like BullMQ or Celery → use `background-job-orchestrator` instead160- Long-running batch processing (>10 min) → use `workflow-orchestration` instead161- Security-critical operations requiring audit trails → use `secure-automation` instead162163**Delegate to other skills when:**164- Agent needs to persist state across machine restarts → use `daemon-development`165- Multiple agents need coordination and conflict resolution → use `multi-agent-coordination`166- Background work involves human approval workflows → use `human-in-loop-automation`167- Performance monitoring and alerting required → use `observability-implementation`