Daemon Development
Build long-running background processes that start reliably, run continuously, recover from failures, and shut down gracefully. Expert-level daemon architecture across macOS launchd, Linux systemd, and AI-powered services.
Decision Points
1. Platform-Specific Init System Choice
Platform Detected?
├─ macOS
│ ├─ Must run at boot (no user login): LaunchDaemon → /Library/LaunchDaemons/
│ ├─ User session required (GUI/files): LaunchAgent → ~/Library/LaunchAgents/
│ └─ System-wide user service: LaunchAgent → /Library/LaunchAgents/
├─ Linux
│ ├─ systemd available: systemd unit file → /etc/systemd/system/
│ ├─ Legacy SysV: init.d script (rare, avoid if possible)
│ └─ Container: s6 or built-in supervision
└─ Cross-platform dev
├─ Node.js app: pm2 for development, systemd/launchd for production
└─ Other languages: Direct systemd/launchd implementation
2. Service Type Configuration
Daemon Startup Behavior?
├─ Simple process (doesn't fork)
│ ├─ systemd: Type=simple
│ └─ launchd: Standard plist (no special keys)
├─ Signals readiness when ready
│ ├─ systemd: Type=notify + sd_notify("READY=1")
│ └─ launchd: N/A (use health check instead)
├─ Forks child process (legacy)
│ ├─ systemd: Type=forking + PIDFile (avoid)
│ └─ launchd: Not supported (rewrite to not fork)
└─ Socket-activated
├─ systemd: Type=simple + [Socket] section
└─ launchd: Sockets dict in plist
3. Restart Policy Design
Failure Recovery Strategy?
├─ Critical service (must always run)
│ ├─ systemd: Restart=always, RestartSec=5
│ └─ launchd: KeepAlive=true, ThrottleInterval=10
├─ Crash recovery only
│ ├─ systemd: Restart=on-failure
│ └─ launchd: KeepAlive={SuccessfulExit=false}
├─ Manual restart preferred
│ ├─ systemd: Restart=no
│ └─ launchd: KeepAlive=false
└─ Rate-limited restart
├─ systemd: StartLimitBurst=5, StartLimitIntervalSec=60
└─ launchd: ThrottleInterval=30 (built-in)
4. AI Daemon Rate Limiting Strategy
LLM API Connection Pattern?
├─ Single provider, token bucket
│ ├─ Token estimation: prompt_tokens + max_completion_tokens
│ ├─ Bucket refill: tokens_per_minute from provider limits
│ └─ Overflow: Queue requests with priority
├─ Multi-provider failover
│ ├─ Circuit breaker per provider (3 failures = 30s timeout)
│ ├─ Rate limit per provider independently
│ └─ Failover order: primary → secondary → queue
├─ Streaming responses
│ ├─ Reserve tokens optimistically
│ ├─ Adjust on actual_tokens in real-time
│ └─ Handle mid-stream rate limits gracefully
└─ Batch processing
├─ Group similar requests to maximize throughput
└─ Split large batches if they hit rate limits
5. Graceful Shutdown Handling
SIGTERM Received?
├─ Web server daemon
│ ├─ 1. server.close() - stop accepting new connections
│ ├─ 2. Wait for active requests (timeout: TimeoutStopSec-5s)
│ ├─ 3. Close database connections
│ └─ 4. exit(0)
├─ Queue worker daemon
│ ├─ 1. Stop polling for new jobs
│ ├─ 2. Finish current job (timeout protection)
│ ├─ 3. Flush any pending state
│ └─ 4. exit(0)
├─ AI daemon
│ ├─ 1. Stop accepting new LLM requests
│ ├─ 2. Drain in-flight requests (respect provider timeouts)
│ ├─ 3. Save rate limit state to disk
│ └─ 4. Close provider connections, exit(0)
└─ Database/stateful daemon
├─ 1. Checkpoint/flush transactions
├─ 2. Close client connections gracefully
├─ 3. Release file locks
└─ 4. exit(0)
Failure Modes
1. Restart Loop Death Spiral
- Symptoms: High CPU usage, rapid log growth, service stuck in "activating" state
- Diagnosis: Daemon crashes immediately on startup, init system restarts too quickly
- Detection:
journalctl -u service shows start/crash/start pattern every few seconds
- Fix: Add
RestartSec=10 (systemd) or ThrottleInterval=15 (launchd), implement startup validation
2. Zombie Process Accumulation
- Symptoms:
ps aux shows <defunct> processes, parent daemon still running but degraded
- Diagnosis: Daemon spawns child processes but doesn't reap them (missing SIGCHLD handler)
- Detection:
ps axo pid,ppid,stat,comm | grep Z shows zombie children
- Fix: Install SIGCHLD handler with
wait() or waitpid(), use signal(SIGCHLD, SIG_IGN) if children are fire-and-forget
3. Connection Leak Cascade
- Symptoms: "Too many open files" errors, service degrades over time, works fine after restart
- Diagnosis: File descriptors or network connections not closed properly
- Detection:
lsof -p <daemon_pid> | wc -l grows continuously, eventual EMFILE errors
- Fix: Add connection pooling with max limits, implement proper cleanup in error paths, set
LimitNOFILE in systemd
4. Log Disk Explosion
- Symptoms: Disk space alerts, daemon crashes with "No space left on device"
- Diagnosis: No log rotation configured, daemon writes unbounded logs
- Detection: Log files in GB range, df shows /var/log at 100% capacity
- Fix: Configure logrotate/newsyslog, use systemd journald with
SystemMaxUse, add log level controls
5. Rate Limit Thrash (AI Daemons)
- Symptoms: High latency, request timeouts, alternating success/failure patterns
- Diagnosis: No backoff after rate limit hits, daemon hammers API repeatedly
- Detection: Provider API returns 429 errors, response times spike periodically
- Fix: Implement exponential backoff, parse Retry-After headers, circuit breaker per provider
Worked Examples
AI Daemon Case Study: LLM Rate Limit Management
Scenario: Building an AI daemon that processes user requests through OpenAI API, needs 99.9% uptime with graceful rate limit handling.
1. Initial Architecture Decision
# Decision: Multi-provider with circuit breakers
# Primary: OpenAI GPT-4, Secondary: Anthropic Claude, Tertiary: Local model
# systemd unit file choice
Type=notify # Daemon signals when fully initialized
Restart=on-failure
RestartSec=10
2. Rate Limiting Implementation
// Token bucket per provider (expert catches: different providers = different limits)
const rateLimiters = {
openai: new TokenBucket({ tokensPerMinute: 40000, burstCapacity: 8000 }),
anthropic: new TokenBucket({ tokensPerMinute: 25000, burstCapacity: 5000 }),
};
// Novice mistake: Request-based limiting
// Expert insight: LLM APIs are token-based, not request-based
async processRequest(req: LLMRequest) {
const estimatedTokens = this.estimateTokens(req.prompt, req.maxTokens);
await this.rateLimiters.openai.acquire(estimatedTokens);
// ... proceed with API call
}
3. Circuit Breaker Configuration
// Expert trade-off: Aggressive vs Conservative failover
const circuitBreaker = new CircuitBreaker({
failureThreshold: 3, // Conservative: 5+ for stable APIs, 3 for flaky ones
timeout: 30000, // Aggressive: 15s, Conservative: 60s
resetTimeout: 60000, // How long to wait before retry
});
// Decision point: When to fail over?
if (error.status === 429) {
// Rate limited: backoff on primary, don't fail over yet
await this.exponentialBackoff(provider, error.retryAfter);
} else if (error.status >= 500) {
// Server error: immediate failover to secondary
this.circuitBreaker.recordFailure('openai');
}
4. Graceful Shutdown Pattern
// Expert catches: Race conditions in shutdown
let shutdownInProgress = false;
process.on('SIGTERM', async () => {
if (shutdownInProgress) return; // Idempotent shutdown
shutdownInProgress = true;
console.log('SIGTERM received, draining connections...');
// 1. Stop accepting new requests
server.close();
// 2. Wait for in-flight requests (with timeout)
const drainTimeout = setTimeout(() => {
console.log('Drain timeout, force exit');
process.exit(1);
}, 25000); // systemd TimeoutStopSec=30, so exit by 25s
await Promise.all([
this.drainActiveRequests(),
this.flushRateLimitState(), // Save token bucket state to disk
]);
clearTimeout(drainTimeout);
process.exit(0);
});
5. Trade-offs and Decision Results
- Aggressive restart:
RestartSec=5 for quick recovery vs RestartSec=30 to avoid thrashing
- Circuit breaker: 3 failures for failover (catches transient issues) vs 5 failures (more stable)
- Drain timeout: 25s (safe margin) vs 29s (maximize request completion)
- Result: 99.95% uptime achieved, average failover time 2.3 seconds
Quality Gates
Not-For Boundaries
This skill is NOT for:
- Container orchestration: For Kubernetes deployments, Docker Swarm, or container-specific supervision → use
devops-automator
- One-shot scheduled tasks: For cron jobs that run and exit, periodic batch processing → use
task-scheduler
- Web application deployment: For nginx/Apache configuration, reverse proxy setup, SSL termination → use
backend-architect
- Queue worker frameworks: For Sidekiq, Celery, Bull queues with built-in supervision → use
background-job-orchestrator
- Development process management: For hot-reloading, file watching, development servers → use standard dev tooling (nodemon, cargo watch)
- Database administration: For MySQL/PostgreSQL service configuration → use
database-architect
Delegate to other skills when:
- Building microservice architecture →
backend-architect handles service mesh, load balancing
- Setting up CI/CD pipelines →
devops-automator handles deployment automation
- Implementing job queues →
background-job-orchestrator handles queue-specific patterns
- Creating always-on AI agents →
always-on-agent-architecture handles AI-specific lifecycle needs
1---2name: daemon-development3description: Build daemon/background processes that start on boot, run continuously, and manage their own lifecycle. Covers macOS launchd (plist files, agents vs daemons), Linux systemd (unit files), Windows services, process supervision, logging, health checks, graceful shutdown, auto-restart, and AI-powered daemons that manage LLM API connections and rate limits. Activate on: "daemon", "background process", "launchd", "systemd", "service file", "plist", "launch agent", "launch daemon", "auto-start", "always running", "process supervisor", "pm2", "background service", "boot service", "AI daemon", "long-running process". NOT for: container orchestration (use devops-automator), cron jobs that run and exit (use task-scheduler), web server deployment (use backend-architect).4license: Apache-2.05---6
7# Daemon Development
8
9Build long-running background processes that start reliably, run continuously, recover from failures, and shut down gracefully. Expert-level daemon architecture across macOS launchd, Linux systemd, and AI-powered services.
10
11## Decision Points
12
13### 1. Platform-Specific Init System Choice
14```
15Platform Detected?
16├─ macOS
17│ ├─ Must run at boot (no user login): LaunchDaemon → /Library/LaunchDaemons/
18│ ├─ User session required (GUI/files): LaunchAgent → ~/Library/LaunchAgents/
19│ └─ System-wide user service: LaunchAgent → /Library/LaunchAgents/
20├─ Linux
21│ ├─ systemd available: systemd unit file → /etc/systemd/system/
22│ ├─ Legacy SysV: init.d script (rare, avoid if possible)
23│ └─ Container: s6 or built-in supervision
24└─ Cross-platform dev
25 ├─ Node.js app: pm2 for development, systemd/launchd for production
26 └─ Other languages: Direct systemd/launchd implementation
27```
28
29### 2. Service Type Configuration
30```
31Daemon Startup Behavior?
32├─ Simple process (doesn't fork)
33│ ├─ systemd: Type=simple
34│ └─ launchd: Standard plist (no special keys)
35├─ Signals readiness when ready
36│ ├─ systemd: Type=notify + sd_notify("READY=1")
37│ └─ launchd: N/A (use health check instead)
38├─ Forks child process (legacy)
39│ ├─ systemd: Type=forking + PIDFile (avoid)
40│ └─ launchd: Not supported (rewrite to not fork)
41└─ Socket-activated
42 ├─ systemd: Type=simple + [Socket] section
43 └─ launchd: Sockets dict in plist
44```
45
46### 3. Restart Policy Design
47```
48Failure Recovery Strategy?
49├─ Critical service (must always run)
50│ ├─ systemd: Restart=always, RestartSec=5
51│ └─ launchd: KeepAlive=true, ThrottleInterval=10
52├─ Crash recovery only
53│ ├─ systemd: Restart=on-failure
54│ └─ launchd: KeepAlive={SuccessfulExit=false}
55├─ Manual restart preferred
56│ ├─ systemd: Restart=no
57│ └─ launchd: KeepAlive=false
58└─ Rate-limited restart
59 ├─ systemd: StartLimitBurst=5, StartLimitIntervalSec=60
60 └─ launchd: ThrottleInterval=30 (built-in)
61```
62
63### 4. AI Daemon Rate Limiting Strategy
64```
65LLM API Connection Pattern?
66├─ Single provider, token bucket
67│ ├─ Token estimation: prompt_tokens + max_completion_tokens
68│ ├─ Bucket refill: tokens_per_minute from provider limits
69│ └─ Overflow: Queue requests with priority
70├─ Multi-provider failover
71│ ├─ Circuit breaker per provider (3 failures = 30s timeout)
72│ ├─ Rate limit per provider independently
73│ └─ Failover order: primary → secondary → queue
74├─ Streaming responses
75│ ├─ Reserve tokens optimistically
76│ ├─ Adjust on actual_tokens in real-time
77│ └─ Handle mid-stream rate limits gracefully
78└─ Batch processing
79 ├─ Group similar requests to maximize throughput
80 └─ Split large batches if they hit rate limits
81```
82
83### 5. Graceful Shutdown Handling
84```
85SIGTERM Received?
86├─ Web server daemon
87│ ├─ 1. server.close() - stop accepting new connections
88│ ├─ 2. Wait for active requests (timeout: TimeoutStopSec-5s)
89│ ├─ 3. Close database connections
90│ └─ 4. exit(0)
91├─ Queue worker daemon
92│ ├─ 1. Stop polling for new jobs
93│ ├─ 2. Finish current job (timeout protection)
94│ ├─ 3. Flush any pending state
95│ └─ 4. exit(0)
96├─ AI daemon
97│ ├─ 1. Stop accepting new LLM requests
98│ ├─ 2. Drain in-flight requests (respect provider timeouts)
99│ ├─ 3. Save rate limit state to disk
100│ └─ 4. Close provider connections, exit(0)
101└─ Database/stateful daemon
102 ├─ 1. Checkpoint/flush transactions
103 ├─ 2. Close client connections gracefully
104 ├─ 3. Release file locks
105 └─ 4. exit(0)
106```
107
108## Failure Modes
109
110### 1. **Restart Loop Death Spiral**
111- **Symptoms**: High CPU usage, rapid log growth, service stuck in "activating" state
112- **Diagnosis**: Daemon crashes immediately on startup, init system restarts too quickly
113- **Detection**: `journalctl -u service` shows start/crash/start pattern every few seconds
114- **Fix**: Add `RestartSec=10` (systemd) or `ThrottleInterval=15` (launchd), implement startup validation
115
116### 2. **Zombie Process Accumulation**
117- **Symptoms**: `ps aux` shows `<defunct>` processes, parent daemon still running but degraded
118- **Diagnosis**: Daemon spawns child processes but doesn't reap them (missing SIGCHLD handler)
119- **Detection**: `ps axo pid,ppid,stat,comm | grep Z` shows zombie children
120- **Fix**: Install SIGCHLD handler with `wait()` or `waitpid()`, use `signal(SIGCHLD, SIG_IGN)` if children are fire-and-forget
121
122### 3. **Connection Leak Cascade**
123- **Symptoms**: "Too many open files" errors, service degrades over time, works fine after restart
124- **Diagnosis**: File descriptors or network connections not closed properly
125- **Detection**: `lsof -p <daemon_pid> | wc -l` grows continuously, eventual EMFILE errors
126- **Fix**: Add connection pooling with max limits, implement proper cleanup in error paths, set `LimitNOFILE` in systemd
127
128### 4. **Log Disk Explosion**
129- **Symptoms**: Disk space alerts, daemon crashes with "No space left on device"
130- **Diagnosis**: No log rotation configured, daemon writes unbounded logs
131- **Detection**: Log files in GB range, df shows /var/log at 100% capacity
132- **Fix**: Configure logrotate/newsyslog, use systemd journald with `SystemMaxUse`, add log level controls
133
134### 5. **Rate Limit Thrash (AI Daemons)**
135- **Symptoms**: High latency, request timeouts, alternating success/failure patterns
136- **Diagnosis**: No backoff after rate limit hits, daemon hammers API repeatedly
137- **Detection**: Provider API returns 429 errors, response times spike periodically
138- **Fix**: Implement exponential backoff, parse Retry-After headers, circuit breaker per provider
139
140## Worked Examples
141
142### AI Daemon Case Study: LLM Rate Limit Management
143
144**Scenario**: Building an AI daemon that processes user requests through OpenAI API, needs 99.9% uptime with graceful rate limit handling.
145
146**1. Initial Architecture Decision**
147```bash
148# Decision: Multi-provider with circuit breakers
149# Primary: OpenAI GPT-4, Secondary: Anthropic Claude, Tertiary: Local model
150
151# systemd unit file choice
152Type=notify # Daemon signals when fully initialized
153Restart=on-failure
154RestartSec=10
155```
156
157**2. Rate Limiting Implementation**
158```typescript
159// Token bucket per provider (expert catches: different providers = different limits)
160const rateLimiters = {
161 openai: new TokenBucket({ tokensPerMinute: 40000, burstCapacity: 8000 }),
162 anthropic: new TokenBucket({ tokensPerMinute: 25000, burstCapacity: 5000 }),
163};
164
165// Novice mistake: Request-based limiting
166// Expert insight: LLM APIs are token-based, not request-based
167async processRequest(req: LLMRequest) {
168 const estimatedTokens = this.estimateTokens(req.prompt, req.maxTokens);
169 await this.rateLimiters.openai.acquire(estimatedTokens);
170 // ... proceed with API call
171}
172```
173
174**3. Circuit Breaker Configuration**
175```typescript
176// Expert trade-off: Aggressive vs Conservative failover
177const circuitBreaker = new CircuitBreaker({
178 failureThreshold: 3, // Conservative: 5+ for stable APIs, 3 for flaky ones
179 timeout: 30000, // Aggressive: 15s, Conservative: 60s
180 resetTimeout: 60000, // How long to wait before retry
181});
182
183// Decision point: When to fail over?
184if (error.status === 429) {
185 // Rate limited: backoff on primary, don't fail over yet
186 await this.exponentialBackoff(provider, error.retryAfter);
187} else if (error.status >= 500) {
188 // Server error: immediate failover to secondary
189 this.circuitBreaker.recordFailure('openai');
190}
191```
192
193**4. Graceful Shutdown Pattern**
194```typescript
195// Expert catches: Race conditions in shutdown
196let shutdownInProgress = false;
197
198process.on('SIGTERM', async () => {
199 if (shutdownInProgress) return; // Idempotent shutdown
200 shutdownInProgress = true;
201
202 console.log('SIGTERM received, draining connections...');
203
204 // 1. Stop accepting new requests
205 server.close();
206
207 // 2. Wait for in-flight requests (with timeout)
208 const drainTimeout = setTimeout(() => {
209 console.log('Drain timeout, force exit');
210 process.exit(1);
211 }, 25000); // systemd TimeoutStopSec=30, so exit by 25s
212
213 await Promise.all([
214 this.drainActiveRequests(),
215 this.flushRateLimitState(), // Save token bucket state to disk
216 ]);
217
218 clearTimeout(drainTimeout);
219 process.exit(0);
220});
221```
222
223**5. Trade-offs and Decision Results**
224- **Aggressive restart**: `RestartSec=5` for quick recovery vs `RestartSec=30` to avoid thrashing
225- **Circuit breaker**: 3 failures for failover (catches transient issues) vs 5 failures (more stable)
226- **Drain timeout**: 25s (safe margin) vs 29s (maximize request completion)
227- **Result**: 99.95% uptime achieved, average failover time 2.3 seconds
228
229## Quality Gates
230
231- [ ] Boot test passes: Service starts automatically after system reboot
232- [ ] Log rotation configured: Logs rotate daily/weekly, old logs purged, no disk space growth
233- [ ] SIGTERM handling verified: Graceful shutdown completes within TimeoutStopSec
234- [ ] Health endpoint responds: HTTP 200 with meaningful status (uptime, queue depth, dependencies)
235- [ ] Restart backoff works: Service doesn't enter tight restart loop on immediate crash
236- [ ] Resource limits enforced: Memory/CPU/file descriptor limits prevent resource exhaustion
237- [ ] Non-root execution: Service runs as dedicated user with minimal privileges
238- [ ] Config reload tested: SIGHUP reloads configuration without full restart
239- [ ] Crash recovery verified: Process crash triggers automatic restart within RestartSec
240- [ ] Security hardening applied: NoNewPrivileges, ProtectSystem, ReadWritePaths configured
241- [ ] For AI daemons: Token-based rate limiting implemented with proper estimation
242- [ ] For AI daemons: Provider failover works when primary returns 5xx errors
243
244## Not-For Boundaries
245
246**This skill is NOT for:**
247- **Container orchestration**: For Kubernetes deployments, Docker Swarm, or container-specific supervision → use `devops-automator`
248- **One-shot scheduled tasks**: For cron jobs that run and exit, periodic batch processing → use `task-scheduler`
249- **Web application deployment**: For nginx/Apache configuration, reverse proxy setup, SSL termination → use `backend-architect`
250- **Queue worker frameworks**: For Sidekiq, Celery, Bull queues with built-in supervision → use `background-job-orchestrator`
251- **Development process management**: For hot-reloading, file watching, development servers → use standard dev tooling (nodemon, cargo watch)
252- **Database administration**: For MySQL/PostgreSQL service configuration → use `database-architect`
253
254**Delegate to other skills when:**
255- Building microservice architecture → `backend-architect` handles service mesh, load balancing
256- Setting up CI/CD pipelines → `devops-automator` handles deployment automation
257- Implementing job queues → `background-job-orchestrator` handles queue-specific patterns
258- Creating always-on AI agents → `always-on-agent-architecture` handles AI-specific lifecycle needs