Pilot AWS Operations (founder box)
Since the S6-lite cutover (2026-07-16, TASK-409) the Pilot daemon runs on an
EC2 box, NOT locally. Local pgrep pilot returning 0 is correct.
Constants
INSTANCE i-0e0c1ca34e7b561f9 # "pilot-founder-box", t3.xlarge, eu-central-1a
PROFILE quantflow # export AWS_PROFILE=quantflow for every aws call
REGION eu-central-1
RUNNER i-0147f5c24d234cdbb # mgmt runner (AdministratorAccess) — IAM/privileged ops only
DAEMON tmux session "pilot" as ec2-user, wrapper /home/ec2-user/start-pilot.sh
SLACK alerts → #pilot-reports · merge asks → founder bot DM · chat → #pointer (3 config keys; restart to apply) — see .agent/system/references/reference_slack_notifications_routing.md
BINARY /var/lib/pilot/bin/pilot # ec2-user-owned (self-upgrade works, #4470); /usr/local/bin/pilot is a SYMLINK to it — never `sudo install` over the symlink; rollback at pilot.prev
STATE /home/ec2-user/.pilot → /var/lib/pilot/pilot-home (200GB data volume)
REPOS /Users/aleks.petrov/Projects → /var/lib/pilot/repos (path shim — ledger keys
on the macOS-era absolute paths; NEVER "fix" these symlinks)
DB /home/ec2-user/.pilot/data/pilot.db
LOG /home/ec2-user/.pilot/logs/daemon.log (+ daemon-stderr.log)
The SSM command pattern (used by everything below)
export AWS_PROFILE=quantflow AWS_DEFAULT_REGION=eu-central-1
CMD=$(aws ssm send-command --instance-ids i-0e0c1ca34e7b561f9 \
--document-name AWS-RunShellScript \
--parameters 'commands=["<shell here>"]' \
--query Command.CommandId --output text)
sleep 10
aws ssm get-command-invocation --command-id $CMD \
--instance-id i-0e0c1ca34e7b561f9 --query StandardOutputContent --output text
For multi-line/quote-heavy payloads write a JSON file and use
--parameters file:///tmp/x.json; for shipping FILES to the box, base64 the
content into the command (echo <b64> | base64 -d > target) — inline heredocs
with \n escapes DO NOT survive SSM JSON (verified failure mode).
Operations
Status (first move, zero GitHub quota)
~/bin/pilot-board # daemon health, queue, autopilot, log tail
~/bin/pilot-board --gh # + GitHub issues/PRs (costs shared user quota — sparingly)
Remote half lives at /usr/local/bin/pilot-board-remote on the box (edit there).
Live TUI dashboard
~/bin/pilot-dash # SSM interactive → tmux attach -t pilot
Detach: Ctrl-B then D. NEVER press q or Ctrl-C inside the TUI — that
stops the daemon. Fallback: plain ssm session → sudo su - ec2-user →
tmux attach -t pilot.
Logs
# tail (adjust -n / add grep):
SSM: tail -50 /home/ec2-user/.pilot/logs/daemon.log
# common filters: 'rate limit', 'claim lost', 'approval', 'ERROR', a task id
Queue / ledger queries (read-only)
SSM: sqlite3 -column /home/ec2-user/.pilot/data/pilot.db \
"SELECT task_id,status,datetime(created_at) FROM executions \
WHERE status IN ('running','queued') ORDER BY created_at;"
Key tables: executions, execution_claims (task_id, project_path, generation),
autopilot_pr_state, autopilot_scope_release, instance_events.
⚠️ Timestamp trap: pre-2026-07-16 rows may carry legacy …+02:00 string format —
never filter by string time ranges across eras; use rowid or exact ids.
Stop / start / restart ⚠️ OPERATOR-CONSENT ACTIONS
Never do these on your own judgment during watch/autonomous modes; in interactive sessions get explicit user go-ahead. In-flight executions die (they retry via generation claims — proven safe, but wasteful).
# STOP (graceful; wait, then verify 0):
SSM: sudo -iu ec2-user tmux send-keys -t pilot C-c # TUI quit = graceful shutdown
sleep 15; ps -eo comm | grep -c '^pilot$' # must be 0
# START:
SSM: sudo -iu ec2-user tmux new-session -d -s pilot -x 220 -y 50 /home/ec2-user/start-pilot.sh
sleep 20; ps -eo comm | grep -c '^pilot$' # must be 1; then check pilot-board
# RESTART = STOP, verify, START, verify (banner version + no Telegram 409 + poller ticks).
After ANY start: verify version (/usr/local/bin/pilot version), exactly 1
process, curl -s localhost:9091/metrics | grep pilot_queue_depth (via SSM).
Rebuild / upgrade the box binary
SSM (as ec2-user, NO sudo — dir is ec2-user-owned since 2026-07-19):
cd /Users/aleks.petrov/Projects/startups/pilot && \
git fetch -q --tags origin main && git checkout -q <tag-or-origin/main> && \
make build && install -m 0755 bin/pilot /var/lib/pilot/bin/pilot && \
git checkout -q main && /var/lib/pilot/bin/pilot version
Then RESTART (above) to activate. Prefer building from a released tag. For releases, daemon self-upgrade also works now (#4470 preflight + writable dir) — letting it upgrade itself on the next train is the default path. Expect ~1 quiet hour post-restart until #4391 ships: startup rescans can burn the GitHub user-aggregate rate pool (see Troubleshooting).
Metrics tunnel (grafterm / local Prometheus tools)
~/bin/pilot-tunnel # box:9091 → localhost:9091, keep running; then `pilot-tui` (grafterm alias)
Box shell (interactive)
aws ssm start-session --target i-0e0c1ca34e7b561f9 --profile quantflow --region eu-central-1
sudo su - ec2-user
Privileged infra ops (IAM, volumes, instance-level)
User aleks lacks iam:PassRole/ec2:CreateVolume etc. Route through the mgmt
runner via SSM RunShellScript on i-0147f5c24d234cdbb (it has admin). Scope
every grant minimally; never print policy docs containing secrets.
Hard rules
- One daemon, ever. Never start pilot locally while the box serves the
repos (dual-serve = cross-machine duplicate class; claims are per-DB).
Rollback procedure lives in
.agent/tasks/TASK-409-s6-lite-aws-cutover.md. - Never touch the path shims (
/Users/aleks.petrov/...symlinks on the box) — ledger + claims key on those exact strings. - GitHub API is one shared per-USER pool (5000/hr) across every token,
session, and the daemon. Prefer sqlite/metrics over
ghfor status. If the daemon logs "rate limit exceeded for user ID …" — it self-recovers on the rolling window; do not thrash retries. (#4391 tracks the durable fix.) - Secrets: never echo tokens/keys; never
psfull args of processes holding tokens; the wrapper script pattern exists precisely to keep tokens out of argv. Config on the box is verbatim-from-laptop; changes = operator consent- restart.
- Trust the ledger over dashboard panels (known mislabel: awaiting_approval rendered as "rebase" until the GH-4383 fix is in the running binary).
- Name your ledger. Every status/diagnosis claim must state which data
source produced it: box DB (via SSM), GitHub, or the laptop archive. The
laptop's
~/.pilot/data/pilot.dbis a FROZEN pre-2026-07-16 archive (S6-lite cutover) — plausible-looking rows, months stale. A 2026-07-27 incident: a session read it and confidently misdiagnosed healthy tasks as "failed". Before reasoning about any DB: verify the path is the box's (/var/lib/pilot/pilot-home/data/pilot.db) and check row freshness (select max(datetime(created_at)) from executions). mem-160 family.
Troubleshooting quick table
| Symptom | Likely cause | Move |
|---|---|---|
| Board says queue empty but daemon log shows executions; log claim generations exceed claims table | split-brain shadow ledger (#4393 class — daemon opened a DB at an unshimmed path) | sudo readlink /proc/$(pgrep -x pilot)/fd/* must include /var/lib/pilot/pilot-home/data/pilot.db; if not: STOP daemon, locate+merge shadow DB, fix shim |
ALL task executions fail unknown: exit status 1 after ~3m, 0 tokens, stream shows api_retry/fetch failed; judge/preflight children work |
RLIMIT_AS cap on executor children (#4401 class — GH-3028 "RSS cap", darwin no-op) | grep 'address space' /proc/<claude-child>/limits must be unlimited; config subprocess_limits is enabled: false since 2026-07-17 (backup config.yaml.bak-4396) — OOM cap off until #4401 |
| Queue frozen, pollers 403 "rate limit … user ID" | user-aggregate GitHub pool exhausted (startup rescans, parallel sessions) | wait for rolling window; stop nonessential gh usage; see #4391 |
| Queue frozen, "dispatch claim lost" every poll, no 403s | dead-owner non-terminal rows holding gen-N claims (post-restart/cutover) | see #4392; workaround = mark orphan rows stalled with audit note (exact-id UPDATE, never string time-ranges) — this is a RECOVERY signal (retry the task), not a cancel; see next row if the goal is to stop a task |
| Need to stop a task from ever being re-picked (dup ticket, wrong scope, abandoned, operator error) | stalled is NOT a cancel verb — it means "dead owner, retry me" and the dispatcher grants it fresh generations exempt from the repick hard cap forever if hand-written for this purpose (GH-4655 incident) |
pilot task cancel <task-id> [--project <path>] [--reason "..."] (GH-4678) — the real terminal cancel; NEVER hand-write status='stalled' to try to stop a task |
Canceled a GitHub-backed task (GH-N) and want it to run again |
Relabel/reopen alone does NOT re-arm it — GH-5127/5129 recovery proved the old cancel hint false. As of GH-5139, re-arm needs a genuine reopen or (re-)label event dated after the cancel (checked via the GitHub issue-events timeline), not merely "issue currently open/labeled"; probed lazily on the next poll tick (repick-backoff throttled, not instant) | Remove the trigger label then re-add it (or close then reopen) — either produces a fresh timestamped event. Confirm re-arm in daemon.log: "GH-5139: canceled task re-armed via GitHub reopen/relabel". Non-GitHub adapters and any GH-N with no such event stay permanently canceled — file a new issue/task instead |
| tmux session gone, no pilot process | wrapper/script error at spawn | check daemon-stderr.log; verify wrapper intact (cat start-pilot.sh); restart per above |
| TUI monochrome | TERM captured at daemon start | ~/.tmux.conf already sets 256color; colors return on next restart — do not bounce for paint |
| PR stuck "rebase N/3" in panel | awaiting_approval mislabel | ledger: SELECT stage FROM autopilot_pr_state WHERE pr_number=N |
| Box unreachable via SSM | agent/instance down | aws ec2 describe-instances --instance-ids i-0e0c… → LOUD escalate to operator; never assume |
Multi-day uptime right after a train — looks like "installed but never restarted" |
NOT a stale process — self-upgrade's restart leg is syscall.Exec (PID preserved, so ps etime/uptime survives the restart; mem hot-restart-preserves-pid-uptime-false-mismatch). Since 2026-08-13 board ver reads the RUNNING process (pilot_build_info metric, #4864/PR#4865; repointed pilot-board-remote) — a (disk!) suffix means the daemon metric was unreachable and the value is the disk binary's |
Trust ver (metric-sourced). Cross-check: curl -s localhost:9091/metrics | grep pilot_build_info via SSM, or grep 'upgrade verified complete' ~/.pilot/logs/daemon.log (boot-reconcile from=/to= logged by the NEW image; failure counterpart previous upgrade did NOT take effect). Do NOT bounce the daemon on uptime evidence alone |
Refs
- Cutover plan + rollback:
.agent/tasks/TASK-409-s6-lite-aws-cutover.md - Restart discipline:
.agent/sops/operations/safe-daemon-restart.md(verify PID+banner ALWAYS) - Open hardening: #4391 (rate-budget client), #4392 (orphan reconciliation)