Daily Security Scanning & Fleet Health Automation (v2)
When to Use
Build or maintain a daily security scan when:
- User requests a "security scan" or "daily security briefing"
- Setting up cron jobs for CTO oversight
- Automating fleet health monitoring
- Checking SSL cert expiry, CVE feeds, supply chain vulns
- Any recurring security audit workflow
Scan Architecture (v2 — 9 Layers)
The canonical script lives at the workspace path ~/clawd/tools/appie-3-daily-security-scan.sh. The actual runnable copy is at ~/.hermes/scripts/appie-3-daily-security-scan.sh (or the profile-specific scripts dir, e.g. ~/.hermes-appie3/scripts/ on some setups). Every layer maps to a function.
Layer 1: Local Machine Health
- Disk usage (df -h /)
- Memory: vm_stat — use grep, NOT awk /pattern/
* WRONG: vm_stat | awk '/pages active/ {print $NF}' → empty!
* RIGHT: vm_stat | grep 'pages active' | awk '{print $NF}'
* On macOS vm_stat output starts with uppercase "Pages", awk /pages/ doesn't match
* If >90% active: dump top 5 processes by RSS + swap info
- Load averages (sysctl vm.loadavg / uptime)
- Hermes agent count (pgrep -f hermes_cli | wc -l)
- Failed logins in 24h (log show --predicate)
Layer 2: Fleet Health (SSH via Tailscale)
Uses 4 category arrays (all indexed, pipe-separated — bash 3.x compat):
| Array | Emoji | Condition | Check Type |
|---|---|---|---|
FLEET |
🟢/🔴 | SSH key works | Full SSH health (disk, load, Hermes, updates) |
BROKEN_KEYS |
🟡 | Online, port 22 open, key rejected | nc -zv port check (<1s) |
TAILNET_ONLINE |
⚪ | Online, no SSH daemon | Tailnet status only |
GHOSTS |
💤 | Offline >7d | Archived, no active check |
Entry format: "name|user@tailscale_ip|ssh_key_path|description"
SSH: -o ConnectTimeout=5 -o BatchMode=yes -o StrictHostKeyChecking=no
Results cached to FLEET_CACHE file (pipe-separated: name|status|desc)
Key v2 fix: FLEET_CACHE is shared between markdown report and Telegram output. NO separate SSH loop for Telegram. This eliminates the v1 Telegram divergence bug.
Tailnet-only entries (BROKEN_KEYS, TAILNET_ONLINE, GHOSTS) also write to FLEET_CACHE so the Telegram output can render all 4 groups in sequence from a single cache read. This is critical for completeness — the Telegram scan should show ALL tailnet nodes, not just SSH-reachable ones.
Layer 3: Local Open Ports
lsof -iTCP -sTCP:LISTEN -P -n | awk 'NR>1 && !seen[$1,$9]++'
Flags unexpected dev services on unprivileged ports. Expected: bun on 37701 (Hermes internal).
Layer 4: Supply Chain Security
# npm audit (high+ only)
npm audit --audit-level=high
# Python safety check
safety check --short
# Gitleaks secret scan — CRITICAL: always use --no-git with .gitleaks.toml
gitleaks detect --source $CLAWD_DIR --no-git --config .gitleaks.toml --verbose
- ALWAYS --no-git: git mode scans 2369 commits (487MB) → 35k false positives
- ALWAYS .gitleaks.toml: suppress example keys, lockfile hashes, test data
- ALWAYS timeout 60: gitleaks --no-git can CPU-spike to 975%
- Extract leak count from "leaks found: N" line, not grep -c
Layer 5: SSL/TLS Certificates
# Use brew OpenSSL — system LibreSSL can't parse x509 output
ossl="/opt/homebrew/bin/openssl"
cert_raw=$(echo "" | "$ossl" s_client -servername "$domain" -connect "$domain":443 2>&1)
enddate=$(echo "$cert_raw" | "$ossl" x509 -noout -enddate 2>/dev/null | cut -d= -f2)
- NO 2>/dev/null on s_client (kills output in subshell)
- NO timeout wrapper (kills mid-handshake)
- Flag <7d 🔴, <30d ⚠️
Layer 6: Security Headers
curl -sI --max-time 5 "https://$domain"
# Check for: HSTS, CSP, X-Frame-Options, X-Content-Type-Options
All 4 required. Score: 4/4 🟢, 2-3 ⚠️, 0-1 🔴.
Caveat: follow redirects with -L if domain uses Cloudflare/redirect chains.
Layer 7: Pending Updates (fleet SSH)
ssh <node> "apt list --upgradable | grep -v 'Listing...' | wc -l"
ssh <node> "apt list --upgradable | grep -i security | wc -l"
- Security count via
grep -i security(not-security— varies by distro) - 0 updates ✅, 1-20 ⚠️, 20+ or any security 🔴
Layer 8: Tailscale Network
tailscale status --json | python3 -c "import sys,json; ..."
Counts peers, finds offline nodes, shows last-seen timestamps.
Layer 9: CVE Watch (v2 — no NVD)
PRIMARY: GitHub Advisory API (no auth, no rate limit issues)
GET https://api.github.com/advisories?type=reviewed&severity=critical&per_page=8
- Returns GHSA advisories sorted by
published_atdesc - Use HTTP status code check (curl -w %{http_code}) — don't rely on python successfully parsing
- Filter by published_at within 48h via Python datetime comparison
- Also fetch high severity for awareness
SECONDARY: OSV.dev (per-package, always works)
POST https://api.osv.dev/v1/query
{"package": {"name": "openssl", "ecosystem": "PyPI"}}
- Query key packages: openssl, node, curl
- Also query agent frameworks via
agent-framework-cve-scan.py(seereferences/agent-framework-cve-scan.md) - Filter by published date within 90d
- No auth needed, no rate limits observed
TERTIARY: Agent Framework & Go Ecosystem CVE Scanner (standalone Python script)
- Covers 20 packages: 17 Python agent frameworks (LangChain, CrewAI, Semantic Kernel, AutoGen, LlamaIndex, LiteLLM, guardrails-ai, giskard, etc.) + 3 Go infra packages (golang.org/x/crypto, github.com/go-chi/chi, github.com/sigstore/rekor)
- Checks
pip listlocally + OSV.dev API per package - Go packages added 2026-06-26 n.a.v. 7 critical SSH crypto CVEs published 2026-06-25
- Run:
python3 ~/clawd/tools/agent-framework-cve-scan.py - See
references/agent-framework-cve-scan.md
NVD API v2.0 is NOT used — requires API key to avoid 5 req/30s limit. Free key tier exists but is unavailable from this environment.
Layer 10: Bot & Client Health Check (ad-hoc, not in daily scan)
Run periodically (weekly or on demand) — check all active Telegram bots and web live services for basic health:
# 1. List deployed bots from project config
# 2. For each bot endpoint, check:
# - curl -m 5 <bot_url>/health (or /) — returns 200?
# - curl -m 5 <bot_url> | grep -i "ok\|alive\|running"
# 3. For Telegram bots, check response via bot API:
# curl -m 5 "https://api.telegram.org/bot<TOKEN>/getMe"
# 4. For YouTube / media bots, check if deploy is still live on platform
Not automated as a daily cron (too many client-specific endpoints, rate-limit risk on Telegram API). Run as an ad-hoc CTO audit.
Layer 11: Continuous Improvement — Research → Scripts
Seyed's standing directive: use the daily AI briefing research to improve the security scripts. After each daily briefing, scan the research output for:
| Signal | Action |
|---|---|
| New CVE class or attack vector | Add a check layer or tool to the scan |
| New tool or best practice | Add install command + verification to the scan or host-init |
| Configuration hardening advice | Add to the security suggestion pipeline |
| Client bot platform deprecation | Flag in bot health check layer |
| New scanning methodology | Replace or augment an existing scan layer |
Implementation checklist after each daily briefing:
- Read the research output (
~/clawd/appie-brain/knowledge/research/daily-research/YYYY-MM-DD/README.md) - Cross-reference against existing scan layers — what's missing?
- For any gap: add a new function to the scan script or update existing logic, or create a standalone script for cross-platform use
- Test the change: run the affected layers manually
- If the improvement is structural (new layer, new tool), update this SKILL.md — add a reference file if the new tool has its own docs
- Log to Mission Control:
mc-log-task.py "Security script improvement: <summary>" --agent Appie-3
Do not batch up improvements. Make them as you discover them. A one-line regex addition or a new cert check costs nothing; deferring it until "next Monday" means it never happens.
Examples of recent improvements from research:
agent-framework-cve-scan.py— created from 2026-06-20 briefing which found CVE-2026-26030 (Semantic Kernel RCE), GHSA-gr75-jv2w-4656 (LangChain path traversal). Expanded 2026-06-25 +3 (litellm, guardrails-ai, giskard). Expanded 2026-06-26 +3 Go infra packages (golang.org/x/crypto after 7 critical SSH CVEs, go-chi/chi IP spoofing, sigstore/rekor OOM). Seereferences/agent-framework-cve-scan.md.headroom-aiv0.26.0 — installed after 2026-06-20 briefing flagged Headroom (60-95% token compression). Has MCP server, pure Python, Apache-2.0.- SkillsGuard evaluated — TypeScript/Node.js project (not installable on Python stack), cloud API available at
https://skillsguard.apiskillsguard.workers.dev/scan.
Tailnet Reconnaissance — Full Fleet Exploration Pattern
A standalone workflow for when you need a complete picture of every machine on Tailscale: what's online, what ports are open, what services run, and whether SSH keys work. Use this before setting up scans, deploying keys, or auditing fleet security posture.
Workflow
Step-by-step exploration, run each command and compile results:
Step 1: List all machines
tailscale status
# Full listing with IP, name, OS, online/offline status
tailscale status | grep -v offline # only online machines
tailscale status --json # programmatic access
Step 2: Ping all online machines
for ip in <all_online_tailscale_ips>; do
result=$(ping -c 2 -W 3 $ip 2>&1 | tail -1)
echo "$ip -> $result"
done
Helps identify reachability and latency. Machines behind DERP relays show higher RTT (100-400ms). Local DERP-free machines show 1-5ms.
Step 3: Port-scan per machine
Use nc -zv for fast TCP port checks. Common ports to probe:
# Linux servers
for port in 22 80 443 3000 5000 8000 8080 8443 9090 2375 2376 6443 7860 8888; do
nc -zv -w 2 $ip $port 2>&1 | grep -v "Connection refused"
done
Additional macOS-specific ports: 5900 (VNC), 7000 (AirPlay), 5353 (mDNS).
Key port signatures:
22→ SSH (verify key auth next)443→ HTTPS web service80→ HTTP (redirect to HTTPS or legacy)3000→ Common dev/Next.js default port5000→ Flask/Express/development5900→ VNC/Screen Sharing (macOS)8443→ Alternative HTTPS
Step 4: Identify web services
For every open HTTP/HTTPS port, curl the root path to identify the service:
# HTTPS
curl -sk https://$ip/ | head -c 200
curl -sk https://$ip/ 2>&1 | grep -i "<title>"
# HTTP on alternate ports
curl -sk http://$ip:$port/ | head -c 200
Look for:
<title>tag — identifies the site/app- Server headers:
curl -sI https://$ip/ | grep -i server - Framework fingerprints (Next.js, WordPress, Flask, etc.)
- JSON responses if it's an API
- Empty response may indicate TLS SNI filtering — try
-H "Host: example.com"
Step 5: Test SSH Connectivity & Diagnose Failures
When SSH fails, methodically diagnose why — the remediation differs by failure mode.
5a. Basic Auth Test
ssh -o ConnectTimeout=5 -o BatchMode=yes -o StrictHostKeyChecking=no \
-i ~/.ssh/id_ed25519 <user>@<ip> "hostname && uname -a"
Outcomes by error message:
| Error | Meaning | Next step |
|---|---|---|
Permission denied (publickey) |
Port open, SSH running, key not accepted | Deploy the key (see refs) |
Connection refused |
Port 22 closed, no SSH daemon | Go to 5b below |
Connection timed out / hangs |
Firewall dropping or host unreachable | Check network (ping) |
unexpected HTTP response: 502 Bad Gateway via tailscale ssh |
Tailscale SSH not configured on remote | Go to 5b below |
Also test alternative keys: id_ed25519_github, id_ed25519_spark, etc.
5b. Port-Closed Diagnosis (SSH daemon not running)
When port 22 is closed, use this systematic flow:
# 1) Verify basic network reachability
ping -c 2 -W 3 <IP>
# 2) Check if the host is on Tailscale and its state
tailscale status | grep <IP-or-name>
# Look for: "active" → reachable, "offline" → unreachable
# 3) Check the remote peer's Tailscale SSH capabilities
tailscale status --json | python3 -c "
import sys, json
d = json.load(sys.stdin)
for pid, p in d.get('Peer', {}).items():
if '<IP>' in str(p.get('TailscaleIPs', [])):
print('OS:', p.get('OS'))
print('Online:', p.get('Online'))
print('Capabilities:', p.get('Capabilities'))
print('CapMap:', p.get('CapMap'))
"
Interpretation:
Capabilities: null/CapMap: null→ Tailscale SSH is NOT enabled on the remote node. SSH won't work viatailscale ssheither."https://tailscale.com/cap/ssh": nullin CapMap → Tailscale SSH IS enabled. If it fails with 502, check the Tailscale admin console.- Node is
activewith direct connection → reachable (even if SSH is closed).
# 4) Port scan common services to see if ANY ports are open
for port in 22 443 5900 8080 8443 3000 5000; do
nc -zv -w 3 <IP> $port 2>&1 | grep -v "Connection refused"
done
Remediation by failure mode:
| Failure | Fix |
|---|---|
| Port 22 closed, macOS | On the Mac: System Settings → General → Sharing → Remote Login. Or in Tailscale Admin Console → approve node for Tailscale SSH. |
| Port 22 closed, Linux | sudo systemctl enable --now sshd or sudo apt install openssh-server |
Port 22 open, Permission denied |
Deploy SSH public key to remote ~/.ssh/authorized_keys (see references/ssh-key-deployment.md / references/mac-mini-coding-harness-recovery.md) |
| Tailscale SSH 502, port 22 works fine | Retry direct ssh via Tailscale IP. The 502 is a Tailscale relay issue, not an SSH error. |
| No ports open at all (22, 443, 5900, etc. all refused) | Machine may be asleep, on a restricted network, or fully locked down. macOS: no services enabled in Sharing prefs. |
For detailed session walkthrough with real error transcripts, see references/ssh-connectivity-diagnostics.md.
macOS coding-harness recovery pattern: if a Mac mini is online, port 22 is open, but common users reject Appie-3's key, stop guessing users/passwords. Ask the local operator or another authorized agent on the box to add Appie-3's public key to the active macOS user's ~/.ssh/authorized_keys, enable Remote Login, and report the actual username. Then verify harness state (~/clawd, ~/.hermes, ~/.codex, ~/.claude, and ps for hermes|claude|codex|copilot|node|python) before changing anything. Full recipe: references/mac-mini-coding-harness-recovery.md.
Step 6: Check local authorized_keys
cat ~/.ssh/authorized_keys
cat ~/.ssh/id_ed25519.pub # your own public key
This tells you:
- Who can SSH into you (the keys in authorized_keys)
- Who you claim to be (the pubkey comment field)
- Whether cross-fleet SSH is unidirectional only
Step 7: Compile and present
Present findings as a compact table:
| Machine | IP | OS | Ping | Open ports | Services |
|---------|----|----|------|------------|---------|
| machine | 100.x.x.x | Linux | Xms | 22, 443, 3000 | SSH, web (app name) |
Include:
- Online/offline status per machine
- Reachable services and what they are
- SSH key acceptance status
- The user's own pubkey (ready to copy-paste for deployment)
Common Findings
- SSH key not deployed: New VPS/machine provisioning — pubkey needs to be added to authorized_keys. Common pattern: Appie-3's key isn't on any fleet machine, but other agents' keys already are.
- VNC open: macOS machines with Screen Sharing enabled. Port 5900 open means anyone on tailnet could attempt authentication. Verify VNC is password-protected.
- Web service on alternate port: Next.js dev server (3000), Flask/Express (5000), alternative HTTPS (8443) — may be development/staging environments not meant for production.
- Offline nodes >30d: Move to GHOSTS array in scan script. Their Tailscale IPs may change on reconnection.
- Multi-OS tailnet: macOS machines have VNC (5900), AirPlay (7000), and different SSH configuration paths. Linux servers are more consistent.
Multi-Agent Cron Inventory — Fleet-Wide Schedule Audit
A systematic method for auditing ALL cron jobs across a multi-agent fleet. Use when Seyed asks for a comprehensive schedule review, "make all cron jobs aesthetic", or before any fleet-wide cron reformatting.
Workflow
Step 1: Enumerate all accessible machines
# SSH key test across ALL tailnet hosts
for host in "root@100.x.x.x" "appie@100.x.x.x" "eva@100.x.x.x" "root@100.x.x.x"; do
result=$(ssh -o ConnectTimeout=3 -o BatchMode=yes -o StrictHostKeyChecking=no \
-i ~/.ssh/id_ed25519 "$host" "hostname" 2>&1)
echo "$host → $result"
done
Note: SSH user varies by machine — try root, appie, the user's name (eva, harry). Don't assume one user works everywhere.
Step 2: For each reachable machine, identify the runtime
# Hermes cron
hermes cron list
# OpenClaw cron (legacy)
ls ~/.openclaw/cron/ 2>/dev/null
# launchd plists (macOS)
ls ~/Library/LaunchAgents/ | grep -v 'com\.google\|homebrew'
Key signals for runtime detection via ps aux:
hermes_cli.main gateway run→ Hermes agentopenclawin process name → OpenClaw agentclaudeorcodexin process name → Claude Code / Codex CLI
Step 3: Check cron managers
| Runtime | Cron Manager | Command |
|---|---|---|
| Hermes | hermes cron |
hermes cron list |
| OpenClaw (legacy) | cron files | ls ~/.openclaw/cron/ |
| macOS (legacy) | launchd plists | launchctl list | grep com.weblyfe |
| System crontab | system cron | crontab -l |
Step 4: Document per-job health
For each job, note:
- Name and schedule
- Last run status (ok/error/failed delivery)
- Delivery target (telegram/origin/none)
- If job is broken: capture the exact error message
Step 5: Identify aesthetic/styling improvements
Check for:
- Naming convention: consistent prefix per agent (
appie3-,appie2-), no mixed case - Output formatting: Markdown + emoji headers vs plain text
- Delivery target: all jobs deliver to Telegram (or all to origin — be consistent)
- Timing deconfliction: no overlapping runs on the same agent
- Broken jobs: provider errors, expired credits, delivery failures — all need fixing
Fleet Topology Map (2026-06-24)
| Machine | IP | Runtime | Cron | Notes |
|---|---|---|---|---|
| appie-3-hermes | 100.69.131.51 | Hermes | hermes cron (6 jobs) |
✅ stable |
| appie-2 | 100.118.143.10 | Hermes | hermes cron (13 jobs) |
🟡 7 broken |
| appie-1 | 100.101.29.56 | OpenClaw | launchd (40+ plists) | 🟢 |
| eva | 100.99.64.92 | Hermes + OpenClaw | launchd | 🟢 |
| eugi | 100.110.58.73 | Hermes | 0 cron jobs | schone lei |
| Deadpool | ? (off-tailnet) | ? | ? | Aparte Hetzner VPS voor Roslan + EKO |
Zie references/fleet-inventory.md voor de complete actuele inventaris.
Tailnet Discovery
Not all fleet machines are on Tailscale. Deadpool (Hetzner VPS voor Roslan + EKO) is NOT on the tailnet and has no SSH config entry. When checking "all machines", distinguish between tailnet-connected and off-tailnet infrastructure.
# Full peer list with online status
tailscale status
# JSON output for programmatic use (Self vs Peer distinction is CRITICAL)
tailscale status --json
The local machine appears in Self, ALL other nodes in Peer dict:
d = json.load(sys.stdin)
self = d.get('Self', {})
self_online = self.get('Online', False) # local machine
peer_online = any(p.get('Online', False)
for p in d.get('Peer', {}).values()) # others
WRONG: Looping over Peer to find "appie-1" → always returns offline.
RIGHT: Check Self for the local machine, Peer for all others.
This matters for watchdog scripts that check local node tailnet health. The tailscale status CLI output shows the local node without a status column (making it look unknown), but JSON Self.Online is always True for the running machine.
Connectivity Checks
tailscale ping -c 1 --timeout 3s <IP> # cross-platform
ssh -o ConnectTimeout=5 -o BatchMode=yes -o StrictHostKeyChecking=no <user>@<ip> <cmd>
Key: Never use direct Hetzner IPs — they change on reprovision. Always use Tailscale IPs (100.x.x.x). SSH key timeout: ConnectTimeout=5 is required; ConnectTimeout=3 gives false positives on DERP relay nodes.
Cross-Machine File Retrieval
When you need to find and retrieve a file from a fleet machine:
Know where canonical assets live — common locations per machine:
- appie-1 (Mac Mini):
~/clawd/projects/openclaw-guide/canonical/(Build-Your-Own-Appie PDFs),~/clawd/assets/agent-avatars/(Telegram bot avatars) - appie-2 (Hetzner VPS):
/root/appie-brain/,/root/.hermes/ - eugi (Hetzner VPS):
/root/.openclaw/,/root/.hermes/
- appie-1 (Mac Mini):
Search remotely — use
findwith name globs:ssh <user>@<tailscale-host> "find <path> -iname '*pattern*v4*' -o -iname '*pdf*' 2>/dev/null"- Always use
find(notls -R) — scales to large directories - Quote paths with spaces:
"/Users/appie/clawd/projects/openclaw-guide/" - Use wildcards for version numbers:
*v4.5*
- Always use
SCP back — file to local
/tmp/:scp <user>@<tailscale-host>:"<remote-path>" /tmp/<local-name>scpvia Tailscale IP/hostname works without extra flags- Use
/tmp/for ephemeral copies; move to~/clawd/for permanent storage
Deliver the file:
- Telegram: include
MEDIA:/tmp/<file>in response — sends as native attachment - Email: NOT set up on this VPS (no himalaya, mailtools, or SMTP creds). Ask Seyed for creds or an alternative
- Public URL: upload somewhere if Seyed needs to share the link
- Telegram: include
Clean up — remove from
/tmp/when no longer needed:rm /tmp/<local-name>
Pitfall: macOS paths start with /Users/, not /root/. Check remote OS before constructing search paths.
SSH Key Inventory
Keys stored in ~/.ssh/. SSH config (~/.ssh/config) has host aliases for fleet machines.
Current reality (2026-06-24): appie-2, eugi, appie-1, and eva all accept Appie-3's key. All other fleet nodes reject it or have no SSH daemon. Deploying the key to remaining BROKEN_KEYS nodes is the next step for fleet-wide SSH management.
| File | Purpose | Status | Last Verified |
|---|---|---|---|
~/.ssh/id_ed25519 |
Default key (appie-3-hermes@tailnet) |
🟢 appie-2, eugi, appie-1, eva / 🔴 others | 2026-06-24 |
~/.ssh/id_ed25519_github |
GitHub-only key | 🟢 Works | — |
~/.ssh/id_ed25519_spark |
Spark Atlas key | 🔴 Stale (node offline) | 2026-06-11 |
Fleet SSH Access Map (2026-06-23)
| Device | OS | Port 22 | SSH Key | Category | Remedy |
|---|---|---|---|---|---|
| appie-2 | Linux | ✅ Open | ✅ Root accepted | FLEET 🟢 | — |
| eugi | Linux | ✅ Open | ✅ Root accepted | FLEET 🟢 | — |
| appie-1 | macOS | ✅ Open | ✅ appie@100.101.29.56 |
FLEET 🟢 | User appie, discovered 2026-06-24 |
| eva | macOS | ✅ Open | ✅ eva@100.99.64.92 |
FLEET 🟢 | User eva, discovered 2026-06-24 |
| appie-mc-1 | Linux | ✅ Open | ❌ Key rejected | BROKEN_KEYS 🟡 | Deploy pubkey to root@100.107.179.3 |
| harry (mac-mini-van-harry) | macOS | ✅ Open | ❌ Key rejected | BROKEN_KEYS 🟡 | Add pubkey to macOS user's authorized_keys |
| rabi (mac-mini-van-rabi) | macOS | ❌ Closed | N/A | TAILNET_ONLINE ⚪ | Enable Remote Login in Sharing prefs |
| techwiz-mbp | macOS | ❌ Closed | N/A | TAILNET_ONLINE ⚪ | Enable Remote Login in Sharing prefs |
| ipad-pro | iOS | N/A | N/A | TAILNET_ONLINE ⚪ | No SSH possible |
| spark-atlas | Linux | ❌ Offline | N/A | GHOSTS 💤 | Wait for node to come online |
| artemis | macOS | ❌ Offline | N/A | GHOSTS 💤 | Wait for node to come online |
| mac-studio-luminaire | macOS | ❌ Offline | N/A | GHOSTS 💤 | Wait for node to come online |
| iphone181 | iOS | ❌ Offline | N/A | GHOSTS 💤 | No SSH possible |
| macbook-air-lorenzo | macOS | ❌ Offline | N/A | GHOSTS 💤 | Wait for node to come online |
| wolf-diddy | macOS | ❌ Offline | N/A | GHOSTS 💤 | External peer, no access |
Key SSH failure patterns (2026-06-19 audit):
- macOS with port 22 open: SSH server running but
appie-3-hermes@tailnetkey not in any macOS user's~/.ssh/authorized_keys. Remedy: deploy the key to the actual macOS user. - Linux with port closed: SSH daemon not installed/started (
systemctl enable --now sshd+apt install openssh-serverif needed). - Linux with
Connection refusedafter initialPermission denied: fail2ban likely blocked the source IP (eugi pattern). - macOS with all ports closed: Remote Login + Screen Sharing both disabled in System Sharing prefs.
See references/ssh-connectivity-diagnostics.md for detailed error transcripts and step-by-step per-failure diagnosis.
Fleet Node Authorized Keys Audit
Check how many authorized keys exist per node and who they belong to:
for host in "root@100.118.143.10" "root@100.110.58.73"; do
echo "=== $host ==="
ssh -o ConnectTimeout=5 -o BatchMode=yes -o StrictHostKeyChecking=no \
-i "$HOME/.ssh/id_ed25519" "$host" \
"cat /root/.ssh/authorized_keys | awk '{print \$3}'"
done
| Node | Keys | Identities |
|---|---|---|
| appie-2 | 4 | appie@weblyfe-ocean, appie@weblyfe, seyed@Techwiz-MacBook-Pro-8, mc@appie-mc-1 |
| eugi | 2 | appie@weblyfe, Appie-2@appie-brain |
Document any unexpected keys. The comment field (3rd column) identifies the key owner. Stale keys (old hostnames, decommissioned users) should be removed.
Fleet Node Deploy Keys
Fleet nodes may carry per-service deploy keys for CI/CD:
ssh <host> "ls -la /root/.ssh/ | grep -v authorized_keys | grep -v known_hosts"
Example from appie-2:
privanotify_admin_deploy → PrivaNotify admin
privanotify_deploy → PrivaNotify
weblyfe_ai_deploy → Weblyfe AI
weblyfe_ai_deploy_new → Weblyfe AI (new)
github_do_appie → GitHub Actions deploy
Document the purpose of each key. Remove stale ones.
SSH Config Audit
ssh <host> "cat /root/.ssh/config"
Check for:
PasswordAuthentication noPubkeyAuthentication yes- No wildcard
Host *patterns that expose sensitive hosts
SSH config aliases can drift over time as Tailscale IPs change during reprovisioning. Periodically run tailscale status and update SSH config.
Hermes Agent Detection
ssh <host> "ps aux | grep -iE '(hermes|claude|python.*bot|node.*bot|discord|telegram)' | grep -v grep"
The process looks like:
/usr/local/lib/hermes-agent/venv/bin/python -m hermes_cli.main gateway run --replace
Hermes proc count on a healthy node: 2-5 (gateway + workers). Count of 0-1 is a warning.
Appie-Brain Locations
| Machine | Path |
|---|---|
| appie-3-hermes (deze VPS) | /root/clawd/appie-brain/ |
| appie-2 (Hetzner) | /root/appie-brain/ |
| appie-1 (Mac Mini) | ~/clawd/appie-brain/ |
Fleet Reality-Check (Periodic Inventory Refresh)
Fleet node status changes over time — offline nodes come back, new nodes join, nodes get decommissioned. All 4 arrays in the scan script drift from reality.
Run this check monthly to verify the script's arrays match real Tailscale status:
# 1. Get actual Tailscale online state
tailscale status | grep -v '^$' | awk '{print $1, $2, $(NF)}'
# 2. Compare against script's FLEET array
grep -A20 '^FLEET=(' ~/clawd/tools/appie-3-daily-security-scan.sh
# 3. Check BROKEN_KEYS — are they still online?
grep -A15 '^BROKEN_KEYS=(' ~/clawd/tools/appie-3-daily-security-scan.sh
# 4. Check TAILNET_ONLINE — still no SSH?
grep -A10 '^TAILNET_ONLINE=(' ~/clawd/tools/appie-3-daily-security-scan.sh
# 5. Check GHOSTS — any back online?
grep -A15 '^GHOSTS=(' ~/clawd/tools/appie-3-daily-security-scan.sh
When to move a node between arrays:
- Ghost comes online → move to FLEET (with current SSH key and Tailscale IP)
- Active node offline >30d → move to GHOSTS
- Node decommissioned → remove from both arrays
2026-06-02 update: spark-atlas and mac-mini-van-eva both returned online after 51+ days offline. Move from GHOSTS back to FLEET. Verify SSH keys still work before re-adding.
Fleet Status Categories
15 tailnet nodes (2026-06-23) grouped into 4 arrays:
FLEET — SSH-reachable (daily full health check)
FLEET=( # dagelijks SSH health → 🟢 (full data) / 🔴 (failed)
"appie-2|root@100.118.143.10|$HOME/.ssh/id_ed25519|Weblyfe VPS"
"eugi|root@100.110.58.73|$HOME/.ssh/id_ed25519|Ubuntu VPS (FSN)"
"appie-1|appie@100.101.29.56|$HOME/.ssh/id_ed25519|Mac Mini Weblyfe (macOS)"
"eva|eva@100.99.64.92|$HOME/.ssh/id_ed25519|Mac Mini Naoufal (macOS)"
)
BROKEN_KEYS — online, port 22 open, key rejected
BROKEN_KEYS=( # daily nc port check only → 🟡
"appie-mc-1|root@100.107.179.3|$HOME/.ssh/id_ed25519|Mission Control v1 (Linux) — key rejected"
"harry|100.79.180.56|none|Mac Mini Harry (macOS) — key rejected"
)
TAILNET_ONLINE — online, no SSH daemon at all
TAILNET_ONLINE=( # tailnet status only → ⚪
"rabi|100.67.184.25|macOS|Mac Mini Rabi — Remote Login off"
"techwiz-mbp|100.87.99.11|macOS|Techwiz MacBook Pro — Remote Login off"
"ipad-pro|100.105.56.2|iOS|iPad Pro — no SSH"
)
GHOSTS — offline >7d (logged but no active check)
GHOSTS=( # archived → 💤
"spark-atlas|100.69.197.43|linux|DGX Spark — offline 9d"
"artemis|100.95.165.116|macOS|Artemis — offline 11d"
"mac-studio-luminaire|100.126.237.96|macOS|Studio Luminaire — offline 11d"
"iphone181|100.98.117.8|iOS|iPhone S3YED — offline 4d"
"macbook-air-lorenzo|100.112.169.28|macOS|Book Air Lorenzo — offline 74d"
"wolf-diddy|100.102.181.116|macOS|Wolf Diddy — offline 32d"
)
Transition rules:
- Ghost comes online → move to BROKEN_KEYS or TAILNET_ONLINE (test SSH first)
- BROKEN_KEYS node gets SSH key deployed → move to FLEET
- Online node offline >7d → move to GHOSTS
- Node decommissioned → remove from all arrays
- iOS/iPad nodes belong in TAILNET_ONLINE or GHOSTS — never in FLEET (no SSH)
Voorkomt: valse 🔴 elke dag, 5s SSH timeouts per ghost, ruis in CTO notes.
CRITICAL: Maintener de arrays actief. Wanneer een ghost node terug online komt (check via tailscale status), verplaats hem direct terug naar BROKEN_KEYS of TAILNET_ONLINE. De SSH key en Tailscale IP kunnen veranderd zijn sinds de node offline ging.
Periodic refresh: zie "Fleet Reality-Check" sectie hierboven.
Fleet Health Report Template
🌐 Weblyfe Tailnet — [DATE]
### ✅ Actieve Bots
appie-1 | 100.101.29.56 | macOS | ✅
appie-2 | 100.118.143.10 | Linux | ✅
### 💤 Gearchiveerd
spark-atlas — offline >30d
eva — offline >30d
### 🔴 Offline (overige nodes)
macbook-air-van-lorenzo — last seen 2026-04-10
Quick Fleet Scan Pattern (Telegram / on-demand)
When Seyed asks for a quick fleet analysis, keep it fast and operational rather than running the full daily scan:
- Local health first:
hostname,date -Is,uptime,df -h /,free -m, Hermes gateway service status, failed SSH logins in 24h, pending apt update count. - Tailnet state via JSON: parse
tailscale status --json; includeSelfandPeerseparately. Summarize online/offline, OS, Tailscale IP, and last-seen age. - Expected-offline context matters: if Seyed says a device is physically offline/in suitcase, mark it
expected/no action, not an incident. Still include last-seen time. - Reachability probes: for online/core nodes, run
tailscale pingand quick TCP checks for22,80,443,3000,5000,8000,8080,8443,5900. - SSH auth check: test key-based SSH separately from port reachability. Report
port 22 open but key rejectedas an access/remediation issue, not node downtime. - Service fingerprints: for open HTTP(S) ports, use
curl -skIand a short body/title probe. On macOS,5000often fingerprints as AirTunes and5900is Screen Sharing/VNC, both tailnet-accessible rather than internet-facing. - Report concise risk classes:
healthy,expected offline,access issue,tailnet-exposed service,ghost/offline >30d. End with the next 1-3 recommended actions.
Avoid over-escalating offline personal devices, iOS nodes, or user-declared travel/suitcase devices. The security signal is unexpected offline core infrastructure, key rejection on nodes that should be manageable, public DNS/service failure, or exposed services that should not be tailnet-wide.
Telegram Fleet Output — Grouped by Status Category
The daily scan Telegram output groups all fleet nodes by their connectivity status for at-a-glance readability. Implemented via FLEET_CACHE pipe-separated file and a case statement that tracks group transitions:
*Fleet*
SSH ✅
🟢 appie-2
🟢 eugi
SSH key rejected
🟡 appie-mc-1
🟡 appie-1
🟡 eva
🟡 harry
Online (geen SSH)
⚪ rabi
⚪ techwiz-mbp
⚪ ipad-pro
Offline
💤 spark-atlas
💤 artemis
...
Implementation pattern in the script's Telegram generation block:
current_group=""
while IFS="|" read -r name status desc; do
case "$status" in
"🟢") group="SSH ✅" ;;
"🔴") group="SSH ❌" ;;
"🟡") group="SSH key rejected" ;;
"⚪") group="Online (geen SSH)" ;;
"💤") group="Offline" ;;
*) group="" ;;
esac
if [ "$group" != "$current_group" ] && [ -n "$group" ]; then
[ -n "$current_group" ] && tg_echo ""
tg_echo " $group"
current_group="$group"
fi
tg_echo " $status $name"
done < "$FLEET_CACHE"
Key constraint: ALL 4 array types (FLEET, BROKEN_KEYS, TAILNET_ONLINE, GHOSTS) must write to the same FLEET_CACHE in array order. The Telegram block reads once sequentially. This guarantees the Telegram output matches the report and no nodes are missed.
Cache format (pipe-separated, no JSON):
appie-2|🟢|Weblyfe VPS (Hetzner)
eugi|🟢|Ubuntu VPS (FSN)
appie-mc-1|🟡|Mission Control v1 — key rejected
rabi|⚪|Mac Mini Rabi — Remote Login off
spark-atlas|💤|DGX Spark — offline 9d
Watchdog Architecture (separate cron jobs, no_agent=True)
These run independently from the daily scan to catch problems early. Script-only mode: zero LLM tokens per tick, silent when healthy.
| Watchdog | Schedule | Script | Trigger |
|---|---|---|---|
| Disk | Every 4h | disk-watchdog.sh |
>80% ⚠️, >90% 🔴 on any fleet node |
| Tailscale | Every 30min | tailscale-watchdog.sh |
Core node (appie-1, appie-2, eugi) goes offline |
| SSH key audit | 1st of month | ssh-key-audit.sh |
Monthly authorized_keys review per node |
| Failed login trend | In daily scan | failed-login-trend.sh |
Spike >5x yesterday or >10/day |
| Agent framework CVEs (Python + Go) | Daily 06:30 UTC | agent-framework-cve-scan.py (via LLM cron job c503abd02155) |
Any CVE reported in last 90d for 20 packages (17 Python agent + 3 Go infra). Uses [SILENT] when clean. |
Cron job setup:
# Scripts must be at ~/.hermes/scripts/ (or the profile-specific scripts dir)
# Symlinks may be blocked by the cron scheduler — use real files
cronjob action=create name="disk-watchdog" schedule="0 */4 * * *" \
no_agent=True script="disk-watchdog.sh" deliver="telegram:1817919454"
Security Governance & Suggestion Pipeline
A multi-agent governance pattern for human-in-the-loop security improvements. Appie-3 (CTO specialist) creates the suggestion plan, Appie-Opus (orchestrator) presents suggestions to Seyed, Seyed approves before execution.
When to Use
- After the daily security scan finds issues that need remediation
- User asks for "daily security suggestions", "action plan for opus", "suggestions to present to seyed"
- Setting up a security governance workflow where changes require human approval
- Any recurring task where the CTO recommends but the user decides
Architecture
Appie-3 (CTO) Appie-Opus (orchestrator) Seyed (human)
│ │ │
│ Creates plan ──────────────►│ │
│ DAILY-SECURITY- │ │
│ SUGGESTIONS.md │ │
│ │ Presents S001-S00X ──►│
│ │ "Wat, waarom, │
│ │ risico, revert" │
│ │◄─────── 👍 / 👎 ──────│
│ │ │
│ (Appie-3 only acts │ Executes approved │
│ when directly asked) │ suggestions │
Suggestion Format
Each suggestion MUST include four fields:
| Field | Description |
|---|---|
| Wat | What the check/change is (1-2 zinnen) |
| Waarom | Why it matters (1 zin) |
| Risico | Risk level: Geen / Laag / Medium |
| Revert | Exact command to undo the change |
Safety Rules (for Appie-Opus)
- Read-only or reversible only — no
rm -rf,chmod -R,iptables -F - STOP on error — report to Seyed, do not self-fix
- No lockouts — always test key-based login before disabling password auth
- No service restarts without explicit Seyed approval
- Every suggestion has a revert step — if missing, do not execute
Suggestion Count
S001-S027 exist (expanded 2026-06-20 met S021-S027: agent framework CVE scan, Headroom install/verify/MCP, SkillsGuard cloud API, CVE-2026-2256 check, MCP cutover check, SSH key deployment plan).
Daily Rotation
27 suggestions exist (S001-S027), rotated by day:
| Dag | Category | Sample Suggestions |
|---|---|---|
| Maandag | Access audit | authorized_keys check (appie-2 + eugi), tailnet status, S027 (SSH key plan) |
| Dinsdag | Config check | .env perms, open ports on fleet, S025 (CVE-2026-2256 check) |
| Woensdag | Dependency scan | gitleaks, npm audit, S021 (agent framework CVE scan), S026 (MCP cutover) |
| Donderdag | Web security | SSL certs, security headers, failed logins, S024 (SkillsGuard) |
| Vrijdag | System health | disk, pending updates, S022-S023 (Headroom check/MCP) |
| Weekend | Fleet check | tailnet status, spark-atlas SSH, gateway uptime |
Plan Location
Canonical plan: ~/clawd/appie-3-cto/DAILY-SECURITY-SUGGESTIONS.md
Zie references/security-suggestions.md voor de volledige, actuele suggestielijst.
Backup Sync (optional, needs approval)
backup-sync.sh — rsyncs appie-brain, Hermes config, SSH config, and secrets to appie-2.
Not active by default. Script exists at ~/.hermes-appie3/scripts/backup-sync.sh.
Cross-Platform Detection (macOS vs Linux)
The daily scan script may run on macOS (Mac Mini) or Linux (Hetzner VPS). Use an uname check at the top to branch OS-specific commands:
OS="$(uname -s)"
case "$OS" in
Darwin) IS_MACOS=true; IS_LINUX=false ;;
Linux) IS_MACOS=false; IS_LINUX=true ;;
esac
Command Equivalents
| macOS Command | Linux Alternative | When to Use |
|---|---|---|
vm_stat |
free or /proc/meminfo |
Memory check |
sysctl -n vm.loadavg |
cat /proc/loadavg |
Load averages |
log show --predicate ... --last 24h |
journalctl --since "24 hours ago" | grep -c "Failed password" |
Failed login count |
date -j -f "%b %d %H:%M:%S %Y %Z" "$date" +%s |
date -d "$date" +%s |
Date parsing (SSL cert expiry) |
date -v-1d +%Y-%m-%d |
date -d "yesterday" +%Y-%m-%d |
…(truncated)