Agent Audit
Scan your entire OpenClaw setup and get actionable cost/performance recommendations.
What This Skill Does
- Scans config — reads OpenClaw config to map models to agents/tasks
- Analyzes cron history — checks every cron job's model, token usage, runtime, success rate
- Classifies tasks — determines complexity level of each task
- Calculates costs — per agent, per cron, per task type using provider pricing
- Recommends changes — with confidence levels and risk warnings
- Generates report — markdown report with specific savings estimates
Running the Audit
python3 {baseDir}/scripts/audit.py
Options:
python3 {baseDir}/scripts/audit.py --format markdown # Full report (default)
python3 {baseDir}/scripts/audit.py --format summary # Quick summary only
python3 {baseDir}/scripts/audit.py --dry-run # Show what would be analyzed
python3 {baseDir}/scripts/audit.py --output /path/to/report.md # Save to file
How It Works
Phase 1: Discovery
- Read OpenClaw config (
~/.openclaw/openclaw.json or similar)
- List all cron jobs and their configurations
- List all agents and their default models
- Detect provider (Anthropic, OpenAI, Google, xAI) from model names
Phase 2: History Analysis
- Pull cron job run history (last 7 days by default)
- Calculate per-job: avg tokens, avg runtime, success rate, model used
- Pull session history where available
- Calculate total token spend by model tier
Phase 3: Task Classification
Classify each task into complexity tiers:
| Tier |
Examples |
Recommended Models |
| Simple |
Health checks, status reports, reminders, notifications |
Cheapest tier (Haiku, GPT-4o-mini, Flash, Grok-mini) |
| Medium |
Content drafts, research, summarization, data analysis |
Mid tier (Sonnet, GPT-4o, Pro, Grok) |
| Complex |
Coding, architecture, security review, nuanced writing |
Top tier (Opus, GPT-4.5, Ultra, Grok-2) |
Classification signals:
- Simple: Short output (<500 tokens), low thinking requirement, repetitive pattern, status/health tasks
- Medium: Medium output, some reasoning needed, creative but templated, research tasks
- Complex: Long output, multi-step reasoning, code generation, security-critical, tasks that previously failed on weaker models
Phase 4: Recommendations
For each task where the model tier doesn't match complexity:
⚠️ RECOMMENDATION: Downgrade "Knox Bot Health Check" from opus to haiku
Current: anthropic/claude-opus-4 ($15/M input, $75/M output)
Suggested: anthropic/claude-haiku ($0.25/M input, $1.25/M output)
Reason: Simple status check averaging 300 output tokens
Estimated savings: $X.XX/month
Risk: LOW — task is simple pattern matching
Confidence: HIGH
Safety Rules — NEVER Recommend Downgrading:
- Coding/development tasks
- Security reviews or audits
- Tasks that have previously failed on weaker models
- Tasks where the user explicitly chose a higher model
- Complex multi-step reasoning tasks
- Anything the user flagged as critical
Phase 5: Report Generation
Output a clean markdown report with:
- Overview — total agents, crons, monthly spend estimate
- Per-agent breakdown — model, usage, cost
- Per-cron breakdown — model, frequency, avg tokens, cost
- Recommendations — sorted by savings potential
- Total potential savings — monthly estimate
- One-liner config changes — exact model strings to swap
Model Pricing Reference
See references/model-pricing.md for current pricing across all providers.
Update this file when prices change.
Task Classification Details
See references/task-classification.md for detailed heuristics
on how tasks are classified into complexity tiers.
Important Notes
- This skill is read-only — it never changes your config automatically
- All recommendations include risk levels and confidence scores
- When unsure about a task's complexity, it defaults to keeping the current model
- The audit should be re-run periodically (monthly) as usage patterns change
- Token counts are estimates based on cron history — actual costs depend on your provider's billing
1---2name: agent-audit3description: Audit your AI agent setup for performance, cost, and ROI. Scans OpenClaw config, cron jobs, session history, and model usage to find waste and recommend optimizations. Works with any model provider (Anthropic, OpenAI, Google, xAI, etc.). Use when: (1) user says "audit my agents", "optimize my costs", "am I overspending on AI", "check my model usage", "agent audit", "cost optimization", (2) user wants to know which cron jobs are expensive vs cheap, (3) user wants model-task fit recommendations, (4) user wants ROI analysis of their agent setup, (5) user says "where am I wasting tokens".4---5
6# Agent Audit
7
8Scan your entire OpenClaw setup and get actionable cost/performance recommendations.
9
10## What This Skill Does
11
121. **Scans config** — reads OpenClaw config to map models to agents/tasks
132. **Analyzes cron history** — checks every cron job's model, token usage, runtime, success rate
143. **Classifies tasks** — determines complexity level of each task
154. **Calculates costs** — per agent, per cron, per task type using provider pricing
165. **Recommends changes** — with confidence levels and risk warnings
176. **Generates report** — markdown report with specific savings estimates
18
19## Running the Audit
20
21```bash
22python3 {baseDir}/scripts/audit.py
23```
24
25Options:
26```bash
27python3 {baseDir}/scripts/audit.py --format markdown # Full report (default)
28python3 {baseDir}/scripts/audit.py --format summary # Quick summary only
29python3 {baseDir}/scripts/audit.py --dry-run # Show what would be analyzed
30python3 {baseDir}/scripts/audit.py --output /path/to/report.md # Save to file
31```
32
33## How It Works
34
35### Phase 1: Discovery
36- Read OpenClaw config (`~/.openclaw/openclaw.json` or similar)
37- List all cron jobs and their configurations
38- List all agents and their default models
39- Detect provider (Anthropic, OpenAI, Google, xAI) from model names
40
41### Phase 2: History Analysis
42- Pull cron job run history (last 7 days by default)
43- Calculate per-job: avg tokens, avg runtime, success rate, model used
44- Pull session history where available
45- Calculate total token spend by model tier
46
47### Phase 3: Task Classification
48Classify each task into complexity tiers:
49
50| Tier | Examples | Recommended Models |
51|------|----------|-------------------|
52| **Simple** | Health checks, status reports, reminders, notifications | Cheapest tier (Haiku, GPT-4o-mini, Flash, Grok-mini) |
53| **Medium** | Content drafts, research, summarization, data analysis | Mid tier (Sonnet, GPT-4o, Pro, Grok) |
54| **Complex** | Coding, architecture, security review, nuanced writing | Top tier (Opus, GPT-4.5, Ultra, Grok-2) |
55
56Classification signals:
57- **Simple**: Short output (<500 tokens), low thinking requirement, repetitive pattern, status/health tasks
58- **Medium**: Medium output, some reasoning needed, creative but templated, research tasks
59- **Complex**: Long output, multi-step reasoning, code generation, security-critical, tasks that previously failed on weaker models
60
61### Phase 4: Recommendations
62For each task where the model tier doesn't match complexity:
63
64```
65⚠️ RECOMMENDATION: Downgrade "Knox Bot Health Check" from opus to haiku
66 Current: anthropic/claude-opus-4 ($15/M input, $75/M output)
67 Suggested: anthropic/claude-haiku ($0.25/M input, $1.25/M output)
68 Reason: Simple status check averaging 300 output tokens
69 Estimated savings: $X.XX/month
70 Risk: LOW — task is simple pattern matching
71 Confidence: HIGH
72```
73
74### Safety Rules — NEVER Recommend Downgrading:
75- Coding/development tasks
76- Security reviews or audits
77- Tasks that have previously failed on weaker models
78- Tasks where the user explicitly chose a higher model
79- Complex multi-step reasoning tasks
80- Anything the user flagged as critical
81
82### Phase 5: Report Generation
83Output a clean markdown report with:
841. **Overview** — total agents, crons, monthly spend estimate
852. **Per-agent breakdown** — model, usage, cost
863. **Per-cron breakdown** — model, frequency, avg tokens, cost
874. **Recommendations** — sorted by savings potential
885. **Total potential savings** — monthly estimate
896. **One-liner config changes** — exact model strings to swap
90
91## Model Pricing Reference
92
93See [references/model-pricing.md](references/model-pricing.md) for current pricing across all providers.
94Update this file when prices change.
95
96## Task Classification Details
97
98See [references/task-classification.md](references/task-classification.md) for detailed heuristics
99on how tasks are classified into complexity tiers.
100
101## Important Notes
102
103- This skill is **read-only** — it never changes your config automatically
104- All recommendations include risk levels and confidence scores
105- When unsure about a task's complexity, it defaults to keeping the current model
106- The audit should be re-run periodically (monthly) as usage patterns change
107- Token counts are estimates based on cron history — actual costs depend on your provider's billing