Skill Auditor v3.1.0
The definitive security scanner for OpenClaw/ClawHub skills. Best-in-class detection across 18 security checks including prompt injection detection — the first scanner to catch agent manipulation attacks in skill documentation. 5-dimension trust scoring, trend tracking, diff analysis, and benchmarking. Zero false positives on legitimate skills.
When to Activate
- Installing a new skill from ClawHub - run
inspect.sh for full pre-install validation
- Auditing existing skills - use
audit.sh to scan any skill directory
- Generating trust scores - use
trust_score.py for 0-100 rating across 5 dimensions
- Comparing skills - use
trust_score.py --compare for side-by-side analysis
- Tracking improvements - use
trust_score.py --save-trend to monitor score over time
- Reviewing updates - use
diff-audit.sh to compare before/after versions
- Batch scanning - use
audit-all.sh or benchmark.sh for fleet-wide analysis
Quick Start
# Audit a single skill
bash audit.sh /path/to/skill
# Trust score (0-100 across 5 dimensions)
python3 trust_score.py /path/to/skill
# Compare two skills side by side
python3 trust_score.py /path/to/skill1 --compare /path/to/skill2
# Track score over time
python3 trust_score.py /path/to/skill --save-trend
python3 trust_score.py /path/to/skill --trend
# Diff audit (before/after update)
bash diff-audit.sh /path/to/old-version /path/to/new-version
# Benchmark against a corpus
bash benchmark.sh /path/to/skills-dir
# Inspect a ClawHub skill before installing
bash inspect.sh skill-slug
# Audit all installed skills
bash audit-all.sh
# Generate a markdown report
bash report.sh
# Run test suite (28 assertions)
bash test.sh
Guardrails / Anti-Patterns
DO:
- ✓ Always audit skills before installing from untrusted sources
- ✓ Review trust scores - reject skills scoring below 60 (D grade)
- ✓ Use
diff-audit.sh when updating skills to catch regressions
- ✓ Use
--json output for CI/CD pipeline integration
- ✓ Run
--save-trend periodically to track skill health
DON'T:
- ✗ Install skills scoring below 40 (F grade) without extensive manual review
- ✗ Ignore CRITICAL findings - they indicate potential security threats
- ✗ Blindly add skills to allowlist without understanding why they access credentials
- ✗ Skip audit because a skill is "popular" or "official"
Security Checks (18 total)
| # |
Check |
Severity |
Description |
| 1 |
credential-harvest |
CRITICAL |
Scripts reading API keys/tokens AND making network calls |
| 2 |
exfiltration-url |
CRITICAL |
webhook.site, requestbin, ngrok URLs in scripts |
| 3 |
obfuscated-payload |
CRITICAL |
Base64-encoded URLs or shell commands |
| 4 |
sensitive-fs |
CRITICAL |
/etc/passwd, ~/.ssh, ~/.aws/credentials access |
| 5 |
crypto-wallet |
CRITICAL |
Hardcoded ETH/BTC wallet addresses (drain attacks) |
| 6 |
dependency-confusion |
CRITICAL |
Internal/private-scoped packages in public deps |
| 7 |
typosquatting |
CRITICAL |
Misspelled package names (lodahs, requets, etc.) |
| 8 |
symlink-attack |
CRITICAL |
Symlinks targeting sensitive system paths |
| 9 |
code-execution |
WARNING |
eval(), exec(), subprocess patterns |
| 10 |
time-bomb |
WARNING |
Date/time comparisons that could trigger delayed payloads |
| 11 |
telemetry-detected |
WARNING |
Analytics SDKs, tracking pixels, phone-home behavior |
| 12 |
excessive-permissions |
WARNING |
>15 bins/env/config items requested |
| 13 |
unusual-ports |
WARNING |
Network calls to non-standard ports |
| 14 |
prompt-injection |
CRITICAL |
Agent manipulation in docs: "ignore instructions", role hijacking, hidden HTML directives |
| 15 |
download-execute |
CRITICAL |
curl|bash, wget|sh, eval $(curl), unsafe pip/npm installs |
| 16 |
hidden-file |
WARNING |
Suspicious dotfiles that may hide malicious content |
| 17 |
env-exfiltration |
CRITICAL |
Reading sensitive env vars + outbound network calls |
| 18 |
privilege-escalation |
CRITICAL |
sudo, chmod 777/setuid, writes to system paths |
Context-aware: credential mentions in documentation are INFO, not CRITICAL.
Trust Score (5 Dimensions)
| Dimension |
Max |
What's Measured |
| Security |
35 |
Audit findings (criticals = -18, warnings = -4) |
| Quality |
22 |
Description, version, usage docs, examples, metadata, changelog |
| Structure |
18 |
File organization, tests, README, reasonable scope |
| Transparency |
15 |
License, no minified code, code comments |
| Behavioral |
10 |
Rate limiting, error handling, input validation |
Grades: A (90+), B (75+), C (60+), D (40+), F (<40)
Comparative Scoring
python3 trust_score.py /path/to/skill-a --compare /path/to/skill-b
Shows per-dimension deltas and overall score difference.
Trend Tracking
python3 trust_score.py /path/to/skill --save-trend # Record score
python3 trust_score.py /path/to/skill --trend # View history
Stores up to 50 entries per skill in trust_trends.json.
Tools
| File |
Purpose |
| audit.sh |
Single skill security audit (18 checks) |
| audit-all.sh |
Batch scan all installed skills |
| trust_score.py |
Trust score calculator (5-dimension, 0-100) |
| diff-audit.sh |
Compare skill versions for security regressions |
| benchmark.sh |
Corpus-wide audit with aggregate statistics |
| inspect.sh |
ClawHub pre-install workflow |
| report.sh |
Markdown report generator |
| test.sh |
Automated test suite (28 assertions, 12 test skills) |
| allowlist.json |
Known-good credential skills |
Test Suite
12 test skills (8 malicious, 4 clean) with 28 automated assertions:
bash test.sh
Malicious fixtures: credential harvest, obfuscated payload, sensitive fs reads, crypto wallets, time bombs, symlink attacks, prompt injection, download-execute, privilege escalation.
Clean fixtures: basic skill, credential docs (false positive check), network skill, dotfiles skill.
Exit Codes
- 0: PASS / safe to install
- 1: REVIEW / warnings found
- 2: FAIL / critical issues
- 3: Error / bad input
Changelog
See CHANGELOG.md for full version history.
1---2name: skill-auditor-43description: The definitive security scanner for OpenClaw skills. 18 security checks including prompt injection detection, download-and-execute, privilege escalation, credential harvesting, supply chain attacks, crypto drains, and more. 5-dimension trust scoring with trend tracking.4---5
6# Skill Auditor v3.1.0
7
8The definitive security scanner for OpenClaw/ClawHub skills. Best-in-class detection across 18 security checks including **prompt injection detection** — the first scanner to catch agent manipulation attacks in skill documentation. 5-dimension trust scoring, trend tracking, diff analysis, and benchmarking. Zero false positives on legitimate skills.
9
10## When to Activate
11
121. **Installing a new skill** from ClawHub - run `inspect.sh` for full pre-install validation
132. **Auditing existing skills** - use `audit.sh` to scan any skill directory
143. **Generating trust scores** - use `trust_score.py` for 0-100 rating across 5 dimensions
154. **Comparing skills** - use `trust_score.py --compare` for side-by-side analysis
165. **Tracking improvements** - use `trust_score.py --save-trend` to monitor score over time
176. **Reviewing updates** - use `diff-audit.sh` to compare before/after versions
187. **Batch scanning** - use `audit-all.sh` or `benchmark.sh` for fleet-wide analysis
19
20## Quick Start
21
22```bash
23# Audit a single skill
24bash audit.sh /path/to/skill
25
26# Trust score (0-100 across 5 dimensions)
27python3 trust_score.py /path/to/skill
28
29# Compare two skills side by side
30python3 trust_score.py /path/to/skill1 --compare /path/to/skill2
31
32# Track score over time
33python3 trust_score.py /path/to/skill --save-trend
34python3 trust_score.py /path/to/skill --trend
35
36# Diff audit (before/after update)
37bash diff-audit.sh /path/to/old-version /path/to/new-version
38
39# Benchmark against a corpus
40bash benchmark.sh /path/to/skills-dir
41
42# Inspect a ClawHub skill before installing
43bash inspect.sh skill-slug
44
45# Audit all installed skills
46bash audit-all.sh
47
48# Generate a markdown report
49bash report.sh
50
51# Run test suite (28 assertions)
52bash test.sh
53```
54
55## Guardrails / Anti-Patterns
56
57**DO:**
58- ✓ Always audit skills before installing from untrusted sources
59- ✓ Review trust scores - reject skills scoring below 60 (D grade)
60- ✓ Use `diff-audit.sh` when updating skills to catch regressions
61- ✓ Use `--json` output for CI/CD pipeline integration
62- ✓ Run `--save-trend` periodically to track skill health
63
64**DON'T:**
65- ✗ Install skills scoring below 40 (F grade) without extensive manual review
66- ✗ Ignore CRITICAL findings - they indicate potential security threats
67- ✗ Blindly add skills to allowlist without understanding why they access credentials
68- ✗ Skip audit because a skill is "popular" or "official"
69
70## Security Checks (18 total)
71
72| # | Check | Severity | Description |
73|---|-------|----------|-------------|
74| 1 | credential-harvest | CRITICAL | Scripts reading API keys/tokens AND making network calls |
75| 2 | exfiltration-url | CRITICAL | webhook.site, requestbin, ngrok URLs in scripts |
76| 3 | obfuscated-payload | CRITICAL | Base64-encoded URLs or shell commands |
77| 4 | sensitive-fs | CRITICAL | /etc/passwd, ~/.ssh, ~/.aws/credentials access |
78| 5 | crypto-wallet | CRITICAL | Hardcoded ETH/BTC wallet addresses (drain attacks) |
79| 6 | dependency-confusion | CRITICAL | Internal/private-scoped packages in public deps |
80| 7 | typosquatting | CRITICAL | Misspelled package names (lodahs, requets, etc.) |
81| 8 | symlink-attack | CRITICAL | Symlinks targeting sensitive system paths |
82| 9 | code-execution | WARNING | eval(), exec(), subprocess patterns |
83| 10 | time-bomb | WARNING | Date/time comparisons that could trigger delayed payloads |
84| 11 | telemetry-detected | WARNING | Analytics SDKs, tracking pixels, phone-home behavior |
85| 12 | excessive-permissions | WARNING | >15 bins/env/config items requested |
86| 13 | unusual-ports | WARNING | Network calls to non-standard ports |
87| 14 | prompt-injection | CRITICAL | Agent manipulation in docs: "ignore instructions", role hijacking, hidden HTML directives |
88| 15 | download-execute | CRITICAL | curl\|bash, wget\|sh, eval $(curl), unsafe pip/npm installs |
89| 16 | hidden-file | WARNING | Suspicious dotfiles that may hide malicious content |
90| 17 | env-exfiltration | CRITICAL | Reading sensitive env vars + outbound network calls |
91| 18 | privilege-escalation | CRITICAL | sudo, chmod 777/setuid, writes to system paths |
92
93Context-aware: credential mentions in documentation are INFO, not CRITICAL.
94
95## Trust Score (5 Dimensions)
96
97| Dimension | Max | What's Measured |
98|-----------|-----|-----------------|
99| Security | 35 | Audit findings (criticals = -18, warnings = -4) |
100| Quality | 22 | Description, version, usage docs, examples, metadata, changelog |
101| Structure | 18 | File organization, tests, README, reasonable scope |
102| Transparency | 15 | License, no minified code, code comments |
103| Behavioral | 10 | Rate limiting, error handling, input validation |
104
105Grades: A (90+), B (75+), C (60+), D (40+), F (<40)
106
107### Comparative Scoring
108```bash
109python3 trust_score.py /path/to/skill-a --compare /path/to/skill-b
110```
111Shows per-dimension deltas and overall score difference.
112
113### Trend Tracking
114```bash
115python3 trust_score.py /path/to/skill --save-trend # Record score
116python3 trust_score.py /path/to/skill --trend # View history
117```
118Stores up to 50 entries per skill in `trust_trends.json`.
119
120## Tools
121
122| File | Purpose |
123|------|---------|
124| audit.sh | Single skill security audit (18 checks) |
125| audit-all.sh | Batch scan all installed skills |
126| trust_score.py | Trust score calculator (5-dimension, 0-100) |
127| diff-audit.sh | Compare skill versions for security regressions |
128| benchmark.sh | Corpus-wide audit with aggregate statistics |
129| inspect.sh | ClawHub pre-install workflow |
130| report.sh | Markdown report generator |
131| test.sh | Automated test suite (28 assertions, 12 test skills) |
132| allowlist.json | Known-good credential skills |
133
134## Test Suite
135
13612 test skills (8 malicious, 4 clean) with 28 automated assertions:
137
138```bash
139bash test.sh
140```
141
142Malicious fixtures: credential harvest, obfuscated payload, sensitive fs reads, crypto wallets, time bombs, symlink attacks, prompt injection, download-execute, privilege escalation.
143Clean fixtures: basic skill, credential docs (false positive check), network skill, dotfiles skill.
144
145## Exit Codes
146- 0: PASS / safe to install
147- 1: REVIEW / warnings found
148- 2: FAIL / critical issues
149- 3: Error / bad input
150
151## Changelog
152
153See [CHANGELOG.md](CHANGELOG.md) for full version history.