/skill-ops v1.2
Skill/agent ops hub — 3 modes: snapshot, usage, frequency. Merges skill-versioning + skill-health-report.
Dominant Variable
Snapshot integrity + invocation log completeness — if a snapshot doesn't match the original, rollback is meaningless. Without logs, usage analysis is impossible.
Key Assumptions
- Permission to create
~/.claude/.harness/snapshots/— if broken: report permission issue + provide manual mkdir command. - Target file is
~/.claude/skills/*/SKILL.mdor~/.claude/agents/*.md— if broken: ask for the skill name directly. - Retention policy: keep last 5 — 6th and older are deleted oldest-first. Ignore if deletion fails.
- Invocation logs: session-checkpoint Phase 3.7 appends to
invocations/YYYY-MM.jsonl— if broken: state "no logs". - SKILLS/AGENTS_INVENTORY.md is the source of truth — if broken: analyze from log-derived names only (mark incomplete).
Trigger
/skill-ops(snapshot default)/skill-ops health/skill-ops invocations- "skill version", "skill usage frequency", "harness score"
Discard If
- Target file is outside
~/.claude/skills/or~/.claude/agents/(project code → git handles it) - Target file doesn't exist (nothing to snapshot)
- A same-day snapshot with an identical SHA-256 content hash already exists (duplicate is pointless)
- Health mode:
invocations/directory itself doesn't exist → report "no logs" - Health mode: 0 JSONL files within scan range → report "no logs in range"
Mode
| Mode | Role | Trigger |
|---|---|---|
| snapshot | Pre-change snapshot + regression-detection restore command | /skill-ops (default) |
| health | Usage/Dead/Unused/Discard report | /skill-ops health |
| invocations | Per-skill frequency rollup from session JSONL | /skill-ops invocations |
Snapshot Mode
Phase 0: Parse Input
- Extract target file path from user input (absolute path > skill name > user question)
- Extract skill name:
~/.claude/skills/<name>/SKILL.mdor~/.claude/agents/<name>.md
Phase 1: Read Original
[READ] {TARGET_FILE}→ compute its SHA-256 content hash (ORIGINAL_HASH)- Failure (file missing) → "target file not found" + end with BROKEN status
Phase 2: Prepare Directory + Same-Day Duplicate Check
TIMESTAMP=$(date +%Y-%m-%d-%H%M%S)
SKILL_SNAP_DIR=~/.claude/.harness/snapshots/{skill-name}
ORIGINAL_HASH=$(sha256sum "${TARGET_FILE}" | cut -d' ' -f1)
mkdir -p ${SKILL_SNAP_DIR}/${TIMESTAMP}
- Same-day snapshot exists + its SHA-256 hash equals
ORIGINAL_HASH→ "identical-content snapshot already exists" → go to Phase 6 (equal line counts alone do NOT count as identical — always compare hashes)
Phase 3: Save + Verify (Invariant #2)
[WRITE]snapshot →[READ]re-verify → compare its SHA-256 hash againstORIGINAL_HASH- Mismatch →
⚠️ Snapshot verification failed+ end with PARTIAL
Phase 4: Clean Up Old Snapshots (delete beyond 5)
SNAP_DIR=~/.claude/.harness/snapshots/{skill-name}
COUNT=$(find "${SNAP_DIR}" -mindepth 1 -maxdepth 1 -type d 2>/dev/null | wc -l)
if [ "${COUNT}" -gt 5 ]; then
find "${SNAP_DIR}" -mindepth 1 -maxdepth 1 -type d | sort | head -n "$((COUNT-5))" | while IFS= read -r path; do
[ -n "${path}" ] && rm -rf -- "${path}"
done
fi
COUNTis computed explicitly (previously undefined) and cleanup is skipped entirely whenCOUNT≤ 5, sohead -nnever receives a zero/negative argument.- The
while IFS= read -r pathloop replacesxargs—xargs' default whitespace-delimited splitting mishandles snapshot paths containing spaces, whileread -rconsumes each line whole. [ -n "${path}" ]guards against an empty line reachingrm -rf.
Phase 5: Show Prior Score Store Score
- Extract
harness_score(0-100 scale — check-harness's project/user-level aggregate score; a different schema from this skill's own 0-10S_Qmetric in Quality Mode below) +datefrom the latest~/.claude/.harness/scores/*.jsonfile - If none: "No Score Store — run a quality audit first to have something to compare against"
Phase 6: Output Rollback Command
Rollback: cp ~/.claude/.harness/snapshots/{skill-name}/{TIMESTAMP}/SKILL.md {TARGET_FILE}
List: ls ~/.claude/.harness/snapshots/{skill-name}/
Health Mode
Phase 0: Verify
- Confirm
invocations/directory exists - Load SKILLS_INVENTORY.md / AGENTS_INVENTORY.md
Phase 1: Parse Invocation Logs + Correction History
Per-skill rollup from JSONL over the last 30 days (default):
skills[]→ invocation countdiscarded[]→ Discard If trigger countlast_seen→ last invocation date- Retained snapshot count: number of snapshot directories currently under
~/.claude/.harness/snapshots/{skill}/— NOT the skill's true cumulative edit count, since Invariant 4 caps retention at 5 (a skill edited more than 5 times still shows at most 5 here). A count at or near the 5-snapshot cap is a stability-watch signal (frequent recent edits)
Phase 2: Classify Status (deterministic)
Don't eyeball this against the criteria table — call the bucket classifier per skill instead:
python scripts/skill_health_bucket.py bucket --count {invocation_count_30d} --last-seen {last_seen_date} --discard-rate {discard_rate}
--last-seen and --discard-rate come from the Phase 1 rollup. The script is the source of truth for the label; the table below is reference only, for reading the output — not for manually re-deriving it.
| Status | Criterion | Label |
|---|---|---|
| Active | ≥ threshold (2x/30d) within range | 🟢 |
| Low | ≥1x, below threshold | 🟡 |
| Unused | 0x, under 90 days | 🔴 |
| Dead | 0x, 90+ days | 💀 |
| Discarded | Discard If triggered only | ⚪ |
| Unknown | No logs | ❓ |
Unused + no recent edits is a retire-candidate signal.
Skill-bank alignment signal: a skill bank that has drifted out of alignment with your current goals or workflow can underperform having no skill bank at all. Treat Unused and clearly-misaligned skills as a stronger retire-candidate signal than either alone.
Retirement-judge audit gate: before wiring Health mode's Dead/Unused/Discard classifications into any automated delete/archive pipeline, validate the classifier itself — deliberately include a few known-good (still-needed) skills in the candidate pool and check whether the classifier still flags them (false positives). A high false-positive rate means the retirement mechanism looks like it's working but silently isn't. Until that validation exists, this mode stays report-only — deletion is always the user's call (Invariant 5).
Phase 3: Generate + Save Report
📊 Skill Health — {YYYY-MM-DD}
🟢 Active {N} | 🟡 Low {N} | 🔴 Unused {N} | 💀 Dead {N}
[Full status table]
[💀 Dead — recommend immediate review]
[⚪ Discard ratio >30% warning]
Save: ~/.claude/.harness/reports/skill-health-{date}.md
Invocations Mode
Phase 7: Invocation Frequency Scan
Aggregate Skill tool calls from session JSONL to measure per-skill monthly invocation frequency.
- tool_use metadata only — never read prompt text
- Windows: use
python/python3on PATH; if neither resolves, check common install locations before failing - Save:
~/.claude/.harness/invocations/YYYY-MM.json
Output:
📊 Skill Invocation Report (YYYY-MM)
Top 5: [most invoked]
Zero-invocation: [never-invoked list — SHARPEN candidates]
Quality Mode
Trigger:
/skill-ops qualityor/skill-ops --quality
Purpose
Calculate a per-skill quality score (S_Q, 0-10 scale — distinct from the 0-100 harness_score in Snapshot Mode's Score Store) and identify the bottom quartile as optimization targets.
⚠️ Boundary: S_Q is an operational signal for "keep vs. retire this skill" — not a quality oracle. It measures usage plus a handful of structural checklist items, not whether the skill's content is actually good. Don't read a low S_Q as "this skill is badly written" — it may simply be under-used. Deep content-quality review of a skill's actual reasoning/instructions is a separate activity outside this skill's scope (see
not_forabove).
Phase 8: Quality Score Scan (deterministic)
Structure score, usage score, and their sum are computed by the script — never re-derive them by reading the checklist and eyeballing points. The bullets below are what each score means, not steps to apply by hand.
- Load skill list:
~/.claude/skills/*/SKILL.md+SKILLS_INVENTORY.md - Structure score (0-5):
Checks, +1 each: Dominant Variable present · Discard If present · Invariants has a violation-consequence clause · Scope Boundary has 2+ rows on each side · Rationalization Table has 3+ rows.python scripts/skill_health_bucket.py structural --file <path to SKILL.md> - Usage score (0-5):
Weights: 5+ invocations in 30 days (+2) / 1-4 (+1) / 0 (0) · Discard If trigger rate < 30% (+1) · last modified within 30 days (+1) or within 90 days (+0.5) · related lesson exists (correction history = usage evidence) (+0.5).python scripts/skill_health_bucket.py usage --invocation-count-30d {N} --discard-rate {F} \ --days-since-modified {N} [--has-related-lesson] - S_Q = structure + usage (0-10):
python scripts/skill_health_bucket.py sq --structural {F} --usage {F} - Bottom 25% = optimization targets. Top 75% = keep as-is.
Output
📊 Skill Quality Report (YYYY-MM-DD)
S_Q ≥ 7: {N} (STRONG)
S_Q 4-6: {N} (ADEQUATE)
S_Q < 4: {N} (OPTIMIZE) ← bottom quartile
[OPTIMIZE target table: skill name | structure | usage | S_Q | 1-line improvement direction]
Save: ~/.claude/.harness/reports/skill-quality-{date}.md
Scope Boundary
| Does | Does NOT |
|---|---|
| [READ] Read the original snapshot target file | Directly modify skill/agent files |
| [WRITE] Save timestamped snapshot file | Execute automatic restoration (proposal only) |
| [BASH] Delete old snapshots (beyond 5) | Upload to external storage/cloud |
| [READ] Check prior Score Store score | Run a quality audit itself |
| [READ] Parse invocations JSONL (tool_use only) | Read session prompt text |
| [WRITE] health report / invocations JSON | Judge skill quality or decide deletion |
| [BASH] Scan session JSONL for frequency rollup | Access project code or databases |
[BASH] Call scripts/skill_health_bucket.py for bucket/structural/usage/S_Q scoring |
Manually re-derive those scores by eye |
Targets only
~/.claude/global skills/agents. Project code version control is git's job.
Safety Layers
| Risky Action | Reversibility | Applied Layers |
|---|---|---|
Delete old snapshots (rm -rf) |
medium | L1+L3 |
| Roll back a skill file (Write overwrite) | medium | L1+L3 |
- L1 (Invariants): mandatory SHA-256 hash re-verification after save. No automatic restoration.
- L3 (User Approval): deletion only after explicit user request. Rollback only after stating "current→rollback" and getting user confirmation.
Error Recovery
| Failure Type | Detection | Recovery |
|---|---|---|
tool_failure |
Write/Read failure | State "snapshot save failed". Never proceed with comparison without a snapshot |
logic_inconsistency |
harness_score DELTA (0-100 scale, from Phase 5's Score Store — not the 0-10 S_Q scale below) ≤ -5 but content actually improved |
State "possible false positive" + ask user to re-review |
missing_data |
Target file missing / invocations log missing | Discard that mode + state the reason |
input_error |
Target skill unclear | Default to full-list scan. If specific target intended, ask 1 clarifying question |
Invariants (never violate)
- Confirm original exists before snapshotting: Write only after successful Read. Abort if original is missing. Violation → empty snapshot.
- Re-verify Read after Write: SHA-256 hash mismatch → PARTIAL. Violation → reporting a corrupted snapshot as "done".
- No automatic restoration: only output the restore
cpcommand. Execution is the user's job. Violation → unintended file overwrite. - Keep last 5: delete 6th and beyond. Violation → unbounded directory growth.
- No automatic deletion (Health): never delete/move files even at 0 usage. Report only. Violation → No Action default violation.
- No logs ≠ unused (Health): sessions that skipped session-checkpoint may still have been used despite missing logs. Treat as Unknown. Violation → truthful-reporting violation.
- Below threshold ≠ Dead (Health): Low (below threshold) and Dead (0x for 90+ days) are distinct. Violation → misclassifying an in-use skill.
- Bucket/structure/usage/S_Q scores are computed via
scripts/skill_health_bucket.py, never eyeballed: counting is a job for the script, judgment (retire or not) stays with the user/LLM. Violation → scores drift silently between runs and stop being comparable.
Truthful Reporting
- no mock deception: never say "save complete" without a post-Write Read re-verification. Never assume "used" from absent logs.
- no test façade: SHA-256 hash mismatch = PARTIAL. Never assume "it probably worked".
- no silent brokenness: final status must be labeled
WORKING/PARTIAL/BROKEN.
Rationalization Table
| Rationalization | Rebuttal |
|---|---|
| "Skipping the re-verify after Write is fine if it succeeded" | Violates Invariant 2. A silent Write failure means rollback is attempted without a real snapshot |
| "Auto-restore would be more convenient" | Violates Invariant 3. If the user restores without understanding the regression cause, the root cause remains |
| "Snapshots older than 90 days can just stay" | Slows Glob traversal + wastes space. 90-day cleanup happens via session-checkpoint guidance, after user approval |
| "Skills at 0 usage can be auto-deleted" | Violates Invariant 5. Could be emergency-only, seasonal, or recently added. User decides |
| "Months with no logs can just be treated as 0 invocations" | Violates Invariant 6. Must be treated as Unknown |
| "High Discard If ratio → recommend immediate retirement" | Related to Invariant 7. The safeguard may simply be working correctly. Propose re-review only |
| "The criteria table is simple enough to just eyeball" | Violates Invariant 8. Manual application drifts from the script's exact thresholds and regex logic — the same skill can score differently run to run |