Session Efficiency Audit
You are the reasoning layer (L2/L3) of a layered audit engine. Bundled deterministic scripts (L0/L1) digest raw session logs into metadata-only artifacts; you reason over aggregates and escalate to content only through a budget-capped fetch interface.
{base_directory} below refers to this skill's directory (the convention used by the runtime to inject script paths). Set the workdir once and use it everywhere:
export AUDIT_WORKDIR=<session scratchpad>/audit_workdir # working artifacts — ephemeral
Two lifetimes. The JSON artifacts are intermediates and belong in the ephemeral workdir. The report is the deliverable and must outlive the session — it goes to a dated file in a stable, project-independent archive:
mkdir -p ~/.claude/audit-reports
# final report -> ~/.claude/audit-reports/<YYYY-MM-DD>.md
Keeping past reports is also the only way to answer "am I improving?": cross-run persistence is not implemented (I7), so trend comes from comparing today's report to the archive, not from the data layer.
Invariants (non-negotiable)
The list below is canonical. Each item is prefixed with its invariant id (I1–I8, with I7 intentionally omitted — cross-run persistence is not implemented; do not promise trends across runs).
- (I1) Never read a raw session JSONL directly (no Read/cat/grep on
~/.claude/projects/**/*.jsonl). All access goes throughrunandfetch. - (I2) Reason over metadata by default; fetch content only to resolve a specific hypothesis,
--max-bytes ≤ 2000, at most ~10 fetches per audit. - (I3) Waste math never trusts raw
input_tokens/output_tokenssums (streaming placeholders / thinking excluded). Anchor oncache_read_input_tokens,cache_creation_input_tokens, and content bytes. - (I4) All aggregation is deduped by
requestId(the scripts do this — do not hand-roll parsing). - (I5) The active session is self-excluded by the runner; never audit the session you are running in.
- (I6) Report the audit's own self-cost from
fetch_log.jsonlat the end — and label it a lower bound: it counts escalation fetches only, not the Phase 1 landscape or your own reasoning turns. - (I8) Write only to
$AUDIT_WORKDIRand the report. Never edit this skill's own files (SKILL.md,bin/,src/,data/) or the user's repo. You are an installed bundle: a relative path likesrc/views.jspoints at your own running code, not at the project that built you. An audit that patches its own renderer measures itself with something that changed mid-run, and the next install destroys the change without a word.
Phase 0 — Digest (L0 + L1, zero LLM)
node {base_directory}/bin/audit.js run # full directory
node {base_directory}/bin/audit.js run --max 50 # quick pass, newest 50 sessions
Writes to $AUDIT_WORKDIR: manifest.json (thresholds + pricing: priced models, unpriced models seen, and the share of waste that carries a price), l1_findings.json (rule findings: {rule, severity, sessionId, turnPointers, evidenceStats, estWasteTokens, project}), overview.json (aggregates: projects, tools, models, gapBuckets, dates, skills, sessions sorted by waste). Per-session and per-project rollups carry wasteUsd alongside wasteTokens, plus models, usdPerMTok, and pricedShare.
Rules emitted: CACHE_TTL_EXPIRY, DUP_TOOL_CALL, BIG_TOOL_OUTPUT, RETRY_STORM, CACHE_MISS_RATE, CONTEXT_GROWTH, NO_SUBAGENT.
Phase 1 — Read the aggregate landscape
node {base_directory}/bin/audit.js views
One bounded block: totals (tokens and dollars), findings-by-rule with a sessions-affected count, cache hit-ratio distribution, projects and worst sessions by waste (full session ids, with the fetch command to inspect one), idle-gap cost curve, tools by bytes, peak-context and delegation, per-date trend, skill usage.
Do not read overview.json / l1_findings.json wholesale (>100KB), and do not hand-roll node one-liners for anything views already prints — every query and its output lands in your transcript, which is the audit's real self-cost.
If a stat you need is missing, say so in the report — do not add it to views.js. You are running from an installed bundle, so src/views.js is your own code, and editing it mid-audit means the run measures itself with a renderer that changed partway through; the next install silently wipes the change anyway (I8). Record the gap under a "stats this report wanted and could not get" note, and a human ports it to the repo. Where a missing stat blocks a specific claim, a single narrow query is acceptable — but state in the report that you ran it, because the transcript cost is real and unlogged.
Two traps the view exists to prevent:
- Finding counts are not population counts. Every rule has entry gates —
CACHE_MISS_RATEneeds ratio < 0.5 and ≥5 turns — so "1 finding" never means "1 session with that property". Read the distribution block, never the rule count. - Raw vs. cost-equivalent tokens.
gapBuckets.cacheCreationis raw; rule waste is cost-equivalent (creation × 1.15,rules.js:118). The view prints the gap curve in both units. Compare like with like before claiming two methods corroborate.
Phase 2 — L2 hypothesis loop
- From the landscape, form candidate patterns. Recurring high-value ones: TTL expiry after idle gaps (check
gapBuckets— cost per resume vs thelt_1mbaseline), large-file re-read loops (DUP + BIG onReadin the same sessions), marathon sessions near the context ceiling, image-heavy low-cache sessions, snapshot-happy browser loops, scan-heavy sessions withhasSubagents: false. - Confirm or dismiss with narrower queries over per-session
findingsByRuleand per-session finding details. - Where metadata cannot resolve intent, escalate with fetch:
node {base_directory}/bin/audit.js fetch <session-id> --kind <kind> [--limit N] [--max-bytes B] [--uuid U] [--radius K]
| kind | returns | use to judge |
|---|---|---|
user_text |
user messages | intent vs agent behavior |
error_head |
head of errored results | transient vs deterministic failure |
tool_input |
tool call params | which file/command was duplicated |
assistant_head |
head of assistant text | verbosity / narration |
turn_window |
metadata around --uuid |
sequence reconstruction |
- Write
$AUDIT_WORKDIR/l2_hypotheses.json:
[{ "pattern": "NAME", "evidence": ["stat or fetch ref", "..."], "confidence": "high|medium|low",
"escalation_fetches": ["..."], "resolution": "resolved|escalate", "interpretation": "one sentence" }]
Phase 3 — L3 attribution
Reason over the hypotheses plus fetched snippets only. Write $AUDIT_WORKDIR/attribution.json, ranked by waste:
[{ "rank": 1, "finding": "...", "attribution": "habit|skill_file|config",
"estWasteShare": "~N% (xK of yK)", "trend": "improving|worsening|steady",
"fix": "concrete behavior change, named skill edit, or config change" }]
Attribution guide: habit = working style (resuming big sessions after breaks, no /clear at pivots, inline scanning, whole-file re-reads); skill_file = a named skill drives the pattern (say which file and what to change); config = MCP bloat, model default, thinking budget, CLAUDE.md size.
Phase 4 — Report
Write ~/.claude/audit-reports/<YYYY-MM-DDTHHMMSSZ>.md (create the directory if absent). Generate the timestamp with date -u +%Y-%m-%dT%H%M%SZ (UTC, second-precision). No sequence suffix: the timestamp disambiguates, and a same-second collision would mean two simultaneous audits — outside this skill's pacing.
The report is read by a developer deciding whether to spend an afternoon on this, and on what. A finding they cannot price, verify, or check off later is a finding they will not act on. Every requirement below exists to close one of those three gaps.
Header. projectsDir, session count, and date window, so a later comparison knows its scope. Nothing else — keep tool-internal notes (artifact disagreements, re-renders, workdir paths) out of the deliverable; they read as "this tool contradicts itself" directly above the numbers. Put them in the workdir if you need them.
Required content:
Totals in dollars and tokens.
viewsprints both. The dollar figure is what makes the report decidable; a token count alone is a unit nobody budgets in. Carry through both disclosuresviewsprints — the unpriced-model share, and that thinking tokens and compaction are unmeasured — and label the totals a heuristic floor.Findings-by-rule table with the sessions-affected column. 200 findings across 4 sessions is one bad week; across 60 it is a habit. The finding count alone cannot distinguish them, and the fix differs.
Per-project breakdown (
views→ Projects by waste). A dev acts on one repo at a time; the directory total tells them nothing about where to start. Note where the dollar and token rankings disagree — that means a pricier model, and it changes the priority.Ranked fixes, each carrying all five of:
- Evidence — real file names, real behaviors, the user's own quoted words where fetched.
- Full session ids, never truncated, plus the command to check one:
node {base_directory}/bin/audit.js fetch <session-id> --kind user_text --limit 3 --max-bytes 500An 8-char prefix cannot be resumed or fetched, so it makes the claim unverifiable. - Cost — dollars and tokens, with the share of headline.
- Attribution — habit / skill_file / config.
- A target metric — the exact row and number that should move by the next audit ("
CACHE_TTL_EXPIRY4,202K / $21 → under 1,500K"). Without one, "step away less" is unfalsifiable and next month's report has nothing to compare.
A "do this first" pick, chosen on effort × permanence — not on waste rank. A
skill_fileorconfigfix is one edit that keeps paying; ahabitfix is indefinite discipline with no enforcement. When the top-ranked item is a habit and a smaller one is a one-line skill edit, say plainly that the skill edit goes first and why. For anyskill_filefix, name the file and the text to change — "instruct it to read less" is not executable.Trend. Before writing,
ls ~/.claude/audit-reports/. If prior reports exist, read the most recent totals table and state the direction — which rules grew, which shrank, whether a previously recommended fix stuck. If the archive is empty say "first audit — no baseline yet"; do not infer a trend from within-run date buckets alone.Audit self-cost — fetch count + bytes from
fetch_log.jsonl, labelled a lower bound (I6).ASCII charts — four visualizations placed inline with the table each summarizes in the report (findings-by-rule table → rule chart immediately after, per-project breakdown → project chart, per-date trend → date chart, cache hit-ratio section → distribution histogram). Each chart is authored in two versions:
- Wide (~100 cols) in the report file: full labels and session counts.
- Compact (~40 cols) in the chat summary: shortened labels, no session counts.
Chart data is plotted only from values already pulled from
viewsoutput (I8) — never re-derived or hand-rolled. If a stat is missing, skip that chart and note it under "stats this report wanted and could not get".
The chat summary carries all four compact charts (rule, project, date, cache-hit), placed right after the headline paragraph and before the trend story, so the visual precedes the prose explanation.
Chart format
Wide version (~100 cols, label ≤ 25 chars, bar 50 chars, numbers ~20 chars):
NO_SUBAGENT ████████████████████████████████████████████ 243 sess / $73.04
CACHE_TTL_EXPIRY ██████████████████████ 119 sess / $35.26
BIG_TOOL_OUTPUT ████ 62 sess / $7.58
DUP_TOOL_CALL ████████ 98 sess / $14.01
Compact version (~40 cols, label ≤ 15 chars, bar 20 chars, numbers ~8 chars):
NO_SUBAGENT ████████████████████ $73
CACHE_TTL ██████████ $35
DUP_TOOL_CALL ████ $14
BIG_TOOL ██ $8
- Wrap each chart in a fenced code block so monospace alignment survives rendering.
- Bar character:
█(U+2588). Round bars to whole units — no fractional blocks. - Layout: right-padded label · bar · numbers. Fixed label width so bars line up.
- Scale: auto-scale to the maximum value so the largest bar fills the bar column.
- Sort: rows ordered by value descending — the worst item leads.
Then summarize the top 3 changes in chat, in the user's own context, and say where the report was saved.
Interpretation notes
CONTEXT_GROWTHcarries zero direct waste — treat it as an amplifier of everything else in that session.CACHE_TTL_EXPIRYfindings carryevidenceStats.gapKind. Onlyuser_idlegaps are priced as waste;tool_runtimegaps are reported at zero waste by design — the wait was a long-running command, not a habit, and attributing it to behaviour produces a fix the user cannot act on. CheckgapKindbefore calling any TTL finding a habit.NO_SUBAGENTis excluded from headline waste and reported in its own section. It prices a different counterfactual (delegate the phase) than the rules it overlaps (read less, read once), so summing it into the total double-counts bytes. Never add it to the headline figure.- Rank by prefix persistence: early, long-riding costs beat late one-offs of the same size.
- Dollars are priced per session, from its own model mix — so the token ranking and the dollar ranking can legitimately disagree (a project on a pricier model costs more per wasted token). Where they do, the dollar order is the one to act on, and worth calling out. Unpriced models contribute tokens but no dollars;
manifest.pricing.pricedWasteSharesays how much of the total is covered. - Peak context can exceed 200K on 1M-context models; don't call it a bug.
- A healthy directory (hit ratio >0.95, near-linear growth) deserves a short report saying so — do not manufacture findings.