Harness Eval

Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps. HIGH PRIORITY questionnaires at top: Q1 optional docs, Q2 B/C budget before Track A (certainty/tokens). A always runs after Q2; B/C opt-in. ADRs/RFCs excluded from T2. Mixed apply uses 11-mixed-apply.md (KEEP/CUT). Use when the user says harness eval, harness-eval, harness debug, audit AGENTS.md, audit skills/rules, instruction audit, redundancy of agent instructions, usefulness of skills, Ship/Review/Hold/Slim/Keep-core for harness, or wants Track A/B/C harness evaluation. Do NOT use for harness setup or init, feature spec-driven work (pbs-spec-driven), or applying Ship/Slim trims unless the user explicitly asks after the report.

Peterson-Benhame Updated

File contents

Peterson-Benhame/agent-skills/tree/main/packages/skills-catalog/skills/(development)/harness-eval commit 5ba67f4ed5

Frequently asked questions

npx skillmds@latest add peterson-benhame/harness-eval