Jailbreak harness engineering
Minimal harness
seeds/prompts by taxonomy tagrunners/target.pymodel/agent adapterscore.pyrubric (refuse/partial/full + keyword checks)results/jsonl transcriptsreport.mdgenerator
Good scorers
- Refusal detector (string + judge)
- Technical completeness checklist per task type
- Safety-filter hit taxonomy
- Latency/cost
CI ideas
- Nightly against local agent builds
- Diff refusal strings after Claude/Codex updates → feed Unleash re-anchor
- Track regression: yesterday's wins failing today
Unleash integration
- On local CC/Codex/OMP: setup → patch → verify before prompt campaigns
- Store new refusal strings under research notes for pool patches
- Skills pack path:
contrib/skills/installed to~/.agents/skills+~/.claude/skills