Capture Evidence

Use when a developer wants to build an eval from their real LLM app before changing anything — "measure how my app is doing today", "build an eval from my workload", "we have no baseline", "is my current model actually good". Turns the workload into auditable local artifacts (harness, metric, frozen splits, baseline); has a public-benchmark on-ramp when no traces exist.

understudylabs Updated

File contents

understudylabs/understudy-agent-tools/tree/main/skills/capture-evidence commit fe63d1b74c

Frequently asked questions

npx skillmds@latest add understudylabs/capture-evidence