Eval

Use this skill when running pre-release validation, detecting calibration regression, or tuning a skill — to run plugin evaluation tiers (Tier 1 linters and Tier 2 judge-calibration drift smoke) by wrapping the eval/runner.py harness. Tier 1 is free (no LLM); Tier 2 budgets tokens per `eval/config.json`. Tier 3 (full behavioral suites) is planned but not yet shipped — runner returns error code 3 if invoked.

avav25 Updated

File contents

avav25/ai-assets/tree/main/plugin/skills/eval commit 18852a3d82

Frequently asked questions

npx skillmds@latest add avav25/eval