Evaluation And Quality Harness

Measures exactitude, quality, and stability of runs produced by the teaching-agent system using programmatic checks, rubric-based grading, and benchmark cases. Use when validating workflows, benchmarking outputs, or comparing runs.

alainlebret Updated

File contents

alainlebret/claude-agents/tree/main/higher-ed-teaching-agents/skills/evaluation-and-quality-harness commit f97051de4b

Frequently asked questions

npx skillmds@latest add alainlebret/evaluation-and-quality-harness