Eval Result Interpreter

Analyzes AI agent evaluation results - primarily from Copilot Studio (the worked example here, via its CSV export) but also from custom harnesses or any evaluator that produces per-case pass/fail rows - using Microsoft's Triage & Improvement Playbook. Returns a SHIP / ITERATE / BLOCK verdict with root cause classification, diagnostic triage, prioritized remediation, and pattern analysis.

varunk130 735bff6 34.5 KB Updated

File contents

varunk130/AI-Eval-Skills/tree/main/skills/eval-result-interpreter commit 735bff67c1

Frequently asked questions

npx skillmds@latest add varunk130/eval-result-interpreter