Eval Benchmark Runner

Use when running the automated daily evaluation suite that measures the legal AI system's output quality across all benchmark datasets. Orchestrates the full eval pipeline — loading datasets, calling the production model, scoring with LLM-as-judge rubrics, detecting regressions, and publishing results to the leaderboard and observability dashboards.

sboghossian Updated

File contents

sboghossian/mini-claude-for-legal/tree/main/skills/eval/eval-benchmark-runner commit 7a5f2572ed

Frequently asked questions

npx skillmds@latest add sboghossian-mini-claude-for-legal/eval-benchmark-runner