Agents CLI Benchmark Eval

Designs, runs, and scores local evaluation datasets against an ADK agent or command, computes exact-match, token-overlap (ROUGE-like), semantic-similarity proxy, rubric scores, and routing precision/recall, then writes JUnit XML and Markdown reports. TRIGGER when users ask to 'run evaluation datasets', 'evaluate agent accuracy', 'calculate trigger precision/recall', 'generate a JUnit test report for my agent', 'set up an agent benchmark suite', or score an ADK agent locally. DO NOT TRIGGER for live production load testing, training/fine-tuning a model, or general unit tests without an agent output/dataset evaluation objective.

tommywagz Updated

File contents

tommywagz/Gemini-Enterprise-Skills/tree/main/skills/agents-cli-benchmark-eval commit bce29b6772

Frequently asked questions

npx skillmds@latest add tommywagz/agents-cli-benchmark-eval