Run repeatable agent evaluation suites with trajectory and simulator coverage using Strands Evals

Build repeatable evaluation experiments for agents and LLM apps with output checks, trajectory scoring, simulators, and trace-based review.

agentskillexchange Updated 28 repo stars

File contents

agentskillexchange/skills/tree/main/skills/run-repeatable-agent-evaluation-suites-with-trajectory-and-simulator-coverage-using-strands-evals commit ca465a8040

Frequently asked questions

npx skillmds@latest add agentskillexchange/run-repeatable-agent-evaluation-suites-with-trajectory-and-s