Run repeatable agent evaluation suites with trajectory and simulator coverage using Strands Evals
Build repeatable evaluation experiments for agents and LLM apps with output checks, trajectory scoring, simulators, and trace-based review.
Prerequisites
Python 3.10+, pip, optional judge-model access
Installation
Use the upstream install or setup path that matches your environment:
- pip install strands-agents-evals
- pip install -e .
- pip install -e ".[test]"
- pip install -e ".[test,dev]"
Requirements and caveats from upstream:
- ◆ Python SDK
- python
Basic usage or getting-started notes:
Multiple Evaluation Types: Output evaluation, trajectory analysis, tool usage assessment, and interaction evaluation
bash
from strands import Agent
Extracted from upstream docs: https://raw.githubusercontent.com/strands-agents/evals/HEAD/README.md