# Run repeatable agent evaluation suites with trajectory and simulator coverage using Strands Evals

> Build repeatable evaluation experiments for agents and LLM apps with output checks, trajectory scoring, simulators, and trace-based review.

- Skill: `agentskillexchange/run-repeatable-agent-evaluation-suites-with-trajectory-and-s` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/run-repeatable-agent-evaluation-suites-with-trajectory-and-s`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/run-repeatable-agent-evaluation-suites-with-trajectory-and-s/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/run-repeatable-agent-evaluation-suites-with-trajectory-and-s

---


# Run repeatable agent evaluation suites with trajectory and simulator coverage using Strands Evals

Build repeatable evaluation experiments for agents and LLM apps with output checks, trajectory scoring, simulators, and trace-based review.

## Prerequisites

Python 3.10+, pip, optional judge-model access

## Installation

Use the upstream install or setup path that matches your environment:
- pip install strands-agents-evals
- pip install -e .
- pip install -e ".[test]"
- pip install -e ".[test,dev]"

Requirements and caveats from upstream:
- <a href="https://python.org"><img alt="Python versions" src="https://img.shields.io/pypi/pyversions/strands-agents-evals"/></a>
- ◆ <a href="https://github.com/strands-agents/sdk-python">Python SDK</a>
- python

Basic usage or getting-started notes:
- **Multiple Evaluation Types**: Output evaluation, trajectory analysis, tool usage assessment, and interaction evaluation
- bash
- from strands import Agent

- Source: https://github.com/strands-agents/evals
- Extracted from upstream docs: https://raw.githubusercontent.com/strands-agents/evals/HEAD/README.md

## Documentation

- https://github.com/strands-agents/evals

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/run-repeatable-agent-evaluation-suites-with-trajectory-and-simulator-coverage-using-strands-evals/)

