Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation

Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.

agentskillexchange Updated 28 repo stars

File contents

Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation

Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.

Prerequisites

Python environment, target agent endpoint or integration, optional AWS services such as Bedrock or SageMaker

Installation

No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.

Documentation

Source

agentskillexchange/skills/tree/main/skills/benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation commit d483583ea4

Frequently asked questions

npx skillmds@latest add agentskillexchange/benchmark-virtual-agents-with-scripted-multi-turn-conversati