Agent Trajectory Evaluator

Evaluate a multi-step AI agent's whole run — tool calls, intermediate steps, and final result — not just final-answer correctness, so you can pinpoint WHERE it went wrong. Use when building or debugging a tool-using or multi-step agent, when final-answer-only evals can't explain failures, or when a prompt/model change quietly makes the agent less efficient or more error-prone even though the answer still looks right.

imtiazrayhan 1db1acb 7.6 KB Updated

File contents

imtiazrayhan/agentscamp-library/tree/main/skills/agent-trajectory-evaluator commit 1db1acbf08

Frequently asked questions

npx skillmds@latest add imtiazrayhan/agent-trajectory-evaluator