Autonomous Agent Trajectory Adjudication

Use this skill when evaluating an autonomous agent's or multi-agent system's full multi-step trajectory in any environment (CLI, Web, GUI, or team-based role pipelines), focusing on terminal outcomes, step-level correctness, and environmental grounding. Trigger it when users say things like 'check if the server repair worked', 'trace where the error started in the pipeline', 'assign blame for the team failure', 'verify the GUI agent's clicks', or 'score the step-by-step progress of the browser bot'. It covers verifying system states (metrics, status codes, screenshots), detecting harmful side effects, penalizing inefficient reasoning loops, and tracing 'Error Origins' or 'Repair' events across sequential multi-agent or environment handoffs.

dingxingdi Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/skill_bank_evolved/eval/skills/autonomous-agent-trajectory-adjudication commit d5df6dd0a3

Frequently asked questions

npx skillmds@latest add dingxingdi/autonomous-agent-trajectory-adjudication