Grade agent trajectories and tool-use decisions with AgentEvals

Score whether an agent took a sensible intermediate path, called tools correctly, and reached the outcome without relying only on final-answer checks.

agentskillexchange Updated 28 repo stars

File contents

Grade agent trajectories and tool-use decisions with AgentEvals

Score whether an agent took a sensible intermediate path, called tools correctly, and reached the outcome without relying only on final-answer checks.

Prerequisites

Python or TypeScript runtime, agent run outputs or trajectories, optional LLM judge provider

Installation

Use the upstream install or setup path that matches your environment:

  • pip install agentevals
  • npm install agentevals @langchain/core
  • pip install openai
  • npm install openai

Requirements and caveats from upstream:

Basic usage or getting-started notes:

Documentation

Source

agentskillexchange/skills/tree/main/skills/grade-agent-trajectories-and-tool-use-decisions-with-agentevals commit 674dd22af7

Frequently asked questions

npx skillmds@latest add agentskillexchange/grade-agent-trajectories-and-tool-use-decisions-with-agentev