Evaluations

Set up comprehensive evaluations for your AI agent with LangWatch — experiments (batch testing), evaluators (scoring functions), datasets, online evaluation (production monitoring), and guardrails (real-time blocking). Supports both code (SDK) and platform (CLI) approaches. Use when the user wants to evaluate, test, benchmark, monitor, or safeguard their agent.

langwatch Updated

File contents

langwatch/tokenmaxxer/tree/main/.agents/skills/evaluations commit cf6ddf3b27

Frequently asked questions

npx skillmds@latest add langwatch/evaluations-3