Agent Evals And Observability

Design, run, review, or release framework- and vendor-neutral evaluations and observability for AI agents. Use when defining agent evals, datasets, graders, trajectory review, regression analysis, release gates, production traces, or privacy-aware telemetry. Covers task and trajectory contracts, statistical comparisons, and incident-to-case learning; route framework implementation to pydanticai or langgraph when needed.

theheavenlyd3mon 8f17c39 19 files · 35.1 KB Updated 28 repo stars

File contents

theheavenlyd3mon/hermes-profiles/tree/main/profiles/mlops/skills/magnus/agent-evals-and-observability commit 8f17c394f1

Frequently asked questions

npx skillmds add theheavenlyd3mon/agent-evals-and-observability