Eval And Improve

Run this repo's eval suite (python -m evals), diagnose every failing case, and fix in scope (agent instructions, or the eval case when its assertion was wrong) until all cases pass. Use whenever the user wants to run the evals, check for regressions, or fix a red eval suite ("run the evals", "why is a case failing", "make the eval suite green"). Targets the committed cases in evals/cases.py; for ad-hoc probing of one agent's live behavior use improve-agent instead.

agno-agi 4743d0d 10.7 KB Updated

File contents

agno-agi/context/tree/main/.agents/skills/eval-and-improve commit 4743d0d665

Frequently asked questions

npx skillmds@latest add agno-agi/eval-and-improve