Evaluate Agent Quality

Establish whether an agent actually works, and whether it still works, using a recorded eval set rather than ad hoc chats. Use when an agent is about to ship and the only testing was somebody typing a few questions into the test pane, when an agent that used to answer correctly now does not and nobody can say when it broke, when a knowledge source or model version changed and you need to know what it affected, when a stakeholder asks how accurate it is and there is no number to give them, when you cannot decide what "correct" even means for a generative answer, or when the agent works for the builder and fails for a user with different permissions. For a one-off read of a single agent's configuration, review-copilot-studio-agent is the better fit; this one is about measurement over time.

RagnarPitla 970e14c 19.5 KB Updated

File contents

RagnarPitla/microsoft-agent-skills-private-archive/tree/main/.github/skills/evaluate-agent-quality commit 970e14c130

Frequently asked questions

npx skillmds@latest add ragnarpitla/evaluate-agent-quality-2