Auditing Agent Behavior

Audit an agent for risky behavior before it ships, the way you would red-team it, but automated. An auditor agent runs many multi-turn scenarios against your target agent from seed instructions you write, and a judge model scores the transcripts for deception, sycophancy, oversight subversion, power-seeking, and cooperation with misuse. Covers writing seed instructions that probe your agent's real risks, adapting the scoring rubric to your domain, and reading flagged transcripts. Use this when someone wants to red-team or safety-test an agent, asks how to find deceptive or manipulable behavior, audits a model or agent before deployment, or compares behavior across model versions. Trigger on "red-team my agent," "audit agent behavior," "test for deception," "alignment testing," and similar. This is auditing behavior for safety; measuring task success is build-agent-evals, and stopping objective-gaming is reward-hacking-and-inoculation.

Hoja-Solutions c192d6e 3.2 KB Updated

File contents

Hoja-Solutions/agent-stdlib/tree/main/skills/auditing-agent-behavior commit c192d6e999

Frequently asked questions

npx skillmds@latest add hoja-solutions/auditing-agent-behavior