Reward Safety

Route and execute reward-seeking safety work for AI agents. Use when a request mentions rewards, graders, verifiers, hidden tests, monitors, oversight, reward hacking, evaluation gaming, metagaming, authority conflicts, honesty versus task completion, or whether an agent is optimizing the score instead of the intended outcome. Also use for /reward-safety.

jpoindexter Updated

File contents

jpoindexter/reward-seeking-safety-skills/tree/main/skills/reward-safety commit 1ca831c52c

Frequently asked questions

npx skillmds@latest add jpoindexter/reward-safety