Reward Process Integrity

Audit and harden the reward path around an AI agent. Use when an agent might read or modify graders, hidden tests, verifier code, score files, budgets, model selection, monitors, audit logs, release thresholds, or other evaluation machinery; also use for reward hacking and grader gaming controls.

jpoindexter Updated

File contents

jpoindexter/reward-seeking-safety-skills/tree/main/skills/reward-process-integrity commit 4d47b0d483

Frequently asked questions

npx skillmds@latest add jpoindexter/reward-process-integrity