Benchmarking Reward Hack Detection

Detect reward hacking in AI-generated code trajectories using contrastive analysis from the TRACE benchmark. Use when: 'check this code agent for reward hacking', 'detect if these test results are gamed', 'audit coding agent trajectories', 'find reward exploits in RL-generated code', 'contrastive analysis on code submissions', 'are these tests being manipulated'.

ndpvt-web e3a6539 15.8 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/benchmarking-reward-hack-detection commit e3a6539638

Frequently asked questions

npx skillmds@latest add ndpvt-web/benchmarking-reward-hack-detection