Trace Reward Hack Detection Eval

This benchmark evaluates an LLM's ability to detect and classify reward hacking behaviors in multi-turn code generation trajectories. It specifically probes contrastive anomaly detection capabilities by presenting clusters of mixed benign and malicious trajectories, testing whether models can disentangle subtle semantic and syntactic exploit patterns without prior taxonomy exposure. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports Detection Rate.

qhjqhj00 681f7be 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/trace-reward-hack-detection-eval commit 681f7be050

Frequently asked questions

npx skillmds add qhjqhj00/trace-reward-hack-detection-eval