Backbench Eval

This benchmark probes an agent's ability to recover from harmful states in real-world computer use environments by backtracking or remediating to a safe operational state. It evaluates how well agents align with human preferences during recovery under varying resource constraints (step limits). Use when the user wants to benchmark on BackBench, or asks about evaluating this task. Reports Bradley-Terry rating.

qhjqhj00 d75105a 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/backbench-eval commit d75105acdb

Frequently asked questions

npx skillmds add qhjqhj00/backbench-eval