Twinviews Bias Eval

Evaluates whether reward models exhibit political bias by measuring the average reward scores assigned to politically left-leaning versus right-leaning statements on the same topics. The protocol compares mean reward differences across model sizes and training runs to detect systematic left-leaning skew. Use when the user wants to benchmark on TwinViews-13k, or asks about evaluating this task. Reports average_reward.

qhjqhj00 36371fb 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/twinviews-bias-eval commit 36371fb59d

Frequently asked questions

npx skillmds add qhjqhj00/twinviews-bias-eval