Model Written Evaluations

Tests whether language models exhibit specific emergent or undesirable behaviors (e.g., sycophancy, self-preservation, political bias) by measuring their preference for behavior-matching versus behavior-mismatching labels. It evaluates how well smaller or preference-trained models predict the behavioral tendencies of larger or RLHF-trained counterparts. Use when the user wants to benchmark on Model-Written Evaluations (133 behaviors), or asks about evaluating this task. Reports accuracy.

qhjqhj00 f72d8b3 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/model-written-evaluations commit f72d8b3f24

Frequently asked questions

npx skillmds add qhjqhj00/model-written-evaluations