Inverse Constitutional AI Eval

Tests a framework's ability to compress pairwise preference data into interpretable natural language principles (constitutions) and use them to reconstruct original annotations. It probes the model's adaptability to aligned, unaligned, individual, and demographic group preferences, as well as its capacity for bias detection. Use when the user wants to benchmark on Synthetic data, AlpacaEval, Chatbot Arena Conversations, PRISM, or asks about evaluating this task. Reports agreement.

qhjqhj00 b12ae07 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/inverse-constitutional-ai-eval commit b12ae07c3a

Frequently asked questions

npx skillmds add qhjqhj00/inverse-constitutional-ai-eval