Ravel Eval

Evaluates interpretability methods' ability to disentangle polysemantic language model representations by isolating causal attributes through activation interventions on residual stream features. Use when the user wants to benchmark on RAVEL, or asks about evaluating this task. Reports Disentanglescore.

qhjqhj00 614302d 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ravel-eval commit 614302d338

Frequently asked questions

npx skillmds add qhjqhj00/ravel-eval