Intervention Scoring

Evaluates the fidelity of natural language explanations for sparse autoencoder (SAE) features by measuring how well an explanation predicts the downstream effects of directly intervening on the feature's activation, rather than just correlating with input contexts. Use when the user has predictions and gold and needs to compute intervention_scoring.

qhjqhj00 505c1d4 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/intervention_scoring commit 505c1d48e8

Frequently asked questions

npx skillmds add qhjqhj00/intervention-scoring