Transformer Interpretability Eval

Evaluates the faithfulness and class-specificity of Transformer interpretability methods by measuring how well highlighted input features align with model predictions. It probes explanation quality through pixel/token masking, segmentation overlap, and rationale extraction accuracy. Use when the user wants to benchmark on ImageNet Validation (ILSVRC 2012), ImageNet-Segmentation, Movie Reviews, or asks about evaluating this task. Reports AUC (Positive/Negative Perturbation).

qhjqhj00 a639432 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/transformer-interpretability-eval commit a6394325a1

Frequently asked questions

npx skillmds add qhjqhj00/transformer-interpretability-eval