Fewstab Eval

This benchmark evaluates the robustness of few-shot image classifiers to spurious class-attribute correlations. It constructs evaluation tasks where support sets contain images with engineered spurious attributes, and query sets lack these attributes or contain attributes from other classes, then measures classification accuracy compared to standard random task sampling. Use when the user wants to benchmark on miniImageNet, tieredImageNet, CUB-200, or asks about evaluating this task. Reports wAcc-A.

qhjqhj00 ad91f9a 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/fewstab-eval commit ad91f9a934

Frequently asked questions

npx skillmds add qhjqhj00/fewstab-eval