Subgroup Benchmarking Eval

Evaluates the precision of statistical estimators (Empirical Bayes, Synthetic Regression, Direct Training) for estimating model performance on data subgroups with limited observations. It probes how well these methods reduce mean squared error and produce reliable confidence intervals when benchmarking LLMs, vision models, and tabular classifiers on niche tasks. Use when the user wants to benchmark on LLM MC QA tasks, Computer Vision tasks (LAION CLIP benchmark), COCO Captions, Tabular Fairness Datasets (ACS, COMPAS, Student), or asks about evaluating this task. Reports MSE.

qhjqhj00 6b7fb2c 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/subgroup-benchmarking-eval commit 6b7fb2cbc2

Frequently asked questions

npx skillmds add qhjqhj00/subgroup-benchmarking-eval