Aide Benchmark Eval

Evaluates the performance of LLMs fine-tuned on synthetic data generated by AIDE across a suite of standard knowledge and reasoning benchmarks. It probes zero-shot and few-shot generalization capabilities compared to models fine-tuned on human-curated gold data. Use when the user wants to benchmark on MMLU, FinBen, ARC-Challenge, GSM8K, TruthfulQA, MedQA, BIG-Bench, or asks about evaluating this task. Reports zero-shot accuracy.

qhjqhj00 a76c999 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/aide-benchmark-eval commit a76c999ee5

Frequently asked questions

npx skillmds add qhjqhj00/aide-benchmark-eval