Craft Eval

Evaluates the ability of instruction-tuned LLMs to perform domain-specific multiple-choice question answering and text generation tasks. It measures how well models fine-tuned on synthetic, corpus-retrieved data generalize to held-out human-annotated benchmarks in biology, medicine, commonsense, recipe generation, and summarization. Use when the user wants to benchmark on ScienceQA (BioQA), MedMCQA (MedQA), CommonsenseQA 2.0 (CSQA), RecipeNLG (RecipeGen), CNN-DailyMail (Summarization), or asks about evaluating this task. Reports accuracy.

qhjqhj00 3d86256 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/craft-eval commit 3d86256195

Frequently asked questions

npx skillmds add qhjqhj00/craft-eval