eval-gauntlet-and-lima-judge-eval
LIMIT: Less Is More for Instruction Tuning Across Evaluation Paradigms — Aditi Jha et al. (2023) (arXiv:2311.13133, 2023)
What this evaluates
This protocol evaluates instruction-tuned LLMs across two distinct paradigms: traditional closed-domain NLP benchmarks and open-ended generation quality. It probes whether performance on standard accuracy-based tasks aligns with preference judgments from a large language model judge, highlighting the tension between task-diverse versus style-aligned training data.
Datasets
- MosaicML Eval Gauntlet — total ?; splits: test (-1)
- LIMA test set — total 300; splits: test (300)
Metrics
accuracy(primary) — range: [0, 1]- Proportion of correctly answered questions per category, averaged across all categories.
preference rate— range: percent- Percentage of prompts where the judge model (GPT-4) prefers the target model's response over the baseline model's response.
Input / output format
Input: Instruction prompts from the MosaicML Eval Gauntlet or the 300-sample LIMA test set.
Output: Model-generated text responses to each prompt.
Scoring recipe
# For Eval Gauntlet
gauntlet_scores = []
for category in eval_gauntlet_categories:
correct = sum(1 for q, a in category if model_answer(q) == gold(a))
gauntlet_scores.append(correct / len(category))
final_accuracy = sum(gauntlet_scores) / len(gauntlet_scores)
# For LIMA Judge Eval
preferences = 0
for prompt in lima_test_set:
response_target = model_target.generate(prompt)
response_baseline = model_baseline.generate(prompt)
if gpt4_judge.prefers(response_target, response_baseline):
preferences += 1
preference_rate = (preferences / len(lima_test_set)) * 100
Common pitfalls
- Assuming higher accuracy on traditional benchmarks correlates with better open-ended generation quality; the paper shows a strong dichotomy between the two.
- Overestimating the impact of dataset size over dataset composition; random subsets of diverse tasks match full-dataset performance on the gauntlet.
- Using in-domain test sets for judge evaluation may bias preference rates toward the training data's stylistic conventions rather than general helpfulness.
Evidence (verbatim from paper)
Each plot shows the accuracy (between 0–1) of models on a given category of the MosaicML Eval Gauntlet, and the average score across all categories is shown in the first subplot.
Citation
@misc{jha2023limit,
title={LIMIT: Less Is More for Instruction Tuning Across Evaluation Paradigms},
author={Aditi Jha et al. (2023)},
year={2023},
note={arXiv:2311.13133}
}
- arXiv: 2311.13133