Mmlu Bbh Gsm8k Eval

Evaluates how fine-tuning large language models on syntactically or semantically perturbed instructions impacts downstream generalization across factual knowledge, complex reasoning, and mathematical problem-solving. It also measures potential side effects on model toxicity and truthfulness under varying noise levels during both training and evaluation. Use when the user wants to benchmark on MMLU, BBH, GSM8K, ToxiGen, TruthfulQA, or asks about evaluating this task. Reports average test accuracy.

qhjqhj00 02caf8d 5.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmlu-bbh-gsm8k-eval commit 02caf8df9e

Frequently asked questions

npx skillmds add qhjqhj00/mmlu-bbh-gsm8k-eval