Off Policy Sft Eval

Evaluates the trade-off between improving downstream mathematical reasoning capabilities and mitigating catastrophic forgetting on general-domain knowledge benchmarks after off-policy supervised fine-tuning. It measures how well a model retains pre-trained general knowledge while learning a new specialized task. Use when the user wants to benchmark on Math500, MinervaMath, AMC23, AGIEval-Math, IMO-Bench, MMLU, MMLU-Pro, AGIEval, or asks about evaluating this task. Reports OverallAvg.

qhjqhj00 339f278 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/off-policy-sft-eval commit 339f278d19

Frequently asked questions

npx skillmds add qhjqhj00/off-policy-sft-eval