Sft Generalization Eval

Probes whether language models trained via supervised fine-tuning (SFT) truly learn reasoning and planning capabilities or merely memorize instruction templates. It tests generalization across unseen action mappings (instruction variations) and increased grid/card complexity (difficulty variations). Use when the user wants to benchmark on Sokoban, General Points, or asks about evaluating this task. Reports exact-match accuracy.

qhjqhj00 ceeff9d 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sft-generalization-eval commit ceeff9d4dd

Frequently asked questions

npx skillmds add qhjqhj00/sft-generalization-eval