Ttsr Eval

Evaluates the ability of LLMs to improve reasoning performance at test time through self-reflection and targeted variant question synthesis, without external supervision. It probes how well a model can adapt its policy to difficult mathematical and general reasoning problems by diagnosing its own failures and generating corrective training signals. Use when the user wants to benchmark on AMC23, MATH-500, Minerva, OlympiadBench, AIME 2024, AIME 2025, GPQA-Diamond, MMLU-Pro, or asks about evaluating this task. Reports accuracy.

qhjqhj00 9855330 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ttsr-eval commit 985533093d

Frequently asked questions

npx skillmds add qhjqhj00/ttsr-eval