Phi 4 Reasoning Eval

Evaluates large language models on reasoning-specific capabilities including mathematics, scientific QA, coding, algorithmic planning, and spatial reasoning. It probes the model's ability to generate step-by-step solution traces and produce correct final answers under varying decoding temperatures and run counts. Use when the user wants to benchmark on AIME, GPQA Diamond, OmniMATH, LiveCodeBench, Codeforces, or asks about evaluating this task. Reports pass@1 accuracy.

qhjqhj00 3466f7f 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/phi-4-reasoning-eval commit 3466f7f475

Frequently asked questions

npx skillmds add qhjqhj00/phi-4-reasoning-eval