Phibench Eval

Evaluates diverse reasoning and coding capabilities using an internal benchmark designed to minimize data contamination and LLM-judge bias. It probes a model's ability to debug, extend, and explain code, as well as identify errors in mathematical proofs and generate related problems. Use when the user wants to benchmark on PhiBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 64c1792 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/phibench-eval commit 64c17921df

Frequently asked questions

npx skillmds add qhjqhj00/phibench-eval