Scireplicate Bench Eval

Evaluates an LLM's ability to comprehend algorithmic descriptions from academic papers and translate them into executable code. It probes the model's capacity for algorithmic reasoning, dependency resolution, and practical implementation within a repository context. Use when the user wants to benchmark on SciReplicate-Bench, or asks about evaluating this task. Reports Execution Accuracy.

qhjqhj00 d41a140 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scireplicate-bench-eval commit d41a140aa0

Frequently asked questions

npx skillmds add qhjqhj00/scireplicate-bench-eval