Scinav Eval

Evaluates an autonomous agent's capability to generate executable scientific and data-science code by measuring execution reliability and task-goal satisfaction. It probes the model's ability to navigate constrained search budgets while producing outputs that match predefined success criteria across diverse disciplinary benchmarks. Use when the user wants to benchmark on ScienceAgentBench, DA-Code, or asks about evaluating this task. Reports Success Rate (SR).

qhjqhj00 06f9c2b 4.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scinav-eval commit 06f9c2bbb9

Frequently asked questions

npx skillmds add qhjqhj00/scinav-eval