Astrovisbench Eval

Evaluates large language models' ability to act as coding assistants for astronomy-specific scientific workflows. It probes domain-specific API usage, data manipulation, and the generation of research-standard visualizations from natural language queries. Use when the user wants to benchmark on AstroVisBench, or asks about evaluating this task. Reports execution-based evaluation.

qhjqhj00 06c9444 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/astrovisbench-eval commit 06c9444920

Frequently asked questions

npx skillmds add qhjqhj00/astrovisbench-eval