Astrovlbench Eval

This benchmark evaluates the ability of vision-language models to perform multi-modal astronomical reasoning across five distinct observational modalities, including optical imaging, radio interferometry, photometry, light curves, and spectroscopy. It probes whether models can correctly classify celestial objects and interpret physical features, while also testing the impact of prompt guidance and input representation (visual vs. numerical) on classification accuracy and reasoning quality. Use when the user wants to benchmark on AstroVLBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 688542f 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/astrovlbench-eval commit 688542f513

Frequently asked questions

npx skillmds add qhjqhj00/astrovlbench-eval