Sphinx Eval

Probes visual perception and reasoning capabilities of vision-language models across 25 distinct task types, including symmetry, spatial transformations, chart interpretation, and sequence prediction. Uses a synthetic environment with verifiable ground truth to measure model accuracy against human baselines. Use when the user wants to benchmark on Sphinx, or asks about evaluating this task. Reports accuracy.

qhjqhj00 864f87b 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sphinx-eval commit 864f87b22c

Frequently asked questions

npx skillmds add qhjqhj00/sphinx-eval