Vla Generality Benchmark Eval

Evaluates the cross-domain generality of vision-language-action models across vision-language understanding, discrete multi-agent control, and continuous robot manipulation. Probes strict format compliance, semantic alignment, action prediction accuracy, and failure modes such as output collapse or modality misalignment. Use when the user wants to benchmark on PIQA, SQA3D, RoboVQA, ODINW, BFCL, Overcooked, Open-X, or asks about evaluating this task. Reports EMR.

qhjqhj00 8318352 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vla-generality-benchmark-eval commit 83183523c8

Frequently asked questions

npx skillmds add qhjqhj00/vla-generality-benchmark-eval