C2rope 3d Vqa Eval

Evaluates a 3D large multimodal model's ability to perform spatial reasoning and visual question answering on multi-view 3D scene data. It probes the model's capacity to retain early visual context, understand spatial relationships, and generate accurate text responses to complex 3D scene queries. Use when the user wants to benchmark on ScanQA, SQA3D, or asks about evaluating this task. Reports EM@1.

qhjqhj00 1159891 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/c2rope-3d-vqa-eval commit 11598915d9

Frequently asked questions

npx skillmds add qhjqhj00/c2rope-3d-vqa-eval