Pts3d LLM Eval

Evaluates multimodal large language models on 3D scene understanding tasks, including visual grounding, dense captioning, and spatial/situated question answering. It specifically probes how different 3D token structures (point-based vs. video-based) and feature fusion strategies impact performance on indoor RGB-D scans. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports NS.

qhjqhj00 43327bd 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pts3d-llm-eval commit 43327bdc90

Frequently asked questions

npx skillmds add qhjqhj00/pts3d-llm-eval