Cov Chain Of View Spatial Reasoning

Enable vision-language models to perform embodied question answering in 3D environments through active camera exploration. CoV uses training-free test-time reasoning to iteratively select relevant viewpoints and adjust camera angles until sufficient context is gathered, achieving 11-13% accuracy improvements across spatial reasoning benchmarks.

adu2021 Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/cov-chain-of-view-spatial-reasoning commit a1899d7ff8

Frequently asked questions

npx skillmds@latest add adu2021/cov-chain-of-view-spatial-reasoning