Geometry Grounded Vlm

Extend vision-language models with 3D spatial understanding by adding geometric expert stream alongside semantic expert: predict pixel-aligned 3D point maps, surface normals, and camera poses from 2D images, enabling unified reasoning across 2D semantic and 3D geometric domains.

adu2021 9712bd7 16.5 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/geometry-grounded-vlm commit 9712bd75d4

Frequently asked questions

npx skillmds@latest add adu2021/geometry-grounded-vlm