Vision Deepresearch Incentivizing Deepresearch Cap

Multi-turn, multi-entity, multi-scale visual and textual deep research agent for answering complex questions about images. Implements the Vision-DeepResearch paradigm: iterative reasoning-then-search with progressive visual cropping and text retrieval. Use when: 'research this image', 'identify everything in this photo', 'what is the history behind this building', 'find information about objects in this picture', 'deep research this visual scene', 'multi-hop visual question answering'.

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/vision-deepresearch-incentivizing-deepresearch-cap commit 1c57ca4f95

Frequently asked questions

npx skillmds@latest add ndpvt-web/vision-deepresearch-incentivizing-deepresearch-cap