Tao Generate Image Grounding

Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, or run the image_grounding pipeline. Triggers include 'image grounding', 'phrase grounding', 'ground captions', 'auto-label image grounding', 'image_grounding'.

NVIDIA-TAO Updated

File contents

NVIDIA-TAO/tao-skills-bank/tree/main/skills/data/tao-generate-image-grounding commit e3ab8c404b

Frequently asked questions

npx skillmds@latest add nvidia-tao/tao-generate-image-grounding