Cross Modal Alignment And Dialogue Grounding

Use this skill when the user asks 'what is visually present when this sound happens?', 'the video shows one thing but the person says another, which is right?', 'correct the spelling of this name based on the screen', 'find the exact time for this speech', or 'match the voice to the person'. It is triggered for tasks requiring explicit temporal grounding involving both audio and video, implicit cross-modal retrieval, or resolving cross-modal contradictions (like using sharp on-screen text to fix noisy or muffled audio transcripts).

dingxingdi Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/skill_bank_evolved/video/skills/cross-modal-alignment-and-dialogue-grounding commit 4c2a387d96

Frequently asked questions

npx skillmds@latest add dingxingdi/cross-modal-alignment-and-dialogue-grounding