Cross Modal Alignment And Dialogue Grounding

Use this skill when the user wants questions about 'who said it', 'which sound matches this moment', 'connect the voice to the gesture', 'use both the video and the audio', or 'ground the answer in dialogue or subtitle timing.' Trigger it whenever neither the video alone nor the audio alone is enough to answer reliably.

dingxingdi Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/examples/evol_ability/20260325_170549/profiles/video/skills/cross-modal-alignment-and-dialogue-grounding commit 94e0f4e3dc

Frequently asked questions

npx skillmds@latest add dingxingdi/cross-modal-alignment-and-dialogue-grounding-2