Modality Gap Driven Subspace Alignment

Align multimodal embeddings (vision-language) by correcting the modality gap using the ReAlign/ReVision technique. Fixes geometric misalignment between image and text embeddings from CLIP-like encoders via mean-shift, trace-scaling, and centroid correction — without retraining. Use when: 'align CLIP embeddings', 'fix modality gap', 'bridge vision-language representations', 'train MLLM with unpaired text data', 'improve cross-modal retrieval accuracy', 'reduce hallucination in multimodal LLM'.

ndpvt-web cdae785 15.9 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/modality-gap-driven-subspace-alignment commit cdae785fcf

Frequently asked questions

npx skillmds@latest add ndpvt-web/modality-gap-driven-subspace-alignment