Multimodal Ml

World-class guidance for building production multimodal AI systems spanning vision, language, audio/speech, and video. Use when training or serving models that combine modalities — contrastive vision-language (CLIP/SigLIP), Vision-Language Models / multimodal LLMs (ViT encoder + projector/connector + LLM, à la LLaVA/Flamingo), ASR/TTS (Whisper), audio encoders, video temporal modeling/frame sampling, image/video generation (latent diffusion, DiT, flow matching), cross-modal embeddings and multimodal RAG. Covers the shared-representation mental model, early/late fusion, native-multimodal vs bolt-on, staged training (pretrain → align → instruction-tune), the serving cost of variable-length visual tokens on context/KV-cache, multimodal evaluation/hallucination/grounding, and the multimodal anti-patterns. Reach for it whenever a model takes images, audio, or video as input or output — not pure-text LLMs.

sanjeevrg89 55d6e17 5 files · 41.9 KB Updated

File contents

sanjeevrg89/arete/tree/main/skills/multimodal-ml commit 55d6e171d3

Frequently asked questions

npx skillmds@latest add sanjeevrg89/multimodal-ml