What I do
- Build vision-language models
- Work with CLIP, BLIP, LLaVA
- Perform visual question answering
- Create multimodal chatbots
When to use me
When building systems that combine vision and language.
Key Concepts
- CLIP
- BLIP and BLIP-2
- LLaVA
- Visual instruction tuning
- Cross-modal attention
- Image-text matching
- Multimodal reasoning