Llava

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis. Use when this capability is needed.

tomevault-io f98ae07 2 files · 8.3 KB Updated

File contents

tomevault-io/skills-registry/tree/main/davila7--claude-code-templates--multimodal-llava commit f98ae079b1

Frequently asked questions

npx skillmds@latest add tomevault-io/llava