Media Tag

AI image and video tagging, captioning, zero-shot classification, and semantic search with open-source + commercial-safe vision-language models: CLIP (MIT, OpenAI, zero-shot classification + text-image similarity, ViT-B/32 to ViT-L/14), SigLIP (Apache 2.0, Google, sigmoid loss, outperforms CLIP on most benchmarks), BLIP-2 (BSD-3-Clause, Salesforce, strong image captioning), LLaVA / LLaVA-NeXT / LLaVA-OneVision (Apache 2.0, open vision-language model for detailed description + VQA + video). Use when the user asks to auto-tag photos, generate alt text, caption an image, describe a video scene-by-scene, build a CLIP semantic-search index, classify images by free-form text labels ("cat vs dog vs car"), bulk-label a folder, generate WCAG alt text, do zero-shot classification, ask a VLM to describe what is happening in a video, or pick between CLIP/SigLIP/BLIP-2/LLaVA.

damionrashford Updated

File contents

damionrashford/media-os/tree/main/skills/media-tag commit 9389f88afe

Frequently asked questions

npx skillmds@latest add damionrashford/media-tag