AI Multimodal

Analyze images/audio/video with Gemini API (better vision than Claude). Generate images (Imagen 4), videos (Veo 3). Use for vision analysis, transcription, OCR, design extraction, multimodal AI. Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/ngxtm--devkit--ai-multimodal commit ba10a09374

Frequently asked questions

npx skillmds@latest add tomevault-io/ai-multimodal