Results for “visual-generation-language”
12 skillsMore results
Vgl
Generates structured VGL JSON for Bria FIBO models, giving deterministic control over objects, lighting, camera, composition, and style instead of natural language prompts.
1 · bundle
Tao Generate Image Grounding
Generates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
2.2k · bundle
Blip 2 Vision Language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
Video Claw
Generates complete AI videos through a 6-stage pipeline (script, character/scene design, storyboard, reference images, video generation, post-production) or one-shot pipelines for short videos, action transfer, and digital human dubbing, all running on local servers.
17 · bundle
Llava
Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
10.4k · bundle
Geminigen AI
Unified multimedia generation API for images, videos, and text-to-speech, replacing separate providers for a single workflow.
10
Video Gen
Generates short AI videos from text or images across Runway, Kling, Sora, Pika, Seedance, Grok Imagine, and Veo, with provider-specific APIs and prompt guidance.
10
AI Video Generation
Generate videos from text, images, or references using 40+ AI models via the inference.sh CLI.
584
Visual Consistency
Mantém a coerência visual entre peças geradas por IA usando modelo fixo, prompt base, seed e referência de estilo, com teste de coerência e biblioteca de prompts.
2
AI Media Generator
Generates high-quality prompts for AI image, video, and music generation platforms, with optional browser automation to submit them to target sites.
17 · bundle
Llava
Runs the open-source LLaVA vision-language model for image understanding, captioning, visual question answering, and multi-turn image conversations, including setup, inference, and training guidance.
2