Image & Video Generation Agent Skills
Image & Video Generation
189 skillsscientific-schematics
Create publication-quality scientific diagrams using AI generation with smart iterative refinement and quality review.
30.2k · bundle
infsh-cli
Run 250+ AI apps from the command line: generate images and videos, call LLMs, search the web, create 3D models, and automate Twitter posts.
584 · bundle
agent-tools
Run 250+ AI apps from the command line: generate images and videos, call LLMs, search the web, create 3D models, and automate Twitter posts.
584 · bundle
p-image
Generate images using Pruna's optimized P-Image models via the inference.sh CLI, supporting text-to-image, LoRA styles, image editing, and multi-image compositing.
584
p-video
Generate videos using Pruna's optimized P-Video and WAN models via the inference.sh CLI, supporting text-to-video, image-to-video, audio input, and multiple resolutions.
584
gpt-image
Generate and edit images using OpenAI's GPT-Image-2 model via the inference.sh CLI, supporting text-to-image, image editing, inpainting, and batch generation.
584
flux-image
Generate images using FLUX models via the inference.sh CLI, supporting text-to-image, image-to-image, and LoRA fine-tuning.
584
google-veo
Generate videos using Google Veo models via the inference.sh CLI, with support for multiple Veo versions and cinematic prompts.
584
happyhorse
Generate and edit videos using Alibaba HappyHorse 1.0 models via the inference.sh CLI, supporting text-to-video, image-to-video, reference-to-video, and video editing with natural language.
584
qwen-image-2
Generate and edit images using Alibaba Qwen-Image-2.0 models via the inference.sh CLI, with support for text-to-image, multi-image editing, and text rendering.
584
ai-podcast
Generate multi-person talking head podcast videos from scratch using AI — character creation, TTS, avatar animation, and video stitching.
584
nano-banana-2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
image-to-video
Convert still images to animated videos using the inference.sh CLI, with guidance on model selection, motion prompting, and camera movement.
584
p-video-avatar
Generate talking head avatar videos from a portrait image using the inference.sh CLI, with built-in TTS, multilingual support, and competitive pricing.
584
image-upscaling
Upscale and enhance images using Real-ESRGAN, Thera, FLUX Upscaler, and Topaz via the inference.sh CLI.
584
qwen-image-2-pro
Generate images with Alibaba Qwen-Image-2.0-Pro via inference.sh CLI, with professional text rendering and fine-grained realism for posters, banners, and text-heavy designs.
584
ai-image-generation
Generate images with 50+ AI models including GPT-Image-2, FLUX, Gemini, Grok, Seedream, and Reve via the inference.sh CLI.
584
ai-video-generation
Generate videos from text, images, or references using 40+ AI models via the inference.sh CLI.
584
ai-content-pipeline
Build multi-step AI content creation pipelines combining image, video, audio, and text using the inference.sh CLI.
584
talking-head-production
Create talking head videos with AI avatars, lipsync, and voiceover using the inference.sh CLI.
584
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
llava
Enables visual instruction tuning and image-based conversations using open-source vision-language models. Supports multi-turn image chat, visual question answering, and image understanding tasks.
10.4k · bundle
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
segment-anything-model
Segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks with zero-shot transfer.
10.4k · bundle
stable-diffusion-image-generation
Generate images from text prompts, perform image-to-image translation, inpainting, and build custom diffusion pipelines using Stable Diffusion models via HuggingFace Diffusers.
10.4k · bundle
higgsfield-soul-id
Train a personalized identity model on a person's face for use in Higgsfield's image and video generation.
518 · bundle
higgsfield-generate
Generate images, videos, 3D assets, and audio via the Higgsfield AI CLI, including Marketing Studio ads and Virality Predictor analysis.
518 · bundle
muapi-media-editing
Edit and enhance images and videos with AI via muapi.ai — prompt-based editing, upscaling, background removal, face swap, lipsync, video effects, and more.
3.7k · bundle
muapi-media-generation
Generate AI images, videos, music, and audio from the terminal via muapi.ai — supports 100+ models including Flux, Midjourney v7, Kling 3.0, Veo3, and Suno V5.
3.7k · bundle
muapi-workflow
Build, run, and visualize multi-step AI generation workflows by chaining image, video, and audio nodes into automated pipelines.
3.7k · bundle
muapi-ai-clipping
Turn a long video into viral-ready short clips with a single API call, including transcription, highlight ranking, and auto-crop.
3.7k · bundle
muapi-seedance-2
Generate cinematic videos with Seedance 2.0 using director-level prompts, camera grammar, and multi-mode generation across Chinese, Global, and VIP tiers.
3.7k · bundle
muapi-storyboard
Generate a sequence of keyframe images from a story premise using the MuAPI image generation service.
3.7k
muapi-music-video
Generates a short music video from a song theme by creating keyframes, animating them, and producing a matching soundtrack.
3.7k
muapi-nano-banana
Generates high-fidelity images using reasoning-driven prompts and structured creative briefs via muapi.ai.
3.7k · bundle
muapi-ai-fight-scene
Generate a high-cut-density action/fight scene by composing a 16-cell storyboard image and driving Seedance 2.0 image-to-video from that storyboard.
3.7k