Image & Video Generation Agent Skills
Image & Video Generation
189 skillstransformers-js
Run state-of-the-art machine learning models directly in JavaScript/TypeScript across browsers and server-side runtimes using Transformers.js.
10.8k · bundle
azure-ai-openai-dotnet
Integrate Azure OpenAI and OpenAI services in .NET applications for chat completions, embeddings, image generation, audio transcription, and assistants.
2.7k
azure-ai-vision-imageanalysis-java
Analyze images using Azure AI Vision SDK for Java, enabling captioning, OCR, object detection, tagging, and smart cropping.
2.7k · bundle
azure-ai-vision-imageanalysis-py
Analyze images using Azure AI Vision SDK: generate captions, tags, detect objects, extract text (OCR), detect people, and suggest smart crops.
2.7k
nv-reason-cxr
Runs chest X-ray reasoning smoke tests using the NV-Reason-CXR-3B model via local inference or a public Hugging Face Space API.
2.2k · bundle
nv-generate-mr
Generates synthetic body MRI volumes using NVIDIA's NV-Generate-CTMR rflow-mr model. Wraps the upstream diffusion inference pipeline with config staging, output validation, and NIfTI volume summarization.
2.2k · bundle
vss-summarize-video
Summarize recorded video clips using the LVS microservice with a VLM fallback, producing a narrative summary with timestamped events.
2.2k · bundle
nv-generate-ct-rflow
Generates synthetic CT volumes and masks using NVIDIA's rectified-flow pipeline for medical imaging research.
2.2k · bundle
nv-generate-mr-brain
Generates synthetic brain MRI volumes using NVIDIA's NV-Generate-CTMR workflow, with configurable modality and random seed.
2.2k · bundle
tao-train-visual-changenet
Trains, evaluates, exports, and runs inference for Visual ChangeNet models used in AOI defect detection, comparing image pairs for PASS/NO_PASS classification or change-segmentation masks.
2.2k · bundle
vss-deploy-dense-captioning
Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
2.2k · bundle
vss-deploy-detection-tracking-2d
Deploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
2.2k · bundle
vss-deploy-detection-tracking-3d
Deploy and operate the RTVI-CV-3D microservice for multi-camera 3D detection and tracking, supporting sample datasets, custom videos, and RTSP streams.
2.2k · bundle
gpt-image-2
Generates and edits images using GPT Image 2 across three modes: direct generation via OpenAI-compatible API, prompt engineering for host-native image tools, or pure prompt advisory. Includes 80+ structured templates for posters, UI mockups, product visuals, maps, slides, and more.
9.2k · bundle
videodb
Ingest, index, search, edit, and generate video and audio content from files, URLs, live streams, or desktop capture.
226k · bundle
fal-ai-media
Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
226k
imagen
Generates images using Google Gemini's image generation model for UI placeholders, documentation, and design assets.
42.4k
daily-gift
Decides whether a personalized gift should be created each day, then generates it as an interactive H5 page, AI image, or AI video through a five-stage creative pipeline.
42.4k
gemini-omni-flash-api
Generate and edit videos using the Gemini Omni Flash model: text-to-video, image-to-video, video editing, and turn-by-turn refinement via the official google-genai SDK.
3.8k · bundle
generate-image
Generate images using AI from OpenAI or Google Gemini, with support for textures, icons, sprites, and visual assets.
36.2k
resemble-detect
Detect AI-generated audio, images, video, and text, trace synthesis sources, apply watermarks, verify speaker identity, and analyze media intelligence using the Resemble AI platform.
36.2k · bundle
nano-banana-pro-openrouter
Generate or edit images via OpenRouter using the Gemini 3 Pro Image model, with support for prompt-only generation, single-image edits, and multi-image compositing at 1K/2K/4K resolutions.
36.2k · bundle
vision-analysis
Analyze, describe, and extract information from images using the MiniMax vision MCP tool, with modes for general description, OCR, UI review, chart data extraction, and object detection.
12.9k
gif-sticker-maker
Convert user photos into 4 animated GIF stickers in Funko Pop / Pop Mart style with customizable captions.
12.9k · bundle
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform from the terminal.
12.9k
heygen-automation
Automate AI video generation workflows: browse avatars and templates, generate personalized videos, track processing status, and retrieve shareable URLs via the Composio MCP integration.
66.9k
openai-automation
Automate OpenAI API operations: generate text and multimodal responses with structured output, create embeddings, generate images, and list models via the Composio MCP integration.
66.9k
claid-ai-automation
Automate Claid AI image editing and enhancement tasks through Composio's toolkit via Rube MCP.
66.9k
replicate-automation
Automate Replicate AI model operations: run predictions, upload files, inspect model schemas, list versions, and manage prediction history via the Composio MCP integration.
66.9k
dreamstudio-automation
Automate Dreamstudio image generation tasks through Composio's toolkit via Rube MCP, with dynamic tool discovery and connection management.
66.9k
all-images-ai-automation
Automate image generation, editing, and analysis tasks using All Images AI through the Rube MCP toolkit.
66.9k
baoyu-image-gen
Generates images using multiple AI providers including OpenAI, Google, Azure, and others. Supports text-to-image, reference images, aspect ratios, and batch generation from prompt files.
23.1k · bundle
baoyu-cover-image
Generates article cover images with customizable type, palette, rendering, text, and mood dimensions, supporting multiple aspect ratios and backends.
23.1k · bundle
baoyu-danger-gemini-web
Generates images and text via reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.
23.1k · bundle
baoyu-article-illustrator
Analyzes article structure, identifies positions requiring visual aids, and generates illustrations with consistent type, style, and palette.
23.1k · bundle
generate-image
Generate and edit high-quality images using OpenRouter's AI models including FLUX.2 Pro and Gemini 3.1 Flash Image Preview.
30.2k · bundle