Results for “image-text”
112 skillsfal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
1
fal-ai-media
Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
226k
fal
Generate images, videos, audio, and more using fal.ai AI models, with support for text-to-image, image-to-video, text-to-speech, speech-to-text, image editing, upscaling, model search, workflow creation, and cost estimation.
567 · bundle
qwen-image-2
Generate and edit images using Alibaba Qwen-Image-2.0 models via the inference.sh CLI, with support for text-to-image, multi-image editing, and text rendering.
584
flux-image
Generate images using FLUX models via the inference.sh CLI, supporting text-to-image, image-to-image, and LoRA fine-tuning.
584
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
1
More results
baoyu-image-gen
Generates images via OpenAI, Google, DashScope, and Replicate APIs, supporting text prompts, reference images, aspect ratios, and quality presets.
559 · bundle
p-image
Generate images using Pruna's optimized P-Image models via the inference.sh CLI, supporting text-to-image, LoRA styles, image editing, and multi-image compositing.
584
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
0
nano-banana
Generate images with Google Gemini native image models via the inference.sh CLI, supporting text-to-image, image editing, multi-image input, and various output options.
584
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
32 · bundle
coyo-700m-image-text-pair-dataset-github-kakaobrain-coyo-700
COYO-700M: Image-Text Pair Dataset
6
stable-diffusion-image-generation
Generate images from text prompts, perform image-to-image translation, inpainting, and build custom diffusion pipelines using Stable Diffusion models via HuggingFace Diffusers.
10.4k · bundle
gpt-image
Generate and edit images using OpenAI's GPT-Image-2 model via the inference.sh CLI, supporting text-to-image, image editing, inpainting, and batch generation.
584
nano-banana-2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
qwen-image-2-pro
Generate images with Alibaba Qwen-Image-2.0-Pro via inference.sh CLI, with professional text rendering and fine-grained realism for posters, banners, and text-heavy designs.
584
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including queue management and file uploads.
1 · bundle
stable-diffusion-image-generation
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
1 · bundle
stable-diffusion-image-generation
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
0 · bundle
stable-diffusion-image-generation
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
0 · bundle
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
baoyu-danger-gemini-web
Generates images and text via reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.
23.1k · bundle
generate-image
Generate images using AI from OpenAI or Google Gemini, with support for textures, icons, sprites, and visual assets.
36.2k
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
flux-kontext
Edit images with Flux 1 Kontext Pro (Black Forest Labs' precise local image-edit model) on RunComfy — bundled with the model's documented prompting patterns so the skill gets sharper output than naive prompting against the same model. Documents Flux Kontext's strengths (single-reference precise local edits, strong prompt control, consistent high-fidelity outputs), the schema (single image + prompt), and when to route to Nano Banana Edit / GPT Image 2 edit / Flux 2 Klein instead. Calls `runcomfy run blackforestlabs/flux-1-kontext/pro/edit` through the local RunComfy CLI. Triggers on "flux kontext", "flux-kontext", "flux 1 kontext", "kontext", "BFL kontext", or any explicit ask to edit with this model.
12
imagen
Generates images using Google Gemini's image generation model for UI placeholders, documentation, and design assets.
42.4k
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design projects.
3
imagegen
Generates or edits images for projects using the OpenAI Image API, with support for batch runs and structured prompt augmentation.
61
flamingo-a-visual-language-model-for-few-shot-learning-arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
baoyu-image-gen
Generates images using multiple AI providers including OpenAI, Google, Azure, and others. Supports text-to-image, reference images, aspect ratios, and batch generation from prompt files.
23.1k · bundle
vgl
Generates structured VGL JSON prompts for Bria's FIBO image generation models, covering text-to-image, editing, inpainting, outpainting, and captioning with a deterministic schema.
1 · bundle
image-gen
Generates AI images via Ideogram, Leonardo, or Flux using plugin scripts, or delivers an optimized prompt when no API key is configured. Requires explicit cost confirmation before each paid generation.
2
fal-ai-media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
1
image-gen
Generates images from text prompts using diffusion models, covering prompt engineering, inpainting/outpainting, ControlNet, and API integration for production workflows.
10