Results for “image-text”

112 skills
More results
ecnu-icalk
baoyu-image-gen
Generates images via OpenAI, Google, DashScope, and Replicate APIs, supporting text prompts, reference images, aspect ratios, and quality presets.
559 · bundle
inference-sh
p-image
Generate images using Pruna's optimized P-Image models via the inference.sh CLI, supporting text-to-image, LoRA styles, image editing, and multi-image compositing.
584
sakamoto-family-smile
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
0
inference-sh
nano-banana
Generate images with Google Gemini native image models via the inference.sh CLI, supporting text-to-image, image editing, multi-image input, and various output options.
584
dvcrn
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
32 · bundle
jiachen-t-wang
coyo-700m-image-text-pair-dataset-github-kakaobrain-coyo-700
COYO-700M: Image-Text Pair Dataset
6
orchestra-research
stable-diffusion-image-generation
Generate images from text prompts, perform image-to-image translation, inpainting, and build custom diffusion pipelines using Stable Diffusion models via HuggingFace Diffusers.
10.4k · bundle
inference-sh
gpt-image
Generate and edit images using OpenAI's GPT-Image-2 model via the inference.sh CLI, supporting text-to-image, image editing, inpainting, and batch generation.
584
inference-sh
nano-banana-2
Generate images using Google Gemini 3.1 Flash Image Preview via the inference.sh CLI, with support for text-to-image, image editing, multi-image input, and Google Search grounding.
584
inference-sh
qwen-image-2-pro
Generate images with Alibaba Qwen-Image-2.0-Pro via inference.sh CLI, with professional text rendering and fine-grained realism for posters, banners, and text-heavy designs.
584
kbarbel640-del
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including queue management and file uploads.
1 · bundle
tianhao909
stable-diffusion-image-generation
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
1 · bundle
qcmuu
stable-diffusion-image-generation
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
0 · bundle
ichichuang
stable-diffusion-image-generation
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
0 · bundle
upayanghosh
synapse-image-describe
Provides detailed, structured image descriptions covering objects, people, colors, text, and scene context, with an overview and interpretation.
14
jimliu
baoyu-danger-gemini-web
Generates images and text via reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.
23.1k · bundle
github
generate-image
Generate images using AI from OpenAI or Google Gemini, with support for textures, icons, sprites, and visual assets.
36.2k
orchestra-research
clip
Enables zero-shot image classification, image-text matching, and cross-modal retrieval using OpenAI's CLIP model.
10.4k · bundle
runcomfy-com
flux-kontext
Edit images with Flux 1 Kontext Pro (Black Forest Labs' precise local image-edit model) on RunComfy — bundled with the model's documented prompting patterns so the skill gets sharper output than naive prompting against the same model. Documents Flux Kontext's strengths (single-reference precise local edits, strong prompt control, consistent high-fidelity outputs), the schema (single image + prompt), and when to route to Nano Banana Edit / GPT Image 2 edit / Flux 2 Klein instead. Calls `runcomfy run blackforestlabs/flux-1-kontext/pro/edit` through the local RunComfy CLI. Triggers on "flux kontext", "flux-kontext", "flux 1 kontext", "kontext", "BFL kontext", or any explicit ask to edit with this model.
12
antigravity
imagen
Generates images using Google Gemini's image generation model for UI placeholders, documentation, and design assets.
42.4k
peteedoo
clip
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
0 · bundle
phoroth
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design projects.
3
comeonoliver
imagegen
Generates or edits images for projects using the OpenAI Image API, with support for batch runs and structured prompt augmentation.
61
jiachen-t-wang
flamingo-a-visual-language-model-for-few-shot-learning-arxiv
Flamingo: A Visual Language Model for Few-Shot Learning
6
orchestra-research
blip-2-vision-language
Generate image captions, answer visual questions, and perform image-text retrieval using BLIP-2's Q-Former architecture with frozen vision encoders and LLMs.
10.4k · bundle
jimliu
baoyu-image-gen
Generates images using multiple AI providers including OpenAI, Google, Azure, and others. Supports text-to-image, reference images, aspect ratios, and batch generation from prompt files.
23.1k · bundle
kbarbel640-del
vgl
Generates structured VGL JSON prompts for Bria's FIBO image generation models, covering text-to-image, editing, inpainting, outpainting, and captioning with a deterministic schema.
1 · bundle
pwdev-solucoes
image-gen
Generates AI images via Ideogram, Leonardo, or Flux using plugin scripts, or delivers an optimized prompt when no API key is configured. Requires explicit cost confirmation before each paid generation.
2
tangchunwu
fal-ai-media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
1
oyi77
image-gen
Generates images from text prompts using diffusion models, covering prompt engineering, inpainting/outpainting, ControlNet, and API integration for production workflows.
10