Results for “image-text”

230 skills
doany-ai
gpt-image-2
Generate and edit images with OpenAI GPT Image 2 (ChatGPT Images 2.0) on RunComfy. Documents GPT Image 2's strengths (embedded text, logos, multilingual typography, instruction precision), its 3 fixed sizes, edit-with-preservation language, and when to route to a sibling (Flux 2 / Nano Banana Pro / Seedream) instead. Calls `runcomfy run openai/gpt-image-2/text-to-image` or `/edit` through the local RunComfy CLI. Triggers on "gpt image 2", "gpt-image-2", "ChatGPT Images 2", "image 2", or any explicit ask to generate or edit with this model.
5
qhjqhj00
visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
lingxling
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design assets.
253
jorcan
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design assets.
0 · bundle
jimliu
baoyu-image-gen
Generates images using multiple AI providers including OpenAI, Google, Azure, and others. Supports text-to-image, reference images, aspect ratios, and batch generation from prompt files.
23.1k · bundle
kbarbel640-del
web-search
Searches the web for text, news, images, videos, and books using the duckse CLI, returning clean results in text or JSON format.
1 · bundle
scoheart
xhs-publish
Publishes image-text, video, and long-form content to Xiaohongshu (Little Red Book) via a local CLI script, with support for scheduling, tags, and visibility settings.
2
mhassan0000
fal-3d
Generates 3D models from text or images via fal.ai, useful for game assets, AR previews, product mockups, and concept sculpting.
1
lucaspmarie-a11y
imagen
Generates images from text prompts using Google Gemini's image generation model, saving them as PNG files for use in UI, documentation, and design projects.
5
oimiragieo
threejs-loaders
Three.js asset loading - GLTF, textures, images, models, async patterns. Use when loading 3D models, textures, HDR environments, or managing loading progress.
0
samuraigpt
muapi-one-shot-video
Generate a single continuous cinematic shot video with no cuts, using text or image prompts and configurable style, duration, and aspect ratio.
3.7k
oyi77
video-gen
Generates short AI videos from text or images across Runway, Kling, Sora, Pika, Seedance, Grok Imagine, and Veo, with provider-specific APIs and prompt guidance.
10
samuraigpt
muapi-keyboard-art-maker
Generate artistic top-down photos of keyboard keycaps arranged to spell out custom text messages.
3.7k
nvidia
tao-finetune-clip
Fine-tune and deploy CLIP vision-language models for zero-shot classification, image-text retrieval, and embedding extraction with ONNX and TensorRT support.
2.2k · bundle
demerzels-lab
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation directly from the command line.
10 · bundle
microsoft
azure-ai-contentsafety-java
Analyze text and images for harmful content using Azure AI Content Safety SDK for Java. Supports hate, violence, sexual content, and self-harm detection with blocklist management.
2.7k · bundle
oyi77
image-gen
Generates images from text prompts using diffusion models, covering prompt engineering, inpainting/outpainting, ControlNet, and API integration for production workflows.
10
ecnu-icalk
baoyu-post-to-wechat
Publishes articles and image-text posts to a WeChat Official Account via browser automation or API, converting markdown to HTML and validating metadata.
559 · bundle
iterationlayer
generate-email-banner
Generate a personalized email banner image with text, logo, and brand colors.
2
ssrjkk
gradio
Creates ML demo interfaces with Gradio, supporting images, text, audio, and video inputs/outputs.
2 · bundle
sinhoneyy
textme
Text Claude from your phone — set up the njerschow/textme daemon so inbound iMessages drive a Claude Code session on your laptop, with voice notes, image input, code execution, and a phone-number whitelist.
11
m00sp
generate-image
Generate images using AI. Use when asked to generate, create, or make images, textures, icons, sprites, artwork, visual assets, or mockups. Supports OpenAI (gpt-image-2) and Google Gemini (Nano Banana). Requires an API key for the chosen provider.
0
inference-sh
og-image-design
Design Open Graph and social sharing images with platform-specific specs, text placement, and branding guidelines. Generate images via HTML-to-image or AI, and configure OG meta tags for Facebook, Twitter, LinkedIn, and more.
584
seaworld008
gpt-image2
Use when the user asks Codex to directly generate images with gpt-image-2 using inherited OpenAI/Codex-compatible environment credentials or local GPT_IMAGE2_* overrides, including text-to-image, reference-image guided generation, ratios, resolution, quality, variants, and saved local image files; run the bundled Node CLI and keep URL/sk configuration private.
65 · bundle
johnalbertini14-glitch
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
1 · bundle
inference-sh
happyhorse
Generate and edit videos using Alibaba HappyHorse 1.0 models via the inference.sh CLI, supporting text-to-video, image-to-video, reference-to-video, and video editing with natural language.
584
pwdev-solucoes
alt-text
Escreve texto alternativo acessível para peças de redes sociais, transmitindo a informação da imagem em vez de descrevê-la, com regras por tipo e formato de saída.
2
iterationlayer
generate-front-book-cover
Generate a front cover image with custom artwork, title text, and author attribution.
2
microsoft
azure-ai-vision-imageanalysis-py
Analyze images using Azure AI Vision SDK: generate captions, tags, detect objects, extract text (OCR), detect people, and suggest smart crops.
2.7k
composiohq
openai-automation
Automate OpenAI API operations: generate text and multimodal responses with structured output, create embeddings, generate images, and list models via the Composio MCP integration.
66.9k
tianhao909
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
1 · bundle
qcmuu
blip-2-vision-language
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
0 · bundle
inference-sh
p-video
Generate videos using Pruna's optimized P-Video and WAN models via the inference.sh CLI, supporting text-to-video, image-to-video, audio input, and multiple resolutions.
584
mineru98
imagine
Generate or edit images with Codex. Use this skill whenever the user says "imagine ...", asks to create an image from a text description, transform or restyle an existing image, produce artwork / illustrations / logos / concept art, make image variations, or asks for any kind of AI image generation or image-to-image editing. All outputs are saved inside the current project's `./images/` folder by default.
13 · bundle
prime-skills
ai-image-generation
Generate and edit images on RunComfy via the `runcomfy` CLI — a smart router across the full image-model catalog: FLUX 2 (Klein 9B/4B, Pro, Dev, Flash, Turbo, Max), Google Nano Banana 2 / Pro, OpenAI GPT Image 2, ByteDance Seedream 5 / 4-5 / 4-0 and Dreamina 4-0, Alibaba Qwen Image and Z-Image Turbo, Wan 2-7. Covers both text-to-image (t2i) and image-to-image / edit (i2i) endpoints — the skill picks the right model for the user's actual intent (typography precision, photoreal portraits, sub-second iteration, multi-reference brand styling, open-weights workflow) and ships each model's documented prompting patterns plus the minimal `runcomfy run` invoke. Triggers on "generate image", "make a picture", "text to image", "AI image", "make an image of …", "image to image", "i2i", or any explicit ask to create or restyle an image.
33
runcomfy-com
ai-image-generation
Generate and edit images on RunComfy via the `runcomfy` CLI — a smart router across the full image-model catalog: FLUX 2 (Klein 9B/4B, Pro, Dev, Flash, Turbo, Max), Google Nano Banana 2 / Pro, OpenAI GPT Image 2, ByteDance Seedream 5 / 4-5 / 4-0 and Dreamina 4-0, Alibaba Qwen Image and Z-Image Turbo, Wan 2-7. Covers both text-to-image (t2i) and image-to-image / edit (i2i) endpoints — the skill picks the right model for the user's actual intent (typography precision, photoreal portraits, sub-second iteration, multi-reference brand styling, open-weights workflow) and ships each model's documented prompting patterns plus the minimal `runcomfy run` invoke. Triggers on "generate image", "make a picture", "text to image", "AI image", "make an image of …", "image to image", "i2i", or any explicit ask to create or restyle an image.
12